Donatella Sciuto

dblp:31/1913 · DBLP profile ↗
← Back
228ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0001-9030-6940ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 196 · 6 first-author · 10 since 2021Software engineering, systems software and programming languages · 31 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 7Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Computer networks · 4 · 2 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Cognitive oracles: On-chain explainable machine learning
abstract
Blockchains offer a mechanism for executing code in a credible, transparent and uncensorable way. These characteristics lend themselves to the implementation of applications that guarantee different actors regarding the credible execution of mechanisms, including insurance and impartial certification. A significant challenge in this domain is enabling smart contracts to reliably access real-world state information. We propose a “cognitive” oracle architecture that makes complex and ambiguous conditions—such as identifying specific elements within an image—verifiable on-chain. Unlike traditional oracles that focus on scalar data feeds, our oracle integrates machine learning classifiers with SNARK proofs to provide trustworthy semantic evaluations. To mitigate inaccuracies and faults, the oracle also embeds a jury-based dispute resolution layer supported by verifiable machine learning explanations. We evaluate this oracle system in the context of a carbon credit application based on small-size neural network inferences. The results and cost analyses indicate that the economic viability of such an oracle depends on minimizing dispute frequencies and selecting optimal configurations. The inclusion of explanation methods adds a layer of game-theoretic complexity but also enhances the quality of the jury’s work, as well as overall transparency and trust.
Francesco Bruschi, Donatella Sciuto
Comput. Commun.3
2025 DDRoute: a Novel Depth-Driven Approach to the Qubit Routing Problem
abstract
In the Noisy Intermediate-Scale Quantum (NISQ) era, the topological constraints present in many of the currently available quantum devices pose a physical limit on the feasible interactions between qubits. To comply with such limitations, the compilation of quantum circuits requires solving the Qubit Routing Problem (QRP), by inserting SWAP operations among qubits. The State of the Art provides heuristic algorithms addressing this task, yet the depth of the output circuits is often incompatible with the current limits of quantum hardware. Therefore, we propose DDRoute, a novel heuristic algorithm to solve QRP, designed to reduce the depth overhead introduced by the routing process in the compiled circuits. Our experimental evaluation proves the efficiency of our approach, with a depth reduction of up to 70% with respect to the state-of-the-art routing procedures.
Alessandro Annechini, Marco Venere, Donatella Sciuto, Marco D. Santambrogio
DAC3
2025 A Decentralized Approach to Award Game Achievements
abstract
Blockchain technology has the potential to transform the gaming industry by enabling players to own in-game assets and receive NFTs or tokens as rewards for their accomplishments. However, one critical issue is determining how to ensure that achievement conditions are met (e.g., that the player completed level number 10). Current approaches either open cheating backdoors (e.g., if the client checks the conditions) or introduce centralization points and limitations (if a backend checks the condition). To overcome these problems a potential solution may seem to execute games directly on-chain, but the high computational cost (especially on Ethereum) and latency has made it impossible so far; however, the development of new technologies like proofs of computation enables new approaches. In this article we employ succinct proofs of computation to allow the player to prove to a smart contract on a blockchain that he reached (achieved) a certain state in an action game, that can then be played directly in the client; this allows to implement decentralized systems that can ensure fair and transparent rewarding. Moreover, we address the problem of proving that a certain game on the client has taken no more than a specific amount of time.
Francesco Bruschi, Donatella Sciuto, Tommaso Paulon, Andrea Marchesi
Distributed Ledger Technol. Res. Pract.2
2025 QUEKUF: An FPGA Union Find Decoder for Quantum Error Correction on the Toric Code
abstract
Quantum computing represents an exciting computing paradigm that promises to solve problems untractable for a classical computer. The main limiting factor for quantum devices is the noise impacting qubits, which hinders the superpolynomial speedup promise. Thus, although Quantum Error Correction (QEC) mechanisms are paramount, QEC demands high speed and low latency to scale quantum computations to real-life-sized problems. Within this context, hardware accelerators, such as Field Programmable Gate Arrays (FPGAs), represent a valuable approach to fulfilling QEC requirements. Nevertheless, the literature falls short in proposing solutions targeting the toric code, a type of quantum Low-Density Parity Check code capable of encoding two logical qubits, thus requiring fewer physical qubits. This manuscript presents QUEKUF , an FPGA-based QEC dataflow architecture dealing with the toric code. QUEKUF disposes of parallel processing units to spatially parallelize QEC, which a centralized controller orchestrates for data movement and operation decisions. We also provide a latency-oriented resource optimization model to identify the best theoretical configuration of QUEKUF that minimizes latency and optimizes resource requirements based upon high-level quantum parameters. Experimental results show that QUEKUF attains up to \(7.30\times\) speedup and \(81.51\times\) improvement in energy efficiency over a C++ implementation with error-free syndromes while keeping high accuracy.
Federico Valentino, Beatrice Branchini, Davide Conficconi, Donatella Sciuto, Marco D. Santambrogio
ACM Trans. Reconfigurable Technol. Syst.4
2025 Rock the QASBA: Quantum Error Correction Acceleration via the Sparse Blossom Algorithm on FPGAs
abstract
Quantum computing is a new paradigm of computation that exploits principles from quantum mechanics to achieve an exponential speedup compared to classical logic. However, noise strongly limits current quantum hardware, reducing achievable performance and limiting the scaling of the applications. For this reason, current noisy intermediate-scale quantum devices require Quantum Error Correction (QEC) mechanisms to identify errors occurring in the computation and correct them in real time. Nevertheless, the high computational complexity of QEC algorithms is incompatible with the tight time constraints of quantum devices. Thus, hardware acceleration is paramount to achieving real-time QEC. This work presents QASBA, an FPGA-based hardware accelerator for the Sparse Blossom Algorithm (SBA), a state-of-the-art decoding algorithm. After profiling the state-of-the-art software counterpart, we developed a design methodology for hardware development based on the SBA. We also devised an automation process to help users without expertise in hardware design in deploying architectures based on QASBA. We implement QASBA on different FPGA architectures and experimentally evaluate resource usage, execution time, and energy efficiency of our solution. Our solution attains up to \(25.05\times\) speedup and \(304.16\times\) improvement in energy efficiency compared to the software baseline.
Marco Venere, Beatrice Branchini, Davide Conficconi, Donatella Sciuto, Marco D. Santambrogio
ACM Trans. Reconfigurable Technol. Syst.4
2024 Decentralized Updates of IoT and Edge Devices
Francesco Bruschi, Marco Zanghieri, Michele Terziani, Donatella Sciuto
AINA (5)4
2024 An Untraceable Credential Revocation Approach Based on a Novel Merkle Tree Accumulator
abstract
As digital identity gains increasing importance, current centralized digital identity systems face significant limitations. Recent years have seen a shift towards decentralized and self-sovereign identity models, underpinned by verifiable credentials and cryptographic signatures. A critical challenge in these systems is credential revocation, presenting unique issues absent in centralized systems. In this context, blockchain (BC) systems offer a viable support for implementing revocation mechanisms without compromising the advantages of decentralized identity. However, they introduce storage and cost challenges. This work proposes a credential revocation approach utilizing a Merkle tree-based accumulator. Our accumulator enables a trade-off between anonymity and proof complexity, offers resistance to credential tracking, and relies solely on hash function collision resistance, making it quantum-safe. Additionally, we have integrated the accumulator within a broader identity management library and reported its performance on standard hardware.
Nacereddine Sitouah, Francesco Bruschi, Francesco Lorenzo Pallotta, Riccardo Mencucci, Donatella Sciuto
ICBC5
2023 Iris: Automatic Generation of Efficient Data Layouts for High Bandwidth Utilization
abstract
Optimizing data movements is becoming one of the biggest challenges in heterogeneous computing to cope with data deluge and, consequently, big data applications. When creating specialized accelerators, modern high-level synthesis (HLS) tools are increasingly efficient in optimizing the computational aspects, but data transfers have not been adequately improved. To combat this, novel architectures such as High-Bandwidth Memory with wider data busses have been developed so that more data can be transferred in parallel. Designers must tailor their hardware/software interfaces to fully exploit the available bandwidth. HLS tools can automate this process, but the designer must follow strict coding-style rules. If the bus width is not evenly divisible by the data width (e.g., when using custom-precision data types) or if the arrays are not power-of-two length, the HLS-generated accelerator will likely not fully utilize the available bandwidth, demanding even more manual effort from the designer. We propose a methodology to automatically find and implement a data layout that, when streamed between memory and an accelerator, uses a higher percentage of the available bandwidth than a naive or HLS-optimized design. We borrow concepts from multiprocessor scheduling to achieve such high efficiency.
Stephanie Soldavini, Donatella Sciuto, Christian Pilato
ASP-DAC2
2023 Optimizing the Use of Behavioral Locking for High-Level Synthesis
abstract
The globalization of the electronics supply chain requires effective methods to thwart reverse engineering and intellectual property (IP) theft. Logic locking is a promising solution, but there are many open concerns. First, even when applied at a higher level of abstraction, locking may result in significant overhead without improving the security metric. Second, optimizing a security metric is application-dependent and designers must evaluate and compare alternative solutions. We propose a metaframework to optimize the use of behavioral locking during the high-level synthesis (HLS) of IP cores. Our method operates on chip’s specification (before HLS) and it is compatible with all HLS tools, complementing industrial EDA flows. Our metaframework supports different strategies to explore the design space and to select points to be locked automatically. We evaluated our method on the optimization of differential entropy, achieving better results than random or topological locking: 1) we always identify a valid solution that optimizes the security metric, while topological and random locking can generate unfeasible solutions; 2) we minimize the number of bits used for locking up to more than 90% (requiring smaller tamper-proof memories); and 3) we make better use of hardware resources since we obtain similar overheads but with higher security metric.
Christian Pilato, Luca Collini, Luca Cassano, Donatella Sciuto, Siddharth Garg, Ramesh Karri
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 Faber: A Hardware/SoftWare Toolchain for Image Registration
abstract
Image registration is a well-defined computation paradigm widely applied to align one or more images to a target image. This paradigm, which builds upon three main components, is particularly compute-intensive and represents many image processing pipelines’ bottlenecks. State-of-the-art solutions leverage hardware acceleration to speed up image registration, but they are usually limited to implementing a single component. We present Faber, an open-source HW/SW CAD toolchain tailored to image registration. The Faber toolchain comprises HW/SW highly-tunable registration components, supports users with different expertise in building custom pipelines, and automates the design process. In this direction, Faber provides both default settings for entry-level users and latency and resource models to guide HW experts in customizing the different components. Finally, Faber achieves from 1.5× to 54× in speedup and from 2× to 177× in energy efficiency against state-of-the-art tools on a Xeon Gold.
Eleonora D'Arnese, Davide Conficconi, Emanuele Del Sozzo, Luigi Fusco, Donatella Sciuto, Marco D. Santambrogio
IEEE Trans. Parallel Distributed Syst.5
2022 High-level design methods for hardware security: is it the right choice? invited
abstract
Due to the globalization of the electronics supply chain, hardware engineers are increasingly interested in modifying their chip designs to protect their intellectual property (IP) or the privacy of the final users. However, the integration of state-of-the-art solutions for hardware and hardware-assisted security is not fully automated, requiring the amendment of stable tools and industrial toolchains. This significantly limits the application in industrial designs, potentially affecting the security of the resulting chips. We discuss how existing solutions can be adapted to implement security features at higher levels of abstractions (during high-level synthesis or directly at the register-transfer level) and complement current industrial design and verification flows. Our modular framework allows designers to compose these solutions and create additional protection layers.
Christian Pilato, Donatella Sciuto, Benjamin Tan 0001, Siddharth Garg, Ramesh Karri
DAC2
2022 A scalable decentralized system for fair token distribution and seamless users onboarding
abstract
Tokens are digital, transferable, and programmable assets and one of the most promising tools offered by blockchains. They could enable a wide range of applications, from down to earth to futuristic. One of the main issues in achieving wide adoption of tokens is onboarding: main platforms require users to deal with specific tools such as wallets, transaction fees, key generation, and storage. The most common solution to offer a familiar experience to naive users are custodial intermediaries, which have the important drawback of centralizing the process, and keeping users away from the advantages of self-sovereign assets control. In this paper, we present a process for the distribution of digital tokens to end users, exploiting “physical” objects for initial distribution. The process is aimed at making the onboarding of users not yet accustomed to blockchain tools and concepts as easy as possible, in a secure and decentralized way.
Francesco Bruschi, Manuel Tumiati, Vincenzo Rana, Mattia Bianchi, Donatella Sciuto
Blockchain Res. Appl.5
2022 A Comprehensive Methodology to Optimize FPGA Designs via the Roofline Model
abstract
With reconfigurable fabrics delivering increasing performance over the years, Field-Programmable Gate Arrays (FPGAs) are becoming an appealing solution for next-generation High-Performance Computing (HPC) systems. However, in order to gain traction among traditional von Neumann architectures, the optimization process of Field-Programmable Gate Array (FPGA) designs should be further abstracted to a higher level. In fact, while High-Level Synthesis (HLS) already provides a handy way to write FPGA code with common high-level languages, substantial effort and expertise are still required to optimize the resulting FPGA design for the underlying hardware. To overcome this problem, we propose a semi-automated performance optimization methodology based on a Hierarchical Roofline model for FPGAs. System-wide and applications-specific optimizations such as off-chip memory transfer and data locality optimizations are guided by the FPGA Roofline model whereas FPGA-specific optimizations are automatically searched by a Design Space Exploration (DSE) engine. We demonstrate the way this methodology allows to easily analyze and optimize to peak system performance a wide set of applications ranging from particle methods, wavefront algorithms, and sparse arithmetic computations. In addition, we prove that the integrated Design Space Exploration (DSE) engine achieves a 14.36x maximum speedup if compared to previous automated solutions in the literature.
Marco Siracusa, Emanuele Del Sozzo, Marco Rabozzi, Lorenzo Di Tucci, Samuel Williams 0001, Donatella Sciuto, Marco D. Santambrogio
IEEE Trans. Computers6
2022 On the Automation of Radiomics-Based Identification and Characterization of NSCLC
abstract
Proper detection and accurate characterization of Non-Small Cell Lung Cancer (NSCLC) are an open challenge in the imaging field. Biomedical imaging is fundamental in lung cancer assessment and offers the possibility of calculating predictive biomarkers impacting patients' management. Within this context, radiomics, which consists of extracting quantitative features from digital images, shows encouraging results for clinical applications, but the sub-optimal standardization of the procedure and the lack of definitive results are still a concern in the field. For these reasons, this work proposes the design and development of LuCIFEx, a fully-automated pipeline for non-invasive in-vivo characterization of NSCLC, aiming to speed up the analysis process and enable an early diagnosis of the tumor.LuCIFEx pipeline relies on routinely acquired [18F]FDG-PET/CT images for the automatic segmentation of the cancer lesion, allowing the computation of accurate radiomic features, then employed for cancer characterization through Machine Learning algorithms. The proposed multi-stage segmentation process can identify the lesion with a mean accuracy of 94.2±5.0%. Finally, the proposed data analysis pipeline demonstrates the potential of PET/CT features for the automatic recognition of lung metastases and NSCLC histological subtypes, while highlighting the main current limitations of the radiomic approach.
Eleonora D'Arnese, Guido Walter Di Donato, Emanuele Del Sozzo, Martina Sollini, Donatella Sciuto, Marco D. Santambrogio
IEEE J. Biomed. Health Informatics5
2021 A Framework for Customizable FPGA-based Image Registration Accelerators
abstract
Image Registration is a highly compute-intensive optimization procedure that determines the geometric transformation to align a floating image to a reference one. Generally, the registration targets are images taken from different time instances, acquisition angles, and/or sensor types. Several methodologies are employed in the literature to address the limiting factors of this class of algorithms, among which hardware accelerators seem the most promising solution to boost performance. However, most hardware implementations are either closed-source or tailored to a specific context, limiting their application to different fields. For these reasons, we propose an open-source hardware-software framework to generate a configurable architecture for the most compute-intensive part of registration algorithms, namely the similarity metric computation. This metric is the Mutual Information, a well-known calculus from the Information Theory, used in several optimization procedures. Through different design parameters configurations, we explore several design choices of our highly-customizable architecture and validate it on multiple FPGAs. We evaluated various architectures against an optimized Matlab implementation on an Intel Xeon Gold, reaching a speedup up to 2.86x, and remarkable performance and power efficiency against other state-of-the-art approaches.
Davide Conficconi, Eleonora D'Arnese, Emanuele Del Sozzo, Donatella Sciuto, Marco D. Santambrogio
FPGA4
2021 A privacy preserving identification protocol for smart contracts
abstract
If on one hand the possibility of using pseudonymous identities is an important feature of blockchains and smart contracts, on the other hand official identity can be required in some applications to comply with regulations such as Know Your Customer and Anti Money Laundering. These regulatory compliance issues are usually dealt with either through “custo-dial” approaches, which however neutralize the decentralization of these systems, or with “whitelisting” mechanisms, that present critical issues with regard to privacy. In this paper we propose a protocol that allows only potentially identifiable users to use a given decentralized applications. The system ensures that users are identifiable only by a competent authority, under certain conditions, and that in any other case they remain pseudonyms. To evaluate performance and costs, we present an Ethereum implementation of the protocol.
Francesco Bruschi, Tommaso Paulon, Vincenzo Rana, Donatella Sciuto
ISCC4
2021 ASSURE: RTL Locking Against an Untrusted Foundry
abstract
Semiconductor design companies are integrating proprietary intellectual property (IP) blocks to build custom integrated circuits (ICs) and fabricate them in a third-party foundry. Unauthorized IC copies cost these companies billions of dollars annually. While several methods have been proposed for hardware IP obfuscation, they operate on the gate-level netlist, i.e., after the synthesis tools embed most of the semantic information into the netlist. We propose ASSURE to protect hardware IP modules operating on the register-transfer level (RTL) description. The RTL approach has three advantages: 1) it allows designers to obfuscate IP cores generated with many different methods (e.g., hardware generators, high-level synthesis tools, and preexisting IPs); 2) it obfuscates the semantics of an IC before logic synthesis; and 3) it does not require modifications to EDA flows. We perform a cost and security assessment of ASSURE against state-of-the-art oracle-less attacks.
Christian Pilato, Animesh Basak Chowdhury, Donatella Sciuto, Siddharth Garg, Ramesh Karri
IEEE Trans. Very Large Scale Integr. Syst.3
2020 BNNsplit: Binarized Neural Networks for embedded distributed FPGA-based computing systems
abstract
In the past few years, Convolutional Neural Networks (CNNs) have seen a massive improvement, outperforming other visual recognition algorithms. Since they are playing an increasingly important role in fields such as face recognition, augmented reality or autonomous driving, there is the growing need for a fast and efficient system to perform the redundant and heavy computations of CNNs. This trend led researchers towards heterogeneous systems provided with hardware accelerators, such as GPUs and FPGAs. The vast majority of CNNs is implemented with floating-point parameters and operations, but from research, it has emerged that high classification accuracy can be obtained also by reducing the floating-point activations and weights to binary values. This context is well suitable for FPGAs, that are known to stand out in terms of performance when dealing with binary operations, as demonstrated in FINN, the state-of-the-art framework for building Binarized Neural Network (BNN) accelerators on FPGAs. In this paper, we propose a framework that extends FINN to a distributed scenario, enabling BNNs implementation on embedded multi-FPGA systems.
Giorgia Fiscaletti, Marco Speziali, Luca Stornaiuolo, Marco D. Santambrogio, Donatella Sciuto
DATE5
2020 A Decentralized System for Fair Token Distribution and Seamless Users Onboarding
abstract
Tokens are digital, transferable and programmable assets and one of the most promising tools offered by blockchains. They could enable a wide range of applications, from down to earth to futuristic. One of the main issues in achieving wide adoption of tokens is onboarding: main platforms require users to deal with specific tools such as wallets, transaction fees, key generation and storage.The most common solution to offer a familiar experience to naive users are custodial intermediaries, which have the important drawback of centralizing the process, and keeping users away from the intrinsic advantages of blockchains.In this paper we present a process for the distribution of digital tokens to end users, exploiting "physical" objects for initial distribution. The process is aimed at making onboarding of users not yet accustomed with blockchain tools and concepts maximally easy, in a secure and decentralized way.
Francesco Bruschi, Manuel Tumiati, Vincenzo Rana, Mattia Bianchi, Donatella Sciuto
ISCC5
2018 A Scalable FPGA Design for Cloud N-Body Simulation
abstract
The N-Body simulation process describes the evolution of a system of forces composed of N bodies, which may represent celestial objects, molecules, and so on. The most accurate algorithm for N-Body simulation, the All-Pairs method, is particularly compute intensive and software implementations on CPUs are inefficient in terms of performance and power consumption. An implementation on a hardware accelerator, such as an FPGA, would benefits in both these terms, exploiting a parallel execution at a relative low power profile. Moreover, it would also benefit faster methods with lower computational complexity, since many of them rely on the All-Pairs approach to approximate the calculation of forces. This work proposes a highly scalable, power efficient and high performance hardware architecture for the N-Body All-Pairs simulation problem. Our final implementation is able to scale up to systems with an arbitrary number of bodies thanks to a tiling approach that allows performance in the order of 13,441 MPairs/s, outperforming state of the art implementations on FPGA in terms of both pure performance, as well as performance per watt ratio. Finally, our design results to be more power efficient than Grape-8 ASIC.
Emanuele Del Sozzo, Marco Rabozzi, Lorenzo Di Tucci, Donatella Sciuto, Marco D. Santambrogio
ASAP4
2018 HLS Support for Polymorphic Parallel Memories
abstract
The importance of High-Level Languages in abstracting machine language to enhance productivity has been proved in many sectors, and has recently encouraged the spread of reconfigurable hardware for general purpose computing. At the same time, Field Programmable Gate Arrays (FPGAs) become popular for data-intensive applications, because they promise customized hardware accelerators and achieve high-performance with low power consumption. However, taking advantage of parallel accesses to the local memories of FPGAs remains difficult, as it currently requires application re-engineering. A solution to this challenge is PolyMem, an easy-to-use parallel memory. In this work, we investigate the implementation, integration, and performance of PolyMem for HLS applications. To this end, we present a novel open-source implementation of PolyMem, optimized for the Xilinx Design Suite. We further demonstrate the use of PolyMem for three different case studies, implemented using both the Vivado workflow with a Virtex-7 VC707, and the SDx workflow with a Kintex Ultrascale 3 ADM-PCIE. Finally, we provide a thorough empirical analysis of these three cases studies in terms of latency, hardware resources, and productivity. Our results demonstrate that PolyMem delivers the expected performance, while enhancing productivity at the cost of a small increase in resources.
Luca Stornaiuolo, Marco Rabozzi, Donatella Sciuto, Marco D. Santambrogio, Giulio Stramondo, Catalin Bogdan Ciobanu, Ana Lucia Varbanescu
VLSI-SoC3
2018 MARC: A Resource Consumption Modeling Service for Self-Aware Autonomous Agents
abstract
Autonomicity is a golden feature when dealing with a high level of complexity. This complexity can be tackled partitioning huge systems in small autonomous modules, i.e., agents. Each agent then needs to be capable of extracting knowledge from its environment and to learn from it, in order to fulfill its goals: this could not be achieved without proper modeling techniques that allow each agent to gaze beyond its sensors. Unfortunately, the simplicity of agents and the complexity of modeling do not fit together, thus demanding for a third party to bridge the gap. Given the opportunities in the field, the main contributions of this work are twofold: (1) we propose a general methodology to model resource consumption trends and (2) we implemented it into MARC, a Cloud-service platform that produces Models-as-a-Service, thus relieving self-aware agents from the burden of building their custom modeling framework. In order to validate the proposed methodology, we set up a custom simulator to generate a wide spectrum of controlled traces: this allowed us to verify the correctness of our framework from a general and comprehensive point of view.
Matteo Ferroni, Andrea Corna, Andrea Damiani, Rolando Brondolin, John Kubiatowicz, Donatella Sciuto, Marco D. Santambrogio
ACM Trans. Auton. Adapt. Syst.6
2018 BuildingRules: A Trigger-Action-Based System to Manage Complex Commercial Buildings
abstract
Modern Building Management Systems (BMSs) have been designed to automate the behavior of complex buildings, but unfortunately they do not allow occupants to customize it according to their preferences, and only the facility manager is in charge of setting the building policies. To overcome this limitation, we present BuildingRules, a trigger-action programming-based system that aims to provide occupants of commercial buildings with the possibility of specifying the characteristics of their office environment through an intuitive interface. Trigger-action programming is intuitive to use and has been shown to be effective in meeting user requirements in home environments. To extend this intuitive interface to commercial buildings, an essential step is to manage the system scalability as large number of users will express their policies. BuildingRules has been designed to scale well for large commercial buildings as it automatically detects conflicts that occur among user specified policies and it supports intelligent grouping of rules to simplify the policies across large numbers of rooms. We ensure the conflict resolution is fast for a fluid user experience by using the Z3 SMT solver. BuildingRules backend is based on RESTful web services so it can connect to various BMSs and scale well with large number of buildings. We have tested our system with 23 users across 17 days in a virtual office building, and the results we have collected prove the effectiveness and the scalability of BuildingRules.
A. A. Nacci, Vincenzo Rana, Bharathan Balaji, Paola Spoletini, Rajesh K. Gupta 0001, Donatella Sciuto, Yuvraj Agarwal
ACM Trans. Cyber Phys. Syst.6
2016 A polyhedral model-based framework for dataflow implementation on FPGA devices of iterative stencil loops
abstract
Iterative Stencil Loops (ISLs) are a specific class of algorithms of great importance for their substantial presence in a lot of industrial and scientific computing applications, such as in numerical methods for solving partial differential equation - e.g. reverse time migration and heat distribution simulation - or in cellular automata - used for instance for random number generation and error correction. In this work, we propose a hardware acceleration methodology based on the polyhedral model and implement the related framework to automatically accelerate ISLs on a multi-FPGA system. The experimental evaluation shows that the throughput obtained by our solution scales linearly with the amount of resources used on the FPGAs, the power efficiency increases proportionally to the amount of instantiated computation, and outperforms the power efficiency figure of state of the art ISL implementations running on an Intel Xeon CPU by at most 10×. A key aspect of this approach is also that no knowledge of the underlying architecture is requested to the application designer, as no code refactoring is needed to make the application suitable to be processed by our framework.
Giuseppe Natale, Giulio Stramondo, Pietro Bressana, Riccardo Cattaneo, Donatella Sciuto, Marco D. Santambrogio
ICCAD5
2016 On How to Accelerate Iterative Stencil Loops: A Scalable Streaming-Based Approach
abstract
In high-performance systems, stencil computations play a crucial role as they appear in a variety of different fields of application, ranging from partial differential equation solving, to computer simulation of particles’ interaction, to image processing and computer vision. The computationally intensive nature of those algorithms created the need for solutions to efficiently implement them in order to save both execution time and energy. This, in combination with their regular structure, has justified their widespread study and the proposal of largely different approaches to their optimization. However, most of these works are focused on aggressive compile time optimization, cache locality optimization, and parallelism extraction for the multicore/multiprocessor domain, while fewer works are focused on the exploitation of custom architectures to further exploit the regular structure of Iterative Stencil Loops (ISLs), specifically with the goal of improving power efficiency. This work introduces a methodology to systematically design power-efficient hardware accelerators for the optimal execution of ISL algorithms on Field-programmable Gate Arrays (FPGAs). As part of the methodology, we introduce the notion of Streaming Stencil Time-step (SST), a streaming-based architecture capable of achieving both low resource usage and efficient data reuse thanks to an optimal data buffering strategy, and we introduce a technique called SSTs queuing that is capable of delivering a pseudolinear execution time speedup with constant bandwidth. The methodology has been validated on significant benchmarks on a Virtex-7 FPGA using the Xilinx Vivado suite. Results demonstrate how the efficient usage of the on-chip memory resources realized by an SST allows one to treat problem sizes whose implementation would otherwise not be possible via direct synthesis of the original, unmanipulated code via High-Level Synthesis (HLS). We also show how the SSTs queuing effectively ensures a pseudolinear throughput speedup while consuming constant off-chip bandwidth.
Riccardo Cattaneo, Giuseppe Natale, Carlo Sicignano, Donatella Sciuto, Marco D. Santambrogio
ACM Trans. Archit. Code Optim.4
2016 State of the Journal
abstract
Discusses the current state of the journal, reports on current and future areas of exploration and research, and presents new editors.
Paolo Montuschi, Edward J. McCluskey, Samarjit Chakraborty, Jason Cong, Ramón M. Rodríguez-Dagnino, Fred Douglis, Lieven Eeckhout, Gernot Heiser, Sushil Jajodia, Ruby B. Lee, Dinesh Manocha, Tomás F. Pena, Isabelle Puaut, Hanan Samet, Donatella Sciuto
IEEE Trans. Computers15
2016 Efficient Hardware Design of Iterative Stencil Loops
abstract
A large number of algorithms for multidimensional signals processing and scientific computation come in the form of iterative stencil loops (ISLs), whose data dependencies span across multiple iterations. Because of their complex inner structure, automatic hardware acceleration of such algorithms is traditionally considered as a difficult task. In this paper, we introduce an automatic design flow that identifies, in a wide family of bidimensional data processing algorithms, subportions that exhibit a kind of parallelism close to that of ISLs; these are mapped onto a space of highly optimized ad-hoc architectures, which is efficiently explored to identify the best implementations with respect to both area and throughput. Experimental results show that the proposed methodology generates circuits whose performance is comparable to that of manually optimized solutions, and orders of magnitude higher than those generated by commercial high-level synthesis tools.
Vincenzo Rana, Ivan Beretta, Francesco Bruschi, A. A. Nacci, David Atienza 0001, Donatella Sciuto
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2016 Parallelizing the Chambolle Algorithm for Performance-Optimized Mapping on FPGA Devices
abstract
The performance and the efficiency of recent computing platforms have been deeply influenced by the widespread adoption of hardware accelerators, such as graphics processing units (GPUs) or field-programmable gate arrays (FPGAs), which are often employed to support the tasks of general-purpose processors (GPPs). One of the main advantages of these accelerators over their sequential counterparts (GPPs) is their ability to perform massive parallel computation. However, to exploit this competitive edge, it is necessary to extract the parallelism from the target algorithm to be executed, which generally is a very challenging task. This concept is demonstrated, for instance, by the poor performance achieved on relevant multimedia algorithms, such as Chambolle, which is a well-known algorithm employed for the optical flow estimation. The implementations of this algorithm that can be found in the state of the art are generally based on GPUs but barely improve the performance that can be obtained with a powerful GPP. In this article, we propose a novel approach to extract the parallelism from computation-intensive multimedia algorithms, which includes an analysis of their dependency schema and an assessment of their data reuse. We then perform a thorough analysis of the Chambolle algorithm, providing a formal proof of its inner data dependencies and locality properties. Then, we exploit the considerations drawn from this analysis by proposing an architectural template that takes advantage of the fine-grained parallelism of FPGA devices. Moreover, since the proposed template can be instantiated with different parameters, we also propose a design metric, the expansion rate, to help the designer in the estimation of the efficiency and performance of the different instances, making it possible to select the right one before the implementation phase. We finally show, by means of experimental results, how the proposed analysis and parallelization approach leads to the design of efficient and high-performance FPGA-based implementations that are orders of magnitude faster than the state-of-the-art ones.
Ivan Beretta, Vincenzo Rana, Abdulkadir Akin, A. A. Nacci, Donatella Sciuto, David Atienza 0001
ACM Trans. Embed. Comput. Syst.5
2015 Occupancy detection via iBeacon on Android devices for smart building management
Andrea Corna, L. Fontana, A. A. Nacci, Donatella Sciuto
DATE4
2015 Thermal-aware floorplanning for partially-reconfigurable FPGA-based systems
Davide Pagano, Mikel Vuka, Marco Rabozzi, Riccardo Cattaneo, Donatella Sciuto, Marco D. Santambrogio
DATE5
2015 Experimental Evaluation and Modeling of Thermal Phenomena on Mobile Devices
abstract
In the context of mobile devices, the thermal problem is an emerging one, as it affects the user experience and involves factors that are both internal and external with respect to the device. In this paper, we present an evaluation of these factors, that consists of two parts. The first one is the analysis of thermal interactions between the internal components of the system, performed with an infrared camera. The second part consists in the analysis of the impact of external temperature on the performance of CPUs and batteries. We finally propose the VirtIRCamera app, a thermal simulator for Android devices, able to generate thermal maps relying on the thermal model proposed within this paper. As a characterization of the thermal phenomenon, this work is the first step in the creation of thermal management techniques that are specifically designed for mobile devices.
Matteo Ferroni, A. A. Nacci, Matteo Turri, Marco D. Santambrogio, Donatella Sciuto
DSD5
2015 Methods and Algorithms for the Interaction of Residential Smart Buildings with Smart Grids
abstract
Smart buildings are essential for future smart grids. We propose a general method and algorithms to integrate residential smart buildings with smart grids. Such integration has to deal with the ability for the smart building to forecast its energy consumption, the proposed method is able to learn the building occupants habits and to use such information to forecast the energy consumption. Moreover, we present a method for selecting the most appropriate appliance usage scheduling given a energy reduction request coming from the smart grid.
Giovanni Bettinazzi, A. A. Nacci, Donatella Sciuto
EUC3
2015 OpenMPower: An Open and Accessible Database About Real World Mobile Devices
abstract
In the last decade we have witnessed the birth and dramatic growth of mobile devices, from cellular-to smart-phones. Despite the huge amount of information achievable from an always-connected reality, researchers that work in the mobile devices field fight against the impossibility to explore, inspect and test their work on such a vast set of possible environments, use case scenarios, hardware and software platforms the smart mobile world is composed of. This pushed the need of a wide open dataset of real world data coming from devices in their real usage context, properly anonymized and conveniently organized to be searchable and accessible. In this paper, we present a platform that brings such a dataset to researchers of the next generation of mobile devices.
Andrea Corna, Andrea Damiani, Matteo Ferroni, A. A. Nacci, Donatella Sciuto, Marco D. Santambrogio
EUC5
2014 On Power and Energy Consumption Modeling for Smart Mobile Devices
abstract
In nowadays life, mobile phones are becoming a cheaper and smaller alternative to laptops for simple, everyday tasks. They experienced an astonishing growth in functionalities and, because of their constant presence in our life, mobile phones became fundamental for the interaction with information coming from the environment. Nevertheless, their resources are limited, both in terms of performance and power, and their availability can greatly vary over time. Especially when dealing with power consumption, mobile devices cannot disregard environment conditions and user habits. Both internal and external conditions are rapidly changing and may influence the response of the entire system, e.g., switching between network types may causes an unpredictable power consumption. In order to puzzle out all these issues, we regard the definition of a power/energy model for mobile devices as a first mandatory step. In literature, several attempts to do so are present, basing their approaches on techniques coming from different computer science fields. They differ in the way they consider hardware components, in the operating system they are suitable for and in the scope of their tests and experiments. Within this paper, we categorize techniques presented in the major works in the field, in order to be able to compare different methods, highlight open issues and give suggestions on future works.
Matteo Ferroni, Andrea Cazzola, Francesco Trovò, Donatella Sciuto, Marco D. Santambrogio
EUC4
2014 cODA: An Open-Source Framework to Easily Design Context-Aware Android Apps
abstract
Mobile devices take an important part in everyday life. They are now cheaper and widespread, but still a lot of time is spent by the users to configure them: users adapt to their own device, not vice versa. Can our smart phones do something smarter? In this work, we propose a framework to support the development of context aware applications for Android devices: the goal of such applications is to reduce as much as possible the interaction with the user, making use of automatic and intelligent components. Moreover, these components should consume as less power and computational resources as possible, being them part of a mobile ecosystem whose battery and hardware are highly constrained. The work implies the study of a methodology that fits the Android framework and the design of a highly extensible software architecture. An open source framework based on the proposed methodology is then described. Some use cases are finally presented, analyzing the performances and the limitations of the proposed methodology.
Matteo Ferroni, Andrea Damiani, A. A. Nacci, Donatella Sciuto, Marco D. Santambrogio
EUC4
2014 A Perspective Vision on Complex Residential Building Management Systems
abstract
Smart buildings have been proposed as the solution for the creation of comfortable and energy-efficient living and working spaces. In the last few years, the energy-efficiency aspect is becoming more and more important and for this reason a lot of research effort has focused on the optimization of energy consumption, especially for what concerns commercial buildings, since they contribute to the 70% of the total energy consumption of an electrical grid. In addition, commercial buildings are usually already equipped with highly-instrumented distributed systems and infrastructures that simplify the creation of a smart environment. Nevertheless, we believe that also residential buildings should be taken into account, since their energy consumption is not negligible (they account for the remaining 30%) and the flexibility of their occupants in the usage of appliances (generally much higher than the ones of commercial buildings) could be better exploited. Within this context, we analyze the hardware infrastructure, intended as the network of sensors, actuators, appliances and accumulators, of this kind of buildings. Moreover, we discuss which kind of networking infrastructures is needed in order to cope with the particularities of complex residential buildings. Finally, we envision a layered architecture, from the aforementioned hardware layer to an envisioned application layer, providing also examples of what could be done, in the near future, to increase the capabilities of our living spaces.
A. A. Nacci, Vincenzo Rana, Donatella Sciuto
EUC3
2014 On How to Efficiently Implement Regular Expression Matching on FPGA-Based Systems
abstract
This work proposes a reconfigurable system able to perform - through a parallel and pipelined core, called ReCPU - regular expression matching. The system can configure on the programmable device, such as a FPGA, a set of ReCPUs, each one exploiting a single instance of the regular expression matching task on the given input string. These cores work in parallel on the same string analyzing different possible matching of the regular expression. Since the system is able to exploit dynamic partial reconfigurations, it can adapt at run-time the number of cores configured on the device, accordingly with the complexity of the regular expression. The adoption of the proposed solution makes it also possible to parallelize the regular expression matching process with a multiple cores architecture drastically reducing the time required for the completion of the task. Finally, run-time reconfiguration capabilities also allow to reduce the amount of resources required by the proposed approach.
Vincenzo Rana, Francesco Bruschi, Marco Paolieri, Donatella Sciuto, Marco D. Santambrogio
EUC4
2014 On How to Design Smart Energy-Efficient Buildings
abstract
Smart spaces are environments such as apartments, offices, museums, hospitals, schools, malls, university campuses, and outdoor areas that are enabled for the cooperation of objects (e.g., sensors, devices, appliances) and systems that have the capability to self-organize themselves, based on given policies.Since they can be used for an efficient management of the energy consumption of buildings, there is a growing interest for them, both in academia and industry.Unfortunately, nowadays, these systems are still designed manually with ad-hoc solutions.As a consequence, a huge effort has to be spent for each new smart building.Within this context, aim of this work is to propose a methodology to automate the design process of such smart spaces.The paper presents an overview of the design flow implemented to support the design of scalable architectures for energy aware smart spaces.
Donatella Sciuto, A. A. Nacci
EUC1
2014 Improving the security and the scalability of the AES algorithm (abstract only)
abstract
Although the reliability and robustness of the AES protocol have been deeply proved through the years, recent research results and technology advancements are rising serious concerns about its solidity in the (quite near) future. In fact, smarter brute force attacks and new computing systems are expected to drastically decrease the security of the AES protocol in the coming years (e.g., quantum computing will enable the development of search algorithms able to perform a brute force attack of a 2n-bit key in the same time required by a conventional algorithm for a n-bit key). In this context, we are proposing an extension of the AES algorithm in order to support longer encryption keys (thus increasing the security of the algorithm itself). In addition to this, we are proposing a set of parametric implementations of this novel extended protocols. These architectures can be optimized either to minimize the area usage or to maximize their performance. Experimental results show that, while the proposed implementations achieve a throughput higher than most of the state-of-the-art approaches and the highest value of the Performance/Area metric when working with 128-bit encryption keys, they can achieve a 84× throughput speedup when compared to the approaches that can be found in literature working with 512-bit encryption keys.
A. A. Nacci, Vincenzo Rana, Marco D. Santambrogio, Donatella Sciuto
FPGA4
2014 An Open-Source, Efficient, and Parameterizable Hardware Implementation of the AES Algorithm
abstract
Although the reliability and robustness of the AES protocol have been deeply proved through the years, recent research results and technology advancements are rising serious concerns about its solidity in the (quite near) future. In this context, we are proposing an extension of the AES algorithm in order to support longer encryption keys (thus increasing the security of the algorithm itself). In addition to this, we are proposing a set of parametric implementations of this novel extended protocols. These architectures can be optimized either to minimize the area usage or to maximize their performance. Experimental results show that, while the proposed implementations achieve a throughput higher than most of the state-of-the-art approaches and the highest value of the Performance/Area metric when working with 128-bit encryption keys, they can achieve a 84× throughput speed-up when compared to the approaches that can be found in literature working with 512-bit encryption keys.
A. A. Nacci, Vincenzo Rana, Donatella Sciuto, Marco D. Santambrogio
ISPA3
2014 A Survey on Recent Hardware and Software-Level Cache Management Techniques
abstract
Multi and many-core processors have emerged as the dominant solution for processing in the whole range of computer system, from small devices to large-scale installations. Chip multi-processors, which are homogeneous, multi and manycore processors, offer an unprecedented amount of on-chip, shared resources and brings a unique set of challenges. Given the importance of the Last-Level Cache management techniques to achieve near-perfect isolation, we survey the state of the art and propose research directions to address the most pressing issues in modern computer systems. To better understand the various research directions in the field, we propose a classification of the presented techniques. Finally, we discuss possible research directions.
Alberto Scolari, Filippo Sironi, Donatella Sciuto, Marco D. Santambrogio
ISPA3
2014 FPGA-Based Design Using the FASTER Toolchain: The Case of STM Spear Development Board
abstract
Even though FPGAs are becoming more and more popular as they are used in many different scenarios like communications and HPC, the steep learning curve needed to work with this technology is still the major limiting factor to their full success. Many works proposed to mitigate this problem by creating a companion of tools to support the designer during the development phase for this technology. The EU FASTER Project aims at realizing an integrated toolchain that assists the designer in the steps of the design flow that are necessary to port a given application onto an FPGA device. The novelty of the framework relies in the fact that the partial dynamic reconfiguration, which FPGA devices can exploit, is seen as a first class citizen throughout the whole design flow. This work reports a case study in which the FASTER toolchain has been used to port a raytracer application onto the STM Spear prototyping embedded platform. The paper discusses the steps done for the realization of the prototype and the results obtained on the target device. It finally reports some improvements that can be exploited to improve the performance of the hardware implementation that has been realized.
Fabrizio Spada, Alberto Scolari, Gianluca Durelli, Riccardo Cattaneo, Marco D. Santambrogio, Donatella Sciuto, Dionisios N. Pnevmatikatos, Georgi Gaydadjiev, Oliver Pell, Andreas Brokalakis, Wayne Luk, Dirk Stroobandt, Danilo Pau
ISPA6
2014 Automated Fine-Grained CPU Provisioning for Virtual Machines
abstract
Ideally, the pay-as-you-go model of Infrastructure as a Service (IaaS) clouds should enable users to rent just enough resources (e.g., CPU or memory bandwidth) to fulfill their service level objectives (SLOs). Achieving this goal is hard on current IaaS offers, which require users to explicitly specify the amount of resources to reserve; this requirement is nontrivial for users, because estimating the amount of resources needed to attain application-level SLOs is often complex, especially when resources are virtualized and the service provider colocates virtual machines (VMs) on host nodes. For this reason, users who deploy VMs subject to SLOs are usually prone to overprovisioning resources, thus resulting in inflated business costs. This article tackles this issue with AutoPro : a runtime system that enhances IaaS clouds with automated and fine-grained resource provisioning based on performance SLOs. Our main contribution with AutoPro is filling the gap between application-level performance SLOs and allocation of a contended resource, without requiring explicit reservations from users. In this article, we focus on CPU bandwidth allocation to throughput-driven, compute-intensive multithreaded applications colocated on a multicore processor; we show that a theoretically sound, yet simple, control strategy can enable automated fine-grained allocation of this contended resource, without the need for offline profiling. Additionally, AutoPro helps service providers optimize infrastructure utilization by provisioning idle resources to best-effort workloads, so as to maximize node-level utilization. Our extensive experimental evaluation confirms that AutoPro is able to automatically determine and enforce allocations to meet performance SLOs while maximizing node-level utilization by supporting batch workloads on a best-effort basis.
Davide B. Bartolini, Filippo Sironi, Donatella Sciuto, Marco D. Santambrogio
ACM Trans. Archit. Code Optim.3
2014 A Mapping-Scheduling Algorithm for Hardware Acceleration on Reconfigurable Platforms
abstract
Reconfigurable platforms are a promising technology that offers an interesting trade-off between flexibility and performance, which many recent embedded system applications demand, especially in fields such as multimedia processing. These applications typically involve multiple ad-hoc tasks for hardware acceleration, which are usually represented using formalisms such as Data Flow Diagrams (DFDs), Data Flow Graphs (DFGs), Control and Data Flow Graphs (CDFGs) or Petri Nets. However, none of these models is able to capture at the same time the pipeline behavior between tasks (that therefore can coexist in order to minimize the application execution time), their communication patterns, and their data dependencies. This article proves that the knowledge of all this information can be effectively exploited to reduce the resource requirements and the timing performance of modern reconfigurable systems, where a set of hardware accelerators is used to support the computation. For this purpose, this article proposes a novel task representation model, named Temporal Constrained Data Flow Diagram (TCDFD), which includes all this information. This article also presents a mapping-scheduling algorithm that is able to take advantage of the new TCDFD model. It aims at minimizing the dynamic reconfiguration overhead while meeting the communication requirements among the tasks. Experimental results show that the presented approach achieves up to 75% of resources saving and up to 89% of reconfiguration overhead reduction with respect to other state-of-the-art techniques for reconfigurable platforms.
Juan Antonio Clemente, Ivan Beretta, Vincenzo Rana, David Atienza 0001, Donatella Sciuto
ACM Trans. Reconfigurable Technol. Syst.5
2013 ThermOS: System support for dynamic thermal management of chip multi-processors
abstract
Constraining the temperature of computing systems has become a dominant aspect in the design of integrated circuits. The supply voltage decrease has lost its pace even though the feature size is shrinking constantly. This results in an increased number of transistors per unit of area and hence a growing power density. Researchers started investigating dynamic thermal management techniques to address the tradeoff between performance and temperature. Hardware dynamic thermal management can guarantee safety but, at the same time, can negatively affect established service-level agreements. On the other hand, software solutions rely on hardware for safety but does not indiscriminately trade-off performance for temperature. We propose ThermOS, an extension for commodity operating systems that harnesses formal feedback control and idle cycle injection to decrease thermal emergencies while showing better efficiency than commodity and cutting edge techniques.
Filippo Sironi, Martina Maggio, Riccardo Cattaneo, Giovanni F. Del Nero, Donatella Sciuto, Marco D. Santambrogio
PACT5
2013 Towards a performance-as-a-service cloud
abstract
Motivation While the pay-as-you-go model of Infrastructure-as-a-Service (IaaS) clouds is more flexible than an in-house IT infrastructure, it still has a resource-based interface towards users, who can rent virtual computing resources over relatively long time scales. There is a fundamental mismatch between this resource-based interface and what users really care about: performance.
Davide B. Bartolini, Filippo Sironi, Martina Maggio, Gianluca Durelli, Donatella Sciuto, Marco D. Santambrogio
SoCC5
2013 Coloring the cloud for predictable performance
abstract
Motivation and Contribution The commodity multicores that power cloud infrastructures hide memory latency through deep memory hierarchies, with the last-level cache (LLC) usually shared among cores. While a shared LLC improves utilization of on-chip resources, it may also lead to unpredictable performance of colocated virtual machines (VMs) as a result of unanticipated contention. Past research showed that the operating system page allocator can favor performance predictability on a physically-addressed shared LLC through page coloring [4, 8, 9]: a software technique that can work on commodity multicores, unlike hardware approaches [2, 7]. The main drawback of page coloring is the high cost of modifying allocations (i.e., recoloring), making this technique almost impractical for applications with varying memory footprints [6].
Alberto Scolari, Filippo Sironi, Davide B. Bartolini, Donatella Sciuto, Marco D. Santambrogio
SoCC4
2013 A high-level synthesis flow for the implementation of iterative stencil loop algorithms on FPGA devices
abstract
The automatic generation of hardware implementations for a given algorithm is generally a difficult task, especially when data dependencies span across multiple iterations such as in iterative stencil loops (ISLs). In this paper, we introduce an automatic design flow to extract parallelism from an ISL algorithm and perform a design space exploration to identify its best FPGA hardware implementation, in terms of both area and throughput. Experimental results show that the proposed methodology generates hardware designs whose performance is comparable to the one of manually-optimized solutions, and orders of magnitude higher than the implementations generated by commercial high-level synthesis tools.
A. A. Nacci, Vincenzo Rana, Francesco Bruschi, Donatella Sciuto, Ivan Beretta, David Atienza 0001
DAC4
2013 Morphone.OS: Context-Awareness in Everyday Life
abstract
Mobile devices, due to their wide distribution and to their increasing smartness and availability of computational power, can become the interaction point between users and their surrounding environments. However, current mobile devices OSes lack of the ability to anticipate and overcome internal and external changes. Integrating mechanisms of self-awareness and self-adaptability in nowadays smartphones is an attractive perspective to match with these requirements. Moreover, adaptive behaviors can enhance the management by the mobile device itself, of the available resources at its best, e.g., the battery life. This paper envisions various situations in which a self-aware mobile device can interact with the surrounding environment and support the user in performing everyday actions. A prototype of such an adaptive device, called morphone.os and based on the Android OS, has been designed and implemented to verify the reaction of the device in different situations providing convincing and promising preliminary results.
A. A. Nacci, Matteo Mazzucchelli, Martina Maggio, Alessandra Bonetto, Donatella Sciuto, Marco D. Santambrogio
DSD5
2013 SMASH: A heuristic methodology for designing partially reconfigurable MPSoCs
abstract
The exploitation of the capabilities offered by reconfigurable architectures is traditionally a demanding task due to the intrinsic time consuming and error prone customization of these systems around the specific application. Moreover, existing approaches are not able to integrate the notion of partial and dynamic reconfiguration (PDR) from the early stages of the decision phases, potentially leading to sub-optimal solutions. In this work, we propose SMASH (Simultaneous Mapping and Scheduling with Heuristics), a highly automated design methodology focused on explicitly taking into account PDR during the design of reconfigurable designs. It combines heuristics for both the design of the architecture and the mapping and scheduling of the partitioned application. We show how this additional degree of freedom leads to architectures whose performance are improved with respect to the baseline.
Riccardo Cattaneo, Christian Pilato, Gianluca Durelli, Marco D. Santambrogio, Donatella Sciuto
RSP5
2013 Adaptive and Flexible Smartphone Power Modeling
A. A. Nacci, Francesco Trovò, Filippo Maggi, Matteo Ferroni, Andrea Cazzola, Donatella Sciuto, Marco D. Santambrogio
Mob. Networks Appl.6
2012 B2IRS: A Technique to Reduce BAN-BAN Interferences in Wireless Sensor Networks
abstract
Wireless Sensor Networks (WSNs) are particular networks characterized by limited energy and computational resources and, if their transmission range is limited to a person's area, they are known as Body Area Networks (BANs). When two or more BANs are co-located and operate on the same channel, active periods can overlap and transmissions can conflict. This phenomenon, that drastically reduces performances and reliability of BANs, is known as BAN-BAN interference. In order to solve this issue, it is possible to employ techniques such as channel switching. However, channel switching is not suitable if the amount of channels is lower than the amount co-located BANs and, considering that interferences of Zigbee/802.15.4 networks with other technologies like 802.11 or Bluetooth reduce the amount of channels available for communication, an alternative approach is required. This paper introduces a BAN-BAN Interference Reduction System (B2IRS) which reschedules beacon packets in order to avoid active period overlap, reducing the interferences between distinct BANs. This approach is complementary to channel switching since it works with single-channel interferences, that arise when no free channels are available anymore. Experimental results, conducted comparing our methodology with the original IEEE 802.15.4, show that B2IRS is able to effectively reduce BAN-BAN interference, making it possible to almost maintain the same performance and energy consumption of an ideal situation (without interferences).
Paolo Roberto Grassi, Vincenzo Rana, Ivan Beretta, Donatella Sciuto
BSN4
2012 Metronome: operating system level performance management via self-adaptive computing
abstract
In this paper, we present Metronome: a framework to enhance commodity operating systems with self-adaptive capabilities. The Metronome framework features two distinct components: Heart Rate Monitor (HRM) and Performance--Aware Fair Scheduler (PAFS). HRM is an active monitoring infrastructure implementing the observe phase of a self--adaptive computing system Observe--Decide--Act (ODA) control loop, while PAFS is an adaptation policy implementing the decide and act phases of the control loop. Metronome was designed and developed looking towards multi--core processors; therefore, its experimental evaluation has been carried on with the PARSEC 2.1 benchmark suite.
Filippo Sironi, Davide B. Bartolini, Simone Campanoni, Fabio Cancare, Henry Hoffmann, Donatella Sciuto, Marco D. Santambrogio
DAC6
2012 An adaptive approach for online fault management in many-core architectures
abstract
This paper presents a dynamic scheduling solution to achieve fault tolerance in many-core architectures. Triple Modular Redundancy is applied on the multi-threaded application to dynamically mitigate the effects of both permanent and transient faults, and to identify and isolate damaged units. The approach targets the best performance, while balancing the use of the healthy resources to limit wear-out and aging effects, which cause permanent damages. Experimental results on synthetic case studies are reported, to validate the ability to tolerate faults while optimizing performance and resource usage.
Cristiana Bolchini, Antonio Miele, Donatella Sciuto
DATE3
2012 On the Development of a Runtime Reconfigurable Multicore System-on-Chip
abstract
Over the last years, several research groups have built reconfigurable systems to obtain high performance at low cost by specializing the computing engine to the computation task. Nowadays, FPGA-based multi-core architectures and reconfigurable computing are widely used for embedded systems, even if the development of complete and efficient solutions on this kind of devices is still quite a complex task. Within this context, what seems to be neglected so far is the combination of a multicore architecture with reconfigurable abilities to vary at runtime not only the hardware components but also the number of the available processors. The variation of the number of processors available on the device can be performed in a dynamic way by using the proposed solution, based on a partial bitstream, characterized by the presence of a reconfigurable system in which both components and component memories can be reconfigured at run-time. This paper presents a study of the viability of making a scalable and flexible multicore System-on-Chip (MPSoC) based on customizable reconfigurable processors, called Multi-Adaptive Reconfigurable Core (MARC), providing the communication infrastructures and the memory management required to create such a complex system-on-chip.
Andrea Cazzaniga, Gianluca Durelli, Christian Pilato, Donatella Sciuto, Marco D. Santambrogio
DSD4
2012 Tacit Consent: A Technique to Reduce Redundant Transmissions from Spatially Correlated Nodes in Wireless Sensor Networks
abstract
This paper introduces Tacit Consent (TaCo), a technique that exploits spatial correlation in Wireless Sensor Networks in order to reduce energy consumption while maintaining a very high accuracy on the measured data. In fact, nodes densely deployed in a field of interest are proved to sense highly correlated data, thus this intrinsic redundancy can be exploited to predict the neighbors' measurements by defining custom estimation functions, which aim at replicating the relationship between the data sensed by two nodes. To exploit spatial correlation, TaCo splits the nodes into two groups: representative and member nodes. Representative nodes directly transmit their measurements, while member nodes use overhearing to understand whether additional information is required. If the estimation function correctly predicts the measurements of the member nodes, they tacitly consent the estimation without performing any transmission. Differently from the other state-of-the-art approaches, such as the YEAST algorithm, TaCo can be used on top of existing routing and clustering protocols. Experimental results prove that TaCo is able to drastically reduce the energy consumption of the network when a high precision of the measured data is required.
Paolo Roberto Grassi, Vincenzo Rana, Ivan Beretta, Donatella Sciuto
DSD4
2012 Energy-Aware FPGA-based Architecture for Wireless Sensor Networks
abstract
Wireless sensor networks (WSNs) are networks of battery-powered sensing devices connected with wireless interfaces. Energy consumption and processing efficiency are relevant characteristics for these systems, thus energy-efficient architectures are required. Recent works show that FPGAs are suitable candidates for efficient data signal processing in WSNs. In this work, we evaluate Flash-based FPGA technology for WSNs' applications, and we present FPGA-based architecture for data signal processing in WSN's nodes. Low static power consumption of Flash-FPGAs is a fundamental characteristic in WSNs systems, which are characterized by long idle periods. SRAM-based and Flash-based FPGAs were compared, and an architecture to manage dynamic energy consumption in Flash-based FPGAs is presented. The architecture has been validated and evaluated on a real testbed, and power consumption results are presented. In our experiments, the overall power consumption of the FPGA is kept below 4mW.
Paolo Roberto Grassi, Donatella Sciuto
DSD2
2012 FASTER: Facilitating Analysis and Synthesis Technologies for Effective Reconfiguration
abstract
The FASTER project aims to ease the definition, implementation and use of dynamically changing hardware systems. Our motivation stems from the promise reconfigurable systems hold for achieving better performance and extending product functionality and lifetime via the addition of new features that work at hardware speed. This is a clear advantage over the more straightforward software component adaptivity. However, designing a changing hardware system is both challenging and time consuming. The FASTER project will facilitate the use of reconfigurable technology by providing a complete methodology that enables designers to easily specify, analyse, implement and verify applications on platforms with general-purpose processors and acceleration modules implemented in the latest reconfigurable technology. To better adapt to different application requirements, the tool-chain will support both region-based and micro-reconfiguration and provide a flexible run-time system that will efficiently manage the reconfigurable resources. We will use applications from the embedded, high performance computing, and desktop domains to demonstrate the potential benefits of the FASTER tools on metrics such as performance, power consumption and total ownership cost.
Dionisios N. Pnevmatikatos, Tobias Becker, Andreas Brokalakis, Karel Bruneel, Georgi Gaydadjiev, Wayne Luk, Kyprianos Papademetriou, Ioannis Papaefstathiou, Oliver Pell, Christian Pilato, M. Robart, Marco D. Santambrogio, Donatella Sciuto, Dirk Stroobandt, Tim Todman
DSD13
2012 An open-source design and validation platform for reconfigurable systems
abstract
Reconfigurable computing is a hot topic for research, as the possibilities and the technology offered by the reconfigurable devices improve year after year both in terms of available configurable logic resources and the possibilities offered to exploit them. This has led CAD tools to grow both in complexity and effectiveness. The expertise required to develop and test a complete system-on-chip using vendors tools has subsequently increased, forcing some designers to create their own tools as support to official development flows. Within this field quite few works have been developed, with respect to the huge effort that has been spent in the exploitation of architectural designs. ReBit is an open-source tool able to help the designer in exploring different placement solutions in the architecture refinement process and in testing the correct execution of an application on a real device.
Alessandra Bonetto, Andrea Cazzaniga, Gianluca Durelli, Christian Pilato, Donatella Sciuto, Marco D. Santambrogio
FPL5
2012 On the automatic integration of hardware accelerators into FPGA-based embedded systems
abstract
This paper proposes an automatic framework for the seamless integration of hardware accelerators, starting from an OpenMP-based application and an XML file describing the HW/SW partitioning. It extends a fully software architecture by generating and integrating the cores, along with the proper interfaces, and the code for scheduling and synchronization. Experimental results show that it is possible to validate different solutions only by varying the input code.
Christian Pilato, Andrea Cazzaniga, Gianluca Durelli, Andrés Otero, Donatella Sciuto, Marco D. Santambrogio
FPL5
2012 On the Evolution of Hardware Circuits via Reconfigurable Architectures
abstract
Traditionally, hardware circuits are realized according to techniques that follow the classical phases of design and testing. A completely new approach in the creation of hardware circuits has been proposed---the Evolvable Hardware (EHW) paradigm, which bases the circuit synthesis on a goal-oriented evolutionary process inspired by biological evolution in Nature. FPGA-based approaches have emerged as the main architectural solution to implement EHW systems. Various EHW systems have been proposed by researchers but most of them, being based on outdated chips, do not take advantage of the interesting features introduced in newer FPGAs. This article describes a project named Hardware Evolution over Reconfigurable Architectures (HERA), which aims at creating a complete and performance-oriented framework for the evolution of digital circuits, leveraging the reconfiguration technology available in FPGAs. The project is described from its birth to its current state, presenting its evolutionary technique tailored for FPGA-based circuits and the most recent enhancements to improve the scalability with respect to problem size. The developed EHW system outperforms the state of the art, proving its effectiveness in evolving both standard benchmarks and more complex real-world applications.
Fabio Cancare, Davide B. Bartolini, Matteo Carminati, Donatella Sciuto, Marco D. Santambrogio
ACM Trans. Reconfigurable Technol. Syst.4
2011 An efficient Quantum-Dot Cellular Automata adder
abstract
This paper presents a ripple-carry adder module that can serve as a basic component for Quantum Dot Automata arithmetic circuits. The main methodological design innovation over existing state of the art solutions was the adoption of so called minority gates in addition to the more traditional majority voters. Exploiting this widened basic block set, we obtained a more compact, and thus less expensive circuit. Moreover, the layout was designed in order to comply with the rules for robustness again noise paths [6].
Francesco Bruschi, Francesco Perini, Vincenzo Rana, Donatella Sciuto
DATE4
2011 A Hybrid Mapping-Scheduling Technique for Dynamically Reconfigurable Hardware
abstract
Reconfigurable computing is a promising technology that offers an interesting trade-off between flexibility and performance, which many recent multi-core embedded system applications demand. In order to achieve these objectives, it is necessary to optimize the deployment of the hardware cores on the FPGA platform, trying to reduce the reconfiguration overhead while meeting the desired performance. In this paper, we propose a hybrid mapping and scheduling technique for multi-core applications on reconfigurable devices, which exploits the information about the relationships among the application cores to minimize the overhead due to reconfiguration.
Juan Antonio Clemente, Vincenzo Rana, Donatella Sciuto, Ivan Beretta, David Atienza 0001
FPL3
2011 A Mapping Flow for Dynamically Reconfigurable Multi-Core System-on-Chip Design
abstract
Nowadays, multi-core systems-on-chip (SoCs) are typically required to execute multiple complex applications, which demand a large set of heterogeneous hardware cores with different sizes. In this context, the popularity of dynamically reconfigurable platforms is growing, as they increase the ability of the initial design to adapt to future modifications. This paper presents a design flow to efficiently map multiple multi-core applications on a dynamically reconfigurable SoC. The proposed methodology is tailored for a reconfigurable hardware architecture based on a flexible communication infrastructure, and exploits applications similarities to obtain an effective mapping. We also introduce a run-time mapper that is able to introduce new applications that were not known at design-time, preserving the mapping of the original system. We apply our design flow to a real-world multimedia case study and to a set of synthetic benchmarks, showing that it is actually able to extract similarities among the applications, as it achieves an average improvement of 29% in terms of reconfiguration latency with respect to a communication-oriented approach, while preserving the same communication performance.
Ivan Beretta, Vincenzo Rana, David Atienza 0001, Donatella Sciuto
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2011 Applying dynamic reconfiguration in the mobile robotics domain: A case study on computer vision algorithms
abstract
Mobile robots are widely used in industrial environments and are expected to be widely available in human environments in the near future, for example, in the area of care and service robots. This article proposes an implementation for a highly customizable color recognition module based on Field Programmable Gate Array (FPGA) hardware to accomplish tasks like real-time frame processing for image streams. In comparison to a pure software solution on a CPU, an attached FPGA-based hardware accelerator enables real-time image processing and significantly reduces the required computing power of the CPU. Instead, the CPU can be used for tasks that cannot be efficiently implemented on FPGAs, for example, because of a large control overhead. We concentrate on a multirobot scenario where a group of robots follows a human team member by keeping a specific formation in order to support the human in exploration and object detection. Additionally, the robots provide a communication infrastructure to maintain a stable multihop communication network between the human and a base station recording all actions and evaluating the captured images and transmitted data. Depending on the current operating conditions, the robot system has to be able to execute a wide variety of different tasks. Since only a small number of tasks have to be executed concurrently, dynamic reconfiguration of the FPGA can be used to avoid the parallel implementation of all tasks on the FPGA. Within this context, this article discusses application fields where dynamic reconfiguration of FPGA-based coprocessors significantly reduces the CPU load and presents examples of how dynamic reconfiguration can be used in exploration.
Federico Nava, Donatella Sciuto, Marco D. Santambrogio, Stefan Herbrechtsmeier, Mario Porrmann, Ulf Witkowski, Ulrich Rückert 0001
ACM Trans. Reconfigurable Technol. Syst.2
2010 Mapping and scheduling of parallel C applications with ant colony optimization onto heterogeneous reconfigurable MPSoCs
abstract
Efficient mapping and scheduling of partitioned applications are crucial to improve the performance on today's reconfigurable multiprocessor systems-on-chip (MPSoCs) platforms. Most of existing heuristics adopt the directed acyclic (task) Graph as representation, that unfortunately, is not able to represent typical embedded applications (e.g., real-time and loop-partitioned). In this paper we propose a novel approach, based on Ant Colony Optimization, that explores different alternative designs to determine an efficient hardware-software partitioning, to decide the task allocation and to establish the execution order of the tasks, dealing with different design constraints imposed by a reconfigurable heterogeneous MPSoC. Moreover, it can be applied to any parallel C application, represented through Hierarchical Task Graphs. We show that our methodology, addressing a realistic target architecture, outperforms existing approaches on a representative set of embedded applications.
Fabrizio Ferrandi, Christian Pilato, Donatella Sciuto, Antonino Tumeo
ASP-DAC3
2010 A reconfigurable multiprocessor architecture for a reliable face recognition implementation
abstract
Face Recognition techniques are solutions used to quickly screen a huge number of persons without being intrusive in open environments or to substitute id cards in companies or research institutes. There are several reasons that require to systems implementing these techniques to be reliable. This paper presents the design of a reliable face recognition system implemented on Field Programmable Gate Array (FPGA). The proposed implementation uses the concepts of multiprocessor architecture, parallel software and dynamic reconfiguration to satisfy the requirement of a reliable system. The target multiprocessor architecture is extended to support the dynamic reconfiguration of the processing unit to provide reliability to processors fault. The experimental results show that, due to the multiprocessor architecture, the parallel face recognition algorithm can achieve a speed up of 63% with respect to the sequential version. Results regarding the overhead in maintaining a reliable architecture are also shown.
Antonino Tumeo, Francesco Regazzoni 0001, Gianluca Palermo, Fabrizio Ferrandi, Donatella Sciuto
DATE5
2010 A Compact Transactional Memory Multiprocessor System on FPGA
abstract
In this paper we present a rapid prototyping platform on a single Field Programmable Gate Array (FPGA) with support for software transactional memory. The system is composed only by off-the-shelf cores and is useful for porting and early validation of programs to the transactional memory programming model. We discuss the implementation of the software layer of this platform, propose an analysis of the system and compare it to a hardware lock based multiprocessor architecture, showing the trade-offs in terms of performance and programming complexity.
Matteo Pusceddu, Simone Ceccolini, Gianluca Palermo, Donatella Sciuto, Antonino Tumeo
FPL4
2010 Multiprocessor systems-on-chip synthesis using multi-objective evolutionary computation
abstract
In this paper, we apply multi-objective evolutionary computation to the synthesis of real-time, embedded, heterogeneous, multiprocessor systems (briefly, Multiprocessor Systems-on-Chip or MP-SoCs). Our approach simultaneously explores the architecture, the mapping and the scheduling of the system, by using multi-objective evolution. In particular, we considered three approaches: a multi-objective genetic algorithm, multi-objective Simulated Annealing, and multi-objective Tabu Search. The algorithms search for optimal architectures, in terms of processing elements (processors and hardware accelerators) and communication infrastructure, and for the best mappings and schedules of multi-rate real-time applications given objectives such as: system area, hard and soft dead-lines violations, dimensions of memory buffers. We formalize the problem, describe our flow and compare the three algorithms, dis- cussing which one performs better with respect to different classes of applications.
Marco Ceriani, Fabrizio Ferrandi, Pier Luca Lanzi, Donatella Sciuto, Antonino Tumeo
GECCO4
2010 A novel design framework for the design of reconfigurable systems based on NoCs
abstract
In the last years, the embedded systems market is considerably grown, even though techniques and methodologies for the design of embedded systems, both from the hardware and the software point of view, have not been able to fully support this growth. Within this context, novel methodologies and design flows are required in order both to improve the quality and to shorten the time-to-market of very complex embedded systems.
Vincenzo Rana, Donatella Sciuto
ACM Great Lakes Symposium on VLSI2
2010 Run-time mapping of applications on FPGA-based reconfigurable systems
abstract
The role of Field-Programmable Gate Arrays (FPGAs) in System-on-Chip (SoC) design considerably increased in the last few years. Their established importance is due to the large amount of hardware resources they offer, as well as to their increasing performance, and furthermore to the support for reconfigurability. Even though FPGAs seem to have reached their maturity, there is still a lack of Computer-Aided Design (CAD) tools able to deal with dynamic reconfiguration. Existing algorithms aim at optimizing the performance of a set of applications, basing the computation on classic metrics (such as communication overhead), while reconfiguration-related issues are not taken into consideration. This work proposes a design methodology to map several applications on the FPGA area at run-time. Starting from a basic solution found at design-time for the initial set of applications, the proposed algorithm makes it possible to map a new application (not known at design-time), both minimizing the number of synthesis processes and optimizing the on-chip performance of the new application. Experimental results show that the proposed approach is able to achieve up to a 18% reduction in the number of reconfigurations with respect to an off-line static-mapping approach, while generally preserving the performance of the executed applications on the FPGA.
Ivan Beretta, Vincenzo Rana, David Atienza 0001, Donatella Sciuto
ISCAS4
2010 A direct bitstream manipulation approach for Virtex4-based evolvable systems
abstract
This work proposes a new Evolvable Hardware (EHW) system able to exploit two-dimensional dynamic reconfigurability and direct bitstream manipulation. These features enhance performance by enabling the parallelism between the evaluation and the reconfiguration phase and by speeding-up the reconfiguration process. The system is hierarchically structured, and can thus be used to evolve circuits mixing the search capabilities offered by a fine grained evolution and the exploitation of functional building blocks typical of functional level evolution. Previous EHW systems were able to evolve simple analog and digital circuits like counters, multiplexers, etc. Our system can be used to quickly evolve circuits able to solve complex problems, like the inverse pendulum problem.
Fabio Cancare, Marco D. Santambrogio, Donatella Sciuto
ISCAS3
2010 A design workflow for dynamically reconfigurable multi-FPGA systems
abstract
Multi-FPGA systems (MFS's) represent a promising technology for various applications, such as the implementation of supercomputers and parallel and computational intensive emulation systems. On the other hand, dynamic reconfigurability expands the possibilities of traditional FPGAs by providing them the capability of adapting their functionality while still running to cope with runtime environment changes. These two research directions are merged together in this work, that describes a methodology for designing dynamic reconfigurable MFS's. In this paper a novel MFS design flow has been described, which makes use of blocks reuse through dynamic reconfigurability to make the implementation of large systems feasible even on multi-FPGA architectures with strict physical constraints. Functional to this goal is the development of an algorithm for the extraction of the isomorphic structures of a circuit that extensively exploits the hierarchy of the design.
Alessandro Panella, Marco D. Santambrogio, Francesco Redaelli, Fabio Cancare, Donatella Sciuto
VLSI-SoC5
2010 Guest Editors' Introduction: Special Section on System-Level Design of Reliable Architectures
abstract
IT is with great pleasure that we introduce this special section on System-Level Design of Reliable Architectures to the audience of the IEEE Transactions on Computers. Six papers have been selected covering a wide spectrum of topics ranging from architectural fault-tolerant techniques to formal methodologies for reliability analysis. These papers are authored by relevant researchers in the field and cover theoretical and experimental topics. The widespread use of electronics in our life is directing more and more attention to the reliability properties of such systems in order to preserve both user’s and environmental safety; therefore, the design of reliable architectures is today a necessity rather than an option, even in not-critical application domains. At the same time, these systems are reaching high complexity levels, thus leading the designer to both develop specific components and to use and compose existing ones to achieve the desired overall functionality. In the former case, ad hoc techniques may be devised, acting on either the hardware or the software to cope with the occurrence of faults. In this latter situation, when combining independently designed modules, the enhancement and assessment of reliability becomes particularly important; for instance, specific approaches are required to be able both to apply fault detection/tolerance techniques from the initial steps of the design flow and to evaluate the effects of faults in a component while interacting with the other ones composing the overall system. As a result, the entire design flow needs to be enhanced to support reliability: from the initial modelling of the system together with the desired properties/requirements, to the fault model, from the hardware/ software partitioning step to the subsequent design exploration phase, where the more traditional metrics covering performance, costs, and power consumption need to be modified to also weight fault detection/tolerance capabilities. Functional verification and reliability analysis constitute two other aspects of this scenario to assess the quality of the designed system in terms of correctness and its ability to deal with failures. In this scenario, new advances have been achieved in all the relevant issues pertaining the system-level design of reliable systems, to support the designers in the development of innovative architectures able to cope with the occurrence of failures. Such advances lead to the definition of both new methodologies, as well as, of new architectures. Furthermore, based on the application environment in which the system will be adopted, different classes of reliability might be necessary; in some situations it is possible to achieve an autonomous fault detection capability, whereas, in critical environments, fault effects need to be completely masked, thus providing fault tolerance properties. The six papers presented in this special section were selected to address the different aspects of the important challenges related to the system level design of reliable systems. They cover all various facets of the issue, offering interesting solutions to tackle the specific problems. The first two papers deal with reliability analysis, which has become a fundamental tool to computer engineers for the validation of the design of hardened system architectures, in particular in safety and mission critical domains, such as medicine, military and transportation. The first paper is entitled “Formal Reliability Analysis Using Theorem Proving” by Osman Hasan, Sofiene Tahar, and Naeem Abbasi. This paper addresses an important aspect of reliability analysis, attempting to introduce formal verification instead of simulation-based and probabilistic approaches to assess the fault tolerance characteristics of the designed systems. The authors propose to conduct a formal reliability analysis of systems within the framework of a higher-order-logic theorem prover. In this paper, they present the higher-orderlogic formalization of some fundamental reliability theory concepts, which can be built upon to precisely analyze the reliability of various engineering systems. The proposed formalization is then applied to analyze the repairability conditions for a reconfigurable memory array in the presence of stuck-at and coupling faults. Still within the context of reliability analysis, the second paper, entitled “Efficient Microarchitectural Vulnerabilities Prediction Using Boosted Regression Trees and Patient Rule Inductions,” by Bin Li, Lide Duan, and Lu Peng, deals with Architectural Vulnerability Factor (AVF) analysis, which reflects the possibility that a transient fault eventually causes a visible error in the program output, and it indicates a system’s susceptibility to transient faults. This metric is increasingly being adopted to evaluate microprocessor’s architectures, due to their high vulnerability to transient faults, derived from shrinking feature sizes, threshold voltage, and increasing frequency. The authors propose an innovative way to predict the architectural vulnerability factor using Boosted Regression Trees, a nonparametric tree-based predictive modeling scheme, to identify the correlation across workloads, execution phases, and processor configurations, between the estimated AVF of a key processor structure and various performance metrics. The next two papers deal with fault detection techniques for different architectural components. The first paper is entitled “Concurrent Structure-Independent Fault Detection Schemes for the Advanced Encryption Standard,” authored by Mehran Mozaffari-Kermani and Arash Reyhani-Masoleh. IEEE TRANSACTIONS ON COMPUTERS, VOL. 59, NO. 5, MAY 2010 577
Cristiana Bolchini, Donatella Sciuto
IEEE Trans. Computers2
2010 Decision-Theoretic Design Space Exploration of Multiprocessor Platforms
abstract
This paper presents an efficient technique to perform design space exploration of a multiprocessor platform that minimizes the number of simulations needed to identify a Pareto curve with metrics like energy and delay. Instead of using semi-random search algorithms (like simulated annealing, tabu search, genetic algorithms, etc.), we use the domain knowledge derived from the platform architecture to set-up the exploration as a discrete-spaceMarkov decision process. The system walks the design space changing its parameters, performing simulations only when probabilistic information becomes insufficient for a decision. A learning algorithm updates the probabilities of decision outcomes as simulations are performed. The proposed technique has been tested with two multimedia industrial applications, namely the ffmpeg transcoder and the parallel pigz compression algorithm. Results show that the exploration can be performed with 5% of the simulations necessary for the most used algorithms (Pareto simulated annealing, nondominated sorting genetic algorithm, etc.), increasing the exploration speed by more than one order of magnitude.
Giovanni Beltrame, Luca Fossati, Donatella Sciuto
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2010 Ant Colony Heuristic for Mapping and Scheduling Tasks and Communications on Heterogeneous Embedded Systems
abstract
To exploit the power of modern heterogeneous multiprocessor embedded platforms on partitioned applications, the designer usually needs to efficiently map and schedule all the tasks and the communications of the application, respecting the constraints imposed by the target architecture. Since the problem is heavily constrained, common methods used to explore such design space usually fail, obtaining low-quality solutions. In this paper, we propose an ant colony optimization (ACO) heuristic that, given a model of the target architecture and the application, efficiently executes both scheduling and mapping to optimize the application performance. We compare our approach with several other heuristics, including simulated annealing, tabu search, and genetic algorithms, on the performance to reach the optimum value and on the potential to explore the design space. We show that our approach obtains better results than other heuristics by at least 16% on average, despite an overhead in execution time. Finally, we validate the approach by scheduling and mapping a JPEG encoder on a realistic target architecture.
Fabrizio Ferrandi, Pier Luca Lanzi, Christian Pilato, Donatella Sciuto, Antonino Tumeo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2010 Placement and Floorplanning in Dynamically Reconfigurable FPGAs
abstract
The aim of this article is to describe a complete partitioning and floorplanning algorithm tailored for reconfigurable architectures deployable on FPGAs and considering communication infrastructure feasibility. This article proposes a novel approach for resource- and reconfiguration- aware floorplanning. Different from existing approaches, our floorplanning algorithm takes specific physical constraints such as resource distribution and the granularity of reconfiguration possible for a given FPGA device into account. Due to the introduction of constraints typical of other problems like partitioning and placement, the proposed approach is named floorplacer in order to underline the great differences with respect to traditional floorplanners. These physical constraints are typically considered at the later placement stage. Different aspects of the problems have been described, focusing particularly on the FPGAs resource heterogeneity and the temporal dimension typical of reconfigurable systems. Once the problem is introduced a comparison among related works has been provided and their limits have been pointed out. Experimental results proved the validity of the proposed approach.
Alessio Montone, Marco D. Santambrogio, Donatella Sciuto, Seda Ogrenci Memik
ACM Trans. Reconfigurable Technol. Syst.3
2009 An application-centered design flow for self reconfigurable systems implementation
abstract
Up to now every proposed methodology for implementing dynamic self reconfigurable systems is architecture-centered. In most cases the system development process is time consuming and requires a very specific technical background. Aim of this work is to provide a fast brain to bit design flow whose goal is to simplify the dynamic reconfigurable system development process by shifting the designer focus from the architecture point of view to the application point of view: designers will not need to possess Dynamic Reconfigurability expertise but just to be skilled with the application domain.
Fabio Cancare, Marco D. Santambrogio, Donatella Sciuto
ASP-DAC3
2009 Prototyping pipelined applications on a heterogeneous FPGA multiprocessor virtual platform
abstract
Multiprocessors on a chip are the reality of these days. Semiconductor industry has recognized this approach as the most efficient in order to exploit chip resources, but the success of this paradigm heavily relies on the efficiency and widespread diffusion of parallel software. Among the many techniques to express the parallelism of applications, this paper focuses on pipelining, a technique well suited to data-intensive multimedia applications. We introduce a prototyping platform (FPGA-based) and a methodology for these applications. Our platform consists of a mix of standard and custom heterogeneous cores. We discuss several case studies, analyzing the interaction of the architecture and applications and we show that multimedia and telecommunication applications with unbalanced pipeline stages can be easily deployed. Our framework eases the development cycle and enables the developers to focus directly on the problems posed by the programming model in the direction of the implementation of a production system.
Antonino Tumeo, Marco Branca, Lorenzo Camerini, Marco Ceriani, Matteo Monchiero, Gianluca Palermo, Fabrizio Ferrandi, Donatella Sciuto
ASP-DAC8
2009 A real-time application design methodology for MPSoCs
abstract
This paper presents a novel technique for the modeling, simulation, and analysis of real-time applications on Multi-Processor Systems-on-Chip (MPSoCs). This technique is based on an application-transparent emulation of OS primitives, including support for RTOS elements. The proposed methodology enables a quick evaluation of the real-time performance of an application in front of different design choices, including the study of system's behavior as tasks' deadlines become stricter or looser. The approach has been verified on a large set of multi-threaded benchmarks. Results show that our methodology (a) enables accurate realtime and responsiveness analysis of parallel applications running on MPSOCs, (b) allows the designer to devise an optimal interrupt distribution mechanism for the given application, and (c) helps dimensioning the system to meet performance and real-time needs.
Giovanni Beltrame, Luca Fossati, Donatella Sciuto
DATE3
2009 HW/SW methodologies for synchronization in FPGA multiprocessors
abstract
odern Field Programmable Gate Arrays (FPGA) can be programmed with multiple soft-core processors. These solutions can be used for MultiProcessor Systems-on-Chip (MPSoCs) prototyping or even for final implementation. Nevertheless, efficient synchronization is required to guarantee performance in multiprocessing environments with the simple cores that do not support atomic instructions and are normally used in the standard FPGA toolchains. In this paper, we introduce two hardware synchronization modules for Xilinx MicroBlaze systems, with local polling or queuing mechanisms for locks and barriers, and present a comparison of these solutions to alternative designs.
Antonino Tumeo, Christian Pilato, Gianluca Palermo, Fabrizio Ferrandi, Donatella Sciuto
FPGA5
2009 A runtime relocation based workflow for self dynamic reconfigurable systems design
abstract
A self, partial and dynamic approach to reconfiguration makes it possible to obtain higher flexibility and better performance with respect to simpler approaches; however, the price for this improvement lies in the increased difficulties in the reconfigurable system creation and management, which become significantly more complex. An automated or semiautomated way to support this kind of systems would simplify the problem by raising the level of abstraction at which the designer has to operate. The aim of this work is the creation of a complete workflow to help the designer in the creation and management of self partially and dynamically reconfigurable systems: the designer should only specify the application, the reconfigurable device, the reconfiguration model (1D vs 2D) and the type of communication infrastructure, and the automated flow will deal with the subsequent steps down to the final architecture implementation. Among other aspects, the provided support includes the definition of area constraint for cores, the creation of an efficient runtime solution for core allocation management and the generation of a solution to obtain internal and fast relocation of cores.
Marco D. Santambrogio, Massimo Morandi, Marco Novati, Donatella Sciuto
FPL4
2009 Evolutionary algorithms for the mapping of pipelined applications onto heterogeneous embedded systems
abstract
In this paper, we compare four algorithms for the mapping of pipelined applications on a heterogeneous multiprocessor platform implemented using Field Programmable Gate Arrays (FPGAs) with customizable processors. Initially, we describe the framework and the model of pipelined application we adopted. Then, we focus on the problem of mapping a set of pipelined applications onto a heterogeneous multiprocessor platform and consider four search algorithms: Tabu Search, Simulated Annealing, Genetic Algorithms, and the Bayesian Optimization Algorithm. We compare the performance of these four algorithms on a set of synthetic problems and on two real-world applications (the JPEG image encoding and the ADPCM sound encoding). Our results show that on our framework the Bayesian Optimization Algorithm outperforms all the other three methods for the mapping of pipelined applications.
Marco Branca, Lorenzo Camerini, Fabrizio Ferrandi, Pier Luca Lanzi, Christian Pilato, Donatella Sciuto, Antonino Tumeo
GECCO6
2009 Reconfigurable NoC design flow for multiple applications run-time mapping on FPGA devices
abstract
Dynamic reconfiguration capabilities exploited by modern FPGA devices improve the flexibility and the reliability of embedded systems. The increasing complexity demands for a design-paradigm shift towards a communication-centric approach. Networks-on-Chip are a promising design paradigm for both homogeneous and heterogeneous systems in which communication is represented in a network-like manner, even if they cannot directly be applied to the dynamic reconfiguration scenario. While in literature there are different approaches to design communication infrastructures able to support the reconfiguration of its functionalities, what seems to be neglected is the definition of a complete design flow for a dynamic reconfigurable communication infrastructure able to adapt itself at runtime to the current working scenario. This paper proposes a design flow to automatically create a reconfigurable architecture that consists of a grid of homogeneous tiles that can be filled with either computational (master or slave cores with their network interfaces) or communication (switches) elements.
Dario Cozzi, Claudia Farè, Alessandro Meroni, Vincenzo Rana, Marco D. Santambrogio, Donatella Sciuto
ACM Great Lakes Symposium on VLSI6
2009 On-line task management for a reconfigurable cryptographic architecture
abstract
The increasing amount of programmable logic provided by modern FPGAs makes it possible to execute multiple hardware applications on the same device. This approach is reinforced by dynamic reconfiguration, which allows a single part of the device to be configured with a single hardware module. The proposed solution is a Linux-based operating system to manage on-demand module configuration on an FPGA while providing a set of high-level abstractions to user applications. The proposed approach has been validated in a cryptographic context using the DES and the AES algorithms.
Ivan Beretta, Vincenzo Rana, Marco D. Santambrogio, Donatella Sciuto
IPDPS4
2009 A multiprocessor self-reconfigurable JPEG2000 encoder
abstract
This paper presents a multiprocessor architecture prototype on a field programmable gate arrays (FPGA) with support for hardware and software multithreading. Thanks to partial dynamic reconfiguration, this system can, at run time, spawn both software and hardware threads, sharing not only the general purpose soft-cores present in the architecture but also area on the FPGA. While on a standard single processor architecture the partial dynamic reconfiguration requires the processor to stop working to instantiate the hardware threads, the proposed solution hides most of the reconfiguration latency through the parallel execution of software threads. We validate our framework on a JPEG 2000 encoder, showing how threads are spawned, executed and joined independently of their hardware or software nature. We also show results confirming that, by using the proposed approach, we are able to hide the reconfiguration time.
Antonino Tumeo, Simone Borgio, Davide Bosisio, Matteo Monchiero, Gianluca Palermo, Fabrizio Ferrandi, Donatella Sciuto
IPDPS7
2009 A Transform-Parametric Approach to Boolean Matching
abstract
In this paper, we address the problem of P-equivalence Boolean matching. We outline a formal framework that unifies some of the spectral- and canonical-form-based approaches to the problem. As a first major contribution, we show how these approaches are particular cases of a single generic algorithm, parametric with respect to a given linear transformation of the input function. As a second major contribution, we identify a linear transformation that can be used to significantly speed up Boolean matching with respect to the state of the art. Experimental results show that, on average, over a large set of randomly generated Boolean functions, our approach is up to five times faster than the main competitor on 20-variable input and scales better, allowing to match even larger components. Finally, as a representative set of Boolean functions that arise in practice, we considered multiplexers with three, four, and five selectors and functions extracted from the ISCAS85 benchmarks suite with a number of input variables up to 20. The reported performance results show that our approach allows us to halve the canonizing computation time.
Giovanni Agosta, Francesco Bruschi, Gerardo Pelosi, Donatella Sciuto
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2009 ReSP: A Nonintrusive Transaction-Level Reflective MPSoC Simulation Platform for Design Space Exploration
abstract
This paper presents reflective simulation platform (ReSP), a transaction-level multiprocessor simulation platform based on the integration of SystemC and Python. ReSP exploits the concept of reflection, enabling the integration of SystemC components without source-code modifications and providing full observability of their internal state. ReSP offers fine-grained simulation control and supports the evaluation of different hardware/software configurations of a given application, enabling complete design space exploration. ReSP allows the evaluation of real-time applications on high-level hardware models since it provides the transparent emulation of POSIX-compliant real-time operating systems (RTOS) primitives. A number of experiments have been performed to validate ReSP and its capabilities, using a set of single- and multithreaded benchmarks, with both POSIX Threads (PThreads) and OpenMP programming styles. These experiments confirm that reflection introduces negligible ( <1%) overhead when comparing ReSP to plain SystemC simulation. The results also show that ReSP can be successfully used to analyze and explore concurrent and reconfigurable applications even at very early development stages. In fact, the average error introduced by ReSP's RTOS emulation is below 6.6 plusmn 5% w.r.t. the same RTOS running on an instruction set simulator, while simulation speed increases by a factor of ten. Owing to the integration with a scripted language, simulation management is simplified, and experimental setup effort is considerably reduced.
Giovanni Beltrame, Luca Fossati, Donatella Sciuto
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2009 Partitioning and Scheduling of Task Graphs on Partially Dynamically Reconfigurable FPGAs
abstract
This paper proposes a new model for the partitioning and scheduling of a specification on partially dynamically reconfigurable hardware. Although this problem can be solved optimally only by tackling its subproblems jointly, the exceeding complexity of such a task leads to a decomposition into two phases. The partitioning phase is based on a new graph-theoretic approach, which aims to obtain near optimality even if performed independently from the subsequent phase. For the scheduling phase, a new integer linear programming formulation and a heuristic approach are developed. Both take into account configuration prefetching and module reuse. The experimental results show that the proposed method compares favorably with existing solutions.
Roberto Cordone, Francesco Redaelli, Massimo Redaelli, Marco D. Santambrogio, Donatella Sciuto
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2009 Internal and External Bitstream Relocation for Partial Dynamic Reconfiguration
abstract
The research described in this paper shows how the runtime relocation of a reconfigurable component can be obtained using a system component that is able to update the bitstream information, moving the reconfigurable module in the desired position. This scenario defines the so-called partial bitstream relocation activity. This paper proposes a relocation filter that can be implemented both as a hardware and a software component. The former is hosted in the static part of the reconfigurable architecture, while the latter is made to be run on the processor placed on the field-programmable gate array (FPGA). The proposed approach has also been validated over different FPGAs, i.e., Virtex II Pro, Virtex 4, and Virtex 5, proposing a runtime relocation support that can be customized to meet all the different constraints associated with these different target architectures.
Simone Corbetta, Massimo Morandi, Marco Novati, Marco D. Santambrogio, Donatella Sciuto, Paola Spoletini
IEEE Trans. Very Large Scale Integr. Syst.5
2008 Lightweight DMA management mechanisms for multiprocessors on FPGA
abstract
This paper presents a multiprocessor system on FPGA that adopts Direct Memory Access (DMA) mechanisms to move data between the external memory and the local memory of each processor. The system integrates all standard DMA primitives via a fast Application Programming Interface (API) and relies on interrupts having also the possibility to manage a command list. This interface allows to program the embedded multiprocessor architecture on FPGA with simple DMAs using the same DMA techniques adopted on high performance multiprocessors with complex DMA controllers. Several experiments demonstrate the performance of our solution, allowing 57% improvement on the execution time of a selected set of benchmarks. We furthermore show how some DMA programming techniques (double and multi-buffering) can be effectively used within our platform, thus easing the design and development of the hardware and the software in a reconfigurable DMA-based environment.
Antonino Tumeo, Matteo Monchiero, Gianluca Palermo, Fabrizio Ferrandi, Donatella Sciuto
ASAP5
2008 ReSP: A non-intrusive Transaction-Level Reflective MPSoC Simulation Platform for design space exploration
abstract
This paper presents ReSP (Reflective Simulation Platform), a Transaction-Level multi-processor simulation platform based on SystemC and Python; SystemC is a standard language for system modeling and verification, and Python provides the platform with reflective capabilities. These are employed to give the designer an easy way to specify the architecture of a system, simulate the given configuration and perform automatic analysis on it. ReSP enables SystemC and Python interoperability through automatic Python wrapper generation. We show that the overhead associated with the Python intermediate layer is around 1%, therefore execution speed is not compromised. The advantages of our approach are: (a) easy integration of external IPs (b) fine grain control of the simulation (c) effortless integration of tools for system analysis and design space exploration. A case study shows how the platform can be extended to support system reliability assessment.
Giovanni Beltrame, Cristiana Bolchini, Luca Fossati, Antonio Miele, Donatella Sciuto
ASP-DAC5
2008 The Shining embedded system design methodology based on self dynamic reconfigurable architectures
abstract
Complex design, targeting system-on-chip based on reconfigurable architectures, still lacks a generalized methodology allowing both the automatic derivation of a complete system solution able to fit into the final device, and mixed hardware-software solutions, exploiting partial reconfiguration capabilities. The shining methodology organizes the input specification of a complex system-on-chip design into three different components: hardware, reconfigurable hardware and software, each handled by dedicated sub-flows. A communication model guarantees reliable and seamless interfacing of the various components. The developed system, stand-alone or OS-based, is architecture-independent. The shining flow reduces the time for system development, easing the design of complex hardware/software reconfigurable applications.
Carlo Curino, Luca Fossati, Vincenzo Rana, Francesco Redaelli, Marco D. Santambrogio, Donatella Sciuto
ASP-DAC6
2008 High-level synthesis with multi-objective genetic algorithm: A comparative encoding analysis
abstract
The high-level synthesis process involves three interdependent and NP-complete optimization problems: (i) the operation scheduling, (ii) the resource allocation, and (iii) the controller synthesis. Evolutionary algorithms have been effectively applied to high level synthesis in presence conflicting design objectives for finding good tradeoffs in the design space. However, so far the design space exploration has been performed using single-objective evolutionary algorithms with an ad hoc fitness function to achieve the desired tradeoff between the objectives. Recently we proposed a framework based on multi-objective genetic algorithms to perform a fully automated design space exploration. In this paper we focus on the choice of the solution representations that can be used to perform the design space exploration with multi-objective genetic algorithms. In particular we consider two specific representations and compare them on a set of benchmark problems. Our results suggest that they have different biases on the search space that make them more effective in different problems and design subspaces. Accordingly, we present a preliminary investigation on a new representation that exploits the advantages of both of them.
Christian Pilato, Daniele Loiacono, Fabrizio Ferrandi, Pier Luca Lanzi, Donatella Sciuto
IEEE Congress on Evolutionary Computation5
2008 Task Scheduling with Configuration Prefetching and Anti-Fragmentation techniques on Dynamically Reconfigurable Systems
abstract
Aim of this paper is to define a scheduling of the task graph of an application that minimizes its total execution time on a partially dynamically reconfigurable FPGA. The scheduler has to take into account the reconfiguration overhead of each task, the area constraint of the target FPGA, the precedences between the tasks, configuration prefetching and module reuse. We introduce an ILP formulation to solve the task scheduling problem in the reconfigurable architecture scenario. This formulation has been used to identify interesting features for a possible heuristic scheduler. The results of the ILP solution show how a reconfiguration- aware scheduler exploiting all the reconfiguration features can outperform one with partial knowledge.
Francesco Redaelli, Marco D. Santambrogio, Donatella Sciuto
DATE3
2008 A Dual-Priority Real-Time Multiprocessor System on FPGA for Automotive Applications
abstract
This paper presents the implementation of a dual-priority scheduling algorithm for real-time embedded systems on a shared memory multiprocessor on FPGA. The dual-priority microkernel is supported by a multiprocessor interrupt controller to trigger periodic and aperiodic thread activation and manage context switching. We show how the dual-priority algorithm performs on a real system prototype compared to the theoretical performance simulations with a typical standard workload of automotive applications, underlining where the differences are.
Antonino Tumeo, Marco Branca, Lorenzo Camerini, Marco Ceriani, Matteo Monchiero, Gianluca Palermo, Fabrizio Ferrandi, Donatella Sciuto
DATE8
2008 Fault Models and Injection Strategies in SystemC Specifications
abstract
This paper presents fault models and fault injection strategies designed in a simulation platform with reflection capabilities, used for simulating complex systems specified by using SystemC and by adopting a platform-based design approach. The approach allows the designer to work at different levels of abstraction and to take into account permanent and transient faults, and -- most important -- it features a transparent and dynamic mechanism for both injecting faults and analyzing the produced errors, in order to evaluate possible fault detection and/or tolerance design techniques.
Cristiana Bolchini, Antonio Miele, Donatella Sciuto
DSD3
2008 Operating system support for online partial dynamic reconfiguration management
abstract
One of the main characteristics of reconfigurable embedded systems is their ability to be dynamically modified to be adapted at run-time to the current environment. This feature, that makes it possible to change the functionality of a system while it is up and running, requires a software application that is able to handle the reconfiguration process. The software for the management of reconfiguration can be developed either as a standalone application, that has to be specifically designed for each given system, or within an operating system, in order to fully exploit both code reuse and code portability. This paper proposes a novel methodology for the design of dynamically reconfigurable systems in which the reconfiguration management is completely assigned to an operating system reconfiguration support. Finally, a prototype implementation is presented, where a standard Linux operating system has been extended with the proposed operating system support in order to handle dynamically reconfigurable hardware resources.
Marco D. Santambrogio, Vincenzo Rana, Donatella Sciuto
FPL3
2008 A design flow tailored for self dynamic reconfigurable architecture
abstract
Dynamic reconfigurable embedded systems are gathering, day after day, an increasing interest from both the scientific and the industrial world. The need of a comprehensive tool which can guide designers through the whole implementation process is becoming stronger. In this paper the authors introduce a new design framework which amends this lack. In particular the paper describes the entire low level design flow onto which the framework is based.
Fabio Cancare, Marco D. Santambrogio, Donatella Sciuto
IPDPS3
2008 HARPE: A Harvard-based processing element tailored for partial dynamic reconfigurable architectures
abstract
Aim of this paper is to propose a reconfigurable processing element based on a Harvard architecture, called HARPE. HARPE's architecture includes a MicroBlaze soft-processor in order to make HARPEs deployable also on devices not having processors on silicon die. In such a context, this work also introduces a novel approach for the management of processor data memory. The proposed approach allows the individual management of data and the dynamic update of the memory, thus making it possible to define partially dynamical reconfigurable multi processing element systems, that consist of several master (e.g., soft-processors, hard-processors or HARPE cores) and slave components. Finally, the proposed methodology enables the possibility of creating a system in which both HARPEs and their memories (data and code) can be separately configured at run time with a partial configuration bitstream, in order to make the whole system more flexible with respect to changes occurring in the external environment.
Alessio Montone, Vincenzo Rana, Marco D. Santambrogio, Donatella Sciuto
IPDPS4
2008 Design methodology for partial dynamic reconfiguration: a new degree of freedom in the HW/SW codesign
abstract
Many emerging products in communication, computing and consumer electronics demand that their functionality remains flexible also after the system has been manufactured and that is why the reconfiguration is starting to be considered into the design flow as a new relevant degree of freedom, in which the designer can have the system autonomously modify its functionalities according to the application's changing needs. Therefore, reconfigurable devices, such as FPGAs, introduce yet another degree of freedom in the design workflow: the designer can have the system autonomously modify the functionality carried out by the IP core according to the application's changing needs while it runs. Research in this field is, indeed, being driven towards a more thorough exploitation of the reconfiguration capabilities of such devices, so as to take advantage of them not only at compile-time, i.e. at the time when the system is first deployed, but also at run-time, which allows the reconfigurable device to be reprogrammed without the rest of the system having to stop running. This paper presents emerging methodologies to design reconfigurable applications, providing, as an example the workflow defined at the Politecnico di Milano.
Marco D. Santambrogio, Donatella Sciuto
IPDPS2
2008 Software and Hardware Techniques for SEU Detection in IP Processors
Cristiana Bolchini, Antonio Miele, Fabio Rebaudengo, Fabio Salice, Donatella Sciuto, Luca Sterpone, Massimo Violante
J. Electron. Test.5
2008 Improving evolutionary exploration to area-time optimization of FPGA designs
Christian Pilato, Antonino Tumeo, Gianluca Palermo, Fabrizio Ferrandi, Pier Luca Lanzi, Donatella Sciuto
J. Syst. Archit.6
2008 Static Analysis of Transaction-Level Communication Models
abstract
We propose a methodology for the early estimation of communication implementation choice effects, starting from an abstract transaction-level system model (TLM). The reference version of the TLM considered is the Open SystemC initiative library. The methodology is based on the computation of metrics that abstract useful information from the initial system model. The metrics are precisely defined upon a general formal model of transaction-level system descriptions. A set of design problems of relevant interest, such as shared communication resource assignment, pipelining partitioning, bandwidth, and latency constraint estimation, is considered to show some potential applications of the metrics proposed.
Giovanni Agosta, Francesco Bruschi, Donatella Sciuto
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2007 A Self-Reconfigurable Implementation of the JPEG Encoder
abstract
Dynamic reconfiguration allows to selectively substitute blocks of logic at run-time in order to improve the area efficiency of a FPGA design. This paper presents the design of a JPEG Encoder which exploits this feature. We propose a mixed HW/SW architecture, where most compute-intensive components of the application are mapped to application-specific HW cores. These cores dynamically alternate on the FPGA. Our purpose is to describe a real-world application of reconfigurable computing, illustrating how this approach allows for saving area with negligible performance overhead. We built a fully-working prototype, which demonstrates that the reconfigurable JPEG encoder achieves 29.6% area saving, 1.5% performance loss, and negligible power overhead with respect to a solution which uses statically mapped HW cores.
Antonino Tumeo, Matteo Monchiero, Gianluca Palermo, Fabrizio Ferrandi, Donatella Sciuto
ASAP5
2007 Fitness inheritance in evolutionary and multi-objective high-level synthesis
abstract
The high-level synthesis process allows the automatic design and implementation of digital circuits starting from a behavioral description. Evolutionary algorithms are very widely adopted to approach this problem or just part of it. Nevertheless, some concerns regarding execution times exist. In evolutionary high-level synthesis, design solutions have to be evaluated to extract information about some figures of merit (such as performance, area, etc.) and to allow the genetic algorithm to evolve and converge to Pareto-optimal solutions. Since the execution time of such evaluations increases with the complexity of the specification, the overall methodology could lead to unacceptable execution time. This paper presents a model to exploit fitness inheritance in a multi-objective optimization algorithm (i.e. NSGA-II) by substituting the expensive real evaluations with estimations based on closeness in an hypothetical design space. The estimations are based on the measure of the distance between individuals and a weighted average of the fitnesses of the closest ones. The results shows that the Pareto-optimal set obtained by applying the proposed model well approximates the set obtained without fitness inheritance. Moreover, the overall execution time is reduced up to the 25% in average.
Christian Pilato, Gianluca Palermo, Antonino Tumeo, Fabrizio Ferrandi, Donatella Sciuto, Pier Luca Lanzi
IEEE Congress on Evolutionary Computation5
2007 A Unified Approach to Canonical Form-based Boolean Matching
abstract
In this paper, we face the problem of P-equivalence Boolean matching. We outline a formal framework that unifies some of the canonical form-based approaches to the problem.
Giovanni Agosta, Francesco Bruschi, Gerardo Pelosi, Donatella Sciuto
DAC4
2007 An efficient cost-based canonical form for Boolean matching
abstract
In this paper, we present new canonical forms for P, NP and NPN equivalence relations on boolean functions. The canonical forms are based on the minimization of a cost function. With respect to previous approaches based on cost minimization, our function allows the minimization algorithm to explore a reduced solution space. This reduction is obtained by partitioning the columns of the boolean function through an equivalence relation. NP and NPN canonical forms are obtained by means of a preprocessing step of negligible computational overhead.
Giovanni Agosta, Francesco Bruschi, Donatella Sciuto
ACM Great Lakes Symposium on VLSI3
2007 A design kit for a fully working shared memory multiprocessor on FPGA
abstract
This paper presents a framework to design a shared memory multiprocessor on a programmable platform. We propose a complete flow, composed by a programming model and a template architecture. Our framework permits to write a parallel application by using a shared memory model. It deals with the consistency of shared data, with no need of hardware coherence protocol, but uses a software model to properlyallsynchronize the local copies with the shared memory image. This idea can be applied both to a scratchpad-based architecture or a cache-based one. The architecture is synthesizable with standard IPs, such as the softcores and interconnect elements, which may be found in any commercial FPGA toolset.
Antonino Tumeo, Matteo Monchiero, Gianluca Palermo, Fabrizio Ferrandi, Donatella Sciuto
ACM Great Lakes Symposium on VLSI5
2007 A novel SoC design methodology combining adaptive software and reconfigurable hardware
abstract
Reconfigurable hardware is becoming a prominent component in a large variety of SoC designs. Reconfigurability allows for efficient hardware acceleration and virtually unlimited adaptability. On the other hand, overheads associated with reconfiguration and interfaces with the software component need to be evaluated carefully during the exploration phase. The aim of this paper is to identify the best trade-off considering application-specific features in software, which can lend itself to software-based acceleration and lead to a revision of the view that certain computationally intensive tasks can only be accelerated through hardware. In order to validate the effectiveness of our proposed techniques, we built an extensive development and experimental setup, bringing together the MLTon-based programming environment and physical mapping of the software and hardware onto a real dynamically reconfigurable SoC system.
Marco D. Santambrogio, Seda Ogrenci Memik, Vincenzo Rana, Umut A. Acar, Donatella Sciuto
ICCAD5
2007 Partial Dynamic Reconfiguration in a Multi-FPGA Clustered Architecture Based on Linux
abstract
Dynamically reconfigurable hardware allows for implementing systems that can be adapted at run-time according to the needs of the user. This paper presents an architecture that is composed of multiple FPGAs that are connected to an embedded processor. Thus, the architecture is referred to as a multi-FPGA clustered architecture (MFCA). All FPGAs can be partially and dynamically reconfigured to integrate user-defined IP-cores into the system at run-time. For the resource management and communication management we have implemented a Linux operating system on the embedded processor that can be used to control the reconfiguration of the FPGAs by means of simple function calls. Furthermore, the Linux OS completely hides the physical infrastructure of the MFCA from user applications, offering a consistent interface to utilize partial reconfiguration.
Vincenzo Rana, Marco D. Santambrogio, Donatella Sciuto, Boris Kettelhoit, Markus Köster, Mario Porrmann, Ulrich Rückert 0001
IPDPS3
2007 Dynamic Reconfigurability in Embedded System Design
abstract
Nowadays, dynamic reconfigurable embedded systems are widely used, since they have the capability to modify their functionalities, adding or removing components and modify interconnections among them. The basic idea behind these systems is to have the system autonomously modify its functionalities according to the application's changes. This paper describes the area of reconfigurable embedded systems presenting both architectural and methodological aspects trying to point out common features and needs. After a brief introduction, an overview of the models of the reconfigurable architectures, and of the design methodologies was presented.
Vincenzo Rana, Marco D. Santambrogio, Donatella Sciuto
ISCAS3
2007 An adaptive genetic algorithm for dynamically reconfigurable modules allocation
abstract
This paper aims at defining an adaptive genetic algorithm tailored for the allocation of dynamically reconfigurable modules. This algorithm can be tuned at run-time with a set of parameters to best characterize different architectural scenarios (i.e., single device or multi-FPGAs characterized by several kinds of communication infrastructures) and to adapt the performance of the algorithm itself to the scenario in which it has to operate. The proposed approach has been validated with a large set of meaningful combinations of parameters (i.e. changing the mutation or the crossover probability), in order to demonstrate the possibility of performing either a fast or an accurate allocation phase.
Vincenzo Rana, Chiara Sandionigi, Marco D. Santambrogio, Donatella Sciuto
VLSI-SoC4
2007 Multi-Accuracy Power and Performance Transaction-Level Modeling
abstract
This paper introduces a modeling and simulation technique that extends transaction-level modeling (TLM) to support multi-accuracy models and power estimation. This approach provides different combinations of power and performance models, and the switching of model accuracy during simulation, allowing the designer to trade off between simulation accuracy and speed at runtime. This is particularly useful during the exploration phase of a design, when the designer changes the features or the parameters of the design, trying to satisfy its constraints. Usually, only limited portions of a system are affected by a single parameter change, and therefore, it is possible to fast-simulate uninteresting sections of the application. In particular, we show how to extend the TLM and modify the SystemC kernel to support multi-accuracy features. The proposed methodology has been tested on several benchmarks, among which is an MPEG4 encoder, showing that simulation speed can be increased of one order of magnitude. On the same benchmarks, we also show how it is possible to choose the optimal performance simulation accuracy for a given power model, maximizing simulation speed for the desired accuracy.
Giovanni Beltrame, Donatella Sciuto, Cristina Silvano
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 Guest Editorial [intro. to the special issue on the 2006 IEEE/ACM Design, Automation and Test in Europe Conference]
abstract
The eight articles in this special issue are extended versions of selected papers from the 9th IEEE/ACM Design, Automation and Test in Europe (DATE) Conference, which was held on March 6-10, 2006 in Munich, Germany. The papers address some of the following topics: multiprocessor systems-on-chip (MPSoC) architectural and methodological issues; the problem of accelerating embedded processor execution through instruction set extensions; techniques for optimal bit width allocation in the conversion from floating point to fixed point in arithmetic circuits for low power; an algorithm to optimize circuits with tight sequential cycles; efficient methods to solve more general quantified Boolean formulas (QBFs); and soft error rate (SER) analysis for combinatorial circuits and the design of reconfigurable continuous-time delta-sigma modulator topologies. The selected papers are briefly summarized.
Georges Gielen, Donatella Sciuto
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2006 Using speculative computation and parallelizing techniques to improve scheduling of control based designs
abstract
Recent research results have seen the application of parallelizing techniques to high-level synthesis. In particular, the effect of speculative code transformations on mixed control-data flow designs has demonstrated effective results on schedule lengths. In this paper we first analyze the use of the control and data dependence graph as an intermediate representation that provides the possibility of extracting the maximum parallelism. Then we analyze the scheduling problem by formulating an approach based on Integer Linear Programming (ILP) to minimize the number of control steps given the amount of resources. We improve the already proposed ILP scheduling approaches by introducing a new conditional resource sharing constraint which is then extended to the case of speculative computation. The ILP formulation has been solved by using a Branch and Cut framework which provides better results than standard branch and bound techniques.
Roberto Cordone, Fabrizio Ferrandi, Marco D. Santambrogio, Gianluca Palermo, Donatella Sciuto
ASP-DAC5
2006 Exploiting TLM and object introspection for system-level simulation
abstract
The introduction of transaction level modeling (TLM) allows a system designer to model a complete application, composed of hardware and software parts, at several levels of abstraction. The simulation speed of TLM is orders of magnitude faster than traditional RTL simulation; nevertheless, it can become a limiting factor when considering a multi-processor system-on-chip (MP-SoC), as the analysis of these systems can be very complex. The main goal of this paper is to introduce a novel way of exploiting TLM features to increase simulation efficiency of complex systems by switching TLM models at runtime. Results show that simulation performance can be increased significantly without sacrificing the accuracy of critical application kernels
Giovanni Beltrame, Donatella Sciuto, Cristina Silvano, Damien Lyonnard, Chuck Pilkington
DATE2
2006 Partial Dynamic Reconfiguration: The Caronte Approach. A New Degree of Freedom in the HW/SW Codesign
abstract
The design of embedded systems is rapidly changed during the last decade. It is possible to identify two main factors that are involved in this process: HW/WS codesign and dynamic reconfigurable architecture. This work aims at introducing an innovative methodology that allows to easily implement on an FPGA a system specification, taking as input its high-level description, such as C or SystemC, and exploiting the capabilities of partial dynamic reconfiguration and HW/SW codesign methodologies. In order to meet the software requirements of complex systems, the solution is also provided with the porting of a real-time GNU/Linux OS, CLinux, which allows software processes to exploit a rich set of features, and with a Linux module that simplifies the handling of reconfiguration
Marco D. Santambrogio, Donatella Sciuto
FPL2
2006 Combining hardware reconfiguration and adaptive computation for a novel SoC design methodology
abstract
In the face of dominant communication overheads and reconfiguration cost of programmable hardware often deployed in SoC environments, a new paradigm is necessary to revisit the partitioning and allocation problems. Our aim is to integrate generalized performance models into codesign to explore the gray area between hardware and software effectively. We propose to use the adaptive computation approach. Adaptivity implies that due to input changes the output of the system is updated only re-evaluating those portions of the program affected by the changes. We study the impact of our model onto a SoC architecture consisting of embedded processors and dynamically reconfigurable hardware. We present an image processing application mapped onto this architecture as a case study
Vincenzo Rana, Marco D. Santambrogio, Seda Ogrenci Memik, Donatella Sciuto
FPT4
2006 Automatic Test Pattern Generation with BOA
Tiziana Gravagnoli, Fabrizio Ferrandi, Pier Luca Lanzi, Donatella Sciuto
PPSN4
2006 An Application Mapping Methodology and Case Study for Multi-Processor On-Chip Architectures
abstract
This paper introduces an application mapping methodology and case study for multiprocessor on-chip architectures. Starting from the description of an application in standard sequential code (e.g. in C), first the application is profiled, parallelized when possible, and then its components are moved to hardware implementation when necessary to satisfy performance and power constraints. The key contribution of this work is a methodology for high-level hardware/software partitioning that allows the designer to use the same code for both hardware and software models for simulation, providing nevertheless preliminary estimations for timing and power consumption. The methodology has been applied to the co-exploration of an industrial case study: an MPEG4 VGA realtime encoder
Giovanni Beltrame, Donatella Sciuto, Cristina Silvano, Pierre G. Paulin, Essaid Bensoudane
VLSI-SoC2
2006 A graph-coloring approach to the allocation and tasks scheduling for reconfigurable architectures
abstract
Designing systems mapped onto FPGAs that foresee a dynamic reconfiguration of the application is a difficult task. It requires that the identification of the reconfigurable tasks and their allocation onto the FPGA must be defined during the design phases. Furthermore, also the schedule of dynamic reconfigurations must be defined. This paper presents an improved scheduling and allocation of reconfigurable tasks onto an FPGA, based on the coloring problem. The proposed algorithm stems from the one previously presented (Ferrandi et al., 2005), but introduces backtracking to improve the performance in terms of number of number of colors, that represent FPGAs areas. The new algorithm has been experimented on the Xilinx-based architecture defined to support dynamic reconfigurability (Donato et al., 2005)
Marco Giorgetta, Marco D. Santambrogio, Donatella Sciuto, Paola Spoletini
VLSI-SoC3
2006 Fast IP-Core Generation in a Partial Dynamic Reconfiguration Workflow
abstract
Reconfigurable devices, such as FPGAs, introduce into the design workflow of embedded systems a new degree of freedom: the designer can have the system autonomously modify the functionality carried out by the IP-core according to the application's changing needs while it runs. The Caronte methodology, based on the modular design approach, is a design workflow that allows the creation and the handling of partial dynamic reconfigurable architectures using Xilinx FPGAs. In order to speed up its execution, it is important to succeed in quickly generate the EDK-based systems that the flow requires for the elaboration of the correct partial reconfiguration bitstreams. To achieve this goal, an IP-core generator framework has been developed, it receives as input the VHDL description of the core functionality of a module, automatically produces as output an IP-core suitable to be inserted into an EDK system. This binding can be performed in a faster way than using EDK to re-create each time the entire architecture, exploiting the EDK system creator tool. IP-core generator can be used each time an IP-core has to be created, and not only in a dynamic reconfigurability environment. Several tests are presented to validate the proposed methodology
Matteo Murgida, Alessandro Panella, Vincenzo Rana, Marco D. Santambrogio, Donatella Sciuto
VLSI-SoC5
2006 Affinity-Driven System Design Exploration for Heterogeneous Multiprocessor SoC
abstract
Continuous advances in silicon technology enable the development of complex system-on-chip as cooperation among digital signal processors (DPSs), general purpose processors (GPPs), and specific hardware components. The impact of this choice is not only limited to the target architecture, but also encompasses the overall system specification. It is thus crucial to manage such a complexity using high-level specification languages and a tool chain supporting the designer throughout a set of strategic decisions, such as the identification of a set of possible target architectures, the verification of the correctness of the specification, and the partitioning of the specification onto a set of computational resources. This paper addresses this type of problem by proposing a design flow supporting the system-level design of heterogeneous multiprocessor system-on-chip (MP-SoC), by extracting information from the system description (e.g., SystemC) - statically and in a fast manner - and by providing a set of quantitative measures correlating the type of executor, the functionality, and a timing estimation. Partitioning and architecture selection are built on top of this data and the final analysis of the selected hardware-software solution over the identified candidates is finally submitted to a timing verification via simulation. Note that the possibility of actually performing a comprehensive design space exploration, in general, is tightly influenced by the interaction between partitioning/architecture-selection and timing simulation in the design flow; for this reason, the description of this aspect is particularly emphasized in the presentation of the methodology. To show the applicability of the proposed methodology, two relevant case studies are described in the paper.
Carlo Brandolese, William Fornaciari, Luigi Pomante, Fabio Salice, Donatella Sciuto
IEEE Trans. Computers5
2005 Reliable System Specification for Self-Checking Data-Paths
abstract
The design of reliable circuits has received a lot of attention in the past, leading to the definition of several design techniques introducing fault detection and fault tolerance properties in systems for critical applications/environments. Such design methodologies tackled the problem at different abstraction levels, from switch-level to logic, RT level, and more recently to system level. The aim of this paper is to introduce a novel system-level technique based on the redefinition of the operator functionality in the system specification. This technique provides reliability properties to the system data path, transparently with respect to the designer. Feasibility, fault coverage, performance degradation and overheads are investigated on a FIR circuit.
Cristiana Bolchini, Fabio Salice, Donatella Sciuto, Luigi Pomante
DATE3
2005 Caronte: A Complete Methodology for the Implementation of Partially Dynamically Self-Reconfiguring Systems on FPGA Platforms
abstract
It is common nowadays to employ FPGAS, not only as a means of rapidly prototyping and testing dedicated solutions, but also as a platform on which to implement actual production systems. Although modern FPGAS allow the designer to modify dynamically even only portions of the chip, to this date there is a lack of satisfying design methodologies that using only non-proprietary widely available tools make it possible to optimally implement a high-level specification into a partially dynamically reconfigurable system. The aim of this work is to propose a methodology for solving this problem. The main features of the Caronte methodology are: 1. full exploitation of partial dynamic reconfiguration; 2. the reconfiguration is internal; 3. a real-time Unix-like operating system helps the management of complex systems with multiple tasks, and simplifies reconfiguration through an optimized device driver.
Alberto Donato, Fabrizio Ferrandi, Massimo Redaelli, Marco D. Santambrogio, Donatella Sciuto
FCCM5
2005 Aspect Orientation in System Level Design
Giovanni Agosta, Francesco Bruschi, Donatella Sciuto
FDL3
2005 Mapping Interface Method Calls over OCP Buses
Francesco Bruschi, Federico Moro, Donatella Sciuto
FDL3
2005 Caronte: A methodology for the Implementation of Partially dynamically Self-Reconfiguring Systems on FPGA Platforms
Alberto Donato, Fabrizio Ferrandi, Massimo Redaelli, Marco D. Santambrogio, Donatella Sciuto
VLSI-SoC5
2004 Plug-in of power models in the StepNP exploration platform: analysis of power/performance trade-offs
abstract
In this paper, we propose a power/performance estimation layer designed for StepNP, a system-level architecture simulation and exploration platform for Network Processors and Multi-Processor Systems-on-Chip (MP-SoCs). The first goal of our work is to plug-in PIRATE, a parameterizable Network on-Chip in the StepNP platform, to support a fast exploration of on-chip interconnection networks. Up to now, StepNP does not provide any energy profiling, so our second goal is to dynamically plug-in power models of the different system components to provide power estimates quickly. The proposed power/performance exploration framework is based on a power characterization methodology and a system-level simulator to dynamically profile the given network application. This framework is intended to be used at different levels of the design, considering several levels of accuracy and taking full advantage of the StepNP performance profiling features. Experimental results are provided for the exploration of an ARM-based MP-SOC including a configurable NoC-IP executing an IPv4 forwarding application.
Giovanni Beltrame, Gianluca Palermo, Donatella Sciuto, Cristina Silvano
CASES3
2004 Analysis and Modeling of Energy Reducing Source Code Transformations
abstract
This paper presents a methodology and a set of models supporting energy-driven source-to-source transformations. The most promising code transformation techniques have been isolated and studied leading to accurate analytical and/or statistical models. Experimental results, obtained for some common embedded-system processors over a set of typical benchmarks, are presented, showing the viability of the proposed approach as a support tool for embedded software design.
Carlo Brandolese, William Fornaciari, Fabio Salice, Donatella Sciuto
DATE4
2004 SystemC and SystemVerilog: Where do They Fit? Where are They Going?
abstract
There is tremendous interest in design languages these days - and more particularly, SystemC and SystemVerilog. Sometimes the truth about design languages can be obscured by marketing and the press. This panel is meant to deepen the technical understanding of the DATE audience on the issue of design languages. It contains five technical experts - an academic expert in design languages and SystemC and SystemVerilog in particular; a language expert for each of SystemC and SystemVerilog; and a user expert for these two languages. The language experts have been heavily involved in the specification and evolution of their respective languages. The user experts have been heavily involved in developing use methodologies for these languages within their own design communities, and in applying them to real design problems. The panellists will consider the questions: 1) what are the key capabilities of these languages and what do they offer to users?; 2) which design problems are they best used for? what is their scope?; 3) how has application of these languages to real design problems improved the productivity of designers and the quality of the design results?; and 4) where should the languages develop further capabilities?.
Donatella Sciuto, Grant Martin, Wolfgang Rosenstiel, Stuart Swan, Frank Ghenassia, Peter Flake, Johny Srouji
DATE1
2004 System Level Hardware-Software Design Exploration with XCS
Fabrizio Ferrandi, Pier Luca Lanzi, Donatella Sciuto
GECCO (2)3
2003 Mining interesting patterns from hardware-software codesign data with the learning classifier system XCS
abstract
Embedded systems are composed of both dedicated elements (hardware components) and programmable units (software components), which have to interact with each other for accomplishing a specific task. One of the aims of hardware-software codesign is the choice of a partitioning between elements that will be implemented in hardware and elements that will be implemented in software is one of the important step in design. In this paper, we present an application of the learning classifier system XCS to the analysis of data derived from hardware-software codesign applications. The goal of the analysis is the discovering or explicitation of existing interelationships among system components, which can be used to support the human design of embedded systems. The proposed approach is validated on a specific task involving a digital sound spatializer.
Fabrizio Ferrandi, Pier Luca Lanzi, Donatella Sciuto
IEEE Congress on Evolutionary Computation3
2003 Static analysis of transaction-level models
abstract
The introduction of design languages, such as SystemC 2.0, that allow the modelling of digital systems at the transaction level will impose some major changes to the design flows. Since these formalisms allow for a higher level of abstraction in the systems description, new methodological tools will be needed to support all design phases. The goal of this paper is twofold: first we formalize in an abstract way a significant set of features of a Transaction Level Model, according to the SystemC 2.0 formalism. Then, upon this model we define numerical metrics that can provide useful information in the analysis of the system-level specifications. In particular these metrics are useful in the design exploration phase, to define the main characteristics of the hardware and software architectures
Giovanni Agosta, Francesco Bruschi, Donatella Sciuto
DAC3
2003 Library Functions Timing Characterization for Source-Level Analysis
Carlo Brandolese, William Fornaciari, Fabio Salice, Donatella Sciuto
DATE4
2003 Transaction Based Design: Another Buzzword or the Solution to a Design Problem?
Heinz-Josef Schlebusch, Gary Smith 0001, Donatella Sciuto, Daniel Gajski, Carsten Mielenz, Christopher K. Lennard, Frank Ghenassia, Stuart Swan, Joachim Kunkel
DATE3
2003 An Internal Representation Model for System-Level Co-Design of Heterogeneous Multiprocessor Embedded System
Fabio Salice, William Fornaciari, Luigi Pomante, Donatella Sciuto
FDL4
2003 Identification of design errors through functional testing
abstract
Verification of the functionality of VHDL specifications is one of the primary and most time consuming tasks of design. However, it must necessarily be an incomplete task because it is impossible to completely exercise the specification by exhaustively applying all input patterns. We present a two-step strategy based on symbolic analysis of the VHDL specification, using a behavioral error model. First, we generate a reduced number of functional test vectors for each process of the specification by using a new analysis metric which we call bit coverage. The error model based on this metric allows the identification of possible design errors represented by redundancies in the VHDL code. Then, through the definition of a controllability measure, we verify if these functional test vectors can be applied to the process inputs when it is interconnected to other processes. If this is not the case, the analysis of the nonapplicable inputs provides identification of possible design errors due to erroneous interconnections. The bit-coverage provides complete statement, condition and branch coverage; and we experimentally show that it allows the identification of possible design errors. Identification and removal of design errors improves the global testability of a design.
Fabrizio Ferrandi, Franco Fummi, Graziano Pravadelli, Donatella Sciuto
IEEE Trans. Reliab.4
2002 Energy estimation and optimization of embedded VLIW processors based on instruction clustering
abstract
Aim of this paper is to propose a methodology for the definition of an instruction-level energy estimation framework for VLIW (Very Long Instruction Word) processors. The power modeling methodology is the key issue to define an effective energy-aware software optimisation strategy for state-of-the-art ILP (Instruction Level Parallelism) processors. The methodology is based on an energy model for VLIW processors that exploits instruction clustering to achieve an efficient and fine grained energy estimation. The approach aims at reducing the complexity of the characterization problem for VLIW processors from exponential, with respect to the number of parallel operations in the same very long instruction, to quadratic, with respect to the number of instruction clusters. Furthermore, the paper proposes a spatial scheduling algorithm based on a low-power reordering of the parallel operations within the same long instruction. Experimental results have been carried out on the Lx processor, a 4-issue VLIW core jointly designed by HPLabs and STMicroelectronics. The results have shown an average error of 1.9% between the cluster-based estimation model and the reference design, with a standard deviation of 5.8%. For the Lx architecture, the spatial instruction scheduling algorithm provides an average energy saving of 12%.
Andrea Bona, Mariagiovanna Sami, Donatella Sciuto, Vittorio Zaccaria, Cristina Silvano, Roberto Zafalon
DAC3
2002 An Instruction-Level Methodology for Power Estimation and Optimization of Embedded VLIW Cores
abstract
Summary form only given. The overall goal of this work is to define an instruction-level power macro-modeling and characterization methodology for VLIW embedded processor cores. The approach presented in this paper is a major extension of the work previously proposed, targeting an instruction-level energy model to evaluate the energy consumption associated with a program execution on a pipelined VLIW core. Our ongoing work aims at defining a power optimization technique based on the proposed model. The technique consists of a spatial rescheduling of the operations within the same long instruction to reduce their instruction power overhead.
Andrea Bona, Mariagiovanna Sami, Donatella Sciuto, Vittorio Zaccaria, Cristina Silvano, Roberto Zafalon
DATE3
2002 Error Simulation Based on the SystemC Design Description Language
abstract
Summary form only given. The combined effects of devices increased complexity and reduced design cycle time creates a testing problem: an increasing larger portion of the design time is devoted to testing and verification. Today EDA tools, moving towards higher levels of abstraction, promise greater designer productivity, resulting in increased design complexity and size. In order to reduce the testing and verification time, different high-level approaches have been proposed in literature. Most of these approaches are based on the definition of an error or fault model, applicable at a higher level of abstraction of the description of the system to be implemented. In this paper we concentrate our attention on the evaluation of error models, used in test generation and in functional verification. Evaluation of error models is also an important aspect when fault injection methodologies are used to evaluate the dependability of complex system.
Francesco Bruschi, Michele Chiamenti, Fabrizio Ferrandi, Donatella Sciuto
DATE4
2002 Functional Verification for SystemC Descriptions Using Constraint Solving
abstract
This paper addresses the problem of test vectors generation starting from an high level description of the system under test, specified in SystemC. The verification method considered is based upon the simulation of input sequences. The system model adopted is the classical Finite State Machine model. Then, according to different strategies, a set of sequences can be obtained, where a sequence is an ordered set of transitions. For each of these sequences, a set of constraints is extracted. Test sequences can be obtained by generating and solving the constraints, by using a constraint solver (GProlog). A solution of the constraint solver yields the values, of the input signals for which a sequence of transitions in the FSM is executed. If the constraints cannot be solved, it implies that the corresponding sequence cannot be executed by any test. The presented algorithm is not based on a specific fault model, but aims at reaching the highest possible path coverage.
Fabrizio Ferrandi, Michele Rendine, Donatella Sciuto
DATE3
2002 Reliability Properties Assessment at System Level: A Co-Design Framework
Cristiana Bolchini, Luigi Pomante, Fabio Salice, Donatella Sciuto
J. Electron. Test.4
2002 Behavioral test generation for the selection of BIST logic
Giuseppe Biasoli, Fabrizio Ferrandi, Alessandro Fin, Franco Fummi, Donatella Sciuto
J. Syst. Archit.5
2002 Test Generation and Testability Alternatives Exploration of Critical Algorithms for Embedded Applications
abstract
Presents an analysis of the behavioral descriptions of embedded systems to generate behavioral test patterns that are used to perform the exploration of design alternatives based on testability. In this way, during the hardware/software partitioning of the embedded system, testability aspects can be considered. This paper presents an innovative error model for algorithmic (behavioral) descriptions, which allows for the generation of behavioral test patterns. They are converted into gate-level test sequences by using more-or-less accurate procedures based on scheduling information or both scheduling and allocation information. The paper shows, experimentally, that such converted gate-level test sequences provide a very high stuck-at fault coverage when applied to different gate-level implementations of the given behavioral specification. For this reason, our behavioral test patterns can be used to explore testability alternatives, by simply performing fault simulation at the gate level with the same set of patterns, without regenerating them for each circuit. Furthermore, whenever gate-level ATPGs are applied on the synthesized gate-level circuits, they obtain lower fault coverage with respect to our behavioral test patterns, in particular when considering circuits with hard-to-detect faults.
Fabrizio Ferrandi, Franco Fummi, Donatella Sciuto
IEEE Trans. Computers3
2002 Static power modeling of 32-bit microprocessors
abstract
The paper presents a novel strategy aimed at modeling instruction energy consumption of 32-bit microprocessors. Different from former approaches, the proposed instruction-level power model is founded on a functional decomposition of the activities accomplished by a generic microprocessor. The proposed model has significant generalization capabilities. It allows estimation of the power figures of the entire instruction-set starting from the analysis of a subset, as well as to power characterize new processors by using the model obtained by considering other microprocessors. The model is formally presented and justified and its actual application over five commercial microprocessors is included. This static characterization is the basic information for system-level power modeling of hardware/software architectures.
Carlo Brandolese, Fabio Salice, William Fornaciari, Donatella Sciuto
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2002 An instruction-level energy model for embedded VLIW architectures
abstract
In this paper, an instruction-level energy model is proposed for the data-path of very long instruction word (VLIW) pipelined processors that can be used to provide accurate power consumption information during either an instruction-level simulation or power-oriented scheduling at compile time. The analytical model takes into account several software-level parameters (such as instruction ordering, pipeline stall probability, and instruction cache miss probability) as well as microarchitectural-level ones (such as pipeline stage power consumption per instruction) providing an efficient pipeline-aware instruction-level power estimation, whose accuracy is very close to those given by RT or gate-level simulations. The problem of instruction-level power characterization of a K-issue VLIW processor is O(N/sup 2K/) where N is the number of operations in the ISA and K is the number of parallel instructions composing the very long instruction. One of the advantages of the proposed model consists of reducing the complexity of the characterization problem to O(K/spl times/N/sup 2/). The proposed model has been used to characterize a four-issue VLIW core with a six-stage pipeline.
Mariagiovanna Sami, Donatella Sciuto, Cristina Silvano, Vittorio Zaccaria
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2002 Low-power data forwarding for VLIW embedded architectures
abstract
Proposes a low-power approach to the design of embedded very long instruction word (VLIW) processor architectures based on the forwarding (or bypassing) hardware, which provides operands from interstage pipeline registers directly to the inputs of the function units. The power optimization technique exploits the forwarding paths to avoid the power cost of writing/reading short-lived variables to/from the register file (RF). Such optimization is justified by the fact that, in application-specific embedded systems, a significant number of variables are short-lived, that is, their liveness (from first definition to last use) spans only few instructions. Values of short-lived variables can thus be accessed directly through the forwarding registers, avoiding writeback to the RF by the producer instruction and successive read from the RF by the consumer instruction. The decision concerning the enabling of the RF writeback phase is taken at compile time by the compiler static scheduling algorithm. This approach implies a minimal overhead on the complexity of the processor control logic and, thus, no critical path increase. The application of the proposed solution to a VLIW embedded core has shown an average RF power saving of 7.8% with respect to the unoptimized approach on the given set of target benchmarks.
Mariagiovanna Sami, Donatella Sciuto, Cristina Silvano, Vittorio Zaccaria, Roberto Zafalon
IEEE Trans. Very Large Scale Integr. Syst.2
2001 Functional test generation for behaviorally sequential models
abstract
Functional testing of HDL specifications is one of the most promising approaches for the verification of the functionalities of a design before synthesis. The contribution of this work is the development of a test generation algorithm targeting a new coverage metric (called bit-coverage) that provides full statement coverage, branch coverage, condition coverage and partial path coverage for behaviorally sequential models. The behavioral test sequences can be also the only way to evaluate testability of VHDL model for which a gate-level representation is not available (e.g third-party cores), since the behavioral error model is characterized also by a high correlation with the RT and gate-level stuck-at fault model. Moreover the preciseness of the proposed coverage metric makes the identified test sequences more effective in identifying design errors, than other test patterns developed by following standard coverage metrics.
Fabrizio Ferrandi, G. Ferrara, Donatella Sciuto, Alessandro Fin, Franco Fummi
DATE3
2001 Exploiting data forwarding to reduce the power budget of VLIW embedded processors
abstract
In this paper, a low-power approach to the design of embedded VLIW processor architectures is proposed. To solve the most part of data hazards in the pipeline, processors use forwarding (or bypassing) hardware to provide the required operands from the inter-stage pipeline registers directly to the inputs of the function units. The operands are then stored in the register file during the write-back pipeline stage. In this paper, we propose a power optimization technique based on the exploitation of the forwarding paths in the processor to avoid the power cost of writing/reading short-lived variables to/from the register file. In application-specific embedded systems, experimental evidence has shown that a significant number of variables are short-lived, that is their liveness (front first definition to last use) spans only few instructions. Values of short-lived variables can be accessed directly through the forwarding registers, avoiding write-back. An application example of our solution to a VLIW embedded core, when accessing the register file, has shown a power saving up to 35% with respect to the unoptimized approach on the given set of target benchmarks. The performance overhead is equal to one-gate delay to be added on the processor critical-path.
Mariagiovanna Sami, Donatella Sciuto, Cristina Silvano, Vittorio Zaccaria, Roberto Zafalon
DATE2
2001 An Assembly-Level Execution-Time Model for Pipelined Architectures
abstract
The aim of this work is to provide an elegant and accurate static execution timing model for 32-bit microprocessor instruction sets, covering also inter-instruction effects. Such effects depend on the processor state and the pipeline behavior, and are related to the dynamic execution of assembly code. The paper proposes a mathematical model of the delays deriving from instruction dependencies and gives a statistical characterization of such timing overheads. The model has been validated on a commercial architecture, the Intel486, by means of timing analysis of a set of benchmarks, obtaining an error within 5%. This model can be seamlessly integrated with a static energy consumption model in order to obtain precise software power and energy estimations.
Giovanni Beltrame, Carlo Brandolese, William Fornaciari, Fabio Salice, Donatella Sciuto, Vito Trianni
ICCAD5
2001 An efficient heuristic approach to solve the unate covering problem
abstract
The paper presents a new approach to solve the unate covering problem based on exploitation of information provided by Lagrangean relaxation. In particular, main advantages of the proposed heuristic algorithm are the effective choice of elements to be included in the solution, cost-related reductions of the problem, and a good lower bound on the optimum. The results support the effectiveness of this approach: on a wide set of benchmark problems, the algorithm nearly always hits the optimum and in most cases proves it to be such. On the problems whose optimum is actually unknown, the best known result is strongly improved.
Roberto Cordone, Fabrizio Ferrandi, Donatella Sciuto, Roberto Wolfler Calvo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2000 An instruction-level functionally-based energy estimation model for 32-bits microprocessors
abstract
The paper presents a novel strategy aimed at modeling the instruction energy consumption of 32-bits microprocessors. The proposed instruction-level pow er model is founded on afunctional decomposition of the activities accomplished by a generic microprocessor and exhibits significant generalization capabilities. It allo ws estimation of the pow er figures of the en tire instruction-set starting from the analysis of a subset, as w ell as to po w er characterize new processors using the model obtained by considering other microprocessors.
Carlo Brandolese, William Fornaciari, Fabio Salice, Donatella Sciuto
DAC4
2000 An Efficient Heuristic Approach to Solve the Unate Covering Problem
abstract
The classical solving approach for two-level logic minimisation reduces the problem to a special case of unate covering and attacks the latter with a (possibly limited) branch-and-bound algorithm. We adopt this approach, but we propose a constructive heuristic algorithm that combines the use of Binary Decision Diagrams (BDDs) with the Lagrangian relaxation. This technique permits us to achieve an effective choice of the elements to include in the solution, as well as cost-related reductions of the problem and a good lower bound on the optimum. The results support the effectiveness of this approach: on a wide set of benchmark problems, the algorithm nearly always hits the optimum, and in most cases proves it to be so. On the problems whose optimum is actually unknown, the best known result is strongly improved.
Roberto Cordone, Fabrizio Ferrandi, Donatella Sciuto, Roberto Wolfler Calvo
DATE3
2000 Power Exploration for Embedded VLIW Architectures
abstract
In this paper, we propose a system-level power exploration methodology for embedded VLIW architectures based on an instruction-level analysis. The instruction-level energy model targets a general pipeline scalar processor; several architectural parameters such as number and type of pipeline stages as well as average stall/latency cycles per instruction and inter-instruction effects are taken into account. The application of the proposed model to VLIW processors results intractable from the point of view of both spatial and temporal complexity (which grow exponentially w.r.t. the number of possible operations in the ISA). To reduce this complexity, the basic model has been extended by assuming that the energy associated with a long instruction is given by the sum of the energy associated with the single operations of the long instruction and the single pipeline stages. The instruction-level energy model has been applied to a simplified VLIW architecture to demonstrate the validity of the proposed approach.
Mariagiovanna Sami, Donatella Sciuto, Cristina Silvano, Vittorio Zaccaria
ICCAD2
2000 An Application of Genetic Algorithms and BDDs to Functional Testing
abstract
This paper describes a functional level rest pattern generator, which combines two techniques: genetic algorithms (GAs) and binary decision diagrams (BDDs). The combined execution of such two techniques achieves better results for functional testing, than the single application of each separated technique. The entire set of functional errors is examined in a shorter time and a more compact test set is produced. The reason of this interesting result has been analyzed in the paper. It mainly depends on the fact that hard to detect errors for GA-based testing techniques are easy to detect than errors for BDD-based techniques and vice versa. The two testing approaches are thus complementary and can effectively cooperate.
Fabrizio Ferrandi, Donatella Sciuto, Alessandro Fin, Franco Fummi
ICCD2
2000 Low-power state assignment techniques for finite state machines
abstract
The problem of minimizing the power consumption in synchronous sequential circuits is explored in this paper. We present a general theoretical framework to solve the state assignment problem for Finite State Machines (FSMs). In this framework, the problem has been separated in two different tasks. First, we define heuristic techniques to visit the State Transition Graph (STG) and thus to assign a priority to the symbolic states. Second, we define encoding techniques to assign binary codes to the symbolic states to reduce the switching activity of state registers. Based on this approach, we propose five power-oriented state assignment algorithms. These techniques have been applied to MCNC benchmark circuits and the experimental results have shown an average reduction in transition activity (power) of 8.56% (5.35%) over well-known low-power state encoding schemes.
Paolo Bacchetta, Lidia Daldoss, Donatella Sciuto, Cristina Silvano
ISCAS3
2000 Testability Alternatives Exploration through Functional Testing
abstract
The aim of this paper is to show the effectiveness of a high-level approach to testability analysis and test pattern generation, when analyzing different classes of architectures implementing the same specification. A unique test set is derived on the behavioral specification, based on a functional error model, which shows a high correlation with the single stuck-at-gate-level fault model. Such a test set is then tailored to the particular gate-level implementation by transforming it into a specific test sequence, based on the scheduling adopted by the high-level synthesis. Experimental results show that the application of such test sequences allows one to accurately evaluate the testability of the architecture in terms of gate-level fault coverage, in a fraction of the time required by a gate-level test pattern generator.
Fabrizio Ferrandi, G. Fornara, Donatella Sciuto, G. Ferrara, Franco Fummi
VTS3
2000 An extended-UIO-based method for protocol conformance testing
Giacomo Buonanno, Franco Fummi, Donatella Sciuto
J. Syst. Archit.3
2000 A Hierarchical Test Generation Approach for Large Controllers
abstract
A testing approach targeted at Hardware Description Language (HDL)-based specifications of complex control devices is proposed. For such architectures, gate-level test pattern generators require insertion of scan paths to enable the flat gate-level representations to be efficiently handled. In contrast, we present a testing methodology based on the hierarchical finite state machine model. Our approach allows the generation of compact test sets with very high stuck-at fault coverages, without any design-for-testability logic other than hardware reset. This method can be used any time the functional information is available together with the gate-level structural description. High fault coverages are achieved with smaller test lengths and execution times with respect to state-of-the-art gate-level test pattern generators.
Franco Fummi, Donatella Sciuto
IEEE Trans. Computers2
2000 Symbolic optimization of interacting controllers based onredundancy identification and removal
abstract
This paper presents a binary decision diagram (BDD)-based algorithm for the optimization of the driven machine, M/sub 2/, of a finite-state machine (FSM) network with cascade connection, M/sub 1//spl rarr/M/sub 2/. The technique we propose relies on redundant faults identification and removal. A fault, f, located into machine M/sub 2/, is redundant with respect to the overall network if the driving machine M/sub 1/ is not able to generate any test sequence for such a fault. When the state transition graph (STG) specifications of the network components are available, the standard way for checking the redundancy condition for the considered fault requires one to first construct the product machine M/sub 2//spl times/M/sub 2//sup F/, where M/sub 2//sup F/ is the faulty FSM, then to connect it to the driving machine, and finally to perform reachability analysis on the composed machine M/sub 1//spl rarr/M/sub 2//spl times/M/sub 2//sup F/. Clearly, the size of such machine limits the applicability of the approach above to systems whose components have a few tens of states at most, even when symbolic traversal algorithms are used. Since we are interested in dealing with networks of larger FSM's (i.e., machines whose STGs can not be represented explicitly), we propose to use the product automaton P'=A/sub 1//spl times/A/sub f/, where A/sub 1/' is the finite automaton (FA) accepting all the output sequences of M/sub 1/, and A/sub f/ is the FA accepting all the test sequences for fault f, instead of machine M/sub 1//spl rarr/M/sub 2//spl times/M/sub 2//sup F/. This simplifies sensibly the task of the reachability analysis program, since A/sub f/ has considerably less states and less edges than the product machine M/sub 2//spl times/M/sub 2//sup F/ and, thus, the size of the BDD representation of its transition relation is much more easily manageable. In addition, differently from other approaches, automaton A/sub 1/' is not required to be deterministic and state minimal. This allows us to avoid the application of determinization and state minimization procedures whose complexity is exponential. We present experimental results For examples (i.e., network of interacting controllers) on which existing optimization methods are not applicable, due to the size of the component FSM's. We also provide a comparison to the data produced by state-of-the-art FSM network optimizers on small benchmarks in order to show the effectiveness of our approach.
Fabrizio Ferrandi, Franco Fummi, Enrico Macii, Massimo Poncino, Donatella Sciuto
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2000 Design of VHDL-based totally self-checking finite-state machine and data-path descriptions
abstract
This paper presents a complete methodology to design a totally self-checking (TSC) sequential system based on the generic architecture of finite-state machine and data path (FSMD), such as the one deriving from VHDL specifications. The control part of the system is designed to be self-checking by adopting a state assignment providing a constant Hamming distance between each pair of binary codes. The design of the data path is based on both classical methodologies (e.g., parity, Berger code) and ad hoc strategies (e.g., multiplexer cycle) suited for the specific circuit structure. Self-checking properties and costs are evaluated on a set of benchmark FSM's and on a number of VHDL circuits.
Cristiana Bolchini, R. Montandon, Fabio Salice, Donatella Sciuto
IEEE Trans. Very Large Scale Integr. Syst.4
1999 Symbolic Functional Vector Generation for VHDL Specifications
abstract
Verification of the functional correctness of VHDL specifications is one of the primary and most time consuming tasks of design. However, it must necessarily be an incomplete task since it is impossible to completely exercise the specification by exhaustively applying all input patterns. The paper aims at presenting a two-step strategy based on symbolic analysis of the VHDL specification, using a behavioral fault model. First, we generate a reduced number of functional test vectors for each process of the specification which allows complete code statement coverage and bit coverage, allowing the identification of possible redundancies in the VHDL process. Then, through the definition of a controllability measure, we verify if these functional test vectors can be applied to the process inputs when interconnected to other processes. If this is not the case, the analysis of the nonapplicable inputs provides identification of possible code redundancies and design errors. Experimental results show that bit coverage provides complete statement coverage and a more detailed identification of possible design errors.
Fabrizio Ferrandi, Franco Fummi, Luca Gerli, Donatella Sciuto
DATE4
1999 Influence of Caching and Encoding on Power Dissipation of System-Level Buses for Embedded Systems
abstract
This paper proposes a methodology to evaluate the effects of encodings on the power consumption of system-level buses in the presence of multi-level cache memories. The proposed model can consider any cache configuration in terms of size, associativity and block. It includes also the most widely adopted power oriented encoding techniques for data and address buses. Experimental results show how the proposed model can be effectively adopted to configure the memory hierarchy and the system bus architecture from the power point of view.
William Fornaciari, Donatella Sciuto, Cristina Silvano
DATE2
1999 Power Estimation of System-Level Buses for Microprocessor-Based Architectures: A Case Study
abstract
The processor-to-memory communication on system-level buses dissipates a significant amount of the overall power in microprocessor-based architectures. A methodology has been set up to evaluate the effects of both encoding schemes and multi-level cache memories on the power consumption associated with the system-level address and data buses of a high-end computer system based on the PowerPC604e architecture. The main goal is to evaluate how different values of cache parameters (cache size, block size, associativity write strategy, and block replacement policy) and the introduction of bus encoding techniques, at the different levels of the memory hierarchy, affect the system-level power dissipation.
William Fornaciari, Donatella Sciuto, Cristina Silvano
ICCD2
1999 Synthesis for Testability of Highly Complex Controllers by Functional Redundancy Removal
abstract
This paper presents a testable synthesis methodology applicable to any top-down design method based on hardware-description-language descriptions, or graphical representations. The methodology is targeted on control-dominated applications and it is based on the identification and removal of a new class of redundant faults, called functionally redundant faults. The formal relation between functionally redundant faults and sequentially redundant faults is introduced. Moreover, the relation between functionally redundant faults and logic synthesis algorithms based on local don't cares is shown. Functionally redundant faults are identified and removed by comparing the implemented synchronous sequential circuit, which can be technology dependent, to its specification. The specification can be a single finite state machine (FSM), a set of interacting FSMs, or a hierarchical FSM that allows the description of highly complex controllers. The proposed methodology produces testable circuits, with area reduction, still mapped on the same technology library, and it manages circuits which cannot be handled by other methods presented in the literature.
Franco Fummi, Donatella Sciuto, Micaela Serra
IEEE Trans. Computers2
1998 A Model for System-Level Timed Analysis and Profiling
abstract
Fast evaluation of functional and timing properties is becoming a key factor to enable cost-effective exploration of mixed hw/sw design alternatives for embedded applications. The goal of this paper is to present a modeling strategy to specify functionality and timing properties of uncommitted mixed hw/sw systems. In addition, the paper proposes a simulation algorithm able to perform fast high-level simulation of the system by taking into account the initial hw vs. sw allocation of system modules. The related CAD simulation environment allows the designer to access profiling information which can be useful to remodel the system to meet the functional/timing goals as well as to drive the following hw vs sw partioning activity. Experimental data obtained by reengineering an industrial design are also included in the paper.
Alberto Allara, William Fornaciari, Fabio Salice, Donatella Sciuto
DATE4
1998 Address Bus Encoding Techniques for System-Level Power Optimization
abstract
The power dissipated by system-level buses is the largest contribution to the global power of complex VLSI circuits. Therefore, the minimization of the switching activity at the I/O interfaces can provide significant savings on the overall power budget. This paper presents innovative encoding techniques suitable for minimizing the switching activity of system-level address buses. In particular, the schemes illustrated here target the reduction of the average number of bus line transitions per clock cycle. Experimental results, conducted on address streams generated by a real microprocessor, have demonstrated the effectiveness of the proposed methods.
Luca Benini, Giovanni De Micheli, Donatella Sciuto, Enrico Macii, Cristina Silvano
DATE3
1998 Fault Analysis in Networks with Concurrent Error Detection Properties
abstract
The design of self-checking circuits through output encoding finds a bottleneck in the realization of the network so that each fault produces only errors detectable by the adopted code. An analysis of an expected TSC network is proposed, based on the application of the weighted observability approach. The aim is the verification of the SC property of the encoded circuit (TSC fault simulation) and identification of critical areas for a consequent manipulation to achieve a complete fault coverage.
Cristiana Bolchini, Fabio Salice, Donatella Sciuto
DATE3
1998 VHDL Testability Analysis Based on Fault Clustering and Implicit Fault Injection
abstract
Testability analysis of VHDL sequential models is the main topic of this paper. We investigate the possibility to obtain information about the testability of a sequential VHDL description before its actual synthesis. The analysis is based on an implicit fault model that injects faults into a BDD based description extracted from the VHDL representation. Such an injection is related to the original VHDL representation thus allowing the identification of potential testability problems before RTL and logic synthesis. Fault injection is performed efficiently by exploiting the concept of fault clustering, that is the possibility of grouping faults and analyzing them concurrently. The proposed methodology is applied to benchmarks for efficiency evaluation and to a real VHDL description.
F. S. Bietti, Fabrizio Ferrandi, Franco Fummi, Donatella Sciuto
Great Lakes Symposium on VLSI4
1998 System-level performance estimation strategy for sw and hw
abstract
The design of an embedded system is a process where the tuning of the architecture should take into account both the functionality and the timing performance while considering the heterogeneity of the hw and sw components. The goal of this paper is to present the new model developed during the SEED Esprit project, to estimate the software and hardware characteristics for cosimulation and profiling within the TOSCA codesign framework. The impact on the design space exploration of such an high-level cosimulation strategy has been tested by considering as a benchmark the reengineering of an industrial device.
Alberto Allara, Carlo Brandolese, William Fornaciari, Fabio Salice, Donatella Sciuto
ICCD5
1998 Automatic VHDL restructuring for RTL synthesis optimization and testability improvement
abstract
A methodology for modifying VHDL descriptions is the core of this paper. Modifications are performed on general RTL descriptions composed of a mix of control and computation, that is, the typical type of description used for designing at the RT level. Such VHDL descriptions are automatically partitioned into a reference model composed of a controller driving a data-path. We call this transformation "VHDL restructuring". A set of restructuring steps is presented aiming at partitioning any VHDL description while guaranteeing the semantic equivalence of the restructured description with the original one. The main motivation to restructuring is the identification and separation of the two parts (FSM+data-path) which can thus be analyzed by using "ad hoc" synthesis, testability and design for testability algorithms. Promising results show that restructuring can sensibly impact on synthesis and testability.
D. Corvino, Italo Epicoco, Fabrizio Ferrandi, Franco Fummi, Donatella Sciuto
ICCD5
1998 Implicit test generation for behavioral VHDL models
abstract
This paper proposes a behavioral-level test pattern generation algorithm for behavioral VHDL descriptions. The proposed approach is based on the comparison between the implicit description of the fault-free behavior and the faulty behavior, obtained through a new behavioral fault model. The paper will experimentally show that the test patterns generated at the behavioral level provide a very high stuck-at fault coverage when applied to different gate-level implementations of the given VHDL behavioral specification. Gate-level ATPGs applied on these same circuits obtain lower fault coverage, in particular when considering circuits with hard to detect faults.
Fabrizio Ferrandi, Franco Fummi, Donatella Sciuto
ITC3
1998 Clock skew reduction in ASIC logic design: a methodology for clock tree management
abstract
This paper presents a methodology for the automatic generation of clock trees in an ASIC design at the gate level. New algorithms and heuristics are described: they have been inserted with success in an industrial ASIC design flow, after the logic synthesis and optimization step. Our algorithms, by different heuristic methods, particularly take into account those elements connected as transmitter-receiver couples which represent the most critical configurations for circuit synchronization. Improvement of clock tree performance has also been obtained by means of an interaction strategy between logic and physical design phases. Such a strategy drives the placement of the clock tree elements in an equidistant way, in order to obtain a controlled routing.
Alessandro Balboni, Claudio Costi, Massimo Pellencin, Andrea Quadrini, Donatella Sciuto
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
1998 Testability analysis and behavioral testing of the Hopfield neural paradigm
abstract
Testability analysis and test pattern generation for neural architectures can be performed at a very high abstraction level on the computational paradigm. In this paper, we consider the case of Hopfield's networks, as the simplest example of networks with feedback loops. A behavioral error model based on finite-state machines (FSM's) is introduced. Conditions for controllability, observability and global testability are derived to verify errors excitation and propagation to outputs. The proposed behavioral test pattern generator creates the minimum length test sequence for any digital implementation.
Cesare Alippi, Franco Fummi, Vincenzo Piuri, Mariagiovanna Sami, Donatella Sciuto
IEEE Trans. Very Large Scale Integr. Syst.5
1998 Power estimation of embedded systems: a hardware/software codesign approach
abstract
The need for low-power embedded systems has become very significant within the microelectronics scenario in the most recent years. A power-driven methodology is mandatory during embedded systems design to meet system-level requirements while fulfilling time-to-market. The aim of this paper is to introduce accurate and efficient power metrics included in a hardware/software (HW/SW) codesign environment to guide the system-level partitioning. Power evaluation metrics have been defined to widely explore the architectural design space at high abstraction level. This is one of the first approaches that considers globally HW and SW contributions to power in a system-level design flow for control dominated embedded systems.
William Fornaciari, Paolo Gubian, Donatella Sciuto, Cristina Silvano
IEEE Trans. Very Large Scale Integr. Syst.3
1998 Automatic generation of error control codes for computer applications
abstract
This paper proposes a methodology, implemented in a tool, to automatically generate the main classes of error control codes (ECC's) widely applied in computer memory systems to increase reliability and data integrity. New code construction techniques extending the features of previous single error correcting (SEC)-double error detecting (DED)-single byte error detecting (SBD) codes have been integrated in the tool. The proposed techniques construct systematic odd-weight-column SEC-DED-SBD codes with odd-bit-per-byte error correcting (OBC) capabilities to enhance reliability in high speed memory systems organized as multiple-bit-per-chip or card.
Franco Fummi, Donatella Sciuto, Cristina Silvano
IEEE Trans. Very Large Scale Integr. Syst.2
1997 Asymptotic Zero-Transition Activity Encoding for Address Busses in Low-Power Microprocessor-Based Systems
abstract
In microprocessor-based systems, large power savings can be achieved through reduction of the transition activity of the on- and off-chip buses. This is because the total capacitance being switched when a voltage change occurs on a bus line is usually sensibly larger than the capacitive load that must be charged/discharged when internal nodes toggle. In this paper, we propose an encoding scheme which is suitable for reducing the switching activity on the lines of an address bus. The technique relies on the observation that, in a remarkable number of cases, patterns traveling onto address buses are consecutive. Under this condition it may therefore be possible, for the devices located at the receiving end of the bus, to automatically calculate the address to be received at the next clock cycle; consequently, the transmission of the new pattern can be avoided, resulting in an overall switching activity decrease. We present analytical and experimental analyses showing the improved performance of our encoding scheme when compared to both binary and Gray addressing schemes, the latter being widely accepted as the most efficient method for address bus encoding. We also propose power and timing efficient implementations of the encoding and the decoding logic, and we discuss the applicability of the technique to real microprocessor-based designs.
Luca Benini, Giovanni De Micheli, Enrico Macii, Donatella Sciuto, Cristina Silvano
Great Lakes Symposium on VLSI4
1997 Parity Bit Code: Achieving a Complete Fault Coverage in the Design of TSC Combinational Networks
abstract
A new methodology for designing Totally Self-Checking combinational circuits through the encoding of the primary outputs with the parity code is presented. The parity code requires that each fault modifies an odd number of outputs for providing its detection, that is, each fault has to be oddly observable. The proposed methodology for fulfilling such a constraint consists of a post-synthesis modification of fault observability through either the introduction of an auxiliary output for the examined network node or the replication of the investigated node. A cost evaluation function allows us to select the most convenient solution in terms of overhead and the final 100% TSC circuit.
Cristiana Bolchini, Fabio Salice, Donatella Sciuto
Great Lakes Symposium on VLSI3
1997 How an "Evolving" Fault Model Improves the Behavioral Test Generation
abstract
By considering test costs at the behavioral level, test problems can be pointed out during the first phases of the design flow. Thus, in case either some testability problems are identified or the size (and hence the cost) of the test set results to be too high, the designer or the high level synthesis tool can modify the circuit to reduce such testability problems. The main problem is the correspondence between the behavioral and RT or gate level fault models. To overcome such limitation, the paper presents a design flow based on the behavioral fault model modification ("evolution") depending on the actual RTL implementation.
Giacomo Buonanno, Fabrizio Ferrandi, L. Ferrandi, Franco Fummi, Donatella Sciuto
Great Lakes Symposium on VLSI5
1997 Improving Design Turnaround Time via Two-Levels Hw/Sw Co-Simulation
abstract
The steadily growing demand of fast turnaround time will shift system tuning from physical prototyping to virtual prototyping. The paper proposes a novel approach for mixed HW-SW implementation of embedded systems, allowing high-level simulation of the overall architecture as well as a deeper analysis of timing performance by exploiting commercial VHDL CAD tools. At the higher level, functional debugging and tradeoff analysis is performed on an OCCAM-based system-level model and at the lower level a VHDL-based description for both the HW and SW is built for fine grain verification of the system. The paper introduces the two levels of simulation, showing their impact in terms of design flow management and design time.
Alberto Allara, S. Filipponi, William Fornaciari, Fabio Salice, Donatella Sciuto
ICCD5
1997 Application of a Testing Framework to VHDL Descriptions at Different Abstraction Levels
abstract
The test problem increasingly affects the system design process and related costs and time to market. Requirements from VLSI/WSI manufacturers are for fast and reliable testability tools, with the possibility of their introduction in early phases of design. The paper presents a global toolset architecture for testability analysis and test pattern generation. Three abstraction levels are considered in this design flow, from the behavioral specifications, through RTL descriptions, down to gate level. In all these phases, VHDL is chosen as the referring description language. The paper then presents an application scenario, detailing the results achieved by the proposed methodology.
M. Bacis, Giacomo Buonanno, Fabrizio Ferrandi, Franco Fummi, Luca Gerli, Donatella Sciuto
ICCD6
1997 A TSC Evaluation Function for Combinational Circuits
abstract
The paper presents an innovative evaluation function for circuits with on-line detecting properties, which considers other aspects beyond area overhead. In particular, this function takes into account the probability of detecting a fault, once it occurs, with respect to the network structure and the application of input configurations. Different implementations of the same device designed to have TSC properties are compared with respect to this innovative evaluation function.
Cristiana Bolchini, Donatella Sciuto, Fabio Salice
ICCD2
1997 Implicit test pattern generation constrained to cellular automata embedding
abstract
This paper presents an implicit methodology that constrains a test pattern generator to identify test sequences which can be reproduced by cellular automata (CA). The so identified CA can be synthesized as an autonomous finite state machine and can be attached to the inputs of a circuit under test (e.g., a controller). In this way, the circuit under test preserves its integrity and its performance is not affected by the proposed testing technique. The overall device (controller+CA) is an off-line self-testable circuit which con potentially self-test all stuck-at faults both of the controller and the CA. Thus, the method can be viewed as a BIST strategy based on the embedding of deterministic test sequences.
Franco Fummi, Donatella Sciuto
VTS2
1997 A complete testing strategy based on interacting and hierarchical FSMs
Franco Fummi, Donatella Sciuto
Integr.2
1997 A VHDL-based approach for power estimation of embedded systems
William Fornaciari, Paolo Gubian, Donatella Sciuto, Cristina Silvano
J. Syst. Archit.3
1997 Special section on VHDL
Donatella Sciuto
J. Syst. Archit.1
1997 Functional design for testability of control-dominated architectures
abstract
Control-dominated architectures are usually described in a hardware description language (HDL) by means of interacting FSMs. A VHDL or Verilog specification can be translated into an interacting FSM (IFSM) representation as described here. The IFSM model allows us to approach the testable synthesis problem at the level of each FSM. The functionality is modified by the addition of transparency to data flow. The complete testability of the IFSM implementation is thus achieved by connecting fully testable implementations of each modified FSM. In this way, test sequences separately generated for each FSM are directly applied to the IFSM to achieve complete fault coverage. The addition of test functionality to each FSM description, and its simultaneous synthesis with the FSM functionality, produces a lower area overhead than that necessary for the application of a partial-scan technique. Moreover, the test generation problem is highly simplified since it is reduced to the test generation for each separate FSM.
Franco Fummi, U. Rovati, Donatella Sciuto
ACM Trans. Design Autom. Electr. Syst.3
1996 Symbolic Optimization of FSM Networks Based on Sequential ATPG Techniques
abstract
This paper presents a novel optimization algorithm for FSM networks that relies on sequential test generation and redundancy removal. The implementation of the proposed approach, which is based on the exploitation of input don't care sequences through regular language intersection, is fully symbolic. Experimental results, obtained on a large set of standard benchmarks, improve over the ones of state-of-the-art methods.
Fabrizio Ferrandi, Franco Fummi, Enrico Macii, Massimo Poncino, Donatella Sciuto
DAC5
1996 Test Generation for Networks of Interacting FSMs Using Symbolic Techniques
abstract
This paper presents a new testing strategy for networks of interacting FSMs. The approach allows us to generate test patterns for faults in the network by separately handling the network's components. The proposed algorithms are fully symbolic; therefore, they allow the manipulation of large designs. Experimental results, though preliminary, are promising.
Fabrizio Ferrandi, Franco Fummi, Enrico Macii, Massimo Poncino, Donatella Sciuto
Great Lakes Symposium on VLSI5
1996 FsmTest: Functional test generation for sequential circuits
G. Buonannoa, Franco Fummi, Donatella Sciuto, Fabrizio Lombardi
Integr.3
1996 VHDL( VHSIC Hardware Description Language)
Donatella Sciuto
J. Syst. Archit.1
1995 Synthesis for testability of large complexity controllers
abstract
Specification of large complexity controllers in industrial design environments is performed by means of a top-down methodology leading to a description based on a hierarchy of FSMs. This paper presents a set of algorithms which compare such hierarchical descriptions with their structural implementations to produce irredundant circuits for which test patterns are easily derived. These algorithms can be inserted into any commercial design flow, based on VHDL descriptions, thus creating a synthesis for testability environment which provides testable and optimized gate-level descriptions.
Franco Fummi, Donatella Sciuto, M. Serro
ICCD2
1995 An Output/State Encoding for Self-Checking Finite State Machine
abstract
A new methodology for defining a self-checking sequential architecture is presented in the paper. A m-out-of-n encoding of the juxtaposition of next-state and output, eventually completed with additional output lines, is provided. The goal is guaranteeing detection for single and multiple unidirectional errors while minimizing area overhead.
Cristiana Bolchini, Donatella Sciuto
ISCAS2
1995 Data Path Testability Analysis Based on BDDs
Giacomo Buonanno, Fabrizio Ferrandi, Donatella Sciuto
ISCAS3
1995 Behavior of Self-Checking Checkers for 1-out-of-3 Codes Based on Pass-Transistor Logic
abstract
VLSI circuit implementation using transmission gates allows improvement of both circuit density and performances, without increasing power dissipation. Moreover, in self-checking structures, it allows detection of a larger class of faults than fully CMOS structures. This is true in particular for 1-out-of-3 codes, for which no gate level implementation of a self-checking checker (with respect to stuck-at faults) has been found yet. In this paper a self-checking checker for 1-out-of-3 codes obtained through a formalized synthesis procedure is presented and compared with other checkers to verify its performances.
Giacomo Buonanno, Fabio Salice, Donatella Sciuto
ISCAS3
1995 GECO: A Tool for Automatic Generation of Error Control Codes for Computer Applications
Luca Penzo, Donatella Sciuto, Cristina Silvano
ISCAS2
1995 TIES: A testability increase expert system for VLSI design
Giacomo Buonanno, Franco Fummi, Donatella Sciuto
J. Electron. Test.3
1995 A new DFT methodology for sequential circuits
Claudio Costi, Micaela Serra, Donatella Sciuto
J. Electron. Test.3
1995 Testability of artificial neural networks: A behavioral approach
Vincenzo Piuri, Mariagiovanna Sami, Donatella Sciuto
J. Electron. Test.3
1995 Construction techniques for systematic SEC-DED codes with single byte error detection and partial correction capability for computer memory systems
abstract
Three new techniques are proposed for constructing a class of codes that extends the protection provided by previous single error correcting (SEC)-double error detecting (DED)-single byte error detecting (SBD) codes. The proposed codes are systematic odd-weight-column SEC-DED-SBD codes providing also the correction of any odd number of erroneous bits per byte, where a byte represents a cluster of b bits of the codeword that are fed by the same memory chip or card. These codes are useful for practical applications to enhance the reliability and the data integrity of byte-organized computer memory systems against transient, intermittent, and permanent failures. In particular they represent a good tradeoff between the overhead in terms of additional check bits and the reliability improvement, due to the capability to correct at least 50% of the multiple errors per byte.>
Luca Penzo, Donatella Sciuto, Cristina Silvano
IEEE Trans. Inf. Theory2
1994 HW/SW Codesign for Embedded Telecom Systems
abstract
The aim of this paper is to define an approach, tailored for control-oriented applications, to manage system cospecification, high-level partitioning, hw/sw tradeoffs and cosynthesis. Our research effort focuses on fulfilling the goal of linking high-level specifications to efficient and cost-effective hw/sw implementations by investigating techniques such as synchronous cospecification styles, direct machine code generation as well as exploiting the capability of commercial VHDL synthesis tools.>
Stefano Antoniazzi, Alessandro Balboni, William Fornaciari, Donatella Sciuto
ICCD4
1994 CMOS Reliability Improvements Through a New Fault Tolerant Technique
abstract
A CMOS gate structure tolerating all single transistor stuck-at (TSA) faults and a large set of multiple faults is presented. Such structure is based on the simultaneous implementation of both the natural and the complemented form of the desired output; such implementation is easily modifiable to achieve fault tolerance and to obtain detectability of most of the faults which are not tolerated since they cause the two output lines to share the same value. Usually, production of the natural and complemented form of the output signal does not require one to double the number of transistors, thus resulting in cheaper (in terms of area) approaches.>
Cristiana Bolchini, Giacomo Buonanno, Donatella Sciuto, Renato Stefanelli
ISCAS3
1994 Two-Dimensional Sequential Array Architectures: Design for Testability Approaches
abstract
Testing of array architectures is an important issue because of the relevance that these structures are assuming in VLSI/WSI designs. The DfT techniques presented in this paper represent a possible approach to allow the verification of the cells composing the entire structure. The structural methodologies cope with the accessibility problems by modifying the interconnection network to "isolate" the cell in exam from the others; the functional approach modifies the cell making it transparent with respect to the data flow if the cell is not being tested. Both techniques aim at defining a sequential array architecture whose elements can be tested by applying patterns defined for the single cell and achieving the same coverage notwithstanding the embedding constituted by the array interconnections.>
Cristiana Bolchini, Franco Fummi, Donatella Sciuto
ISCAS3
1994 Constraint Generation & Placement for Automatic Layout Design of Analog Integrated Circuits
abstract
Electrical performances of integrated circuits operating at high frequencies, can be significantly degraded by electrical parasitics non intentionally introduced during the layout design. We present here new tools for the automatic constraint driven placement and routing of analog integrated circuits. Furthermore we introduce a new constraint generation program, based on AC analysis in the frequency domain. The constraints produced by this program have been-employed to drive the automatic layout tools, during the experiments here reported. These programs automatically produce the layout of high performance integrated circuits, significantly reducing the electrically effective parasitics due to finite length interconnections.>
Margherita Pillan, Donatella Sciuto
ISCAS2
1994 Innovative Structures for CMOS Combinational Gates Synthesis
abstract
Design of multiple outputs CMOS combinational gates is studied. Two techniques for minimization of multiple output functions at the switching level are introduced. These techniques are based on innovative transistor interconnection structures named Delta and Lambda networks. The two techniques can be combined together to obtain further area reductions. Different synthesis algorithms are discussed, from exhaustive enumeration to branch and bound to heuristic techniques allowing to speed up the synthesis process. Simulation results for synthesis are introduced to compare the different algorithms. Design examples are also provided. Electrical simulations show that the dynamic behavior of such structures is comparable to the traditional static or domino implementations (obviously the new and traditional structures have the same static behavior).>
Giacomo Buonanno, Donatella Sciuto, Renato Stefanelli
IEEE Trans. Computers2
1994 ALADIN: a multilevel testability analyzer for VLSI system design
abstract
In order to cope with tomorrow's challenges in the microelectronic market, the reliability of the first phases of the design process must be improved. The possibility of applying techniques for testability analysis at these abstract design levels can considerably help in achieving this goal, reducing at the same time system design costs. In this paper we introduce a novel approach for the application of functional testability at system design level and demonstrate the possibility of its application in an industrial environment. Testability conditions referring to both regular and irregular topologies have been defined, formalized and inserted into the knowledge base of the expert system, ALADIN. This tool operates as a testability analyzer able to identify critical areas for testability in designs whose functional modules and local interconnections are known and described in standard VHDL. The architecture of the tool has been defined in order to satisfy the users' requirements including the integrability into a standard CAD design flow through standard I/O interfaces. Then its application to both a regular and an irregular topology are presented in order to show on real examples which testability conditions apply, and how the tool operates in order to reach the testability assessment. From these industrial case studies, figures of merit are derived from which it is possible to evaluate the importance of the application of such a methodology to system level design.>
Massimo Bombana, Giacomo Buonanno, Patrizia Cavalloro, Fabrizio Ferrandi, Donatella Sciuto, Giuseppe Zaza
IEEE Trans. Very Large Scale Integr. Syst.5
1993 Functional Fault Models and Gate Level Coverage for Sequential Architectures
abstract
This paper introduces and evaluates functional fault models for test pattern generation of sequential circuits at the finite state machine level. Evaluation of the proposed fault models against their gate level fault coverage on multi-level implementations is presented. The relationships between functional and gate level fault coverage are discussed.>
Giacomo Buonanno, Franco Fummi, Donatella Sciuto
ICCD3
1993 Functional Testing and Constrained Synthesis of Sequential Architectures
Giacomo Buonanno, Franco Fummi, Donatella Sciuto
ISCAS3
1993 Fault detection in TFCMOS/DFCMOS combinational gates
Giacomo Buonanno, Fabrizio Lombardi, Donatella Sciuto, Yinan N. Shen
Integr.3
1993 Concurrently self-checking structures for Fsms
Mariagiovanna Sami, Donatella Sciuto, Renato Stefanelli
Microprocess. Microprogramming2
1992 Constant testability of combinational cellular tree structures
Fabrizio Lombardi, Donatella Sciuto
J. Electron. Test.2
1992 A behavioral approach to testability analysis for neural networks
Vincenzo Piuri, Mariagiovanna Sami, Donatella Sciuto, Renato Stefanelli
Microprocess. Microprogramming3
1991 Testability conditions for two-dimensional bilateral arrays
Donatella Sciuto
Integr.1
1991 Multiple stuck-at faults detection in CMOS combinational gates
Giacomo Buonanno, Fabrizio Lombardi, Donatella Sciuto, Y.-N. Sken
Microprocessing and Microprogramming3
1991 The Patricia testability analysis tool
M. Hadjinicolaeu, N. Burgess, Donatella Sciuto, G. Buananno, Patrizia Cavalloro, Giuseppe Zaza
Microprocessing and Microprogramming3
1990 A Routing Algorithm for Harvesting Multipipeline Arrays with Small Intercell and Pipeline Delays
abstract
A novel approach is analyzed for reconfiguring multipipeline arrays from two-dimensional arrays. The proposed approach is fully characterized and the conditions for switching and routing are given. A polynomial time complexity algorithm is proposed for the reconfiguration of multipipeline arrays. It is proved that 100% harvesting is possible using the proposed algorithm while achieving very small intercell and pipeline delays.>
Peter Koo, Fabrizio Lombardi, Donatella Sciuto
ICCAD3
1990 Testing of serial input convolvers
Luca Breveglieri, Luigi Dadda, Donatella Sciuto
Microprocessing and Microprogramming3
1990 An approach to a design for testability personal consultant
Giacomo Buonanno, A. Burri, Franco Fummi, Donatella Sciuto
Microprocessing and Microprogramming4
1989 Linear testability conditions for two-dimensional arrays
Fabrizio Lombardi, Donatella Sciuto
Microprocess. Microprogramming2
1988 Array partitioning: a methodology for reconfigurability and reconfiguration problems
abstract
An approach to array fault tolerance is presented. A parametric tool for reconfigurability detection and reconfiguration of arrays is introduced. The underlying methodology is based on a partitioning algorithm to determine reconfigurability. After reconfiguration a compaction algorithm is used to optimize results. The application of this methodology to fault-stealing-based algorithms has shown good results in terms of the added required redundancy, even for fault distributions that were unreconfigurable when the reconfiguration algorithm was applied alone.>
Fausto Distante, Fabrizio Lombardi, Donatella Sciuto
ICCD3
1988 Behavioral testing of multilevel system software
Fausto Distante, Donatella Sciuto
Microprocess. Microprogramming2
1988 On Functional Testing of Array Processors
abstract
This correspondence presents a new testing method for single instruction multiple data (SIMD) VLSI arrays. A new fault model is presented. Faults are defined at the functional level. A systematic test generation procedure is derived. Testing is performed by sequences of instructions. Two criteria are used. The first criterion establishes the external observability and controllability of the instructions. The second criterion uses instruction cardinality as a metric of instruction complexity. An example of the application of the proposed technique to an existing parallel scheme is described.>
Donatella Sciuto, Fabrizio Lombardi
IEEE Trans. Computers1
1988 An algorithm for functional reconfiguration of fixed-size arrays
abstract
A technique for reconfiguring an array of arbitrary rectangular shape from a fixed-size square array is presented. This type of reconfiguration is referred to as functional reconfiguration, as it maps processing functionalities (given by processing) into an array of dimensions different from those of the physical device. Functional reconfiguration is analyzed using index mapping. Locality and interconnection requirements are presented. Examples are given to illustrate the algorithm. It is also proven that the proposed technique achieves a lower intercell delay than previous techniques for certain values of ratio of the dimensions of the reconfigured rectangular array and the original square array.>
Fabrizio Lombardi, Donatella Sciuto, Renato Stefanelli
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1987 A Technique for Reconfiguring Two Dimensional VLSI Arrays
Fabrizio Lombardi, Donatella Sciuto, Renato Stefanelli
RTSS2
1987 A reconfiguration algorithm for wafer-scale integration of systolic arrays
Leonardo Jervis, Donatella Sciuto
Microprocess. Microprogramming2
1987 A reconfiguration algorithm for delay minimization in VLSI/WSI array processors
Donatella Sciuto
Microprocessing and Microprogramming1