VLDB 2026 Research / reviewers in the wild / expert
Antonio Miele
dblp:46/4009
· DBLP profile ↗
66ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0003-3197-0723ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 56 · 2 first-author · 16 since 2021Software engineering, systems software and programming languages · 12 · 1 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Resolving Energy Storage for Intermittent InferenceabstractWe reveal two fundamental insights when deploying inference workloads on intermittent systems that use capacitors as energy buffers. Because of varying structural characteristics and data-independent execution of Deep Neural Network (DNN) models, different capacitor configurations can swing inference performance from the good to the bad. Second, capacitor leakage, which is inevitable in real deployments and yet vastly overlooked in existing literature, makes a whole difference when determining the most efficient capacitor configuration. Both insights hold independently of specific techniques that enable intermittent inference. They form the basis to design LEACS: an off-line technique to determine an efficient energy storage configuration for intermittent inference workloads. LEACS determines the number and size of capacitors based on the nature of the energy source, while accounting for capacitor leakage. We test LEACS together with two state-of-the-art intermittent inference execution techniques and compare its performance with four different baselines across three hardware platforms, seven DNNs models, and three real-world energy sources. Our results indicate that LEACS improves inference throughput by up to 3.4 ×, with an average 1.59 × across the settings we test. Rei Barjami, Antonio Miele, Luca Mottola |
SenSys | 2 |
| 2026 | Benchmark Suite for Resilience Assessment of Deep Learning ModelsabstractThe reliability assessment of systems powered by artificial intelligence (AI) is becoming a crucial step prior to their deployment in safety and mission-critical systems. Recently, many efforts have been made to develop sophisticated techniques to evaluate and improve the resilience of AI models against the occurrence of random hardware faults. However, due to the intrinsic nature of such models, the comparison of the results obtained in state-of-the-art works is crucial, as reference models are missing. Moreover, their resilience is strongly influenced by the training process, the adopted framework and data representation, and so on. To enable a common ground for future research targeting CNN resilience analysis/hardening, this work proposes a first benchmark suite of DL models commonly adopted in this context, providing the models, the training/test data, and the resilience-related information (fault list, coverage, etc.) that can be used as a baseline for fair comparison. To this end, this research identifies a set of axes that have an impact on the resilience and classifies some popular CNN models, in both PyTorch and TensorFlow. Some final considerations are drawn, showing the relevance of a benchmark suite tailored for the resilience context. Cristiana Bolchini, Alberto Bosio, Luca Cassano, Antonio Miele, Salvatore Pappalardo, Dario Passarello, Annachiara Ruospo, Ernesto Sánchez 0001, Matteo Sonza Reorda, Vittorio Turco |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | HiDP: Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge PlatformsabstractEdge inference techniques partition and distribute Deep Neural Network (DNN) inference tasks among multiple edge nodes for low latency inference, without considering the core-level heterogeneity of edge nodes. Further, default DNN inference frameworks also do not fully utilize the resources of heterogeneous edge nodes, resulting in higher inference latency. In this work, we propose a hierarchical DNN partitioning strategy (HiDP) for distributed inference on heterogeneous edge nodes. Our strategy hierarchically partitions DNN workloads at both global and local levels by considering the core-level heterogeneity of edge nodes. We evaluated our proposed HiDP strategy against relevant distributed inference techniques over widely used DNN models on commercial edge devices. On average our strategy achieved 38% lower latency, 46% lower energy, and 56% higher throughput in comparison with other relevant approaches, Zain Taufique, Aman Vyas, Antonio Miele, Pasi Liljeberg, Anil Kanduri |
DATE | 3 |
| 2025 | Twill: Scheduling Compound AI Systems on Heterogeneous Mobile Edge PlatformsabstractCompound AI (cAI) systems chain multiple AI models to solve complex problems. cAI systems are typically composed of deep neural networks (DNNs), transformers, and large language models (LLMs), exhibiting a high degree of computational diversity and dynamic workload variation. Deploying cAI services on mobile edge platforms poses a significant challenge in scheduling concurrent DNN-transformer inference tasks, which arrive dynamically in an unknown sequence. Existing mobile edge AI inference strategies manage multi-DNN or transformer-only workloads, relying on design-time profiling, and cannot handle concurrent inference of DNNs and transformers required by cAI systems. In this work, we address the challenge of scheduling cAI systems on heterogeneous mobile edge platforms. We present Twill, a run-time framework to handle concurrent inference requests of cAI workloads through task affinity-aware cluster mapping and migration, priority-aware task freezing/unfreezing, and Dynamic Voltage/Frequency Scaling (DVFS), while minimizing inference latency within power budgets. We implement and deploy our Twill framework on the Nvidia Jetson Orin NX platform. We evaluate Twill against state-of-the-art edge AI inference techniques over contemporary DNNs and LLMs, reducing inference latency by 54% on average, while honoring power budgets. Zain Taufique, Aman Vyas, Antonio Miele, Pasi Liljeberg, Anil Kanduri |
ICCAD | 3 |
| 2025 | Runtime Energy-Efficient Control Policy for Mobile Robots with Computing Workload and Battery AwarenessabstractEnergy efficiency is a fundamental goal in robotic control. Various components within a robot, such as mechanical systems, computational units, and sensors, consume energy, all powered by the battery unit. Each component features several actuators and individual controllers that optimize energy usage locally, often without regard to one another. In this paper, we highlight a significant phenomenon indicating a considerable dependency between the mechanical and computational parts of the robot as energy consumers and the battery state of charge (SOC) as the energy provider. We demonstrate that as the battery SOC fluctuates, the behavior of energy consumption also varies, necessitating a unified controller with awareness of this relationship. Motivated by this observation, we propose a battery-aware co-optimization strategy for the mechanical and computational units, leveraging configuration space exploration to optimize the motor speed and the CPU frequency under different environmental conditions and battery SOC levels. Experimental results demonstrate the effectiveness of our approach in extending the operational lifetime of a robot under varying battery SOC and workload conditions, enhancing the energy efficiency of a case study rover by up to 53.93% w.r.t. selected baselines and similar past approaches. M. H. Haghbayan, Abdul Malik, Antonio Miele, Juha Plosila |
IROS | 4 |
| 2025 | A Benchmark Suite to Evaluate DNN's ResilienceabstractAssessing AI systems reliability is essential before deploying them in safety-critical applications. While recent efforts have focused on improving model resilience to random hardware faults, meaningful comparison remains difficult due to the lack of standardized reference models. Different authors use different implementations, which makes comparisons unfair and biased: resilience is influenced by the training processes, the software framework, and data representations. To address these issues, this work introduces a benchmark suite of CNN models to test the resilience of DNNs. The benchmark is structured on different axes: software framework, hardware platform, data representation, task and dataset. It is aimed at providing a shared foundation for fair and reproducible resilience evaluation. Cristiana Bolchini, Alberto Bosio, Luca Cassano, Antonio Miele, Salvatore Pappalardo, Dario Passariello, Annachiara Ruospo, Ernesto Sánchez 0001, Matteo Sonza Reorda, Vittorio Turco |
ITC | 4 |
| 2025 | On the Sweet Spot of Intermittent Inference in the Battery-less Internet of ThingsabstractWe present a measurement and performance analysis of system-level settings to improve the energy efficiency of Deep Neural Network (DNN) inference on battery-less Internet of Things (IoT) devices. To do so, we deliberately trade a small, controllable reduction in inference accuracy for energy gains. Battery-less IoT devices are severely resource-constrained platforms powered by energy harvesting, where execution becomes intermittent as it alternates between bursts of computation and periods of energy recharge. To survive frequent energy failures, devices persist their system state into non-volatile memories, incurring significant energy costs. We leverage aggressive current scaling offered by Spin-Transfer Torque Magnetic RandomAccess Memory (STT-MRAM) during state writes to reduce energy consumption, intentionally allowing controlled write errors that affect inference outcomes. Through an extensive experimental campaign comprising over 2.2+ trillion data points across 4 microcontroller units (MCUs) and 8 benchmarks, we demonstrate that by tolerating a limited accuracy loss we can obtain up to $40 \%$ energy savings. We release our framework and toolset to foster further research in this emerging design space. Rei Barjami, Antonio Miele, Luca Mottola |
MASCOTS | 2 |
| 2025 | Exploiting Approximation for Run-time Resource Management of Embedded HMPsabstractRun-time resource management (RTM) of multi-programmed workloads on heterogeneous multi-core platforms is challenging due to (i) fixed power budget of the device, (ii) variable performance requirements of the workloads, and (iii) unknown arrival of the applications. Existing RTM solutions lack power-performance coordination, resulting in performance degradation during power actuation or power violations during performance provisioning. Exploiting inherent error-resilience of the applications can address the performance loss incurred in power actuation, by combining run-time approximation with traditional power knobs (including Dynamic Voltage/Frequency Scaling, Task Migration, Degree of Parallelism, and CPU Quota ). In this work, we present an accuracy-aware resource management framework that jointly actuates run-time approximation and traditional power knobs for efficient power-performance management of multi-programmed and multi-threaded workloads running on heterogeneous mobile platforms. Our strategy configures the accuracy of the applications at run-time to exploit accuracy-performance trade-offs, by considering system-wide power-performance dynamics. We use heuristic estimation models to jointly enforce accuracy configuration and traditional power knobs settings at run-time. We evaluated our framework on real-world embedded mobile platforms, including Odroid XU3 and Asus Tinker Edge R boards to demonstrate the efficiency of our proposed approach across multiple workload scenarios. Our approach achieved 25% lower performance violations against the state-of-the-art run-time resource management policies at the cost of 2.2% accuracy loss across six applications. Zain Taufique, Anil Kanduri, Antonio Miele, Amir-Mohammad Rahmani, Cristiana Bolchini, Nikil Dutt, Pasi Liljeberg |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2025 | A Coordinated Approach to Control Mechanical and Computing Resources in Mobile RobotsabstractEnergy management of mechanical and cyber parts in mobile robots consists of two processes operating concurrently at runtime. Both the two processes can significantly improve the robots' battery lifetime and further extend mission time. In each process, information on energy consumption of one of the two parts is captured and analyzed to manipulate various mechanical/computational actuators in a robot, such as motor speed and CPU voltage/frequency. In this article, we show that considering management of mechanical and computational segments separately does not necessarily result in an energy-optimal solution due to their co-dependence; as a consequence, a runtime co-management scheme is required. We propose a proactive energy optimization methodology in which dynamically trained internal models are utilized to predict the future energy consumption for the mechanical and computational parts of a mobile robot, and based on that, the optimal mechanical speed and CPU voltage/frequency are determined at runtime. The experimental results on a ground wheeled robot show up to 36.34% reduction in the overall energy consumption compared to the state-of-the-art methods. Sajad Shahsavari, M. H. Haghbayan, Antonio Miele, Eero Immonen, Juha Plosila |
IEEE Trans. Robotics | 3 |
| 2024 | Adaptive Workload Distribution for Accuracy-aware DNN Inference on Collaborative Edge PlatformsabstractDNN inference can be accelerated by distributing the workload among a cluster of collaborative edge nodes. Heterogeneity among edge devices and accuracy-performance trade-offs of DNN models present a complex exploration space while catering to the inference performance requirements. In this work, we propose adaptive workload distribution for DNN inference, jointly considering node-level heterogeneity of edge devices, and application-specific accuracy and performance requirements. Our proposed approach combinatorially optimizes heterogeneity-aware workload partitioning and dynamic accuracy configuration of DNN models to ensure performance and accuracy guarantees. We tested our approach on an edge cluster of Odroid XU4, Raspberry Pi4, and Jetson Nano boards and achieved an average gain of 41.52% in performance and 5.2% in output accuracy as compared to state-of-the-art workload distribution strategies. Zain Taufique, Antonio Miele, Pasi Liljeberg, Anil Kanduri |
ASPDAC | 2 |
| 2024 | Cross-Layer Reliability Analysis of NVDLA Accelerators: Exploring the Configuration SpaceabstractInvestigating the effects of Single Event Upset in domain-specific accelerators represents one of the key enablers to deploy Deep Neural Networks (DNNs) in mission-critical edge applications. Currently, reliability analyses related to DNNs mainly focus either on the DNNs model, at application level, or on the hardware accelerator, at architecture level. This paper presents a systematic cross-layer reliability analysis of NVIDIA Deep-Learning Accelerator, a popular family of industry-grade, open and free DNN accelerators. The goals are i) to analyze the propagation of faults from the hardware to the application level, and ii) to compare different architectural configurations. Our investigation delivers new insights into the performance-accuracy-reliability trade-off spanned by the configuration space of Deep Learning accelerators. In particular, the Failure in Time can be reduced up to 4.3x for the same DNN model accuracy and by up to 9.4x for the same performance, while accounting 6.5x inference latency and 1.1% accuracy drop, respectively. Alessandro Veronesi, Alessandro Nazzari, Dario Passarello, Milos Krstic, Michele Favalli, Luca Cassano, Antonio Miele, Davide Bertozzi, Cristiana Bolchini |
ETS | 7 |
| 2024 | Tango: Low Latency Multi-DNN Inference on Heterogeneous Edge PlatformsabstractThere is an increasing demand to run DNN applications on edge platforms for low-latency inference. Executing multi-DNN workloads with diverse compute and latency requirements on resource-constrained heterogeneous edge platforms poses a significant scheduling challenge. In this work, we present Tango framework for orchestrating multi-DNN inference on heterogeneous edge platforms. Our approach uses a Proximal Policy-based Reinforcement Learning agent to jointly optimize cluster selection, accuracy configuration, and frequency scaling to minimize inference latency with a tolerable accuracy loss. We implemented the proposed Tango framework as a portable middleware and deployed it on real hardware of the Jetson TX edge platform. Our evaluation against relevant multi-DNN scheduling strategies demonstrates 61 % lower latency and 48.4 % lower energy consumption at a maximum accuracy loss of 1.59 %. Zain Taufique, Aman Vyas, Antonio Miele, Pasi Liljeberg, Anil Kanduri |
ICCD | 3 |
| 2024 | Intermittent Inference: Trading a 1% Accuracy Loss for a 1.9x Throughput SpeedupabstractWe present INTERCEPT, a compile-time toolchain enabling manifold throughput improvements when running intermittent DNN inference on IoT devices, in exchange of a maximum 1% accuracy loss. Intermittently-computing IoT devices rely on ambient energy harvesting and compute opportunistically, as energy is available. They use NVM to persist intermediate results in anticipation of energy failures. Without requiring changes to existing models and by exploiting the features of STT-MRAM as NVM, INTERCEPT optimizes the placement and configuration of state persistence operations when executing the inference process. This happens off-line with no user intervention, while enforcing a maximum 1% accuracy loss. Our results, obtained across three platforms and six diverse neural networks, indicate that INTERCEPT provides a 40% energy gain in a single inference process, on average. With the same energy budget, this yields a 1.9x throughput speedup. Rei Barjami, Antonio Miele, Luca Mottola |
SenSys | 2 |
| 2023 | Poster Abstract: Energy vs. Quality of Approximate Non-volatile Writes in Intermittent ComputingabstractWe explore how hardware approximation techniques can be used to reduce the overhead introduced by state persistence operations in intermittent computing. We do so by exploring the trade-off between energy consumption and quality of the results. We specifically adjust the energy/quality ratio of write operations in Spin Transfer Torque Magnetic Random Access Memory (STT-MRAM) by modifying the current applied during these operations. Our evaluation on a heterogeneus set of benchmarks demonstrates up to ≈50% reduction in the state persistence overhead while maintaining an acceptable output quality. Rei Barjami, Antonio Miele, Luca Mottola |
SenSys | 2 |
| 2023 | Fast and Accurate Error Simulation for CNNs Against Soft ErrorsabstractThe great quest for adopting AI-based computation for safety-/mission-critical applications motivates the interest towards methods for assessing the robustness of the application w.r.t. not only its training/tuning but also errors due to faults, in particular soft errors, affecting the underlying hardware. Two strategies exist: architecture-level fault injection and application-level functional error simulation. We present a framework for the reliability analysis of Convolutional Neural Networks (CNNs) via an error simulation engine that exploits a set of validated error models extracted from a detailed fault injection campaign. These error models are defined based on the corruption patterns of the output of the CNN operators induced by faults and bridge the gap between fault injection and error simulation, exploiting the advantages of both approaches. We compared our methodology against SASSIFI for the accuracy of functional error simulation w.r.t. fault injection, and against TensorFI in terms of speedup for the error simulation strategy. Experimental results show that our methodology achieves about 99% accuracy of the fault effects w.r.t. SASSIFI, and a speedup ranging from 44x up to 63x w.r.t. TensorFI, that only implements a limited set of error models. Cristiana Bolchini, Luca Cassano, Antonio Miele, Alessandro Toschi |
IEEE Trans. Computers | 3 |
| 2023 | Run-Time Resource Management in CMPs Handling Multiple Aging MechanismsabstractRun-time resource management is fundamental for efficient execution of workloads on Chip Multiprocessors. Application- and system-level requirements (e.g., on performance versus power versus lifetime reliability) are generally conflicting each other, and any decision on resource assignment, such as core allocation or frequency tuning, may positively affect some of them while penalizing some others. Resource assignment decisions can be perceived in few instants of time on performance and power consumption, but not on lifetime reliability. In fact, this latter changes very slowly based on the accumulation of effects of various decisions over a long time horizon. Moreover, aging mechanisms are various and have different causes; most of them, such as Electromigration (EM), are subject to temperature levels, while Thermal Cycling (TC) is caused mainly by temperature variations (both amplitude and frequency). Mitigating only EM may negatively affect TC and vice versa. We propose a resource orchestration strategy to balance the performance and power consumption constraints in the short-term and EM and TC aging in the long-term. Experimental results show that the proposed approach improves the average Mean Time To Failure at least by 17% and 20% w.r.t. EM and TC, respectively, while providing same performance level of the nominal counterpart and guaranteeing the power budget. M. H. Haghbayan, Antonio Miele, Onur Mutlu, Juha Plosila |
IEEE Trans. Computers | 2 |
| 2022 | Dependability of Alternative Computing Paradigms for Machine Learning: hype or hope?abstractToday we observe amazing performance achieved by Machine Learning (ML); for specific tasks it even surpasses human capabilities. Unfortunately, nothing comes for free: the hidden cost behind ML performance stems from its high complexity in terms of operations to be computed and the involved amount of data. For this reasons, custom Artificial Intelligence hardware accelerators based on alternative computing paradigms are attracting large interest. Such dedicated devices support the energy-hungry data movement, speed of computation, and memory resources that MLs require to realize their full potential. However, when ML is deployed on safety-/mission-critical applications, dependability becomes a concern. This paper presents the state of the art of custom Artificial Intelligence hardware architectures for ML, here Spiking and Convolutional Neural Networks, and shows the best practices to evaluate their dependability. Cristiana Bolchini, Alberto Bosio, Luca Cassano, Bastien Deveautour, Giorgio Di Natale, Antonio Miele, Ian O'Connor, Elena I. Vatajelu |
DDECS | 6 |
| 2022 | Fault Impact Estimation for Lightweight Fault Detection in Image FilteringabstractClassical redundancy-based fault detection techniques, such as Duplication with Comparison (DWC), rely on replicating the computation and comparing the replicas’ output at a bit-wise granularity. In many application environments these costs are prohibitive, especially when applications are characterized by an intrinsic level of tolerance. This article presents a novel fault-detection approach for the specific context of image filtering. Peculiarity of the proposed approach is that it estimates the impact of the fault on the processed output, in order to determine whether the image is usable or should be re-processed. To limit overheads, the proposed solution exploits Approximate Computing (AC), allowing the definition of disciplined AC strategies to trade-off between accuracy and costs. Core of our solution is the successful combination of Image Quality Assessment metrics and Machine Learning models to assess the visual impact of the fault in a lightweight manner. Extensive experimental campaigns demonstrate the effectiveness of the solution, achieving achieving a reduction in terms of execution time up to 44 percent with respect to the classical DWC, with a fault detection precision ranging from 94.58 to 96.70 percent, and recall ranging from 88.2 to 97.8 percent, depending on the adopted level of approximation. Cristiana Bolchini, Giacomo Boracchi, Luca Cassano, Antonio Miele, Diego Stucchi |
IEEE Trans. Computers | 4 |
| 2022 | A Runtime Resource Management and Provisioning Middleware for Fog Computing InfrastructuresabstractThe pervasiveness and growing processing capabilities of mobile and embedded systems have enabled the widespread diffusion of the Fog Computing paradigm in the Internet of Things scenario, where computing is directly performed at the edges of the networked infrastructure in distributed cyber-physical systems. This scenario is characterized by a highly dynamic workload and architecture in which applications enter and leave the system, as well as nodes and connections. This article proposes a runtime resource management and provisioning middleware for the dynamic distribution of the applications on the processing resources. The proposed middleware consists of a two-level hierarchy: (i) a global Fog Orchestrator monitoring the architecture status and (ii) a Local Agent on each node, performing a fine-grain tuning of its resources. The co-operation between these components allows one to dynamically adapt and exploit the fine-grain nodes view for fulfilling the defined system-level goals, for example, minimizing power consumption while meeting Quality of Service requirements such as application throughput. This hierarchical architecture and the adopted policies offer a unified optimization strategy that is unique with regard to existing approaches that typically focus on a single aspect of resource management at runtime. A middleware prototype is presented and experimentally evaluated in a Smart Building case study. Antonio Miele, Henry Zárate, Luca Cassano, Cristiana Bolchini, Jorge Eduardo Ortiz Trivino |
ACM Trans. Internet Things | 1 |
| 2021 | Hierarchical Fault Simulation of Deep Neural Networks on Multi-Core SystemsabstractIn this paper, a hierarchical fault simulation technique for neural networks is proposed, supporting both permanent and temporary faults. In the proposed technique, different levels of hierarchy are used, forming a mixed-level simulation environment. In such an environment, the pre-synthesis behavioral specification of the network and the post-synthesis gate-level model are co-simulated. To accelerate the fault simulation process, faults are injected in the gate-level specification of the selected neurons while the behavioral model in different levels of abstraction is used to simulate the remaining neurons. Further speedup is obtained through event-driven simulation and parallelization. Experimental results confirm the time efficiency of the proposed fault simulation technique. Masoomeh Karami, M. H. Haghbayan, Masoumeh Ebrahimi, Antonio Miele, Hannu Tenhunen, Juha Plosila |
ETS | 4 |
| 2021 | Energy-Efficient Mobile Robot Control via Run-time Monitoring of Environmental Complexity and Computing WorkloadabstractWe propose an energy-efficient controller to minimize the energy consumption of a mobile robot by dynamically manipulating the mechanical and computational actuators of the robot. The mobile robot performs real-time vision-based applications based on an event-based camera. The actuators of the controller are CPU voltage/frequency for the computation part and motor voltage for the mechanical part. We show that independently considering speed control of the robot and voltage/frequency control of the CPU does not necessarily result in an energy-efficient solution. In fact, to obtain the highest efficiency, the computation and mechanical parts should be controlled together in synergy. We propose a fast hill-climbing optimization algorithm to allow the controller to find the best CPU/motor configuration at run-time and whenever the mobile robot is facing a new environment during its travel. Experimental results on a robot with Brushless DC Motors, Jetson TX2 board as the computing unit, and a DAVIS-346 event-based camera show that the proposed control algorithm can save battery energy by an average of 50.5%, 41%, and 30%, in low-complexity, medium-complexity, and high-complexity environments, over baselines. Sherif Abdelmonem Sayed Mohamed, M. H. Haghbayan, Antonio Miele, Onur Mutlu, Juha Plosila |
IROS | 3 |
| 2020 | An Approximation-based Fault Detection Scheme for Image Processing ApplicationsabstractImage processing applications expose an intrinsic resilience to faults. In this application field the classical Duplication with Comparison (DWC) scheme, where output images are discarded as soon as the two replicas’ outputs differ for at least one pixel, may be over-conseravative. This paper introduces a novel lightweight fault detection scheme for image processing applications; i) it extends the DWC scheme by substituting one of the two exact replicas with a faster approximated one; and ii) it features a Neural Network-based checker designed to distinguish between usable and unusable images instead of faulty/unfaulty ones. The application of the hardening scheme on a case study has shown an execution time reduction from 27% to 34% w.r.t. the DWC, while guaranteeing a comparable fault detection capability. Matteo Biasielli, Luca Cassano, Antonio Miele |
DATE | 3 |
| 2020 | Thermal-Cycling-aware Dynamic Reliability Management in Many-Core System-on-ChipabstractDynamic Reliability Management (DRM) is a common approach to mitigate aging and wear-out effects in multi- /many-core systems. State-of-the-art DRM approaches apply finegrained control on resource management to increase/balance the chip reliability while considering other system constraints, e.g., performance, and power budget. Such approaches, acting on various knobs such as workload mapping and scheduling, Dynamic Voltage/Frequency Scaling (DVFS) and Per-Core Power Gating (PCPG), demonstrated to work properly with the various aging mechanisms, such as electromigration, and Negative-Bias Temperature Instability (NBTI). However, we claim that they do not suffice for thermal cycling. Thus, we here propose a novel thermal-cycling-aware DRM approach for shared-memory many-core systems running multi-threaded applications. The approach applies a fine-grained control capable at reducing both temperature levels and variations. The experimental evaluations demonstrated that the proposed approach is able to achieve 39% longer lifetime than past approaches. M. H. Haghbayan, Antonio Miele, Zhuo Zou, Hannu Tenhunen, Juha Plosila |
DATE | 2 |
| 2020 | Dynamic Resource-Aware Corner Detection for Bio-Inspired Vision SensorsabstractEvent-based cameras are vision devices that transmit only brightness changes with low latency and ultra-low power consumption. Such characteristics make event-based cameras attractive in the field of localization and object tracking in resource-constrained systems. Since the number of generated events in such cameras is huge, the selection and filtering of the incoming events are beneficial from both increasing the accuracy of the features and reducing the computational load. In this paper, we present an algorithm to detect asynchronous corners form a stream of events in real-time on embedded systems. The algorithm is called the Three Layer Filtering-Harris or TLF-Harris algorithm. The algorithm is based on an events' filtering strategy whose purpose is 1) to increase the accuracy by deliberately eliminating some incoming events, i.e., noise and 2) to improve the real-time performance of the system, i.e., preserving a constant throughput in terms of input events per second, by discarding unnecessary events with a limited accuracy loss. An approximation of the Harris algorithm, in turn, is used to exploit its high-quality detection capability with a low-complexity implementation to enable seamless real-time performance on embedded computing platforms. The proposed algorithm is capable of selecting the best corner candidate among neighbors and achieves an average execution time savings of 59% compared with the conventional Harris score. Moreover, our approach outperforms the competing methods, such as eFAST, eHarris, and FA-Harris, in terms of real-time performance, and surpasses Arc* in terms of accuracy. Sherif Abdelmonem Sayed Mohamed, Jawad Naveed Yasin, M. H. Haghbayan, Antonio Miele, Jukka Heikkonen, Hannu Tenhunen, Juha Plosila |
ICPR | 4 |
| 2020 | Error Modeling for Image Processing Filters accelerated onto SRAM-based FPGAsabstractImage processing is today employed in a variety of application fields, including safety- and mission-critical ones. In these scenarios it is vital to carefully analyse the reliability of the designed system before deployment and, if necessary, to adopt specific hardening techniques. Two are the techniques generally employed: circuit-level fault injection and application-level functional error simulation. In this paper we present a set of functional error models specific for a number of convolution-based filters that are the basic building blocks for a wide range of image processing applications. The presented error models, derived through a number of circuit-level fault injection experiments, may be integrated into application-level functional error simulators, bridging the gap between the two strategies. The presented error models are the first step towards combining the accuracy of fault injection and the flexibility of error simulation into a widely adopted reliability analysis tool. Cristiana Bolchini, Luca Cassano, Andrea Mazzeo, Antonio Miele |
IOLTS | 4 |
| 2020 | A methodology for the design and deployment of distributed cyber-physical systems for smart environments
Giacomo Tanganelli, Luca Cassano, Antonio Miele, Carlo Vallati |
Future Gener. Comput. Syst. | 3 |
| 2020 | A Neural Network Based Fault Management Scheme for Reliable Image ProcessingabstractTraditional reliability approaches introduce relevant costs to achieve unconditional correctness during data processing. However, many application environments are inherently tolerant to a certain degree of inexactness or inaccuracy. In this article, we focus on the practical scenario of image processing in space, a domain where faults are a threat, while the applications are inherently tolerant to a certain degree of errors. We first introduce the concept of usability of the processed image to relax the traditional requirement of unconditional correctness, and to limit the computational overheads related to reliability. We then introduce our new flexible and lightweight fault management methodology for inaccurate application environments. A key novelty of our scheme is the utilization of neural networks to reduce the costs associated with the occurrence and the detection of faults. Experiments on two aerospace image processing case studies show overall time savings of 14.89 and 34.72 percent for the two applications, respectively, as compared with the baseline classical Duplication with Comparison scheme. Matteo Biasielli, Cristiana Bolchini, Luca Cassano, Erdem Koyuncu, Antonio Miele |
IEEE Trans. Computers | 5 |
| 2020 | CAST: Content-Aware STT-MRAM Cache Write Management for Different Levels of ApproximationabstractSpin transfer torque magnetic RAM (STT-MRAM) technology is one of the most promising alternative for static RAM (SRAM) for implementing on-chip memories. Compared with SRAMs, STT-MRAMs benefit from higher density and near-zero leakage power, nonetheless they impose high energy consumption for reliable write operations. However, in many applications, absolute data integrity is not required; thus, acting on the current applied in the write operations may represent a novel knob for disciplined approximate computing to obtain energy saving with a minimal quality loss in applications' outputs. This article proposes CAST, a hardware/software approach to adjust the energy/quality of write operations in STT-MRAM caches in multicore systems based on the content of requested write operations. CAST utilizes fine-grained cache-line-level actuation knobs with different levels of quality for individual write operations. This unique feature of STT-MRAMs allows to avoid interapplication actuation interference suffered by SRAMs, and makes the approach particularly suitable for systems running multiple applications with mixed accuracy sensitivity. Moreover, CAST exploits another peculiarity of STT-MRAMs represented by the asymmetry and transition-dependency of the write error rate, to further tune in a fine-grained manner the write current to achieve an additional energy saving, even in full-accurate applications. Our evaluations on workloads of full-approximate, mixed-criticality, and full-accurate applications demonstrate up to 57%, 34%, and 21% energy savings over a baseline STT-MRAM cache, respectively, with an acceptable quality of the generated outputs. Amir Mahdi Hosseini Monazzah, Amir-Mohammad Rahmani, Antonio Miele, Nikil Dutt |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | A Runtime Resource Management Policy for OpenCL Workloads on Heterogeneous MulticoresabstractNowadays, runtime workload distribution and resource tuning for heterogeneous multicores running multiple OpenCL applications is still an open quest. This paper proposes an adaptive policy capable at identifying an optimal working point for an unknown multiprogrammed OpenCL workload without using any design-time application profiling or analysis. The approach compared against a design-time optimization strategy demonstrates to be effective in converging to an solution guaranteeing required performance while minimizing power consumption and maximum temperature; it achieves on average values 0.085 W (5.15%) and 0.83°C (1.47%) worse than the static optimal solution. Daniele Angioletti, Francesco Bertani, Cristiana Bolchini, Francesco Cerizzi, Antonio Miele |
DATE | 5 |
| 2019 | A Smart Fault Detection Scheme for Reliable Image Processing ApplicationsabstractTraditional fault detection/tolerance techniques exploit multiple instances of the nominal processing and then perform a bit-wise comparison of the outputs to detect the occurrence of faults. In specific application scenarios, e.g., image/signal processing, the elaboration has an inherent degree of fault tolerance because it is possible to use the output even in the presence of slight alterations. In these contexts, the classical bit-wise comparison may be inefficient. Indeed, it may lead to conservatively discard outputs that have been only slightly altered by the fault and that could still be usefully exploited. In this paper, we propose a smart checking scheme based on Convolutional Neural Networks that rather than distinguishing between faulty and not faulty images, discriminates between usable and not usable images according to the ability of the end user to correctly process the output. The experimental evaluation shows that this solution enables an execution time saving of about 6.35% with a 99.42% accuracy, on average. Matteo Biasielli, Cristiana Bolchini, Luca Cassano, Antonio Miele |
DATE | 4 |
| 2019 | Scalable analytical model for reliability measures in aging VLSI by interacting Markovian agents
Davide Cerotti, Antonio Miele, Marco Gribaudo, Andrea Bobbio, Cristiana Bolchini |
Perform. Evaluation | 2 |
| 2018 | Approximation-aware coordinated power/performance management for heterogeneous multi-coresabstractRun-time resource management of heterogeneous multi-core systems is challenging due to i) dynamic workloads, that often result in ii) conflicting knob actuation decisions, which potentially iii) compromise on performance for thermal safety. We present a runtime resource management strategy for performance guarantees under power constraints using functionally approximate kernels that exploit accuracy-performance trade-offs within error resilient applications. Our controller integrates approximation with power knobs - DVFS, CPU quota, task migration - in coordinated manner to make performance-aware decisions on power management under variable workloads. Experimental results on Odroid XU3 show the effectiveness of this strategy in meeting performance requirements without power violations compared to existing solutions. Anil Kanduri, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Cristiana Bolchini, Nikil Dutt |
DAC | 2 |
| 2018 | Trends in On-chip Dynamic Resource ManagementabstractThe Complexity of emerging multi/many-core architectures and diversity of modern workloads demands coordinated dynamic resource management methods. We introduce a classification for these methods capturing the utilized resources and metrics. In this work, we use this classification to survey the key efforts in dynamic resource management. We first cover heuristic and optimization methods used to manage resources such as power, energy, temperature, Quality-of-Service (QoS) and reliability of the system. We then identify some of the machine learning based methods used in tuning architectural parameters in computer systems. In many cases, resource managers need to enforce design constraints during runtime with a certain level of guarantee. Hence, we also study the trend in deploying formal control theoretic approaches in order to achieve efficient and robust dynamic resource management. Kasra Moazzemi, Anil Kanduri, David Juhasz, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Axel Jantsch, Nikil Dutt |
DSD | 4 |
| 2017 | Optimizing streaming stencil time-step designs via FPGA floorplanningabstractStencil computations represent a highly recurrent class of algorithms in various high performance computing scenarios. The Streaming Stencil Time-step (SST) architecture is a recent implementation of stencil computations on Field Programmable Gate Array (FPGA). In this paper, we propose an automated framework for SST-based architectures capable of achieving the maximum performance level for a given FPGA device through 1) the maximization of basic modules instantiated in the design and 2) optimization of the design floorplanning. Experimental results show that the proposed approach reduces the design time up to 15× w.r.t. naive design space exploration approaches, and improves the performance of the 13%. Marco Rabozzi, Giuseppe Natale, Biagio Festa, Antonio Miele, Marco D. Santambrogio |
FPL | 4 |
| 2017 | Performance/Reliability-Aware Resource Management for Many-Cores in Dark Silicon EraabstractAggressive technology scaling has enabled the fabrication of many-core architectures while triggering challenges such as limited power budget and increased reliability issues, like aging phenomena. Dynamic power management and runtime mapping strategies can be utilized in such systems to achieve optimal performance while satisfying power constraints. However, lifetime reliability is generally neglected. We propose a novel lifetime reliability/performance-aware resource co-management approach for many-core architectures in the dark silicon era. The approach is based on a two-layered architecture, composed of a long-term runtime reliability controller and a short-term runtime mapping and resource management unit. The former evaluates the cores' aging status w.r.t. a target reference specified by the designer, and performs recovery actions on highly stressed cores by means of power capping. The aging status is utilized in runtime application mapping to maximize system performance while fulfilling reliability requirements and honoring the power budget. Experimental evaluation demonstrates the effectiveness of the proposed strategy, which outperforms most recent state-of-the-art contributions. M. H. Haghbayan, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen |
IEEE Trans. Computers | 2 |
| 2017 | Floorplanning Automation for Partial-Reconfigurable FPGAs via Feasible Placements GenerationabstractWhen dealing with partially reconfigurable designs on field-programmable gate array, floorplanning represents a critical step that highly impacts system's performance and reconfiguration overhead. However, current vendor design tools still require the floorplan to be manually defined by the designer. Within this paper, we provide a novel floorplanning automation framework, integrated in the Xilinx tool chain, which is based on an explicit enumeration of the possible placements of each region. Moreover, we propose a genetic algorithm (GA), enhanced with a local search strategy, to automate the floorplanning activity on the defined direct problem representation. The proposed approach has been experimentally evaluated with a synthetic benchmark suite and real case studies. We compared the designed solution against both the state-of-the-art algorithms and alternative engines based on the same direct problem representation. Experimental results demonstrated the effectiveness of the proposed direct problem representation and the superiority of the defined GA engine with respect to the other approaches in terms of exploration time and identified solution. Marco Rabozzi, Gianluca Durelli, Antonio Miele, John Lillis, Marco D. Santambrogio |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Reliability-Aware Runtime Power Management for Many-Core Systems in the Dark Silicon EraabstractPower management of networked many-core systems with runtime application mapping becomes more challenging in the dark silicon era. It necessitates considering network characteristics at runtime to achieve better performance while honoring the peak power upper bound. On the other hand, power management has a direct effect on chip temperature, which is the main driver of the aging effects. Therefore, alongside performance fulfillment, the controlling mechanism must also consider the current cores' reliability in its actuator manipulation to enhance the overall system lifetime in the long term. In this paper, we propose a multiobjective dynamic power management technique that uses current power consumption and other network characteristics including the reliability of the cores as the feedback while utilizing fine-grained voltage and frequency scaling and per-core power gating as the actuators. In addition, disturbance rejecter and reliability balancer are designed to help the controller to better smooth power consumption in the short term and reliability in the long term, respectively. Simulations of dynamic workloads and mixed criticality application profiles show that our method not only is effective in honoring the power budget while considerably boosting the system throughput, but also increases the overall system lifetime by minimizing aging effects by means of power consumption balancing. Amir-Mohammad Rahmani, M. H. Haghbayan, Antonio Miele, Pasi Liljeberg, Axel Jantsch, Hannu Tenhunen |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Lifetime-aware load distribution policies in multi-core systems: An in-depth analysis
Cristiana Bolchini, Luca Cassano, Antonio Miele |
DATE | 3 |
| 2016 | A lifetime-aware runtime mapping approach for many-core systems in the dark silicon era
M. H. Haghbayan, Antonio Miele, Amir-Mohammad Rahmani, Pasi Liljeberg, Hannu Tenhunen |
DATE | 2 |
| 2016 | Workload-aware power optimization strategy for asymmetric multiprocessors
Emanuele Del Sozzo, Gianluca Durelli, Ettore M. G. Trainiti, Antonio Miele, Marco D. Santambrogio, Cristiana Bolchini |
DATE | 4 |
| 2016 | A self-adaptive approach to efficiently manage energy and performance in tomorrow's heterogeneous computing systems
Ettore M. G. Trainiti, Gianluca Durelli, Antonio Miele, Cristiana Bolchini, Marco D. Santambrogio |
DATE | 3 |
| 2016 | Lifetime reliability modeling and estimation in multi-core systemsabstractIn the past decade, aggressive CMOS scaling has led to a considerable increase in the power densities, and consequently, in the operating temperature within the device. As ITRS reported in 2011 [1], such temperature increase is the primary cause of an acceleration in device aging and wear-out phenomena (e.g. electromigration or thermal cycling), that are leading to a higher susceptibility of digital systems to degradation effects, such as timing errors, and eventually to breakdowns. The result, is therefore a dramatic decrease of the lifetime of digital systems. Antonio Miele |
VTS | 1 |
| 2016 | A Power-Aware Approach for Online Test Scheduling in Many-Core ArchitecturesabstractAggressive technology scaling triggers novel challenges to the design of multi-/many-core systems, such as limited power budget and increased reliability issues. Today's many-core systems employ dynamic power management and runtime mapping strategies trying to offer optimal performance while fulfilling power constraints. On the other hand, due to the reliability challenges, online testing techniques are becoming a necessity in current and near future technologies. However, state-of-the-art techniques are not aware of the other power/performance requirements. This paper proposes a power-aware non-intrusive online testing approach for many-core systems. The approach schedules software based self-test routines on the various cores during their idle periods, while honoring the power budget and limiting delays in the workload execution. A test criticality metric, based on a device aging model, is used to select cores to be tested at a time. Moreover, power and reliability issues related to the testing at different voltage and frequency levels are also handled. Extensive experimental results reveal that the proposed approach can i) efficiently test the cores within the available power budget causing a negligible performance penalty, ii) adapt the test frequency to the current cores' aging status, and iii) cover available voltage and frequency levels during the testing. M. H. Haghbayan, Amir-Mohammad Rahmani, Antonio Miele, Mohammad Fattah, Juha Plosila, Pasi Liljeberg, Hannu Tenhunen |
IEEE Trans. Computers | 3 |
| 2015 | A System-Level Simulation Framework for Evaluating Resource Management Policies for Heterogeneous System ArchitecturesabstractNowadays, heterogeneous system architectures, integrating CPUs and one or more kinds of accelerators (e.g., GPUs or HW accelerators), are a promising solution to achieve high performance for data-intensive workloads while fulfilling other system-level requirements on the available power/energy budgets. However, heterogeneity comes at the cost of greater design and management complexity leading to an increasing quest for the definition of innovative runtime resource management policies. We propose a system-level simulation framework implemented in SystemC and Transaction Level Modeling for a fast evaluation of resource management policies for such systems to provide a quick feedback to the middleware designer. A set of case studies shows the efficiency of the proposed framework in supporting a fast analysis of the investigated policies. Antonio Miele, Gianluca Durelli, Marco D. Santambrogio, Cristiana Bolchini |
DSD | 1 |
| 2015 | Floorplanning for Partially-Reconfigurable FPGAs via Feasible Placements DetectionabstractThis work presents a novel floor planner tailored for Partially-Reconfigurable FPGAs having an arbitrary distribution of heterogeneous resources. The proposed approach precomputes a set of feasible placements for each of the reconfigurable regions, thus allowing the designer to set a preference on the types and positions of the desired areas. Then, the core of the approach is based on a Mixed-Integer Linear Programming (MILP) formulation which exploits constraints derived from a conflict graph to prevent overlapping between areas. Experimental results have shown that the defined approach leads to an average 11% improvements in the objective function value w.r.t. The state-of-the-art solutions under the same limited time budget. Marco Rabozzi, Antonio Miele, Marco D. Santambrogio |
FCCM | 2 |
| 2015 | An orchestrated approach to efficiently manage resources in heterogeneous system architecturesabstractNowadays, we are witnessing trends in technology, fabrication processes and computing architectures that lead to the design and development of processing systems constituted by a relevant number of independent, heterogeneous execution resources. The aim is to achieve high-performance while leveraging on other aspects, such as energy consumption. Indeed, heterogeneity comes at the cost of greater design and management complexity. To reach an optimal solution, system architects need to take into account the efficiency of systems' units, i.e., general purpose processors eventually with one or more kinds of accelerators (e.g., GPUs or FPGAs), as well as the workload. This often leads to inefficiency in the exploitation of such resources, and therefore in performance/energy. Within this context, we are proposing a runtime resource manager able to observe the system execution and to dynamically optimise its behaviour with respect to one or more identified functional parameters, according to the architectural characteristics, and the users' and the applications' needs. Such an adaptation characteristic is intrinsically embedded in the device as a software layer, called Orchestrator, able to adapt the runtime resource management according to the target objectives and to the inputs from the external environment. Cristiana Bolchini, Gianluca Durelli, Antonio Miele, Gabriele Pallotta, Marco D. Santambrogio |
ICCD | 3 |
| 2014 | Combined DVFS and mapping exploration for lifetime and soft-error susceptibility improvement in MPSoCsabstractEnergy and reliability optimization are two of the most critical objectives for the synthesis of multiprocessor systems-on-chip (MPSoCs). Task mapping has shown significant promise as a low cost solution in achieving these objectives as standalone or in tandem as well. This paper proposes a multi-objective design space exploration to determine the mapping of tasks of an application on a multiprocessor system and voltage/frequency level of each tasks (exploiting the DVFS capabilities of modern processors) such that the reliability of the platform is improved while fulfilling the energy budget and the performance constraint set by system designers. In this respect, the reliability of a given MPSoC platform incorporates not only the impact of voltage and frequency on the aging of the processors (wear-out effect) but also on the susceptibility to soft-errors - a joint consideration missing in all existing works in this domain. Further, the proposed exploration also incorporates soft-error tolerance by selective replication of tasks, making the proposed approach an interesting blend of reactive and proactive fault-tolerance. The combined objective of minimizing core aging together with the susceptibility to transient faults under a given performance/energy budget is solved by using a multi-objective genetic algorithm exploiting tasks' mapping, DVFS and selective replication as tuning knobs. Experiments conducted with reallife and synthetic application graphs clearly demonstrate the advantage of the proposed approach. Anup Das 0001, Akash Kumar 0001, Bharadwaj Veeravalli, Cristiana Bolchini, Antonio Miele |
DATE | 5 |
| 2014 | A lightweight and open-source framework for the lifetime estimation of multicore systemsabstractThis paper presents a Monte Carlo-based framework for the estimation of lifetime reliability of multicore systems. Existing mathematical tools either consider only the time to the first failure, or are limited by their intrinsic complexity and high computational time. The proposed framework allows to compute quasi-exact results with a reasonable computational time, without adopting typical (and possibly misleading) simplifications that characterize the existing tools for computing Mean Time To Failure (MTTF). The paper describes the framework with all its mathematical details, assumptions and simplifications; it proves the correctness of the obtained results, by comparing them against the exact ones, and underlines the differences with the simplistic approaches, also discussing time overhead improvements. Cristiana Bolchini, Matteo Carminati, Marco Gribaudo, Antonio Miele |
ICCD | 4 |
| 2014 | ADaPT: Automatic Data Personalization based on contextual preferencesabstractThis demo presents a framework for personalizing data access on the basis of the users' context and of the preferences they show while in that context. The system is composed of (i) a server application, which “tailors” a view over the available data on the basis of the user's contextual preferences, previously inferred from log data, and (ii) a client application running on the user's mobile device, which allows to query the data view and collects the activity log for later mining. At each change of context detected by the system the corresponding tailored view is loaded on the client device: accordingly, the most relevant data is available to the user even when the connection is unstable or lacking. The demo features a movie database, where users can browse data in different contexts and appreciate the personalization of the data views according to the inferred contextual preferences. Antonio Miele, Elisa Quintarelli, Emanuele Rabosio, Letizia Tanca |
ICDE | 1 |
| 2014 | Runtime Resource Management in Heterogeneous System Architectures: The SAVE ApproachabstractModern computing systems featuring different kinds of processing elements have proven to be efficient in terms of performance/energy trade-offs. Furthermore these systems usually have to execute multiple concurrent tasks without any apriori knowledge on expected arrival times, in an unpredictable and very dynamic environment. This scenario has propelled an interest towards self-adaptive systems that dynamically reorganize the use of system resources to optimize for a given goal. The SAVE project will develop a Heterogeneous System Architecture that will decide at runtime to execute task on the appropriate kind of resources, based on the current requirements. This paper presents a first implementation of a resource allocation policy that dynamically shares heterogeneous resources between multiple running applications. Resource allocation mechanisms are discussed and evaluated in an experimental campaign, showing how the policy helps in attaining users' applications goals. Gianluca Durelli, Marcello Pogliani, Antonio Miele, Christian Plessl, Heinrich Riebler, Marco D. Santambrogio, Gavin Vaz, Cristiana Bolchini |
ISPA | 3 |
| 2013 | Self-Adaptive Fault Tolerance in Multi-/Many-Core Systems
Cristiana Bolchini, Matteo Carminati, Antonio Miele |
J. Electron. Test. | 3 |
| 2013 | Autonomous Fault-Tolerant Systems onto SRAM-based FPGA Platforms
Cristiana Bolchini, Antonio Miele, Chiara Sandionigi |
J. Electron. Test. | 2 |
| 2013 | A data-mining approach to preference-based data ranking founded on contextual information
Antonio Miele, Elisa Quintarelli, Emanuele Rabosio, Letizia Tanca |
Inf. Syst. | 1 |
| 2013 | Reliability-Driven System-Level Synthesis for Mixed-Critical Embedded SystemsabstractThis paper proposes a design methodology that enhances the classical system-level design flow for embedded systems to introduce reliability-awareness. The mapping and scheduling step is extended to support the application of hardening techniques to fulfill the required fault management properties that the final system must exhibit; moreover, the methodology allows the designer to specify that only some parts of the systems need to be hardened against faults. The reference architecture is a complex distributed one, constituted by resources with different characteristics in terms of performance and available fault detection/tolerance mechanisms. The approach is evaluated and compared against the most recent and relevant work, with an in-depth analysis on a large set of benchmarks. Cristiana Bolchini, Antonio Miele |
IEEE Trans. Computers | 2 |
| 2012 | An adaptive approach for online fault management in many-core architecturesabstractThis paper presents a dynamic scheduling solution to achieve fault tolerance in many-core architectures. Triple Modular Redundancy is applied on the multi-threaded application to dynamically mitigate the effects of both permanent and transient faults, and to identify and isolate damaged units. The approach targets the best performance, while balancing the use of the healthy resources to limit wear-out and aging effects, which cause permanent damages. Experimental results on synthetic case studies are reported, to validate the ability to tolerate faults while optimizing performance and resource usage. Cristiana Bolchini, Antonio Miele, Donatella Sciuto |
DATE | 2 |
| 2012 | Increasing autonomous fault-tolerant FPGA-based systems' lifetimeabstractIn this paper we propose an automated design flow for the implementation of autonomous fault-tolerant systems on SRAM-based FPGA platforms, able to cope with the occurrence of both transient and permanent faults. The goal of the proposed methodology is to increase the system's lifetime, by designing it able to detect and mitigate the effects of soft errors, as well as of permanent, non-recoverable ones, by exploiting dynamic reconfiguration. The application of the hardening design flow to a real case study is reported, to validate the methodology. Cristiana Bolchini, Antonio Miele, Chiara Sandionigi |
ETS | 2 |
| 2011 | Automated Resource-Aware Floorplanning of Reconfigurable Areas in Partially-Reconfigurable FPGA SystemsabstractThe floor planning activity is a key step in the design of systems on FPGAs, but the approaches available today rarely consider both the constraints imposed by the heterogeneous distribution of the resources in the devices and the reconfiguration capabilities. In fact, current-generation FPGAs present a complex architecture, but also offer more sophisticated reconfiguration features. The proposed floor planner, based on an accurate model of the devices, takes into account all these elements and finds an optimal solution, suitable for reconfigurable designs. Cristiana Bolchini, Antonio Miele, Chiara Sandionigi |
FPL | 2 |
| 2011 | Combined architecture and hardening techniques exploration for reliable embedded system designabstractThis paper proposes an approach for hardening embedded systems that combines the identification of the most convenient architecture and the set of application-level fault management techniques, fulfilling reliability requirements and maximizing performance. Cristiana Bolchini, Antonio Miele, Christian Pilato |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | A Novel Design Methodology for Implementing Reliability-Aware Systems on SRAM-Based FPGAsabstractThis paper presents a novel design flow for the implementation of digital systems onto SRAM-based FPGAs with soft error mitigation properties. Traditional fault detection/tolerance techniques are coupled with the device dynamic reconfiguration property to achieve soft error mitigation capabilities, and are applied to the single component, to groups of components or to the entire system, based on the most convenient trade-off with respect to a set of parameters. The design flow performs a two-steps multiobjective design space exploration, driven by a cost function taking into account resource utilization, area, and reconfiguration time. A floorplanning based on precise FPGA resource models is introduced to guarantee the feasibility of the hardened solution, identifying a convenient mapping onto the heterogeneous reconfigurable fabric. Experimental results show that the achieved solutions, aimed at achieving a prompt, "on demand” recovery when fault occurs, are characterized by a reduction in reconfiguration time that is higher than 80 percent, a significant improvement with respect to classical solutions. Cristiana Bolchini, Antonio Miele, Chiara Sandionigi |
IEEE Trans. Computers | 2 |
| 2010 | A multi-objective genetic algorithm framework for design space exploration of reliable FPGA-based systemsabstractThis paper presents a framework for the design space exploration of reliable FPGA systems based on a multi-objective genetic algorithm (NSGA-II). The framework takes into account several design metrics and outputs a set of Pareto-optimal design solutions. The framework is compared to the multi-objective version of simulated annealing (AMOSA) and it is empirically studied in terms of scalability using three real-world circuits and a set of synthetic problems of different sizes. Our results show that the proposed approach generates a rich set of Pareto-optimal solutions whereas AMOSA tends to find suboptimal solutions. Our empirical scalability analysis shows that, while the problem space is exponential in the number n of functional units constituting the system, the number of evaluations required by our framework grows as O(n3.6). Cristiana Bolchini, Pier Luca Lanzi, Antonio Miele |
IEEE Congress on Evolutionary Computation | 3 |
| 2010 | An integrated flow for the design of hardened circuits on SRAM-based FPGAsabstractThis paper presents an enhanced design flow for the implementation of hardened systems on SRAM-based FPGAs, able to cope with the occurrence of Single Event Upsets (SEUs). The framework integrates three strategies independently designed to tackle the problem of SEUs; first a systematic methodology is used to harden the circuit exploiting an enhanced TMR-based technique, coupled with partial dynamic reconfiguration. Then, a robustness analysis is performed to identify possible TMR failures, eventually solved by a specific local re-design of the critical portions of the implementation. We present the overall flow and the benefits of the solution, experimentally evaluated on a realistic circuit. Cristiana Bolchini, Antonio Miele, Chiara Sandionigi, Niccolò Battezzati, Luca Sterpone, Massimo Violante |
ETS | 2 |
| 2009 | A methodology for preference-based personalization of contextual dataabstractThe widespread use of mobile appliances, with limitations in terms of storage, power, and connectivity capability, requires to minimize the amount of data to be loaded on user's devices, in order to quickly select only the information that is really relevant for the users in their current contexts: in such a scenario, specific methodologies and techniques focused on data reduction must be applied. We propose an extension to the data tailoring approach of Context-ADDICT, whose aim is to dynamically hook and integrate heterogeneous data to be stored on small, possibly mobile devices. The main goal of our extension is to personalize the context-dependent data obtained by means of the Context-ADDICT methodology, by allowing the user to express preferences that specify which data s/he is more interested in (and which not) in each specific context. This step allows us to impose a partial order among the data, and to load only the top (most preferred) portion of the data chunks. A running example is used to better illustrate the approach. Antonio Miele, Elisa Quintarelli, Letizia Tanca |
EDBT | 1 |
| 2009 | Multi-level fault modeling for transaction-level specificationsabstractFault modeling is a fundamental element for several activities, ranging from off- and on-line testing, to fault tolerance and dependability-aware design. These activities are carried out during various design phases, dealing with specifications at different abstraction levels. Therefore, modeling faults across abstraction levels is of paramount importance to introduce dependability-related issues from the early phases of design. This paper analyzes how faults can be modeled at the different levels of abstraction with respect to Transaction Level Models, and how these models are related across levels. The work focuses on soft errors and aims at providing support to dependability analysis. A case study of a Transaction Level specification of a Network-on-Chip switch is used to evaluate the methodology and its applicability. Giovanni Beltrame, Cristiana Bolchini, Antonio Miele |
ACM Great Lakes Symposium on VLSI | 3 |
| 2008 | ReSP: A non-intrusive Transaction-Level Reflective MPSoC Simulation Platform for design space explorationabstractThis paper presents ReSP (Reflective Simulation Platform), a Transaction-Level multi-processor simulation platform based on SystemC and Python; SystemC is a standard language for system modeling and verification, and Python provides the platform with reflective capabilities. These are employed to give the designer an easy way to specify the architecture of a system, simulate the given configuration and perform automatic analysis on it. ReSP enables SystemC and Python interoperability through automatic Python wrapper generation. We show that the overhead associated with the Python intermediate layer is around 1%, therefore execution speed is not compromised. The advantages of our approach are: (a) easy integration of external IPs (b) fine grain control of the simulation (c) effortless integration of tools for system analysis and design space exploration. A case study shows how the platform can be extended to support system reliability assessment. Giovanni Beltrame, Cristiana Bolchini, Luca Fossati, Antonio Miele, Donatella Sciuto |
ASP-DAC | 4 |
| 2008 | Fault Models and Injection Strategies in SystemC SpecificationsabstractThis paper presents fault models and fault injection strategies designed in a simulation platform with reflection capabilities, used for simulating complex systems specified by using SystemC and by adopting a platform-based design approach. The approach allows the designer to work at different levels of abstraction and to take into account permanent and transient faults, and -- most important -- it features a transparent and dynamic mechanism for both injecting faults and analyzing the produced errors, in order to evaluate possible fault detection and/or tolerance design techniques. Cristiana Bolchini, Antonio Miele, Donatella Sciuto |
DSD | 2 |
| 2008 | Software and Hardware Techniques for SEU Detection in IP Processors
Cristiana Bolchini, Antonio Miele, Fabio Rebaudengo, Fabio Salice, Donatella Sciuto, Luca Sterpone, Massimo Violante |
J. Electron. Test. | 2 |