VLDB 2026 Research / reviewers in the wild / expert
Andrea Acquaviva
dblp:85/5704
· DBLP profile ↗
98ranked-venue papers
12as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 59 · 10 first-author · 15 since 2021Software engineering, systems software and programming languages · 21 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Computer networks · 5 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Distilling Graph Reasoning into Lightweight CNNs for Near-Sensor Point Cloud Corruption DetectionabstractReal-world point cloud corruption on automotive LiDAR lenses can significantly degrade the reliability of down-stream perception, particularly object detection models trained on clean data, which may yield overconfident false positives. To address this, we propose a near-sensor gating module that classifies incoming point clouds as either clean or contaminated using a teacher–student knowledge distillation pipeline. A Graph Attention Network (GAT), trained directly on raw point clouds, serves as the teacher. On real-world contaminated LiDAR data, the distilled student achieves an average F1-score of 0.83, closely matching the GAT teacher’s 0.88, and consistently outperforming other supervised baselines across diverse test environments. Importantly, the student’s 2D-CNN architecture reduces preprocessing complexity from O(n log n) of graph construction to O(n), enabling faster and more efficient point cloud handling. The student model is quantized to 16-bit fixed-point and deployed on a GAP8 (RISC-V) platform. It achieves an inference latency of 26 milliseconds, consumes only 210µJ per inference, and fits within 84KB of L2 memory. This makes the proposed solution a practical and resource-efficient near-sensor gating module for robust, contaminant-aware perception in embedded automotive systems. The implementation will be available at https://gitlab.com/ecs-lab/distilling-pointcloud-corruption Grafika Jati, Martin Molan, Francesco Barchi, Andrea Bartolini, Giuseppe Mercurio, Andrea Acquaviva |
DATE | 6 |
| 2026 | CVA6-CFI: A First Glance at RISC-V Control-Flow Integrity Extensions
Simone Manoni, Emanuele Parisi, Riccardo Tedeschi, Davide Rossi 0001, Andrea Acquaviva, Andrea Bartolini |
ISCAS | 5 |
| 2026 | FG-Tracer: Tracing Information Flow in Multimodal Large Language Models in Free-Form GenerationabstractMultimodal Large Language Models (MLLMs) have achieved impressive performance across a variety of vision–language tasks. However, their internal working mechanisms remain largely underexplored. In this work, we introduce FG-Tracer, a framework designed to analyze the information flow between visual and textual modalities in MLLMs in free-form generation. Notably, our numerically stabilized computational method enables the first systematic analysis of multimodal information flow in underexplored domains such as image captioning and chain-of-thought (CoT) reasoning. We apply FG-Tracer to three state-of-the-art MLLMs—LLaVA 1.5, LLaMA 3.2-Vision, and Qwen 2.5-VL—across three vision–language benchmarks—TextVQA, COCO 2014, and ChartQA—and we conduct a word-level analysis of multimodal integration. Our findings uncover distinct patterns of multimodal fusion across models and tasks, demonstrating that fusion dynamics are both model- and task-dependent. Overall, FG-Tracer offers a robust methodology for probing the internal mechanisms of MLLMs in free-form settings, providing new insights into their multimodal reasoning strategies. Our source code is publicly available at https://github.com/AImageLab-zip/FG-TRACER Alessia Saporita, Vittorio Pipoli, Federico Bolelli, Lorenzo Baraldi 0001, Andrea Acquaviva, Elisa Ficarra |
WACV | 5 |
| 2026 | Targeting language models for compile-time computing resource optimization: A novel approach based on masked graph autoencoders
Federico Cichetti, Emanuele Parisi, Andrea Acquaviva, Francesco Barchi |
Data Knowl. Eng. | 3 |
| 2026 | A transformer-based approach for source code classification for heterogeneous device mappingabstractThe optimization of code allocation for heterogeneous architectures, such as Central Processing Units (CPUs) and Graphics Processing Units (GPUs), remains challenging due to the limitations of traditional compiler heuristics and existing machine learning approaches. This paper presents a systematic evaluation of Large Language Models (LLMs) for classifying source code execution targets in heterogeneous device mapping. We fine-tune and compare six models: Distilled Bidirectional Encoder Representations from Transformers (DistilBERT), Code Bidirectional Encoder Representations from Transformers (CodeBERT), Code Bidirectional Encoder Representations from Transformers with RoBERTa (Robustly Optimized BERT Pretraining Approach) architecture (CodeBERTa), CodeT5, jTrans, and Deep Learning Low Level Virtual Machine (DeepLLVM), trained on Open Computing Language (OpenCL) kernels. Results show that general-purpose LLMs achieve up to 92.8% accuracy, matching or surpassing code-specific models, and outperform the previous state of the art (DeepLLVM) by up to 5%. Our findings indicate that LLMs pre-trained on general text are not necessarily inferior to code-specialized models, with tokenizer design and pre-training objectives impacting performance more than domain specialization. These results demonstrate the effectiveness of Transformer-based LLMs as a state-of-the-art approach for source code classification in heterogeneous computing contexts. • LLMs achieve SOTA in source code classification for heterogeneous device mapping. • Code-specific LLMs don’t always beat general ones–their edge depends on the dataset. • Transformer LLMs beat DeepLLVM in code-to-architecture mapping; CodeBERTa leads overall. Marco Siino, Emanuele Parisi, Francesco Barchi, Andrea Acquaviva, Andrea Bartolini |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Elevating Datacenter Resilience with ThermADNet: A Thermal Anomaly Detection SystemabstractIn the era of digital transformation, datacenters and High Performance Computing (HPC) Systems have emerged as the backbone of global technology infrastructure, powering essential services across various industries, including finance and healthcare. Therefore, ensuring the uninterrupted service of these datacenters has become a critical challenge. Thermal anomalies pose a significant risk to datacenter operation, potentially leading to hardware deterioration, system downtime, and catastrophic failures. This threat is exacerbated by the growing number of datacenters, increased power density, and heat waves fostered by global warming. Detecting thermal anomalies in datacenters involves several challenges. Large-scale data collection is difficult, requiring diverse monitoring signals from thousands of nodes over long periods. The absence of labeled data complicates the identification of normal and abnormal states. Establishing accurate classification thresholds to minimize false positives and negatives is another significant hurdle. Traditional statistical methods often fail to capture temporal dependencies and complex correlations in monitoring signals. Additionally, finding anomalies at both the system and subsystem levels adds to the complexity. Deploying machine learning models in production environments presents technical and operational challenges, making real-time anomaly detection a demanding task. This paper introduces ThermADNet, a Thermal Anomaly Detection framework that combines statistical rules-based methods with Deep Neural Network (DNN) techniques for thermal anomaly detection in datacenters. ThermADNet utilizes a semi-supervised learning approach by training on a ”semi-normal” dataset, addressing the challenges of large-scale data collection, semi-normal dataset identification, and classification threshold establishment. This framework’s efficacy is validated by its success in identifying real physical thermal failure events within a Tier-0 datacenter, pinpointing anomalies at both the system and subsystem levels, including compute nodes and datacenter infrastructure. In the critical evaluation window covering the July 28 failure, ThermADNet achieves precision and recall up to 0.97, with F1-scores as high as 0.97. By providing detailed information about anomalies, the framework clarifies the characteristics and reasoning behind the DNN outputs, thereby building trust in the AI model and ensuring that users can understand and rely on the system’s decisions. By offering a sophisticated method for thermal anomaly detection, ThermADNet significantly contributes to enhancing datacenter reliability and efficiency. This advancement supports the uninterrupted operation of critical HPC systems, averting considerable economic and societal losses. Mohsen Seyedkazemi Ardebili, Andrea Acquaviva, Luca Benini, Andrea Bartolini |
Future Gener. Comput. Syst. | 2 |
| 2026 | An OpenSSL Engine for Secure and Transparent Offloading of Cryptographic Operations to OP-TEE in Embedded IoT PlatformsabstractEmbedded IoT platforms are frequently deployed in hostile or physically exposed environments, where compromise of the operating system is a realistic threat. In conventional deployments, OpenSSL executes entirely in user space, leaving cryptographic keys and intermediate material potentially exposed in the presence of a compromised OS or privileged malware. This work presents a portable OpenSSL engine that enables secure and transparent offloading of cryptographic operations to OP-TEE, a software stack leveraging a Trusted Execution Environment (TEE) to confine cryptographic key material and security-critical computations within secure-world memory without modifying existing applications or altering the OpenSSL EVP interface. The proposed engine uses a GlobalPlatform Client API implementation as an underlying communication layer and supports both Copy-Based (CB) data transfer and a Shared-Memory (SM) mode, enabling performance tuning under the resource constraints typical of embedded IoT platforms. An experimental evaluation on an NXP i.MX7 industrial gateway shows that secure-world execution introduces bounded overhead dominated by REE–TEE transitions and memory-management operations. For bandwidth-intensive primitives such as SHA-256, SM mode improves throughput by approximately 10–15% over CB transfers. For symmetric encryption algorithms such as AES-256-CBC, SM communication provides a substantial throughput improvement of approximately 40%, indicating that data-movement overhead plays a significant role on embedded platforms. In contrast, asymmetric primitives (e.g. RSA-1024) remain largely insensitive to transfer optimisations because they operate on small, fixed-size operands, making the overall execution time dominated by invocation overheads and internal big-number computations rather than data movement. Tina Sayarmoafi, Francesco Barchi, Lorenzo Bottaccioli, Andrea Acquaviva, Edoardo Patti, Paolo Montuschi, Luca Barbierato |
IEEE Internet Things J. | 4 |
| 2025 | Tracing Information Flow in LLaMA Vision: A Step Toward Multimodal Understanding
Alessia Saporita, Vittorio Pipoli, Federico Bolelli, Lorenzo Baraldi 0001, Andrea Acquaviva, Elisa Ficarra |
CAIP (2) | 5 |
| 2025 | SpikeStream: Accelerating Spiking Neural Network Inference on RISC-V Clusters with Sparse Computation ExtensionsabstractSpiking Neural Network (SNN) inference has a clear potential for high energy efficiency as computation is triggered by events. However, the inherent sparsity of events poses challenges for conventional computing systems, driving the development of specialized neuromorphic processors, which come with high silicon area costs and lack the flexibility needed for running other computational kernels, limiting widespread adoption. In this paper, we explore the low-level software design, parallelization, and acceleration of SNNs on general-purpose multicore clusters with a low-overhead RISC-V ISA extension for streaming sparse computations. We propose SpikeStream, an optimization technique that maps weights accesses to affine and indirect register-mapped memory streams to enhance performance, utilization, and efficiency. Our results on the end-to-end Spiking-VGG11 model demonstrate a significant 4.39× speedup and an increase in utilization from 9.28% to 52.3 % compared to a non-streaming parallel baseline. Additionally, we achieve an energy efficiency gain of 3.46× over LSMCore and a performance gain of 2.38× over Loihi. Simone Manoni, Paul Scheffler, Luca Zanatta, Andrea Acquaviva, Luca Benini, Andrea Bartolini |
DATE | 4 |
| 2025 | ANZIL: Attention-based network for zero-risk inspection of LiDAR point cloud in self-driving carsabstractLiDAR is widely used in autonomous vehicle perception, and its performance relies on algorithmic confidence. However, sensor contamination can lead to catastrophic mistakes in downstream tasks, such as incorrect object detections occurring with high confidence. This underscores the need for point cloud contaminant detection to identify data reliability before downstream processing. To overcome this, we propose a model-agnostic approach that integrates contaminant detection with downstream tasks, using graph representations and attention networks. Trained on real contaminated LiDAR, the contaminant detector complements any existing clean-trained downstream model. Point clouds passing through the contaminant detector are discarded if contamination is detected. We propose a novel cost-benefit methodology, evaluating contaminant detectors in downstream processing. The benefit measures the proportion of discarded frames that would have caused high-confidence errors in the downstream task. The cost includes the model misalignment cost, representing frames wrongly discarded despite being correctly processed, and the true cost, which reflects uncontaminated frames falsely identified as contaminated. The contaminant detector was tested on 78,000 frames with water, dust, mud, salt, oil, and foam contamination from tunnel and outdoor locations. It achieved an F1-score of 0.884 and a recall of 0.957 on a static dataset, and a 0.975 F1-score in unseen real-world environments, demonstrating high sensitivity. In model-agnostic evaluation, the method is applicable to any object detector and reduces catastrophic mistakes by at least 97%, successfully identifying 4,046 failure cases. Deployed on the NVIDIA Jetson AGX Xavier, it achieved inference in under 5 ms and point cloud transformation in 200 ms, making it suitable for edge processing. This enhances sensor reliability, enables automatic cleaning, and improves vehicle safety. Grafika Jati, Martin Molan, Francesco Barchi, Andrea Bartolini, Giuseppe Mercurio, Andrea Acquaviva |
Expert Syst. Appl. | 6 |
| 2025 | Unleashing OpenTitan's Potential: a Silicon-Ready Embedded Secure Element for Root of Trust and Cryptographic OffloadingabstractThe rapid advancement and exploration of open-hardware RISC-V platforms are catalyzing substantial changes across critical sectors, including autonomous vehicles, smart-city infrastructure, and medical devices. Within this technological evolution, OpenTitan emerges as a groundbreaking open-source RISC-V design, renowned for its comprehensive security toolkit and role as a stand-alone system-on-chip (SoC). OpenTitan encompasses different SoC implementations such as Earl Grey, 1 fully implemented and silicon proven, and Darjeeling, 2 announced but not yet fully implemented. The former targets a stand-alone SoC implementation; the latter is oriented towards an integrable implementation. Therefore, the literature currently lacks a silicon-ready embedded implementation of an open-source Root of Trust despite the effort made by lowRISC on the Darjeeling implementation of OpenTitan. We address the limitations of existing implementations, focusing on optimizing data transfer latency between memory and cryptographic accelerators to prevent under-utilization and ensure efficient task acceleration. Our contributions include a comprehensive methodology for integrating custom extensions and intellectual properties (IPs) into the Earl Grey architecture, architectural enhancements for system-level integration, support for varied boot modes, and improved data movement across the platform. These advancements facilitate the deployment of OpenTitan in broader SoCs, even in scenarios lacking specific technology-dependent IPs, providing a deployment-ready research vehicle for the community. We integrated the extended Earl Grey architecture into a reference architecture in 22-nm FDX technology node. Then, we benchmarked the enhanced architecture’s performance, analyzing the latency introduced by the external memory hierarchic levels, presenting significant improvements in cryptographic processing speed, achieving up to 2.7 x speedup for SHA-256/HMAC and 1.6 x for AES accelerators compared with baseline Earl Grey architecture. Maicol Ciani, Emanuele Parisi, Alberto Musa, Francesco Barchi, Andrea Bartolini, Ari Kulmala, Rafail Psiakis, Angelo Garofalo, Andrea Acquaviva, Davide Rossi 0001 |
ACM Trans. Embed. Comput. Syst. | 9 |
| 2025 | Foundation Models for Structural Health MonitoringabstractStructural Health Monitoring (SHM) is a critical task for ensuring the safety and reliability of civil infrastructures, typically realized on bridges and viaducts by means of vibration monitoring. In this paper, we propose for the first time the use of Transformer neural networks, with a Masked Auto-Encoder architecture, asFoundation Modelsfor SHM. We demonstrate the ability of these models to learn generalizable representations from multiple large datasets through self-supervised pre-training, which, coupled with task-specific fine-tuning, allows them to outperform state-of-the-art traditional methods on diverse tasks, including Anomaly Detection (AD) and Traffic Load Estimation (TLE). We then extensively explore model size versus accuracy trade-offs and experiment with Knowledge Distillation (KD) to improve the performance of smaller Transformers, enabling their embedding directly into the SHM edge nodes. We showcase the effectiveness of our foundation models using data from three operational viaducts. For AD, we achieve a near-perfect 99.9% accuracy with a monitoring time span of just 15 windows. In contrast, a state-of-the-art method based on Principal Component Analysis (PCA) obtains its first good result (95.03% accuracy), only considering 120 windows. On two different TLE tasks, our models obtain state-of-the-art performance on multiple evaluation metrics (R2score, MAE% and MSE%). On the first benchmark, we achieve an R2score of 0.97 and 0.90 for light and heavy vehicle traffic, respectively, while the best previous approach (a Random Forest) stops at 0.91 and 0.84. On the second one, we achieve an R2score of 0.54 versus the 0.51 of the best competitor method, a Long-Short Term Memory network. Luca Benfenati, Daniele Jahier Pagliari, Luca Zanatta, Yhorman Alexander Bedoya Velez, Andrea Acquaviva, Massimo Poncino, Enrico Macii, Luca Benini, Alessio Burrello |
IEEE Trans. Sustain. Comput. | 5 |
| 2024 | TinyLid: a RISC-V accelerated Neural Network For LiDAR Contaminant Classification in Autonomous VehicleabstractLiDAR plays a critical role in autonomous car perception. Hence, the robustness of LiDAR data is imperative. However, malfunctions resulting from sensor cover contaminants are unavoidable and can lead to erroneous data that slowly degrade performance. Grafika Jati, Martin Molan, Francesco Barchi, Andrea Bartolini, Andrea Acquaviva |
CF | 5 |
| 2024 | Assessing the Performance of OpenTitan as Cryptographic Accelerator in Secure Open-Hardware System-on-ChipsabstractRISC-V open-source systems are emerging in deployment scenarios where safety and security are critical. OpenTitan is an open-source silicon root-of-trust designed to be deployed in a wide range of systems, from high-end to deeply embedded secure environments. Despite the availability of various cryptographic hardware accelerators that make OpenTitan suitable for offloading cryptographic workloads from the main processor, there has been no accurate and quantitative establishment of the benefits derived from using OpenTitan as a secure accelerator. This paper addresses this gap by thoroughly analysing strengths and inefficiencies when offloading cryptographic workloads to OpenTitan. The focus is on three key IPs --- HMAC, AES, and OpenTitan Big Number accelerator (OTBN) --- which can accelerate four security workloads: Secure Hash Functions, Message Authentication Codes, Symmetric cryptography, and Asymmetric cryptography. For every workload, we develop a bare-metal driver for the OpenTitan accelerator and analyze its efficiency when computation is offloaded from a RISC-V application core within a System-on-Chip designed for secure Cyber-Physical Systems applications. Finally, we assess it against a software implementation on the application core. The characterization was conducted on a cycle-accurate RTL simulator of the System-on-Chip (SoC). Our study demonstrates that OpenTitan significantly outperforms software implementations, with speedups ranging from 4.3x to 12.5x. However, there is potential for even greater gains as the current OpenTitan utilizes a fraction of the accelerator bandwidths, which ranges from 16% to 61%, depending on the memory being accessed and the accelerator used. Our results open the way to the optimization of OpenTitan-based secure platforms, providing design guidelines to unlock the full potential of its accelerators in secure applications. Emanuele Parisi, Alberto Musa, Maicol Ciani, Francesco Barchi, Davide Rossi 0001, Andrea Bartolini, Andrea Acquaviva |
CF | 7 |
| 2024 | TitanCFI: Toward Enforcing Control-Flow Integrity in the Root -of- TrustabstractModern RISC-V platforms control and monitor security-critical systems such as industrial controllers and autonomous vehicles. While these platforms feature a Root-of-Trust (RoT) to store authentication secrets and enable secure boot technologies, they often lack Control-Flow Integrity (CFI) enforcement and are vulnerable to cyber-attacks which divert the control flow of an application to trigger malicious behaviours. Recent techniques to enforce CFI in RISC-V systems include ISA modifications or custom hardware IPs, all requiring ad-hoc binary toolchains or design of CFI primitives in hardware. This paper proposes TitanCFI, a novel approach to enforce CFI in the RoT. TitanCFI modifies the commit stage of the protected core to stream control flow instructions to the RoT and it integrates the CFI enforcement policy in the RoT firmware. Our approach enables maximum reuse of the hardware resource present in the System-on-Chip (SoC), and it avoids the design of custom IPs and the modification of the compilation toolchain, while exploiting the RoT tamper-proof storage and cryptographic accelerators to secure CFI metadata. We implemented the proposed architecture on a modern RISC-V SoC along with a return address protection policy in the RoT, and benchmarked area and runtime overhead. Experimental results show that TitanCFI achieves overhead comparable to SoA hardware CFI solutions for most benchmarks, with lower area overhead, resulting in 1 % of additional area occupation. Emanuele Parisi, Alberto Musa, Simone Manoni, Maicol Ciani, Davide Rossi 0001, Francesco Barchi, Andrea Bartolini, Andrea Acquaviva |
DATE | 8 |
| 2024 | DeepCodeGraph: A Language Model for Compile-Time Resource Optimization Using Masked Graph Autoencoders
Federico Cichetti, Emanuele Parisi, Andrea Acquaviva, Francesco Barchi |
NLDB (1) | 3 |
| 2024 | TitanSSL: Towards Accelerating OpenSSL in a Full RISC-V Architecture Using OpenTitan Root-of-Trust
Alberto Musa, Franco Volante, Emanuele Parisi, Luca Barbierato, Edoardo Patti, Andrea Bartolini, Andrea Acquaviva, Francesco Barchi |
SAFECOMP | 7 |
| 2024 | Enhancing workplace safety: A flexible approach for personal protective equipment monitoringabstractWorkplace safety is a prominent concern, motivating researchers across diverse disciplines to investigate valuable ways to address its challenges. However, creating an efficient system to address this issue remains a significant challenge. Since many accidents happen due to improper usage or complete removal of Personal Protective Equipment (PPE), one straightforward method for enhancing workplace security involves monitoring their usage This paper introduces an Operator Area Network (OAN) system which improves the existing solutions by increasing portability across different users and environments, non-intrusiveness and privacy. To enhance robustness in detecting the situations in which PPEs are not used correctly, we take advantage of Machine Learning to analyse the received signal strength indicator (RSSI) between PPEs in the same OAN The novelty of this work is that it does not exploit RSSI as a proxy of the distance but instead recognises a signature of the correct wearing of the PPE By employing this system, employers can effectively ensure the proper usage of PPE devices at their worksites while also minimising any adverse effects on workers’ comfort and reducing the setup burden for employers. The system runs a Support Vector Machine (SVM) model several times per second and employs a post-processing algorithm to enhance its initial accuracy further As a result, the system effectively reduces false positives by about 80% and swiftly detects instances of improper usage of the worker’s PPE, raising the alarm in less than seven seconds. Moreover, the post-processing algorithm can be customised to meet the specific needs of different use cases, allowing for a flexible trade-off between the detection time interval and the overall accuracy of the detection system. Alessia Pisu, Nicola Elia, Livio Pompianu, Francesco Barchi, Andrea Acquaviva, Salvatore Carta |
Expert Syst. Appl. | 5 |
| 2024 | HazardNet: A thermal hazard prediction framework for datacenters
Mohsen Seyedkazemi Ardebili, Andrea Acquaviva, Luca Benini, Andrea Bartolini |
Future Gener. Comput. Syst. | 2 |
| 2024 | Energy efficient and low-latency spiking neural networks on embedded microcontrollers through spiking activity tuningabstractAbstract In this work, we target the efficient implementation of spiking neural networks (SNNs) for low-power and low-latency applications. In particular, we propose a methodology for tuning SNN spiking activity with the objective of reducing computation cycles and energy consumption. We performed an analysis to devise key hyper-parameters, and then we show the results of tuning such parameters to obtain a low-latency and low-energy embedded LSNN (eLSNN) implementation. We demonstrate that it is possible to adapt the firing rate so that the samples belonging to the most frequent class are processed with less spikes. We implemented the eLSNN on a microcontroller-based sensor node and we evaluated its performance and energy consumption using a structural health monitoring application processing a stream of vibrations for damage detection (i.e. binary classification). We obtained a cycle count reduction of 25% and an energy reduction of 22% with respect to a baseline implementation. We also demonstrate that our methodology is applicable to a multi-class scenario, showing that we can reduce spiking activity between 68 and 85% at iso-accuracy. Francesco Barchi, Emanuele Parisi, Luca Zanatta, Andrea Bartolini, Andrea Acquaviva |
Neural Comput. Appl. | 5 |
| 2023 | Directly-trained Spiking Neural Networks for Deep Reinforcement Learning: Energy efficient implementation of event-based obstacle avoidance on a neuromorphic acceleratorabstractSpiking Neural Networks (SNN) promise extremely low-power and low-latency inference on neuromorphic hardware. Recent studies demonstrate the competitive performance of SNNs compared with Artificial Neural Networks (ANN) in conventional classification tasks. In this work, we present an energy-efficient implementation of a Reinforcement Learning (RL) algorithm using SNNs to solve an obstacle avoidance task performed by an Unmanned Aerial Vehicle (UAV), taking a Dynamic Vision Sensor (DVS) as event-based input. We train the SNN directly, improving upon state-of-art implementations based on hybrid (not directly trained) SNNs. For this purpose, we devise an adaptation of the Spatio-Temporal Backpropagation algorithm (STBP) for RL. We then compare the SNN with a state-of-art Convolutional Neural Network (CNN) designed to solve the same task. To this aim, we train both networks by exploiting a photorealistic training pipeline based on AirSim. To achieve a realistic latency and throughput assessment for embedded deployment, we designed and trained three different embedded SNN versions to be executed on state-of-art neuromorphic hardware, targeting state-of-the-art. We compared SNN and CNN in terms of obstacle avoidance performance showing that the SNN algorithm achieves better results than the CNN with a factor of 6× less energy. We also characterise the different SNN hardware implementations in terms of energy and spiking activity. Luca Zanatta, Alfio Di Mauro, Francesco Barchi, Andrea Bartolini, Luca Benini, Andrea Acquaviva |
Neurocomputing | 6 |
| 2022 | Smart Contracts for Certified and Sustainable Safety-Critical Continuous Monitoring Applications
Nicola Elia, Francesco Barchi, Emanuele Parisi, Livio Pompianu, Salvatore Carta, Andrea Bartolini, Andrea Acquaviva |
ADBIS | 7 |
| 2022 | Meet Monte Cimone: exploring RISC-V high performance compute clustersabstractThe new open and royalty-free RISC-V ISA is attracting interest across the whole computing continuum, from microcontrollers to supercomputers. High-performance RISC-V processors and accelerators have been announced, but RISC-V-based HPC systems will need a holistic co-design effort, spanning memory, storage hierarchy interconnects and full software stack. In this paper, we describe Monte Cimone, a fully-operational multi-blade computer prototype and hardware-software test-bed based on U740, a double precision capable multi-core, 64 bit RISC-V SoC. Monte Cimone does not aim to achieve strong floating point performance, but it was built with the purpose of "priming the pipe" and exploring the challenges of integrating a multi-node RISC-V cluster capable of providing an HPC production stack including interconnect, storage and power monitoring infrastructure on RISC-V hardware. We present the results of our hardware/software integration effort, which demonstrate a remarkable level of software and hardware readiness and maturity - showing that the first-generation of RISC-V HPC machines may not be so far in the future. Federico Ficarelli, Andrea Bartolini, Emanuele Parisi, Francesco Beneventi, Francesco Barchi, Daniele Gregori, Fabrizio Magugliani, Marco Cicala, Cosimo Gianfreda, Daniele Cesarini, Andrea Acquaviva, Luca Benini |
CF | 11 |
| 2022 | Artificial versus spiking neural networks for reinforcement learning in UAV obstacle avoidanceabstractSpiking Neural Networks (SNN) are gaining more interest from the scientific community thanks to the promise of greater energy-efficient and greater computational power. This poses several challenges as today's SNN training for RL is based on Artificial Neural Network (ANN) training and then conversion from ANN to SNN, which does not leverage SNN event-based processing inherent capabilities. The present work compares an ANN and an SNN in an event-camera-based obstacle avoidance task, trained with Reinforcement Learning (RL) using the Deep Q-Learning (DQL) algorithm. We create an experimental setup composed of Unreal Engine 4, AirSim, and an event camera that simulates a real-world obstacle avoidance environment. Additionally, we train an SNN with a gradient-based training method enabling the use of all their expressiveness even in the training phase, showing comparable performance between the ANN and the SNN. To the best of our knowledge, we are the first that implements an entire realistic pipeline with a photo-realistic simulator (Airsim) and train an SNN without converting it from a pre-trained ANN. Luca Zanatta, Francesco Barchi, Andrea Bartolini, Andrea Acquaviva |
CF | 4 |
| 2022 | Making the Most of Scarce Input Data in Deep Learning-Based Source Code Classification for Heterogeneous Device MappingabstractDespite its relatively recent history, deep learning (DL)-based source code analysis is already a cornerstone in machine learning for compiler optimization. When applied to the classification of pieces of code to identify the best computational unit in a heterogeneous Systems-on-Chip, it can be effective in supporting decisions that a programmer has otherwise to take manually. Several techniques have been proposed exploiting different networks and input information, prominently sequence-based and graph-based representations, complemented by auxiliary information typically related to payload and device configuration. While the accuracy of DL methods strongly depends on the training and test datasets, so far no exhaustive and statistically meaningful analysis has been done on its impact on the results and on how to effectively extract the available information. This is relevant also considering the scarce availability of source code datasets that can be labeled by profiling on heterogeneous compute units. In this article, we first present such a study, which leads us to devise the contribution of code sequences and auxiliary inputs separately. Starting from this analysis, we then demonstrate that by using the normalization of auxiliary information, it is possible to improve state-of-the-art results in terms of accuracy. Finally, we propose a novel approach exploiting Siamese networks that further improve mapping accuracy by increasing the cardinality of the dataset, thus compensating for its relatively small size. Emanuele Parisi, Francesco Barchi, Andrea Bartolini, Andrea Acquaviva |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Trimming Feature Extraction and Inference for MCU-Based Edge NILM: A Systematic ApproachabstractNonintrusive load monitoring (NILM) enables the disaggregation of the global power consumption of multiple loads, taken from a single smart electrical meter, into appliance-level details. State-of-the-art approaches are based on machine learning methods and exploit the fusion of time- and frequency-domain features from current and voltage sensors. Unfortunately, these methods are compute-demanding and memory-intensive. Therefore, running low-latency NILM on low-cost resource-constrained microcontroller unit (MCU)-based meters is currently an open challenge. This article addresses the optimization of the feature spaces as well as the computational and storage cost reduction needed for executing state-of-the-art (SoA) NILM algorithms on memory- and compute-limited MCUs. We compare four supervised learning techniques on different classification scenarios and characterize the overall NILM pipeline's implementation on an MCU-basedSmart Measurement Node. Experimental results demonstrate that optimizing the feature space enables edge MCU-based NILM with 95.15% accuracy, resulting in a small drop compared to the most accurate feature vector deployment (96.19%) while achieving up to 5.45× speedup and 80.56% storage reduction. Furthermore, we show that low-latency NILM relying only on current measurements reaches almost 80% accuracy, allowing a major cost reduction by removing voltage sensors from the hardware (HW) design. Enrico Tabanelli, Davide Brunelli, Andrea Acquaviva, Luca Benini |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | Prediction of Thermal Hazards in a Real Datacenter Room Using Temporal Convolutional NetworksabstractDatacenters play a vital role in today's society. At large, a datacenter room is a complex controlled environment composed of thousands of computing nodes, which consume kW of power. To dissipate the power, forced air/liquid flow is employed, with a cost of millions of euros per year. Reducing this cost involves using free-cooling and average case design, which can create a cooling shortage and thermal hazards. When a thermal hazard happens, the system administrators and the facility manager must stop the production to avoid IT equipment damage and wear-out. In this paper, we study the thermal hazards signatures on a Tier-0 datacenter room's monitored data during a full year of production. We define a set of rules for detecting the thermal hazards based on the inlet and outlet temperature of all nodes of a room. We then propose a custom Temporal Convolutional Network (TCN) to predict the hazards in advance. The results show that our TCN can predict the thermal hazards with an Fl-score of 0.98 for a randomly sampled test set. When causality is enforced between the training and validation set the F1-score drops to 0.74, demanding for an in-place online re-training of the network, which motivates further research in this context. Mohsen Seyedkazemi Ardebili, Marcello Zanghieri, Alessio Burrello, Francesco Beneventi, Andrea Acquaviva, Luca Benini, Andrea Bartolini |
DATE | 5 |
| 2021 | Source Code Classification for Energy Efficiency in Parallel Ultra Low-Power MicrocontrollersabstractThe analysis of source code through machine learning techniques is an increasingly explored research topic aiming at increasing smartness in the software toolchain to exploit modern architectures in the best possible way. In the case of low-power, parallel embedded architectures, this means finding the configuration, for instance in terms of the number of cores, leading to minimum energy consumption. Depending on the kernel to be executed, the energy optimal scaling configuration is not trivial. While recent work has focused on general-purpose systems to learn and predict the best execution target in terms of the execution time of a snippet of code or kernel (e.g. offload OpenCL kernel on multicore CPU or GPU), in this work we focus on static compile-time features to assess if they can be successfully used to predict the minimum energy configuration on PULP, an ultra-low-power architecture featuring an on-chip cluster of RISC-V processors. Experiments show that using machine learning models on the source code to select the best energy scaling configuration automatically is viable and has the potential to be used in the context of automatic system configuration for energy minimisation. Emanuele Parisi, Francesco Barchi, Andrea Bartolini, Giuseppe Tagliavini, Andrea Acquaviva |
DATE | 5 |
| 2021 | Exploration of Convolutional Neural Network models for source code classification
Francesco Barchi, Emanuele Parisi, Gianvito Urgese, Elisa Ficarra, Andrea Acquaviva |
Eng. Appl. Artif. Intell. | 5 |
| 2021 | Solar radiation forecasting based on convolutional neural network and ensemble learning
Davide Cannizzaro, Alessandro Aliberti, Lorenzo Bottaccioli, Enrico Macii, Andrea Acquaviva, Edoardo Patti |
Expert Syst. Appl. | 5 |
| 2021 | Manufacturing as a Data-Driven Practice: Methodologies, Technologies, and ToolsabstractIn recent years, the introduction and exploitation of innovative information technologies in industrial contexts have led to the continuous growth of digital shop floor environments. The new Industry 4.0 model allows smart factories to become very advanced IT industries, generating an ever-increasing amount of valuable data. As a consequence, the necessity of powerful and reliable software architectures is becoming prominent along with data-driven methodologies to extract useful and hidden knowledge supporting the decision-making process. This article discusses the latest software technologies needed to collect, manage, and elaborate all data generated through innovative Internet-of-Things (IoT) architectures deployed over the production line, with the aim of extracting useful knowledge for the orchestration of high-level control services that can generate added business value. This survey covers the entire data life cycle in manufacturing environments, discussing key functional and methodological aspects along with a rich and properly classified set of technologies and tools, useful to add intelligence to data-driven services. Therefore, it serves both as a first guided step toward the rich landscape of the literature for readers approaching this field and as a global yet detailed overview of the current state of the art in the Industry 4.0 domain for experts. As a case study, we discuss, in detail, the deployment of the proposed solutions for two research project demonstrators, showing their ability to mitigate manufacturing line interruptions and reduce the corresponding impacts and costs. Tania Cerquitelli, Daniele Jahier Pagliari, Andrea Calimera, Lorenzo Bottaccioli, Edoardo Patti, Andrea Acquaviva, Massimo Poncino |
Proc. IEEE | 6 |
| 2021 | Supporting Telecommunication Alarm Management System With Trouble Ticket PredictionabstractFault alarm data emanated from heterogeneous telecommunication network services and infrastructures are exploding with network expansions. Managing and tracking the alarms with trouble tickets using manual or expert rule-based methods have become challenging due to increase in the complexity of alarm management systems and demand for deployment of highly trained experts. As the size and complexity of networks hike immensely, identifying semantically identical alarms, generated from heterogeneous network elements from diverse vendors, with data-driven methodologies, has become imperative to enhance efficiency. In this article, data-driven trouble ticket prediction models are proposed to leverage alarm management systems. To improve performance, feature extraction, using a sliding time window and feature engineering, from related history alarm streams, is also introduced. The models were trained and validated with a data set provided by the largest telecommunication provider in Italy. The experimental results showed the promising efficacy of the proposed approach in suppressing false positive alarms with trouble ticket prediction. Mulugeta Weldezgina Asres, Million Abayneh Mengistu, Pino Castrogiovanni, Lorenzo Bottaccioli, Enrico Macii, Edoardo Patti, Andrea Acquaviva |
IEEE Trans. Ind. Informatics | 7 |
| 2020 | Learning Behavioral Representations of Human MobilityabstractIn this paper, we investigate the suitability of state-of-the-art representation learning methods to the analysis of behavioral similarity of moving individuals, based on CDR trajectories. The core of the contribution is a novel methodological framework, mob2vec, centered on the combined use of a recent symbolic trajectory segmentation method for the removal of noise, a novel trajectory generalization method incorporating behavioral information, and an unsupervised technique for the learning of vector representations from sequential data. mob2vec is the result of an empirical study conducted on real CDR data through an extensive experimentation. As a result, it is shown that mob2vec generates vector representations of CDR trajectories in low dimensional spaces which preserve the similarity of the mobility behavior of individuals. Maria Luisa Damiani, Andrea Acquaviva, Fatima Hachem, Matteo Rossini |
SIGSPATIAL/GIS | 2 |
| 2019 | Code Mapping in Heterogeneous Platforms Using Deep Learning and LLVM-IRabstractModern heterogeneous platforms require compilers capable of choosing the appropriate device for the execution of program portions. This paper presents a machine learning method designed for supporting mapping decisions through the analysis of the program source code represented in LLVM assembly language (IR) for exploiting the advantages offered by this generalised and optimised representation. To evaluate our solution, we trained an LSTM neural network on OpenCL kernels compiled in LLVM-IR and processed with our tokenizer capable of filtering less-informative tokens. We tested the network that reaches an accuracy of 85% in distinguishing the best computational unit. Francesco Barchi, Gianvito Urgese, Enrico Macii, Andrea Acquaviva |
DAC | 4 |
| 2018 | Impact of graph partitioning on SNN placement for a multi-core neuromorphic architecture: work-in-progress
Francesco Barchi, Gianvito Urgese, Enrico Macii, Andrea Acquaviva |
CASES | 4 |
| 2018 | Multiple alignment of packet sequences for efficient communication in a many-core neuromorphic system: work-in-progress
Gianvito Urgese, Luca Peres, Francesco Barchi, Enrico Macii, Andrea Acquaviva |
CASES | 5 |
| 2018 | GIS-based optimal photovoltaic panel floorplanning for residential installationsabstractShading is a crucial issue for the placement of PV installations, as it heavily impacts power production and the corresponding return of investment. Nonetheless, residential rooftop installations still rely on rule-of-thumb criteria and on gross estimates of the shading patterns, while more optimized approaches focus solely on the identification of suitable surfaces (e.g., roofs) in a larger geographic area (e.g., city or district). This work addresses the challenge of identifying an optimal (with respect to the overall energy production) placement of PV panels on a roof. The novel aspect of the proposed solution lies in the possibility of having a sparse, irregular placement of individual modules so as to better exploit the variance of solar data. The latter are represented in terms of the distribution of irradiance and temperature values over the roof, as elaborated from historical traces and Geographical Information System (GIS) data. Experimental results will prove the effectiveness of the algorithm through three real world case studies, and that the generated optimal solutions allow to increase power production by up to 28% with respect to rule-of-thumb solutions. Sara Vinco, Lorenzo Bottaccioli, Edoardo Patti, Andrea Acquaviva, Enrico Macii, Massimo Poncino |
DATE | 4 |
| 2018 | A Compact PV Panel Model for Cyber-Physical Systems in Smart CitiesabstractOne of the ambitious goals of the "Smart city" paradigm is to design zero-energy buildings. Buildings can be considered as connected cyber-physical systems that require the construction of sound methodologies inherited from the Electronic Design Automation (EDA) research. In particular, aiming at autonomous buildings, the effective design of renewable energy sources is a key aspect for which such methodologies have to be developed. In this work, we propose a modeling strategy for the early estimation of the performance of photovoltaic (PV) arrays. Although a plethora of PV panel models there exists, most of these models suffer from accuracy/complexity tradeoffs. On one hand, building fast models forces to ignore either the correlation between temperature and irradiance, or the topology of panels, thus yielding inaccurate estimations. On the other, more accurate models are time consuming and require costly measurements or circuit analysis, that cannot be extracted from the sole datasheet. This paper proposes a compact semi-empirical model, suitable for real time simulation and built solely from information derived from the PV panel datasheet. The model is built by empirically fitting an expression of the panel operating point as a function of both irradiance and temperature, and of the adopted PV system topology. The accuracy and effectiveness of the proposed model have been validated w.r.t. the production traces of the PV systems of a real world industrial building. Sara Vinco, Lorenzo Bottaccioli, Edoardo Patti, Andrea Acquaviva, Massimo Poncino |
ISCAS | 4 |
| 2018 | Directed Graph Placement for SNN Simulation into a multi-core GALS ArchitectureabstractIn this paper, we present a methodology for effi-ciently mapping neural networks over a neuromorphic computing architecture. The target architecture is a globally asynchronous locally synchronous (GALS) multi-core designed for simulating spiking neural networks (SNN) in real-time, that is spike timings should be the same as in the human brain. The SNN is implemented as a set of concurrent tasks modelling the behaviour of biological neurons, which are executed on the processing cores and communicate through spikes travelling on a network-on-chip. The problem of neuron-to-core mapping is relevant as a non-efficient allocation may impact real-time and reliability of the neural network execution. We designed a task placement pipeline capable of analysing the network of neurons and producing a placement configuration that enables a reduction of communication between computational nodes. The neuron-to-core mapping problem has been formalised as a problem of minimisation of synaptic elongation. Intuitively, this metric represents the cumulative distance that spikes generated by neurons running on a specific core have to travel to reach their destination core. The proposed placement methodology allows using different techniques to solve the problem. In this work Spectral Analysis, Multilevel Static Mapping, and Simulated Annealing were compared evaluating the overall post-placement synaptic elongation. Results point out that mapping solutions taking into account the directionality of the SNN provide a better placement and quantify this impact. Between all techniques considered only the Simulated Annealing was able to overcome an improvement of 25% compared to a random placement. Francesco Barchi, Gianvito Urgese, Andrea Acquaviva, Enrico Macii |
VLSI-SoC | 3 |
| 2018 | A Parallel Hardware Architecture For Quantum Annealing Algorithm AccelerationabstractQuantum Annealing (QA) is an emerging technique, derived from Simulated Annealing, providing metaheuristics for multivariable optimisation problems. Studies have shown that it can be applied to solve NP-hard problems with faster convergence and better quality of result than other traditional heuristics, with potential applications in a variety of fields, from transport logistics to circuit synthesis and optimisation. In this paper, we present a hardware architecture implementing a QA-based solver for the Multidimensional Knapsack Problem, designed to improve the performance of the algorithm by exploiting parallelised computation. We synthesised the architecture using as a target an Altera FPGA board and simulated the execution for solving a set of benchmarks available in the literature. Simulation results show that the proposed implementation is about 100 times faster than a single-thread general-purpose CPU without impact on the accuracy of the solution. Evelina Forno, Andrea Acquaviva, Yuki Kobayashi, Enrico Macii, Gianvito Urgese |
VLSI-SoC | 2 |
| 2017 | Building Energy Modelling and Monitoring by Integration of IoT Devices and Building Information ModelsabstractIn recent years, the research about energy waste and CO2 emission reduction has gained a strong momentum, also pushed by European and national funding initiatives. The main purpose of this large effort is to reduce the effects of greenhouse emission, climate change to head for a sustainable society. In this scenario, Information and Communication Technologies (ICT) play a key role. From one side, advances in physical and environmental information sensing, communication and processing, enabled the monitoring of energy behaviour of buildings in real-time. The access to this information has been made easy and ubiquitous thank to Internet-of-Things (IoT) devices and protocols. From the other side, the creation of digital repositories of buildings and districts (i.e. Building Information Models - BIM) enabled the development of complex and rich energy models that can be used for simulation and prediction purposes. As such, an opportunity is emerging of mixing these two information categories to either create better models and to detect unwanted or inefficient energy behaviours. In this paper, we present a software architecture for management and simulation of energy behaviours in buildings that integrates heterogeneous data such as BIM, IoT, GIS (Geographical Information System) and meteorological services. This integration allows: i) (near-) real-time visualisation of energy consumption information in the building context and ii) building performance evaluation through energy modelling and simulation exploiting data from the field and real weather conditions. Finally, we discuss the experimental results obtained in a real-world case-study. Lorenzo Bottaccioli, Alessandro Aliberti, Francesca Maria Ugliotti, Edoardo Patti, Anna Osello, Enrico Macii, Andrea Acquaviva |
COMPSAC (1) | 7 |
| 2017 | A Flexible Distributed Infrastructure for Real-Time Cosimulations in Smart GridsabstractDue to the increasing penetration of distributed generation, storage, electric vehicles, and new information communication technologies, distribution networks are evolving toward the smart grid paradigm. For this reason, new control strategies, algorithms, and technologies need to be tested and validated before their actual field implementation. In this paper, we present a novel modular distributed infrastructure, based on real-time simulation, for multipurpose smart grid studies. The different components of the infrastructure are described, and the system is applied to a case study based on a real urban district located in northern Italy. The presented infrastructure is shown to be flexible and useful for different and multidisciplinary smart grid studies. Lorenzo Bottaccioli, Abouzar Estebsari, Enrico Pons, Ettore Bompard, Enrico Macii, Edoardo Patti, Andrea Acquaviva |
IEEE Trans. Ind. Informatics | 7 |
| 2017 | IoT Software Infrastructure for Energy Management and Simulation in Smart CitiesabstractThis paper presents an Internet-of-Things software infrastructure that enables energy management and simulation of new control policies in a city district. The proposed platform enables the interoperability and the correlation of (near-)real-time building energy profiles with environmental data from sensors as well as building and grid models. In a smart city context, this platform fulfills 1) the integration of heterogeneous data sources at the building and district level, and 2) the simulation of novel energy policies at the district level aimed at the optimization of the energy usage accounting also for its impact on building comfort. The platform has been deployed in a real-world district and a novel control policy for the heating distribution network has been developed and tested. Results are presented and discussed in the paper. Francesco Gavino Brundu, Edoardo Patti, Anna Osello, Matteo Del Giudice, Niccolo Rapetti, Alexandr Krylovskiy, Marco Jahn, Vittorio Verda, Elisa Guelpa, Laura Rietto, Andrea Acquaviva |
IEEE Trans. Ind. Informatics | 11 |
| 2016 | Toolchain integration of runtime variability and aging awareness in multicore platformsabstractThe impact of process variations and wear-out mechanisms in current and next generation technology nodes is becoming relevant and cannot be compensated at the device or architectural level. Intra-die process variations raising at the core level and platform level makes parallel multicore platforms intrinsically heterogeneous, because the various cores are clocked at different operational frequencies. Power consumption becomes heterogeneous too, both considering dynamic and leakage consumption. Wear-out processes add further uncertainty over time. In this context, to fully exploit the computational capability of the platform parallelism, variability and wear-out aware task allocation strategies must be developed. In this work, we discuss techniques to perform task allocation and we show how they can be integrated in a software toolchain. We report results of the implementation of variability and wear-out awareness in state-of-art multicore platforms. Ramakrishna Nittala, Francesco Barchi, Gianvito Urgese, Andrea Acquaviva |
FDL | 4 |
| 2016 | isomiR-SEA: an RNA-Seq analysis tool for miRNAs/isomiRs expression level profiling and miRNA-mRNA interaction sites evaluationabstractBACKGROUND: Massive parallel sequencing of transcriptomes, revealed the presence of many miRNAs and miRNAs variants named isomiRs with a potential role in several cellular processes through their interaction with a target mRNA. Many methods and tools have been recently devised to detect and quantify miRNAs from sequencing data. However, all of them are implemented on top of general purpose alignment methods, thus providing poorly accurate results and no information concerning isomiRs and conserved miRNA-mRNA interaction sites. RESULTS: To overcome these limitations we present a novel algorithm named isomiR-SEA, that is able to provide users with very accurate miRNAs expression levels and both isomiRs and miRNA-mRNA interaction sites precise classifications. Tags are mapped on the known miRNAs sequences thanks to a specialized alignment algorithm developed on top of biological evidence concerning miRNAs structure. Specifically, isomiR-SEA checks for miRNA seed presence in the input tags and evaluates, during all the alignment phases, the positions of the encountered mismatches, thus allowing to distinguish among the different isomiRs and conserved miRNA-mRNA interaction sites. CONCLUSIONS: isomiR-SEA performances have been assessed on two public RNA-Seq datasets proving that the implemented algorithm is able to account for more reliable and accurate miRNAs expression levels with respect to those provided by two compared state of the art tools. Moreover, differently from the few methods currently available to perform isomiRs detection, the proposed algorithm implements the evaluation of isomiRs and conserved miRNA-mRNA interaction sites already in the first alignment phases, thus avoiding any additional filtering stages potentially responsible for the loss of useful information. Gianvito Urgese, Giulia Paciello, Andrea Acquaviva, Elisa Ficarra |
BMC Bioinform. | 3 |
| 2015 | A new distributed framework for integration of district energy data from heterogeneous devices
Francesco Gavino Brundu, Edoardo Patti, Andrea Acquaviva, Michelangelo Grosso, Gaetano Rasconà, Salvatore Rinaudo, Enrico Macii |
DATE | 3 |
| 2015 | A tool-chain to foster a new business model for photovoltaic systems integration exploiting an Energy Community approachabstractNew approaches and business models for the development of renewable sources are needed as an alternative to feed-in tariffs. In this work, we present a tool-chain based on a distributed infrastructure for planning renewable energy systems deployment. This solution aims at fostering new services and business models by promoting energy community actions. Such tool-chain is able to: (i) evaluate the photovoltaic potential of the rooftops of a community; (ii) perform economic assessments of distributed photovoltaic system plants considering a “community based business model”. As case study, we considered a foothill community in north-west of Italy in which the tool-chain performed economic and energetic analyses. In order to integrate the proposed business model in the Italian regulatory framework, we analysed the Italian laws for electricity distribution and operation, highlighting the limitations in integrating such community approach. Lorenzo Bottaccioli, Edoardo Patti, Andrea Acquaviva, Enrico Macii, Matteo Jarre, Michel Noussan |
ETFA | 3 |
| 2014 | Towards a Software Infrastructure for District Energy ManagementabstractNowadays ICT is becoming a key factor to enhance the energy optimization in our cities. At district level, real-time information can be accessed to monitor and control the energy distribution network. Moreover, the fine grain monitoring and control done at building level can provide additional information to develop more efficient control policies for energy distribution in the district. In this paper we present a distributed software infrastructure for district energy management, which aims to provide a digital archive of the city in which energetic information is available. Such information is considered as the input for a decision system, which aims to increase the energy efficiency by promoting local balancing and shaving peak loads. As case study, we integrated in our proposed cloud the heating distribution network in Turin and we present exploitable options based on real-world environmental data to increase the energy efficiency and minimize the peak request. Edoardo Patti, Andrea Acquaviva, Adriano Sciacovelli, Vittorio Verda, Dario Martellacci, Federico Boni Castagnetti, Enrico Macii |
EUC | 2 |
| 2013 | HW-SW integration for energy-efficient/variability-aware computingabstractRecent trends in embedded system architectures brought a rapid shift towards multicore, heterogeneous and reconfigurable platforms. This imposes a large effort for programmers to develop their applications to efficiently exploit the underlying architecture. In addition, process variability issues lead to performance and power uncertainties, impacting expected quality of service and energy efficiency of the running software. In particular, variability may lead to sub-optimal runtime task allocation. Gasser Ayad, Andrea Acquaviva, Enrico Macii, Brahim Sahbi, Romain Lemaire |
DATE | 2 |
| 2013 | Integration of Literature with Heterogeneous Information for Genes Correlation ScoringabstractDetermining the correlation between biomedical terms is a powerful instrument to help scientist research activity, both to understand experimental results and to design new ones. In particular, a great potential comes from the integration of the many heterogeneous information sources currently available on the Web. In this article we focus on the correlation between genes and biological processes. In this context, we present a methodology for integrating information from biomedical literature with other heterogeneous types of structured information. In particular, the information sources integrated in this work are PubMed abstracts, pathway databases, and NCI thesaurus definitions. The integration is performed at the semantic analysis level using a customized approach we developed to modulate the impact of the different sources on the correlation score. We report the results of a study concerning the impact of the information integration on the correlation score and of the user-level parameters we introduced to modulate the impact of pathway data or NCI definitions with respect to biomedical literature information, depending on the context of the search. To evaluate the methodology, we performed correlation measures on six biological processes and nine genes by comparing the results with and without the integration of pathways and NCI definitions. Francesco Abate, Andrea Acquaviva, Elisa Ficarra, Enrico Macii |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2013 | Semi-Automatic Generation of Device Drivers for Rapid Embedded Platform DevelopmentabstractIP core integration into an embedded platform implies the implementation of a customized device driver complying with both the IP communication protocol and the CPU organization (single processor, SMP, AMP). Such a close dependence between driver and platform organization makes reuse of already existing device drivers very hard. Designers are forced to manually customize the driver code to any different organization of the target platform. This results in a very time-consuming and error-prone task. In this paper, we propose a methodology to semi-automatically generate customized device drivers, thus allowing a more rapid embedded platform development. The methodology exploits the testbench provided with the RTL IP module for extracting the formal model of the IP communication protocol. Then, a taxonomy of device drivers based on the CPU organization allows the system to determine the characteristics of the target platform and to obtain a template of the device driver code. This requires some manual support to identify the target architecture and to generate the desired device driver functionality. The template is used then to automatically generate drivers compliant with 1) the CPU organization, 2) the use in a simulated or in a real platform, 3) the interrupt support, 4) the operating system, 5) the I/O architecture, and 6) possible parallel execution. The proposed methodology has been successfully tested on a family of embedded platforms with different CPU organizations. Andrea Acquaviva, Nicola Bombieri, Franco Fummi, Sara Vinco |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2013 | Gelsius: A Literature-Based Workflow for Determining Quantitative Associations between Genes and Biological ProcessesabstractAn effective knowledge extraction and quantification methodology from biomedical literature would allow the researcher to organize and analyze the results of high-throughput experiments on microarrays and next-generation sequencing technologies. Despite the large amount of raw information available on the web, a tool able to extract a measure of the correlation between a list of genes and biological processes is not yet available. In this paper, we present Gelsius, a workflow that incorporates biomedical literature to quantify the correlation between genes and terms describing biological processes. To achieve this target, we build different modules focusing on query expansion and document cononicalization. In this way, we reached to improve the measurement of correlation, performed using a latent semantic analysis approach. To the best of our knowledge, this is the first complete tool able to extract a measure of genes-biological processes correlation from literature. We demonstrate the effectiveness of the proposed workflow on six biological processes and a set of genes, by showing that correlation results for known relationships are in accordance with definitions of gene functions provided by NCI Thesaurus. On the other side, the tool is able to propose new candidate relationships for later experimental validation. The tool is available at >http://bioeda1.polito.it:8080/medSearchServlet/. Francesco Abate, Andrea Acquaviva, Elisa Ficarra, Roberto Piva, Enrico Macii |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2013 | Aging-Aware Energy-Efficient Workload Allocation for Mobile Multimedia PlatformsabstractMulticore platforms are characterized by increasing variability and aging effects that imply heterogeneity in core performance, energy consumption, and reliability. In particular, wear-out effects such as negative-bias-temperature-instability require runtime adaptation of system resource utilization to time-varying and uneven platform degradation, so as to prevent premature chip failure. In this context, task allocation techniques can be used to deal with heterogeneous cores and extend chip lifetime while minimizing energy and preserving quality of service. We propose a new formulation of the task allocation problem for variability affected platforms, which manages per-core utilization to achieve a target lifetime while minimizing energy consumption during the execution of rate-constrained multimedia applications. We devise an adaptive solution that can be applied online and approximates the result of an optimal, offline version. Our allocator has been implemented and tested on real-life functional workloads running on a timing accurate simulator of a next-generation industrial multicore platform. We extensively assess the effectiveness of the online strategy both against the optimal solution and also compared to alternative state-of-the-art policies. The proposed policy outperforms state-of-the-art strategies in terms of lifetime preservation, while saving up to 20 percent of energy consumption without impacting timing constraints. Francesco Paterna, Andrea Acquaviva, Luca Benini |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2012 | A Cloud Infrastructure for Optimization of a Massive Parallel Sequencing WorkflowabstractMassive Parallel Sequencing is a term used to describe several revolutionary approaches to DNA sequencing, the so-called Next Generation Sequencing technologies. These technologies generate millions of short sequence fragments in a single run and can be used to measure levels of gene expression and to identify novel splice variants of genes allowing more accurate analysis. The proposed solution provides novelty on two fields, firstly an optimization of the read mapping algorithm has been designed, in order to parallelize processes, secondly an implementation of an architecture that consists of a Grid platform, composed of physical nodes, a Virtual platform, composed of virtual nodes set up on demand, and a scheduler that allows to integrate the two platforms. Olivier Terzo, Lorenzo Mossucca, Andrea Acquaviva, Francesco Abate, Rosalba Provenzano |
CCGRID | 3 |
| 2012 | Middleware services for network interoperability in smart energy efficient buildingsabstractOne of the major challenges in today's economy concerns the reduction in energy usage and CO2footprint in existing Public buildings and Spaces without significant construction works, by an intelligent ICT-based service monitoring and managing the energy consumption. In particular, interoperability between heterogeneous devices and networks, both existing and to be deployed is a key features to create efficient services and holistic energy control policies. In this paper we describe an innovative software infrastructure to provide a web-service based, hardware independent access to the heterogeneous networks of wireless sensor nodes, such as smart plugs for measuring energy motes for temperature, relative humidity and light monitoring. The proposed infrastructure allows easy extension to other networks, thus representing a contribute to the opening of a market for ICT-based customized solutions integrating numerous products from different vendors and offering services from design of integrated systems to the operation and maintenance phases. Edoardo Patti, Andrea Acquaviva, Francesco Abate, Anna Osello, A. Cocuccio, Marco Jahn, Marc Jentsch, Enrico Macii |
DATE | 2 |
| 2012 | On the automatic synthesis of parallel SW from RTL models of hardware IPsabstractHeterogeneous multicore system-on-chips (MPSoCs) provide many degrees of freedom to map functionalities on either SW and HW components. In this scenario, enabling the remapping of HW IPs as SW routines allows to fully exploit the computation power and flexibility provided by heterogeneous MPSoCs. On the other hand, reuse of existent IP cores is the key strategy to explore this large design space in a reasonable amount of time and to reduce the error risk during the MPSoC design flow. A methodology for automatic generation of parallel SW code taking into account these aspects is currently missing. This paper aims at overcoming this limitation, by presenting a methodology to automatically generate parallel SW IPs starting from existent RTL IP models. Andrea Acquaviva, Nicola Bombieri, Franco Fummi, Sara Vinco |
ACM Great Lakes Symposium on VLSI | 1 |
| 2012 | Bellerophontes: an RNA-Seq data analysis framework for chimeric transcripts discovery based on accurate fusion modelabstractMOTIVATION: Next-generation sequencing technology allows the detection of genomic structural variations, novel genes and transcript isoforms from the analysis of high-throughput data. In this work, we propose a new framework for the detection of fusion transcripts through short paired-end reads which integrates splicing-driven alignment and abundance estimation analysis, producing a more accurate set of reads supporting the junction discovery and taking into account also not annotated transcripts. Bellerophontes performs a selection of putative junctions on the basis of a match to an accurate gene fusion model. RESULTS: We report the fusion genes discovered by the proposed framework on experimentally validated biological samples of chronic myelogenous leukemia (CML) and on public NCBI datasets, for which Bellerophontes is able to detect the exact junction sequence. With respect to state-of-art approaches, Bellerophontes detects the same experimentally validated fusions, however, it is more selective on the total number of detected fusions and provides a more accurate set of spanning reads supporting the junctions. We finally report the fusions involving non-annotated transcripts found in CML samples. AVAILABILITY AND IMPLEMENTATION: Bellerophontes JAVA/Perl/Bash software implementation is free and available at http://eda.polito.it/bellerophontes/. Francesco Abate, Andrea Acquaviva, Giulia Paciello, Carmelo Foti, Elisa Ficarra, Alberto Ferrarini, Massimo Delledonne, Ilaria Iacobucci, Simona Soverini, Giovanni Martinelli, Enrico Macii |
Bioinform. | 2 |
| 2012 | Variability-Aware Task Allocation for Energy-Efficient Quality of Service Provisioning in Embedded Streaming Multimedia ApplicationsabstractMultimedia streaming applications running on next-generation parallel multiprocessor arrays in sub-45 nm technology face new challenges related to device and process variability, leading to performance and power variations across the cores. In this context, Quality of Service (QoS), as well as energy efficiency, could be severely impacted by variability. In this work, we propose a runtime variability-aware workload distribution technique for enhancing real-time predictability and energy efficiency based on an innovative Linear-Programming + Bin-Packing formulation which can be solved in linear time. We demonstrate our approach on the virtual prototype of a next-generation industrial multicore platform running representative multimedia applications. Experimental results confirm that our technique compensates variability, while improving energy-efficiency and minimizing deadline violations in presence of performance and power variations across the cores. The proposed policy can save up to 33 percent of energy with respect to the state-of-the-art policies and 65 percent of energy with respect to one variability-unaware task allocation policy while providing better QoS. Francesco Paterna, Andrea Acquaviva, Alberto Caprara, Francesco Papariello, Giuseppe Desoli, Luca Benini |
IEEE Trans. Computers | 2 |
| 2012 | Variability-tolerant workload allocation for MPSoC energy minimization under real-time constraintsabstractSub-50nm CMOS technologies are affected by significant variability, which causes power and performance variations among nominally similar cores in MPSoC platforms. This undesired heterogeneity threatens execution predictability and energy efficiency. We propose two techniques to allocate sets of barrier-synchronized tasks. The first technique models allocation as an ILP and achieves optimal results, but requires an offline solver. The second technique adopts a two-stage heuristic approach, and it can be adapted to work online. We tested our approach on the virtual prototype of a next-generation industrial multicore platform. Experimental results demonstrate that our approach minimizes deadline violations while increasing energy efficiency. Francesco Paterna, Andrea Acquaviva, Francesco Papariello, Giuseppe Desoli, Luca Benini |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2011 | Motion Artifact Correction in ASL images: An Improved Automated ProcedureabstractArterial Spin Labelling (ASL) is a perfusion MRI technique with tremendous applications in the study of biological markers and prognostic factors of brain tumors and in the assessment of neural diseases, moreover, it is completely non-invasive as it uses the magnetically inverted blood of the patient as an endogenous tracer. Unfortunately this powerful method is only viable in very limited conditions due to its extreme sensitivity to artifacts originated by head motion, that are not effectively addressed by the current software solutions. This paper presents a motion correction procedure that addresses this issue and provides improved solutions to enhance ASL images of the brain in presence of severe head motion. Experimental results run on a motion-affected pCASL dataset show the concept and demonstrate the superiority of our proposed procedure compared to standard 3D registration. Santa Di Cataldo, Elisa Ficarra, Andrea Acquaviva, Enrico Macii |
BIBM | 3 |
| 2011 | System level techniques to improve reliability in high power microcontrollers for automotive applicationsabstractIn high power microcontrollers, a decrease in circuit lifetime is often observed in safetycritical applications where circuitry is subjected to the most severe stresses and reliability has become a major concern. Thus, ad-hoc design solutions become necessary to mitigate the impact of ageing. In this paper we discuss hardware-software approaches that exploit distributed on-chip monitoring of wear-out parameters to perform ageing-aware allocation of computation and recovery periods on the various computational units. Andrea Acquaviva, Massimo Poncino, Marco Otella, Michele Sciolla |
DATE | 1 |
| 2011 | An efficient on-line task allocation algorithm for QoS and energy efficiency in multicore multimedia platformsabstractThe impact of variability on sub-45nm CMOS multimedia platforms makes hard to provide application QoS guarantees, as the speed variations across the cores may cause sub-optimal and sample-dependent utilization of the available resources and energy budget. These effects can be compensated by an efficient allocation of the workload at run-time. In the context of multimedia applications, a critical objective is to compensate core speed variability while matching time constraints without impacting the energy consumption. In this paper we present a new approach to compute optimal task allocations at run-time. The proposed strategy exploits an efficient and scalable implementation to find on-line the best possible solution in a tightly bounded time. Experimental results demonstrate the effectiveness of compensation both in terms of deadline miss rate and energy savings. Results have been compared with those obtained applying state-of-art techniques on a multithreaded MPEG2 decoder. The validation has been performed on a cycle-accurate virtual prototype of a next-generation industrial multicore platform that has been extended with process variability models. Francesco Paterna, Andrea Acquaviva, Alberto Caprara, Francesco Papariello, Giuseppe Desoli, Luca Benini |
DATE | 2 |
| 2011 | miREE: miRNA Recognition Elements EnsembleabstractBACKGROUND: Computational methods for microRNA target prediction are a fundamental step to understand the miRNA role in gene regulation, a key process in molecular biology. In this paper we present miREE, a novel microRNA target prediction tool. miREE is an ensemble of two parts entailing complementary but integrated roles in the prediction. The Ab-Initio module leverages upon a genetic algorithmic approach to generate a set of candidate sites on the basis of their microRNA-mRNA duplex stability properties. Then, a Support Vector Machine (SVM) learning module evaluates the impact of microRNA recognition elements on the target gene. As a result the prediction takes into account information regarding both miRNA-target structural stability and accessibility. RESULTS: The proposed method significantly improves the state-of-the-art prediction tools in terms of accuracy with a better balance between specificity and sensitivity, as demonstrated by the experiments conducted on several large datasets across different species. miREE achieves this result by tackling two of the main challenges of current prediction tools: (1) The reduced number of false positives for the Ab-Initio part thanks to the integration of a machine learning module (2) the specificity of the machine learning part, obtained through an innovative technique for rich and representative negative records generation. The validation was conducted on experimental datasets where the miRNA:mRNA interactions had been obtained through (1) direct validation where even the binding site is provided, or through (2) indirect validation, based on gene expression variations obtained from high-throughput experiments where the specific interaction is not validated in detail and consequently the specific binding site is not provided. CONCLUSIONS: The coupling of two parts: a sensitive Ab-Initio module and a selective machine learning part capable of recognizing the false positives, leads to an improved balance between sensitivity and specificity. miREE obtains a reasonable trade-off between filtering false positives and identifying targets. miREE tool is available online at http://didattica-online.polito.it/eda/miREE/ Paula Helena Reyes-Herrera, Elisa Ficarra, Andrea Acquaviva, Enrico Macii |
BMC Bioinform. | 3 |
| 2010 | An Automated Tool for Scoring Biomedical Terms Correlation Based on Semantic AnalysisabstractThe considerable improvement in the biotechnogical field and the adoption of screen techniques such as high throughput arrays produced a spread of biological and genetical data and scientific papers, mostly diffused on the Web. Even if the huge amount of available information represents a major step forward for the biomedical research field, the main effort for a scientist is to evaluate the correlation among the concepts both in qualitative and in quantitative terms. In the presented work, exploiting the richness of the UMLS Metathesaurus in combination with the novelty of the literature provided by PubMed, an automatic flow aimed at the scoring the semantic correlation among biomedical terms (e. g. biomolecules and biological processes) is proposed. The experiments show that obtained correlations are fully coherent with the information coming from the biological literature. The accuracy of the results remarks the importance of combining text mining techniques with complete and well structured thesaurus, such as UMLS. Francesco Abate, Elisa Ficarra, Andrea Acquaviva, Enrico Macii |
CISIS | 3 |
| 2010 | MicroRNA Target Prediction and Exploration through Candidate Binding Sites GenerationabstractGene regulation is one of the most important processes in the molecular biology, in the last years the microRNA molecule, one of the non-coding RNAs involved in the process, has been the focus of attention for several studies. The computational research on this area has gained a notable importance, considering the low amount of experimental information available and the lack of understanding of the microRNA binding mechanism. This article deals with the microRNA-target prediction and presents an innovative method for it. First it generates a set of promising binding sites for a given microRNA using a Genetic Algorithm, at the same time a set of target genes is selected based on the biological process under study. Secondly the set of promising binding sites is mapped into the selected set of target genes, in order to provide real binding sites and finally the resulting targets are filtered according to a biological or structural property. The objectives are to provide a flexible method that is capable of incorporating easily new knowledge, is independent of availability of the experimental information and is able to give hints on the research towards new characteristics among the microRNA binding sites such as motifs. The results present some of this novel properties and present a comparison with the most frequently used methods in the field. Paula Helena Reyes-Herrera, Andrea Acquaviva, Elisa Ficarra, Enrico Macii |
CISIS | 2 |
| 2010 | An integrated thermal estimation framework for industrial embedded platformsabstractNext generation industrial embedded platforms require the development of complex power and thermal management solutions. Indeed, an increasingly fine and intrusive thermal control is required because of temperature impact on leakage and reliability. To be effective, the implementation of these policies involves decisions that must be taken during various phases along the design process, to enable the development of architectural level countermeasures and the required hardware knobs, such as power modes, power supply regulation granularity and the number of on-chip temperature sensors. As a consequence, a framework allowing thermal estimation exploiting design-time information is desirable. Andrea Acquaviva, Andrea Calimera, Alberto Macii, Massimo Poncino, Enrico Macii, Matteo Giaconia, Claudio Parrella |
ACM Great Lakes Symposium on VLSI | 1 |
| 2010 | Thermal-aware floorplanning exploration for 3D multi-core architecturesabstractThermal effects are becoming increasingly important in today's sub-micron technologies. Thermal issues affect the performance, the reliability and the cooling costs of integrated systems. High peak temperatures are of major concern in modern 3D designs, where the stacking of multiple layers leads to higher power densities. Therefore, the integration of the thermal-aware design during the initial phases of the design can reduce the cost and the time-to-market of the resulting product. An efficient floorplanning in terms of thermal effects will reduce the appearance of critical hotspots and will spread heat across the chip area. David Cuesta, José Luis Ayala, J. Ignacio Hidalgo, Massimo Poncino, Andrea Acquaviva, Enrico Macii |
ACM Great Lakes Symposium on VLSI | 5 |
| 2009 | Flexible energy-aware simulation of heterogenous wireless sensor networksabstractThis paper presents an accurate and scalable implementation of an energy-aware simulator for wireless sensor networks (WSN's). Scalability and accuracy have been achieved through an energy-aware instrumentation of the Instruction Set Simulator of node's microcontroller and a functional SystemC TLM model of the radio module implementing the IEEE 802.15.4 protocol. The framework allows to execute actual software and to evaluate accurately its effect on the network lifetime. We first validate energy estimation results against a working hardware prototype of a wireless sensor node. The methodology, compared against state-of-the-art simulators such as NS-2, represents a flexible and scalable solution for fast and accurate prototyping of WSN software. Franco Fummi, Giovanni Perbellini, Davide Quaglia, Andrea Acquaviva |
DATE | 4 |
| 2009 | Adaptive idleness distribution for non-uniform aging tolerance in MultiProcessor Systems-on-ChipabstractIn deep submicron designs of MultiProcessor Systems-on-Chip (MPSoC) architectures, uncompensated within-die process variations and aging effects will lead to an increasing uncertainty and unbalancing of expected core lifetimes. In this paper we present an adaptive workload allocation strategy for run-time compensation of variations- and aging-induced unbalanced core lifetimes by means of core activity duty cycling. The proposed techniques regulates the percentage of idle time on short-expected-life cores to meet the platform lifetime target with minimum performance degradation. Experiments have been conducted on a multiprocessor simulator of a next-generation industrial MPSoC platform for multimedia applications made of a general purpose processor and programmable accelerators. Francesco Paterna, Luca Benini, Andrea Acquaviva, Francesco Papariello, Giuseppe Desoli, Mauro Olivieri |
DATE | 3 |
| 2009 | OpenMP Support for NBTI-Induced Aging Tolerance in MPSoCs
Andrea Marongiu, Andrea Acquaviva, Luca Benini |
SSS | 2 |
| 2009 | A Feedback-Based Approach to DVFS in Data-Flow ApplicationsabstractRuntime frequency and voltage adaptation has become very attractive for current and next generation embedded multicore platforms because it allows handling the workload variabilities arising in complex and dynamic utilization scenarios. The main challenge of dynamic frequency adaptation is to adjust the processing speed of each element to match the quality-of-service requirements in the presence of workload variations. In this paper, we present a control theoretic approach to dynamic voltage/frequency scaling for data-flow models of computations mapped to multiprocessor systems-on-chip architectures. We discuss, in particular, nonlinear control approaches to deal with general streaming applications containing both pipeline and parallel stages. Theoretical analysis and experiments, carried out by means of a cycle-accurate energy-aware multiprocessor simulation platform, are provided. We have applied the proposed control approach to realistic streaming applications such as Data Encryption Standard and software-based FM radio. Andrea Alimonda, Salvatore Carta, Andrea Acquaviva, Alessandro Pisano, Luca Benini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2009 | Thermal Balancing Policy for Multiprocessor Stream Computing PlatformsabstractDie-temperature control to avoid hotspots is increasingly critical in multiprocessor systems-on-chip (MPSoCs) for stream computing. In this context, thermal balancing policies based on task migration are a promising approach to redistribute power dissipation and even out temperature gradients. Since stream computing applications require strict quality of service and timing constraints, the real-time performance impact of thermal balancing policies must be carefully evaluated. In this paper, we present the design of a lightweight thermal balancing policy MiGra, which bounds on-chip temperature gradients via task migration. The proposed policy exploits run-time temperature as well as workload information of streaming applications to define suitable run-time thermal migration patterns, which minimize the number of deadline misses. Furthermore, we have experimentally assessed the effectiveness of our thermal balancing policy using a complete field-programmable-gate-array-based emulation of an actual three-core MPSoC streaming platform coupled with a thermal simulator. Our results indicate that MiGra achieves significantly better thermal balancing than state-of-the-art thermal management solutions while keeping the number of migrations bounded. Fabrizio Mulas, David Atienza 0001, Andrea Acquaviva, Salvatore Carta, Luca Benini, Giovanni De Micheli |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2008 | Segmentation of nuclei in cancer tissue images: Contrasting active contours with morphology-based approachabstractIn this paper we present a fully automated morphology-based technique for segmentation of nuclei in cancer tissue images and we compare it with a common technique for biomedical image processing, namely active contours. We discuss the limitations of active contours in the processing of immunohistochemical images characterized by heterogeneously stained nuclear region and noise caused by the presence of multiple tissue layers in the sample. We describe the integration of the proposed approach in a fully automated protein activity quantification tool. Finally, we demonstrate and motivate through extensive experiments that our fully automated morphology-based approach provides better accuracy compared to various active contours implementations. Santa Di Cataldo, Elisa Ficarra, Andrea Acquaviva, Enrico Macii |
BIBE | 3 |
| 2008 | Thermal Balancing Policy for Streaming Computing on Multiprocessor ArchitecturesabstractAs feature sizes decrease, power dissipation and heat generation density exponentially increase. Thus, temperature gradients in multiprocessor systems on chip (MPSoCs) can seriously impact system performance and reliability. Thermal balancing policies based on task migration have been proposed to modulate power distribution between processing cores to achieve temperature flattening. However, in the context of MPSoC for multimedia streaming computing, where timeliness is critical, the impact of migration on quality of service must be carefully analyzed. In this paper we present the design and implementation of a lightweight thermal balancing policy that reduces on-chip temperature gradients via task migration. This policy exploits run-time temperature and load information to balance the chip temperature. Moreover, we assess the effectiveness of the proposed policy for streaming computing architectures using a cycle-accurate thermal-aware emulation infrastructure. Our results using a real-life software defined radio multitask benchmark show that our policy achieves thermal balancing while keeping migration costs bounded. Fabrizio Mulas, Michele Pittau, Marco Buttu, Salvatore Carta, Andrea Acquaviva, Luca Benini, David Atienza 0001, Giovanni De Micheli |
DATE | 5 |
| 2008 | Temperature and Leakage Aware Power Control for Embedded Streaming ApplicationsabstractLeakage power has an increasing relevance for recent and future embedded technologies. Moreover, its contribution is heavily dependent on runtime temperature and process variations, that strongly impact the effectiveness of energy management strategies. Furthermore, in the context of streaming embedded systems, aggressive shutdown techniques that are typically applied to reduce leakage contribution are not suitable because of their negative effect on throughput. In this work we propose a leakage-aware approach for streaming applications that: i) Exploits feedback control policy to guarantee throughput by observing output data rate; ii) Levarages runtime temperature information to adjust the control law so that leakage energy is minimized independently from process variations and temperature conditions. Results show how temperature and leakage variations impact energy management strategies and how the proposed approach is effective in reducing total power consumption. Andrea Alimonda, Andrea Acquaviva, Salvatore Carta |
DSD | 2 |
| 2008 | Analysis of Power Management Strategies for a Large-Scale SoC Platform in 65nm TechnologyabstractLeakage power has become a major concern in nanometer technologies (65 nm and beyond), and new strategies are being proposed to overcome the limitations of traditional dynamic voltage and frequency scaling (DVFS) and shutdown (SD) approaches in dealing with leakage power and with its strong dependency on temperature and process variations. Even though many researchers have proposed effective point-solutions to these issues, a detailed analysis of their impact on a commercial large-scale multimedia SoC platform is still missing. This paper presents an explorative and comparative analysis of DVFS and SD power management options on a multi-million-gate SoC in 65 nm technology, and provides methodology directions and design insights. Andrea Marongiu, Luca Benini, Andrea Acquaviva, Andrea Bartolini |
DSD | 3 |
| 2008 | An energy-aware co-simulation framework for the design of wireless sensor networksabstractIn this paper we present an energy-aware framework for the design of wireless sensor networks, based on a hardware/software/network co-simulation strategy. The framework enables fast development and optimization of software power management policies, since it allows to run real software that can be easily ported on target hardware nodes. The capabilities of the simulation framework have been stressed by comparing different power management strategies, exploring their power-performance trade-offs and their effectiveness in various network conditions. Andrea Acquaviva, Franco Fummi, Giovanni Perbellini, Davide Quaglia |
ACM Great Lakes Symposium on VLSI | 1 |
| 2008 | Interfacing human and computer with wireless body area sensor networks: the WiMoCA solution
Elisabetta Farella, Augusto Pieracci, Luca Benini, Laura Rocchi, Andrea Acquaviva |
Multim. Tools Appl. | 5 |
| 2007 | Dynamic reconfiguration in sensor networks with regenerative energy sourcesabstractIn highly power constrained sensor networks, harvesting energy from the environment makes prolonged or even perpetual execution feasible. In such energy harvesting systems, energy sources are characterized as being regenerative. Regenerative energy sources fundamentally change the problem of power scheduling for embedded devices. Instead of the problem being one of maximizing the lifetime of the system given a total amount of energy, as in traditional battery powered devices, the problem becomes one of preventing energy depletion at any given time. Coupling relatively computationally intensive applications, such as video processing applications, with the constrained FPGAs that are feasible on power constrained embedded systems, makes dynamic reconfiguration essential. It provides the speed comparable to a hardware implementation, but it also allows the dynamic reconfiguration to meet the multiple application needs of the system. Different applications can be loaded on the FPGA, as the system's needs change over time. The problem becomes how to schedule the dynamic reconfiguration to appropriately make use of the regenerative energy source, to ensure the proper availability of energy for the system over time. This paper presents a methodology for carrying out dynamic reconfiguration for regenerative energy sources, based on statistical analysis of tasks and supply energy. The approach is evaluated through extensive simulations. Additionally, the implementation has been evaluated on regenerative energy, dynamically reconfigurable prototype, known as the MicrelEye. The approach is shown to miss 57.7% less deadlines on average than the current approach for reconfiguration with regenerative energy sources Ani Nahapetian, Paolo Lombardo, Andrea Acquaviva, Luca Benini, Majid Sarrafzadeh |
DATE | 3 |
| 2007 | Multi-processor operating system emulation framework with thermal feedback for systems-on-chipabstractMulti-Processor System-On-Chip (MPSoC) can provide the performance levels required by high-end embedded applications. However, they do so at the price of an increasing power density, which may lead to thermal runaway if coupled with low-cost packaging and cooling. Hence, mechanisms to efficiently evaluate the effectiveness of advanced thermal-aware operating-system (OS) strategies (e.g. task migration) onto the available MPSoC hardware are needed. In this paper, we propose a new MPSoC OS emulation framework that enables the study of thermal management strategies at the architectural- and OS-levels with the help of a standard FPGA. This framework includes the hardware and software components needed to accurately model complex MPSoCs architectures, and to test the effects of run-time thermal management strategies at the OS/middleware level with real-life inputs. Our results show that migration overhead is negligible w.r.t. temperature timings, enabling the development of thermal-aware migration strategies. Moreover, the effectiveness of the monitoring and feedback mechanism provides an emulation performance only ten times slower than real time. Salvatore Carta, Andrea Acquaviva, Pablo García Del Valle, David Atienza 0001, Giovanni De Micheli, Fernando Rincón Calle, Luca Benini, Jose Manuel Mendias |
ACM Great Lakes Symposium on VLSI | 2 |
| 2007 | A lightweight parallel java execution environment for embedded multiprocessor systems-on-chipabstractJava is a very popular execution environment for embeddedsystems platforms. However, current industrial and research implementationfocus on single-processor platforms. In this paperwe analyze a prototype parallel Java Virtual Machine implementationtargeted to a symmetric multi-core architecture using sharedmemory communication. We focus on performance and energyanalysis, and we quantify the overheads of the virtual executionenvironment. Marco Mantovani 0001, Simone Leardini, Martino Ruggiero, Andrea Acquaviva, Luca Benini |
ACM Great Lakes Symposium on VLSI | 4 |
| 2007 | Energetic sustainability of routing algorithms for energy-harvesting wireless sensor networks
Emanuele Lattanzi, Edoardo Regini, Andrea Acquaviva, Alessandro Bogliolo |
Comput. Commun. | 3 |
| 2007 | A control theoretic approach to energy-efficient pipelined computation in MPSoCsabstractIn this work, we describe a control theoretic approach to dynamic voltage/frequency scaling (DVFS) in a pipelined MPSoC architecture with soft real-time constraints, aimed at minimizing energy consumption with throughput guarantees. Theoretical analysis and experiments carried out on a cycle-accurate, energy-aware, and multiprocessor simulation platform are provided. We give a dynamic model of the system behavior which allows to synthesize linear and nonlinear feedback control schemes for the run-time adjustment of the core frequencies. We study the characteristics of the proposed techniques in both transient and steady-state conditions. Finally, we compare the proposed feedback approaches and local DVFS policies from an energy consumption viewpoint. Salvatore Carta, Andrea Alimonda, Alessandro Pisano, Andrea Acquaviva, Luca Benini |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2006 | A control theoretic approach to run-time energy optimization of pipelined processing in MPSoCsabstractIn this work we take a control-theoretic approach to feedback-based dynamic voltage scaling (DVS) in multi processor system on chip (MPSoC) pipelined architectures. We present and discuss a novel feedback approach based on both linear and non-linear techniques aimed at controlling interprocessor queue occupancy. Theoretical analysis and experiments, carried out on a cycle-accurate multiprocessor simulation platform, show that feedback-based control reduces energy consumption with respect to standard local DVS policies and highlight that non-linear strategies allows a more flexible and robust implementation in presence of variable workload conditions Andrea Alimonda, Andrea Acquaviva, Salvatore Carta, Alessandro Pisano |
DATE | 2 |
| 2006 | Supporting task migration in multi-processor systems-on-chip: a feasibility studyabstractWith the advent of multi-processor systems-on-chip, the interest in process migration is again on the rise both in research and in product development. New challenges associated with the new scenario include increased sensitivity to implementation complexity, tight power budgets, requirements on execution predictability, the lack of virtual memory support in many low-end MPSoCs. As a consequence, effectiveness and applicability of traditional transparent migration mechanisms are put in discussion in this context. Our paper proposes a task management software infrastructure that is well suited for the constraints of single chip multiprocessors with distributed operating systems. Load balancing in the system is maintained by means of intelligent initial placement and task migration. We propose a user-managed migration scheme based on code checkpointing and user-level middleware support as an effective solution for many MPSoC application domains. In order to prove the practical viability of this scheme, we also propose a characterization methodology for task migration overhead. We derive the minimum execution time following a task migration event during which the system configuration should be frozen to make up for the migration cost. Stefano Bertozzi, Andrea Acquaviva, Davide Bertozzi, Antonio Poggiali |
DATE | 2 |
| 2006 | A Wireless Body Area Sensor Network for Posture DetectionabstractBody Area Sensor Networks (BASN) are an emerging technology enabling the design of natural Human Computer Interfaces (HCI) in the context of Ambient Intelligence. This class of interactive applications poses new challenges on sensor network design that are hard to be faced using traditional solutions optimized for environmental monitoringlike applications. In this paper we present a novel solution for wireless and wearable posture recognition based on a custom-designed wireless body area sensor network, called WiMoCA. Nodes of the network, mounted on different parts of the human body, exploit tri-axial accelerometers to detect body postures. Afterwards we discuss results of interactive performance and power consumption optimizations required to match application constraints. Elisabetta Farella, Augusto Pieracci, Luca Benini, Andrea Acquaviva |
ISCC | 4 |
| 2006 | Power-Aware Network Swapping for Wireless Palmtop PCsabstractVirtual memory is considered to be an unlimited resource in desktop or notebook computers with high storage capabilities. However, in wireless mobile devices, like palmtops and personal digital assistants (PDAs), storage memory is limited or absent due to weight, size, and power constraints, so that swapping over remote memory devices can be considered as a viable alternative. However, power-hungry wireless network interface cards (WNICs) may limit the battery lifetime and application performance if not efficiently exploited. In this paper, we study performance and energy of network swapping in comparison with swapping on local microdrives and flash memories. We report the results of extensive experiments conducted on different WNICs and local swapping devices, using both synthetic and natural benchmarks. Our study points out that remote swapping over power-manageable WNICs can be more efficient than local swapping, especially in bursty workload conditions. Such conditions can be forced where possible by reshaping swapping requests to increase energy efficiency and performance. Andrea Acquaviva, Emanuele Lattanzi, Alessandro Bogliolo |
IEEE Trans. Mob. Comput. | 1 |
| 2005 | Application-Specific Power-Aware Workload Allocation for Voltage Scalable MPSoC PlatformsabstractIn this paper, we address the problem of selecting the optimal number of processing cores and their operating voltage/frequency for a given workload, to minimize overall system power under application-dependent QoS constraints. Selecting the optimal system configuration is non-trivial, since it depends on task characteristics and system-level interaction effects among the cores. For this reason, our QoS-driven methodology for power aware partitioning and frequency selection is based on functional, cycle-accurate simulation on a virtual platform environment. The methodology, being application-specific, is demonstrated on the DES (data encryption system) algorithm, representative of a wider class of streaming applications with independent input data frames and regular work-load. Martino Ruggiero, Andrea Acquaviva, Davide Bertozzi, Luca Benini |
ICCD | 2 |
| 2004 | A low-power motion capture system with integrated accelerometers [gesture recognition applications]abstractMotion capture is an emerging technology enabling the design of natural user interfaces for wearable devices based on gestural recognition. However, costs and energy requirements are critical factors to enable their diffusion to low-end wearable systems. Current commercial products do not match these requirements. For this reason, we developed a low-cost/low-power wearable motion tracking system, based on integrated accelerometers, called MOCA (motion capture with accelerometers). Our system is composed of sensing units connected to a control/acquisition board, responsible for reading and preprocessing data, and a mobile terminal running the recognition algorithm. Experiments performed to validate accuracy, power consumption and real-time performance demonstrate low-power and flexibility features of the proposed tracking system as well as its effectiveness as an input interface. Riccardo Barbieri, Elisabetta Farella, Luca Benini, Bruno Riccò, Andrea Acquaviva |
CCNC | 5 |
| 2004 | Power-Aware Network Swapping for Wireless Palmtop PCsabstractVirtual memory is considered to be an unlimited resource in desktop or notebook computers with high storage memory capabilities. However, in wireless mobile devices like palmtops and personal digital assistants (PDA), storage memory is limited or absent due to weight, size and power constraints. As a consequence, swapping over remote memory devices can be considered as a viable alternative. Nevertheless, power hungry wireless network interface cards (WNIC) may limit the battery lifetime and application performance if not efficiently exploited. In this work we explore performance and energy of network swapping in comparison with swapping on local micro-drives and flash memories. Our study points out that remote swapping over power-manageable WNICs can be more efficient than local swapping and that both energy and performance can be optimized through power-aware reshaping of data requests. Experimental results show that our optimization technique can save up to 60% of communication energy while improving performance. Andrea Acquaviva, Emanuele Lattanzi, Alessandro Bogliolo |
DATE | 1 |
| 2004 | Assessing the Impact of Dynamic Power Management on the Functionality and the Performance of Battery-Powered AppliancesabstractIn this paper we provide an incremental methodology to assess the effect of the introduction of a dynamic power manager in a mobile embedded computing device. The methodology consists of two phases. In the first phase, we verify whether the introduction of the dynamic power manager alters the functionality of the system. We show that this can be accomplished by employing standard techniques based on equivalence checking for noninterference analysis. In the second phase, we quantify the effectiveness of the introduction of the dynamic power manager in terms of power consumption and overall system efficiency. This is carried out by enriching the functional model of the system with information about the performance aspects of the system, and by comparing the values of the power consumption and the overall system efficiency obtained from the solution of the performance model with and without dynamic power manager. To this purpose, first we employ a more abstract performance model based on the Markovian assumption, then we use a more realistic performance model - to be validated against the Markovian one - where general probability distributions are considered. The methodology is illustrated by means of its application to the study of a remote procedure call mechanism - through which a battery-powered device is used by some application requesting information - and of a streaming video service - which is accessed by a mobile client equipped with a power-manageable network interface card. Andrea Acquaviva, Alessandro Aldini, Marco Bernardo 0001, Alessandro Bogliolo, Edoardo Bontà, Emanuele Lattanzi |
DSN | 1 |
| 2004 | Design and simulation of power-aware scheduling strategies of streaming data in wireless LANsabstractOne of the major concern for 802.11b wireless local area network is energy efficiency. In fact, mobile devices spend a large amount of power on their radio interface for accessing multimedia services such as audio and video streaming.In this work we address the problem of energy-aware scheduling of treaming data provided by a single server to multiple clients. We propose both open-loop and closed-loop strategies that exploit application level information to perform energy-effiocient traffic reshaping.We evaluate the effectiveness of the proposed trategies by means of accurate power/performance imulations performed on top of Mathworks' Simulink. System component are modeled as generalized semi-markov processes (GSMPs) and characterized by mean of real-world measurements. In particular, the timing and power behavior of wireless network interface cards is accurately captured in order to evaluate the impact of power management strategies on a wireless 802.11b link. Experimental result how that up to 75%of the communication energy can be aved by mean of power-aware traffic scheduling with negligible user-perceived performance degradation. Andrea Acquaviva, Emanuele Lattanzi, Alessandro Bogliolo |
MSWiM | 1 |
| 2002 | Low Power Control Techniques For TFT LCD DisplaysabstractDisplay power consumption is often the most significant contributor to the overall power budget for many portable devices. Traditionally, liquid crystal display (LCD) power minimization has focused on technology and circuit design. In this paper we take an orthogonal approach, and we introduce several software-only techniques for LCD dynamic power management, which do not require any hardware changes on existing LCDs and their controllers. The power savings achieved are significant: from 40% (with no perceivable image degradation) to 60% (with significant, but tolerable degradation) of total system power, measured on a prototype wearable system platform. Franco Gatti, Andrea Acquaviva, Luca Benini, Bruno Riccò |
CASES | 2 |
| 2002 | A low-power, fixed-point, front-end feature extraction for a distributed speech recognition systemabstractThis work describes the optimization of a signal processing front-end for a distributed speech recognition system with the goal of reducing power consumption. Two categories of source code optimizations were used, architectural and algorithmic. Architectural optimizations reduce the power consumption for a particular system, in this case, the HP Labs Smartbadge IV prototype portable system. Algorithmic optimizations are more general and involve changes in the algorithmic implementation of the source code to run faster and consume less power. A cycle accurate energy simulation shows a reduction in power usage by 83.5% with these optimizations. The optimized source code runs 34 times faster than the original code, therefore it can run at lower processor clock speeds and voltages for further reductions in power consumption. This technique, known as dynamic voltage scaling, was implemented on the Smartbadge IV hardware for an overall reduction in power usage of 89.2%. Brian Delaney, Nikil Jayant, Mathieu Hans, Tajana Rosing, Andrea Acquaviva |
ICASSP | 5 |
| 2001 | Dynamic Voltage Scaling and Power Management for Portable SystemsabstractPortable systems require long battery lifetime while still delivering high performance. Dynamic voltage scaling (DVS) algorithms reduce energy consumption by changing processor speed and voltage at run-time depending on the needs of the applications running. Dynamic power management (DPM) policies trade off the performance for the power consumption by selectively placing components into low-power states. In this work we extend the DPM model presented in [2, 3] with a DVS algorithm, thus enabling larger power savings. We test our approach on MPEG video and MP3 audio algorithms running on the SmartBadge portable device [1]. Our results show savings of a factor of three in energy consumption for combined DVS and DPM approaches. Tajana Rosing, Luca Benini, Andrea Acquaviva, Peter W. Glynn, Giovanni De Micheli |
DAC | 3 |
| 2001 | An adaptive algorithm for low-power streaming multimedia processingabstractThis paper addresses the problem of power consumption in multimedia system architectures and presents an algorithmic optimization technique to achieve the goal of power reduction in the context of real time processing. The technique is based on a mixed speed-setting and shutdown policy. We address the problem from both a theoretical and practical point of view, by presenting a power efficient implementation of a MPEG-layer3 real-time decoder algorithm designed for wearable devices as a case study. The target system is the Hewlett-Packard's SmartBadgeIII prototype of wearable system based on the StrongARM1100 processor. Theoretical analysis as well as quantitative results of power measurements are provided to show the effectiveness of this technique. The experimental set-up is also described. Andrea Acquaviva, Luca Benini, Bruno Riccò |
DATE | 1 |
| 2001 | Software-controlled processor speed setting for low-power streamingmultimediaabstractIn this paper, we describe a software-controlled approach for adaptively minimizing energy in embedded systems for real-time multimedia processing. Energy is optimized by clock speed setting: the software controller dynamically adjusts processor clock speed to the frame rate (FR) requirements of the incoming multimedia stream. The speed-setting policy is based on a system model that correlates clock speed with best case, average case, and worst case sustainable FRs, accounting for data dependency in multimedia streams. The technique has been implemented in a energy-efficient MPEG3 real-time decoder algorithm designed for wearable devices as a case study. The target system is the Hewlett-Packard SmartBadgeIII prototype system based on the StrongARM1100 processor. Hardware measurements show that computational energy can be drastically reduced (up to 40%) with respect to fixed-frequency operation. Andrea Acquaviva, Luca Benini, Bruno Riccò |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2000 | A spatially-adaptive bus interface for low-switching communication (poster session)abstractAdaptive encoding has shown to be an effective approach to bus power minimization in situations where characterization of the input statistics is not available. In this paper, we propose a novel technique for adaptive bus encoding that, conversely from existing solutions, exploits spatial correlations in the input data being transmitted to increase the accuracy in the dynamic selection of the encoding function. We discuss the encoding algorithm and we describe an architecture for its implementation as bus interface. We present experimental data collected in a realistic simulation framework on a number of meaningful benchmarks, and we compare them to those obtained through the application of existing encoding schemes. Andrea Acquaviva, Riccardo Scarsi |
ISLPED | 1 |