VLDB 2026 Research / reviewers in the wild / expert
Daniele Jahier Pagliari
dblp:173/7133
· DBLP profile ↗
65ranked-venue papers
20as first author
44since 2021 · last 2026
0000-0002-2872-7071ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 55 · 18 first-author · 35 since 2021Software engineering, systems software and programming languages · 15 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A System-Level Performance Analysis of On-Device Learning on an Ultra-Low-Power Edge SystemabstractTraining Deep Neural Networks (DNNs) on ultra-low-power processing systems is often hindered by the memory-intensive nature of backpropagation. This paper examines the trade-offs between memory and computational costs when mapping DNN training onto an ultra-low-power RISC-V-based System-on-Chip (SoC) with an 8-core accelerator, 1.5 MB on-chip memory, and a 32 MB off-chip memory. To this end, we develop a framework that automates the deployment of complete training graphs across heterogeneous SoCs. The generated code interleaves memory transfer operations with layer-wise processing functions, which are executed either on the host CPU or, for convolutional layers, on the multi-core accelerator. Overall, our work is the first to provide a full system-level analysis of end-to-end DNN training on OS-less, memory-constrained (32 MB of off-chip memory) edge systems, and we show that computation is the primary bottleneck by accounting for up to 91% of the total energy consumption. Manuele Rusci, Mohamed Amine Hamdi, Daniele Jahier Pagliari, Francesco Conti 0001, Alessio Burrello |
CF | 3 |
| 2026 | Late Breaking Results: CHESSY: Coupled Hybrid Emulation with SystemC-FPGA SynchronizationabstractThe growing complexity of cyber-physical systems (CPSs) calls for early prototyping tools that combine accuracy, speed, and usability. Virtual Platforms (VPs) provide fast functional simulation, but hybrid co-emulation solutions, in which key digital components are deployed on FPGA, become necessary when accurate timing modelling is required and RTL simulation is too costly. However, existing hybrid emulation tools are mostly proprietary, and rely on vendor-specific FPGA features. To address this gap, we introduce an open-source framework that connects SystemC-based VPs with FPGA emulation, enabling full-system co-emulation of digital and non-digital components. The FPGA accelerates the execution of main digital subsystems, while a wrapper coordinates timing and communication with the VP through JTAG, maintaining synchronization with simulated peripherals. Evaluations using a RISC-V SoC, with an example in the biosignals processing domain, show up to 2500× speedup compared to RTL simulation, while maintaining less than 2× total simulation time relative to pure FPGA emulation. Lorenzo Ruotolo, Giovanni Pollo, Mohamed Amine Hamdi, Matteo Risso, Yukai Chen, Enrico Macii, Massimo Poncino, Sara Vinco, Alessio Burrello, Daniele Jahier Pagliari |
DATE | 10 |
| 2026 | End-to-end Automated Deep Neural Network Optimization for PPG-based Blood Pressure Estimation on WearablesabstractPhotoplethysmography-based Blood Pressure (BP) estimation is a challenging task, particularly on resource-constrained wearable devices. However, fully on-board processing is desirable to ensure user data confidentiality. Recent Deep Neural Networks (DNNs) have achieved high BP estimation accuracy by reconstructing BP waveforms or directly regressing BP values, but their large memory, computation, and energy requirements hinder deployment on wearables. This work introduces a fully automated DNN design pipeline that combines hardware-aware Neural Architecture Search, pruning, and Mixed-Precision Search to generate accurate yet compact BP prediction models optimized for ultra-low-power multi-core Systems-on-Chip (SoCs). Starting from state-of-the-art baseline models on four public datasets, our optimized networks achieve up to 7.99% lower error with a 7.5 \(\times\) parameter reduction, or up to 83 \(\times\) fewer parameters with negligible accuracy loss. All models fit within 512 kB of memory on our target SoC (GreenWaves’ GAP8), requiring less than 55 kB and achieving an average inference latency of 142 ms and energy consumption of 7.25 mJ. Patient-specific fine-tuning further improves accuracy by up to 64%, enabling fully autonomous, low-cost BP monitoring on wearables. Francesco Carlucci, Giovanni Pollo, Xiaying Wang, Massimo Poncino, Enrico Macii, Luca Benini, Sara Vinco, Alessio Burrello, Daniele Jahier Pagliari |
ACM Trans. Comput. Heal. | 9 |
| 2026 | Replacing Attention With Modality-Wise Convolution for Energy-Efficient PPG-Based Heart Rate Estimation Using Knowledge DistillationabstractContinuous monitoring of Hearth Rate (HR) based on photoplethysmography (PPG) sensors is an essential capability of nearly all wrist-worn devices. However, arm movements lead to the creation of Motion Artifacts (MA), affecting the accuracy of HR tracking using PPG sensors. This problem is commonly tackled by exploiting the recorded accelerometer data to correlate them with the PPG signal and eventually clean it. Thus, automatic fusion techniques based on Deep Learning (DL) algorithms have been proposed, but they are considered too large and complex to be deployed on wearable devices. The current work presents a novel and lightweight DL architecture, PULSE, improving sensor fusion by applying a multi-head cross-attention layer to the extracted temporal features. Moreover, we propose a relation-based knowledge distillation mechanism to pass PULSE's knowledge to a student network that uses modality-wise convolutions to replace the attention module and mimic the teacher's performance with 5× fewer parameters. The teacher and student are evaluated on two datasets: a) PPG-DaLiA the most extensive available dataset, with PULSE achieving close performance to the best state-of-the-art model, and b) WESAD with PULSE reducing the mean absolute error by 22.6%. The student model is further compressed using post-training quantization and deployed on two commercial-off-the-shelf microcontrollers, demonstrating its suitability for real-time execution, having a close-to-state-of-the-art MAE of 4.81 BPM (+0.40 BPM) on the PPG-DaLiA, but a 10.9× lower memory footprint of 37.9 kB, and consuming 45.9× lower energy (0.577 mJ). Panagiotis Kasnesis, Lazaros Toumanidis, Daniele Jahier Pagliari, Alessio Burrello |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Design and Optimization for AI/ML Acceleration on Resource-constrained and Edge SystemsabstractThe rapid advancement of AI (from foundational machine learning to Large Language Models) and edge computing has placed unprecedented demands on computation, memory, and storage on resource-constrained edge devices. As AI models scale, the ability to efficiently manage computing resources, utilize memory and storage, and reduce energy consumption has become critical. This paper introduces contributions on 4 topics related to deploying AI on resource-constrained edge devices: 1) unlocking training of foundational machine learning algorithms on the edge, 2) exploring hardware-aware DNN architecture and mapping co-optimization for inference on heterogeneous systems, 3) scaling RAG by leveraging advanced memory, storage, and energy-efficient designs, and 4) investigating cost-effective and high-performance large-scale graph processing. Jalil Boukhobza, Alessio Burrello, Yuan-Hao Chang 0001, Yawei Li 0001, Daniele Jahier Pagliari, Chun-Feng Wu, Ming-Chang Yang, Tsun-Yu Yang |
CASES | 5 |
| 2025 | POSTER: V-Seek: Optimizing LLM Reasoning on A Server-Class General-Purpose RISC-V PlatformabstractThis paper addresses the problem of deployment of LLMs on RISC-V-based CPU systems by optimizing LLM inference on the Sophon SG2042.We evaluate the inference performance of two state-of-theart LLMs optimised for reasoning: DeepSeek R1 Distill Llama 8B and DeepSeek R1 Distill QWEN 14B.Thanks to our optimizations on top of the llama.cppinference library, we achieve token generation speeds of 4.32/2.29 tokens per second and prompt processing speeds of 6.54/3.68tokens per second, with a significant speedup of up to 2.9×/3.0×compared to a direct porting of the same library. Javier J. Poveda Rodrigo, Mohamed Amine Hamdi, Cyril Koenig, Alessio Burrello, Daniele Jahier Pagliari, Luca Benini |
CF | 5 |
| 2025 | Multi-Partner Project: Advancing the EDA Tools Landscape for the European RISC-V Ecosystem in TRISTANabstractThe TRISTAN project aims to expand and industrialize the European RISC-V ecosystem to compete effectively with existing commercial alternatives. This initiative specifically targets the critical challenges in the development of Electronic Design Automation (EDA) tools, essential for RISC-V-based solutions, by leveraging the synergy between the open-source community and industrial solutions. This paper presents an overview of the current landscape of TRISTAN's EDA flow, highlighting specific tools and methodologies that streamline the early design phases of RISC-V-based systems. We explore the unique features of these tools, emphasizing how they complement each other to strengthen the overall design process. Fatma Jebali, Caaliph Andriamisaina, Mathieu Jan, Wolfgang Ecker, Florian Egert, Bernhard Fischer, Alessio Burrello, Daniele Jahier Pagliari, Sara Vinco, Giuseppe Tagliavini, Ingo Feldner, Andreas Mauderer, Axel Sauer, Arnór Kristmundsson, Alexander Schober, Téo Bernier, Matti Käyrä, Ulf Schlichtmann, Rocco Jonack |
DATE | 8 |
| 2025 | Coupling Neural Networks and Physics Equations For Li-Ion Battery State-of-Charge PredictionabstractEstimating the evolution of the battery's State of Charge (SoC) in response to its usage is critical for implementing effective power management policies and for ultimately improving the system's lifetime. Most existing estimation methods are either physics-based digital twins of the battery or data-driven models such as Neural Networks (NNs). In this work, we propose two new contributions in this domain. First, we introduce a novel NN architecture formed by two cascaded branches: one to predict the current SoC based on sensor readings, and one to estimate the SoC at a future time as a function of the load behavior. Second, we integrate battery dynamics equations into the training of our NN, merging the physics-based and data-driven approaches, to improve the models' generalization over variable prediction horizons. We validate our approach on two publicly accessible datasets, showing that our Physics-Informed Neural Networks (PINNs) outperform purely data-driven ones while also obtaining superior prediction accuracy with a smaller architecture with respect to the state-of-the-art. Giovanni Pollo, Alessio Burrello, Enrico Macii, Massimo Poncino, Sara Vinco, Daniele Jahier Pagliari |
DATE | 6 |
| 2025 | Automatic integration of SystemC in the FMI standard for Software-defined Vehicle designabstractThe recent advancements of the automotive sector demand robust co-simulation methodologies that enable early validation and seamless integration across hardware and software domains. However, the lack of standardized interfaces and the dominance of proprietary simulation platforms pose significant challenges to collaboration, scalability, and IP protection. To address these limitations, this paper presents an approach for automatically wrapping SystemC models by using the Functional Mock-up Interface (FMI) standard. This method combines the modeling accuracy and fast time-to-market of SystemC with the interoperability and encapsulation benefits of FMI, enabling secure and portable integration of embedded components into co-simulation workflows. We validate the proposed methodology on real-world case studies, demonstrating its effectiveness with complex designs. Giovanni Pollo, Andrei Mihai Albu, Alessio Burrello, Daniele Jahier Pagliari, Cristian Tesconi, Loris Panaro, Dario Soldi, Fabio Autieri, Sara Vinco |
FDL | 4 |
| 2025 | Building Damage Assessment in Conflict Zones: A Deep Learning Approach Using Geospatial Sub-Meter Resolution DataabstractVery High Resolution (VHR) geospatial image analysis is crucial for humanitarian assistance in both natural and anthropogenic crises, as it allows to rapidly identify the most critical areas that need support. Nonetheless, manually inspecting large areas is time-consuming and requires domain expertise. Thanks to their accuracy, generalization capabilities, and highly parallelizable workload, Deep Neural Networks (DNNs) provide an excellent way to automate this task. Nevertheless, there is a scarcity of VHR data pertaining to conflict situations, and consequently, of studies on the effectiveness of DNNs in those scenarios. Motivated by this, our work extensively studies the applicability of a collection of state-of-the-art Convolutional Neural Networks (CNNs) originally developed for natural disasters damage assessment in a war scenario. To this end, we build an annotated dataset with pre- and post-conflict images of the Ukrainian city of Mariupol. We then explore the transferability of the CNN models in both zero-shot and learning scenarios, demonstrating their potential and limitations. To the best of our knowledge, this is the first study to use sub-meter resolution imagery to assess building damage in combat zones. Matteo Risso, Alessia Goffi, Beatrice Alessandra Motetti, Alessio Burrello, Jean Baptiste Bove, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari, Giuseppe Maffeis |
IPAS | 8 |
| 2025 | MEbots: Integrating a RISC-V Virtual Platform with a Robotic Simulator for Energy-aware DesignabstractVirtual Platforms (VPs) enable early software validation of autonomous systems’ electronics, reducing costs and time-to-market. While many VPs support both functional and non-functional simulation (e.g., timing, power), they lack the capability of simulating the environment in which the system operates. In contrast, robotics simulators lack accurate timing and power features. This twofold shortcoming limits the effectiveness of the design flow, as the designer can not fully evaluate the features of the solution under development. This paper presents a novel, fully open-source framework bridging this gap by integrating a robotics simulator (Webots) with a VP for RISC-V-based systems (MESSY). The framework enables a holistic, mission-level, energy-aware co-simulation of electronics in their surrounding environment, streamlining the exploration of design configurations and advanced power management policies. Giovanni Pollo, Mohamed Amine Hamdi, Matteo Risso, Lorenzo Ruotolo, Pietro Furbatto, Matteo Isoldi, Yukai Chen, Alessio Burrello, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari, Sara Vinco |
ISLPED | 11 |
| 2025 | MATCH: Model-Aware TVM-Based Compilation for Heterogeneous Edge DevicesabstractStreamlining the deployment of Deep Neural Networks (DNNs) on heterogeneous edge platforms, coupling within the same micro-controller unit (MCU) instruction processors and hardware accelerators for tensor computations, is becoming one of the crucial challenges of the TinyML field. The best-performing DNN compilation toolchains are usually deeply customized for a single MCU family, and porting them to a different one implies labor-intensive redevelopment of almost the entire compiler. On the opposite side, retargetable toolchains, such as TVM, fail to exploit the capabilities of custom accelerators, producing general but unoptimized code. To overcome this duality, we introduce MATCH, a novel TVM-based DNN deployment framework designed for easy agile retargeting across different MCU processors and accelerators, thanks to a customizable model-based hardware abstraction. We show that a general and retargetable mapping framework can compete with, and even outperform custom toolchains on diverse targets while only needing the definition of an abstract hardware cost model and a SoC-specific API. We tested MATCH on two state-of-the-art heterogeneous MCUs, GAP9 and DIANA. On the four DNN models of the MLPerf Tiny suite MATCH reduces inference latency on average by$60.87\times $on DIANA, compared to using the plain TVM, thanks to the exploitation of the on-board HW accelerator. Compared to HTVM, a fully customized toolchain for DIANA, we still reduce the latency by 16.94%. On GAP9, using the same benchmarks, we improve the latency by$2.15\times $compared to the dedicated DORY compiler, thanks to our heterogeneous DNN mapping approach that synergically exploits the DNN accelerator and the eight-cores cluster available on board. Mohamed Amine Hamdi, Francesco Daghero, Giuseppe Maria Sarda, Josse Van Delm, Arne Symons, Luca Benini, Marian Verhelst, Daniele Jahier Pagliari, Alessio Burrello |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2025 | Optimizing DNN Inference on Multi-accelerator SoCs at Training-TimeabstractThe demand for executing deep neural networks (DNNs) with low latency and minimal power consumption at the edge has led to the development of advanced heterogeneous systems-on-chips (SoCs) that incorporate multiple specialized computing units (CUs), such as accelerators. Offloading DNN computations to a specific CU from the available set often exposes accuracy vs efficiency tradeoffs, due to differences in their supported operations (e.g., standard versus depthwise convolution) or data representations (e.g., more/less aggressively quantized). A challenging yet unresolved issue is how to map a DNN onto these multi-CU systems to maximally exploit the parallelization possibilities while taking accuracy into account. To address this problem, we present ODiMO, a hardware-aware tool that efficiently explores fine-grain mapping of DNNs among various on-chip CUs, during the training phase. ODiMO strategically splits individual layers of the neural network and executes them in parallel on the multiple available CUs, aiming to balance the total inference energy consumption or latency with the resulting accuracy, impacted by the unique features of the different hardware units. We test our approach on CIFAR-10, CIFAR-100, and ImageNet, targeting two open-source heterogeneous SoCs, i.e., DIANA and Darkside. We obtain a rich collection of Pareto-optimal networks in the accuracy versus energy or latency space. We show that ODiMO reduces the latency of a DNN executed on the Darkside SoC by up to$8\times $at iso-accuracy, compared to manual heuristic mappings. When targeting energy, on the same SoC, ODiMO produced up to$50.8\times $more efficient mappings, with minimal accuracy drop$(\lt 0.3\%)$. Matteo Risso, Alessio Burrello, Daniele Jahier Pagliari |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Foundation Models for Structural Health MonitoringabstractStructural Health Monitoring (SHM) is a critical task for ensuring the safety and reliability of civil infrastructures, typically realized on bridges and viaducts by means of vibration monitoring. In this paper, we propose for the first time the use of Transformer neural networks, with a Masked Auto-Encoder architecture, asFoundation Modelsfor SHM. We demonstrate the ability of these models to learn generalizable representations from multiple large datasets through self-supervised pre-training, which, coupled with task-specific fine-tuning, allows them to outperform state-of-the-art traditional methods on diverse tasks, including Anomaly Detection (AD) and Traffic Load Estimation (TLE). We then extensively explore model size versus accuracy trade-offs and experiment with Knowledge Distillation (KD) to improve the performance of smaller Transformers, enabling their embedding directly into the SHM edge nodes. We showcase the effectiveness of our foundation models using data from three operational viaducts. For AD, we achieve a near-perfect 99.9% accuracy with a monitoring time span of just 15 windows. In contrast, a state-of-the-art method based on Principal Component Analysis (PCA) obtains its first good result (95.03% accuracy), only considering 120 windows. On two different TLE tasks, our models obtain state-of-the-art performance on multiple evaluation metrics (R2score, MAE% and MSE%). On the first benchmark, we achieve an R2score of 0.97 and 0.90 for light and heavy vehicle traffic, respectively, while the best previous approach (a Random Forest) stops at 0.91 and 0.84. On the second one, we achieve an R2score of 0.54 versus the 0.51 of the best competitor method, a Long-Short Term Memory network. Luca Benfenati, Daniele Jahier Pagliari, Luca Zanatta, Yhorman Alexander Bedoya Velez, Andrea Acquaviva, Massimo Poncino, Enrico Macii, Luca Benini, Alessio Burrello |
IEEE Trans. Sustain. Comput. | 2 |
| 2024 | Model-Driven Feature Engineering for Data-Driven Battery SOH ModelabstractAccurate State of Health (SoH) estimation is indispensable for ensuring battery system safety, reliability, and run-time monitoring. However, as instantaneous runtime measurement of SoH remains impractical when not unfeasible, appropriate models are required for its estimation. Recently, various data-driven models have been proposed, which solve various weaknesses of traditional models. However, the accuracy of data-driven models heavily depends on the quality of the training datasets, which usually contain data that are easy to measure but that are only partially or weakly related to the physical/chemical mechanisms that determine battery aging. In this study, we propose a novel feature engineering approach, which involves augmenting the original dataset with purpose-designed features that better represent the aging phenomena. Our contribution does not consist of a new machine-learning model but rather in the addition of selected features to an existing model. This methodology consistently demonstrates enhanced accuracy across various machine-learning models and battery chemistries, yielding an approximate 25% SoH estimation accuracy improve-ment. Our work bridges a critical gap in battery research, offering a promising strategy to significantly enhance SoH estimation by optimizing feature selection. Khaled Alamin, Daniele Jahier Pagliari, Yukai Chen, Enrico Macii, Sara Vinco, Massimo Poncino |
DATE | 2 |
| 2024 | AMBEATion: Analog Mixed-Signal Back-End Design Automation with Machine Learning and Artificial Intelligence TechniquesabstractFor the competitiveness of the European economy, automation techniques in the design of complex electronic systems are a prerequisite for winning the global chip challenge. Specifically, while the physical design of digital Integrated Circuits (ICs) can be largely automated, the physical design of Analog-Mixed-Signal (AMS) ICs built with an analog-on-top flow, where digital subsystems are instantiated as Intellectual Property (IP) modules, is still carried out predominantly by hand, with a time-consuming methodology. The AMBEATion consortium, including global semiconductor and design automation companies as well as leading universities, aims to address this challenge by combining classic Electronic Design Automation (EDA) algorithms with novel Artificial Intelligence and Machine Learning (ML) techniques. Specifically, the scientific and technical result expected at the end of the project will be a new methodology, implemented in a framework of scripts for AMS placement, internally making use of state-of-the-art AI/ML models, and fully integrated with Industrial design flows. With this methodology, the AMBEATion consortium aims to reduce the design turnaround-time and, consequently, the silicon development costs of complex AMS ICs. Giulia Elena Aliffi, Joao Baixinho, Dalibor Barri, Francesco Daghero, Nicola Di Carolo, Gabriele Faraone, Michelangelo Grosso, Daniele Jahier Pagliari, Jiri Jakovenko, Vladimír Janícek, Dario Licastro, Vazgen Melikyan, Matteo Risso, Vittorio Romano, Eugenio Serianni, Martin Stastný, Patrik Vacula, Giorgia Vitanza |
DATE | 8 |
| 2024 | Adaptive Deep Learning for Efficient Visual Pose Estimation Aboard Ultra-Low-Power Nano-DronesabstractSub-10cm diameter nano-drones are gaining momentum thanks to their applicability in scenarios prevented to bigger flying drones, such as in narrow environments and close to humans. However, their tiny form factor also brings their major drawback: ultra-constrained memory and processors for the onboard execution of their perception pipelines. Therefore, lightweight deep learning-based approaches are becoming in-creasingly popular, stressing how computational efficiency and energy-saving are paramount as they can make the difference between a fully working closed-loop system and a failing one. In this work, to maximize the exploitation of the ultra-limited resources aboard nano-drones, we present a novel adaptive deep learning-based mechanism for the efficient execution of a vision-based human pose estimation task. We leverage two State-of-the-Art (SoA) convolutional neural networks (CNNs) with different regression performance vs. computational costs trade-offs. By combining these CNNs with three novel adaptation strategies based on the output's temporal consistency and on auxiliary tasks to swap the CNN being executed proactively, we present six different systems. On a real-world dataset and the actual nano-drone hardware, our best-performing system, compared to executing only the bigger and most accurate SoA model, shows 28% latency reduction while keeping the same mean absolute error (MAE), 3% MAE reduction while being iso-latency, and the absolute peak performance, i.e., 6% better than SoA model. Beatrice Alessandra Motetti, Luca Crupi, Mustafa Omer Mohammed Elamin Elshaigi, Matteo Risso, Daniele Jahier Pagliari, Daniele Palossi, Alessio Burrello |
DATE | 5 |
| 2024 | HW-SW Optimization of DNNs for Privacy-Preserving People Counting on Low-Resolution Infrared ArraysabstractLow-resolution infrared (IR) array sensors enable people counting applications such as monitoring the occupancy of spaces and people flows while preserving privacy and minimizing energy consumption. Deep Neural Networks (DNNs) have been shown to be well-suited to process these sensor data in an accurate and efficient manner. Nevertheless, the space of DNNs' archi-tectures is huge and its manual exploration is burdensome and often leads to sub-optimal solutions. To overcome this problem, in this work, we propose a highly automated full-stack optimization flow for DNNs that goes from neural architecture search, mixed-precision quantization, and post-processing, down to the realization of a new smart sensor prototype, including a Microcontroller with a customized instruction set. Integrating these cross-layer optimizations, we obtain a large set of Pareto-optimal solutions in the 3D-space of energy, memory, and accuracy. Deploying such solutions on our hardware platform, we improve the state-of-the-art achieving up to 4.2 x model size reduction, 23.8 x code size reduction, and 15.38 x energy reduction at iso-accuracy. Matteo Risso, Francesco Daghero, Alessio Burrello, Seyedmorteza Mollaei, Marco Castellano, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
DATE | 9 |
| 2024 | Deep Recommender Models Inference: Automatic Asymmetric Data Flow OptimizationabstractDeep Recommender Models (DLRMs) inference is a fundamental AI workload accounting for more than 79% of the total AI workload in Meta's data centers. DLRMs' performance bottleneck is found in the embedding layers, which perform many random memory accesses to retrieve small embedding vectors from tables of various sizes. We propose the design of tailored data flows to speedup embedding look-ups. Namely, we propose four strategies to look up an embedding table effectively on one core, and a framework to automatically map the tables asymmetrically to the multiple cores of a SoC. We assess the effectiveness of our method using the Huawei Ascend AI accelerators, comparing it with the default Ascend compiler, and we perform high-level comparisons with Nvidia A100. Results show a speed-up varying from 1.5x up to 6.5x for real workload distributions, and more than 20x for extremely unbalanced distributions. Furthermore, the method proves to be much more independent of the query distribution than the baseline. Giuseppe Ruggeri, Renzo Andri, Daniele Jahier Pagliari, Lukas Cavigelli |
ICCD | 3 |
| 2024 | Dynamic Decision Tree Ensembles for Energy-Efficient Inference on IoT Edge NodesabstractWith the increasing popularity of Internet of Things (IoT) devices, there is a growing need for energy-efficient machine learning (ML) models that can run on constrained edge nodes. Decision tree ensembles, such as random forests (RFs) and gradient boosting (GBTs), are particularly suited for this task, given their relatively low complexity compared to other alternatives. However, their inference time and energy costs are still significant for edge hardware. Given that said costs grow linearly with the ensemble size, this article proposes the use of dynamic ensembles, that adjust the number of executed trees based both on a latency/energy target and on the complexity of the processed input, to tradeoff computational cost and accuracy. We focus on deploying these algorithms on multicore low-power IoT devices, designing a tool that automatically converts a Python ensemble into optimized C code, and exploring several optimizations that account for the available parallelism and memory hierarchy. We extensively benchmark both static and dynamic RFs and GBTs on three state-of-the-art IoT-relevant data sets, using an 8-core ultralow-power System-on-Chip (SoC), GAP8, as the target platform. Thanks to the proposed early stopping mechanisms, we achieve an energy reduction of up to 37.9% with respect to static GBTs (8.82 uJ versus 14.20 uJ per inference) and 41.7% with respect to static RFs (2.86 uJ versus 4.90 uJ per inference), without losing accuracy compared to the static model. Francesco Daghero, Alessio Burrello, Enrico Macii, Paolo Montuschi, Massimo Poncino, Daniele Jahier Pagliari |
IEEE Internet Things J. | 6 |
| 2023 | HTVM: Efficient Neural Network Deployment On Heterogeneous TinyML PlatformsabstractOptimal deployment of deep neural networks (DNNs) on state-of-the-art Systems-on-Chips (SoCs) is crucial for tiny machine learning (TinyML) at the edge. The complexity of these SoCs makes deployment non-trivial, as they typically contain multiple heterogeneous compute cores with limited, programmer-managed memory to optimize latency and energy efficiency. We propose HTVM – a compiler that merges TVM with DORY to maximize the utilization of heterogeneous accelerators and minimize data movements. HTVM allows deploying the MLPerf™ Tiny suite on DIANA, an SoC with a RISC-V CPU, and digital and analog compute-in-memory AI accelerators, at 120x improved performance over plain TVM deployment. Josse Van Delm, Maarten Vandersteegen, Alessio Burrello, Giuseppe Maria Sarda, Francesco Conti 0001, Daniele Jahier Pagliari, Luca Benini, Marian Verhelst |
DAC | 6 |
| 2023 | Energy-efficient Wearable-to-Mobile Offload of ML Inference for PPG-based Heart-Rate EstimationabstractModern smartwatches often include photoplethysmographic (PPG) sensors to measure heartbeats or blood pressure through complex algorithms that fuse PPG data with other signals. In this work, we propose a collaborative inference approach that uses both a smartwatch and a connected smartphone to maximize the performance of heart rate (HR) tracking while also maximizing the smartwatch's battery life. In particular, we first analyze the trade-offs between running on-device HR tracking or offloading the work to the mobile. Then, thanks to an additional step to evaluate the difficulty of the upcoming HR prediction, we demonstrate that we can smartly manage the workload between smartwatch and smartphone, maintaining a low mean absolute error (MAE) while reducing energy consumption. We benchmark our approach on a custom smartwatch prototype, including the STM32WB55 MCU and Bluetooth Low-Energy (BLE) communication, and a Raspberry Pi3 as a proxy for the smartphone. With our Collaborative Heart Rate Inference System (CHRIS), we obtain a set of Pareto-optimal configurations demonstrating the same MAE as State-of-Art (SoA) algorithms while consuming less energy. For instance, we can achieve approximately the same MAE of TimePPG-Small [1] (5.54 BPM MAE vs. 5.60 BPM MAE) while reducing the energy by 2.03×, with a configuration that offloads 80% of the predictions to the phone. Furthermore, accepting a performance degradation to 7.16 BPM of MAE, we can achieve an energy consumption of 179 uJ per prediction, 3.03× less than running TimePPG-Small on the smartwatch, and 1.82× less than streaming all the input data to the phone. Alessio Burrello, Matteo Risso, Noemi Tomasello, Yukai Chen, Luca Benini, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
DATE | 8 |
| 2023 | PLiNIO: A User-Friendly Library of Gradient-Based Methods for Complexity-Aware DNN OptimizationabstractAccurate yet efficient Deep Neural Networks (DNNs) are in high demand, especially for applications that require their execution on constrained edge devices. Finding such DNNs in a reasonable time for new applications requires automated optimization pipelines since the huge space of hyper-parameter combinations is impossible to explore extensively by hand. In this work, we propose PLiNIO, an open-source library implementing a comprehensive set of state-of-the-art DNN design automation techniques, all based on lightweight gradient-based optimization, under a unified and user-friendly interface. With experiments on several edge-relevant tasks, we show that combining the various optimizations available in PLiNIO leads to rich sets of solutions that Pareto-dominate the considered baselines in terms of accuracy vs model size. Noteworthy, PLiNIO achieves up to 94.34% memory reduction for a <1% accuracy drop compared to a baseline architecture. Daniele Jahier Pagliari, Matteo Risso, Beatrice Alessandra Motetti, Alessio Burrello |
FDL | 1 |
| 2023 | Deep Neural Network Architecture Search for Accurate Visual Pose Estimation aboard Nano-UAVsabstractMiniaturized autonomous unmanned aerial vehicles (UAVs) are an emerging and trending topic. With their form factor as big as the palm of one hand, they can reach spots otherwise inaccessible to bigger robots and safely operate in human surroundings. The simple electronics aboard such robots (sub-100 mW) make them particularly cheap and attractive but pose significant challenges in enabling onboard sophisticated intelligence. In this work, we leverage a novel neural architecture search (NAS) technique to automatically identify several Pareto-optimal convolutional neural networks (CNNs) for a visual pose estimation task. Our work demonstrates how reallife and field-tested robotics applications can concretely leverage NAS technologies to automatically and efficiently optimize CNNs for the specific hardware constraints of small UAVs. We deploy several NAS-optimized CNNs and run them in closed-loop aboard a 27-g Crazyflie nano-UAV equipped with a parallel ultra-low power System-on-Chip. Our results improve the State-of-the-Art by reducing the in-field control error of 32% while achieving a real-time onboard inference-rate of ~10Hz@10mW and ~50Hz@90mW. Elia Cereda, Luca Crupi, Matteo Risso, Alessio Burrello, Luca Benini, Alessandro Giusti, Daniele Jahier Pagliari, Daniele Palossi |
ICRA | 7 |
| 2023 | Model-Driven Dataset Generation for Data-Driven Battery SOH ModelsabstractEstimating the State of Health (SOH) of batteries is crucial for ensuring the reliable operation of battery systems. Since there is no practical way to instantaneously measure it at run time, a model is required for its estimation. Recently, several data-driven SOH models have been proposed, whose accuracy heavily relies on the quality of the datasets used for their training. Since these datasets are obtained from measurements, they are limited in the variety of the charge/discharge profiles. To address this scarcity issue, we propose generating datasets by simulating a traditional battery model (e.g., a circuit-equivalent one). The primary advantage of this approach is the ability to use a simulatable battery model to evaluate a potentially infinite number of workload profiles for training the data-driven model. Furthermore, this general concept can be applied using any simulatable battery model, providing a fine spectrum of accuracy/complexity tradeoffs. Our results indicate that using simulated data achieves reasonable accuracy in SOH estimation, with a 7.2 % error relative to the simulated model, in exchange for a 27X memory reduction and a$\approx 2000\mathrm{X}$speedup. Khaled Alamin, Francesco Daghero, Giovanni Pollo, Daniele Jahier Pagliari, Yukai Chen, Enrico Macii, Massimo Poncino, Sara Vinco |
ISLPED | 4 |
| 2023 | Precision-aware Latency and Energy Balancing on Multi-Accelerator Platforms for DNN InferenceabstractThe need to execute Deep Neural Networks (DNNs) at low latency and low power at the edge has spurred the development of new heterogeneous Systems-on-Chips (SoCs) encapsulating a diverse set of hardware accelerators. How to optimally map a DNN onto such multi-accelerator systems is an open problem. We propose ODiMO, a hardware-aware tool that performs a fine-grain mapping across different accelerators on-chip, splitting individual layers and executing them in parallel, to reduce inference energy consumption or latency, while taking into account each accelerator's quantization precision to maintain accuracy. Pareto-optimal networks in the accuracy vs. energy or latency space are pursued for three popular dataset/DNN pairs, and deployed on the DIANA heterogeneous ultra-low power edge AI SoC. We show that ODiMO reduces energy/latency by up to 33%/31% with limited accuracy drop (−0.53%/-0.32%) compared to manual heuristic mappings. Matteo Risso, Alessio Burrello, Giuseppe Maria Sarda, Luca Benini, Enrico Macii, Massimo Poncino, Marian Verhelst, Daniele Jahier Pagliari |
ISLPED | 8 |
| 2023 | Efficient Deep Learning Models for Privacy-Preserving People Counting on Low-Resolution Infrared ArraysabstractUltralow-resolution infrared (IR) array sensors offer a low cost, energy efficient, and privacy-preserving solution for people counting, with applications, such as occupancy monitoring and visitor flow analysis in private and public spaces. Previous work has shown that deep learning (DL) can yield superior performance on this task. However, the literature was missing an extensive comparative analysis of various efficient DL architectures for IR array-based people counting, that considers not only their accuracy but also the cost of deploying them on memory- and energy-constrained Internet of Things (IoT) edge nodes. Such analysis is key for system designers, since it helps them select the most appropriate DL model given the constraints of their target hardware. In this work, we address this need by comparing six different DL architectures on a novel data set composed of IR images collected from a commercial$8\times8$array, which we made openly available. With a wide architectural exploration of each model type, we obtain a rich set of Pareto-optimal solutions, spanning cross-validated balanced accuracy scores in the 55.70%–82.70% range. When deployed on a commercial microcontroller (MCU) by STMicroelectronics, the STM32L4A6ZG, these models occupy 0.41–9.28kB of memory, and require 1.10–7.74 ms per inference, while consuming 17.18–$120.43 \mu \text{J}$of energy. Our models are significantly more accurate than a previous deterministic method (up to +39.9%), while being up to$3.53\times $faster and more energy efficient. So, our work serves also as a demonstration that DL can not only achieve higher accuracy but also higher efficiency compared to classic algorithms for this type of task. Further, our models’ accuracy is comparable to state-of-the-art DL solutions on similar resolution sensors, despite a much lower complexity. All our models enable continuous, real-time inference on an MCU-based IoT node, with years of autonomous operation without battery recharging. Francesco Daghero, Yukai Chen, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
IEEE Internet Things J. | 9 |
| 2023 | Lightweight Neural Architecture Search for Temporal Convolutional Networks at the EdgeabstractNeural Architecture Search (NAS) is quickly becoming the go-to approach to optimize the structure of Deep Learning (DL) models for complex tasks such as Image Classification or Object Detection. However, many other relevant applications of DL, especially at the edge, are based on time-series processing and require models with unique features, for which NAS is less explored. This work focuses in particular on Temporal Convolutional Networks (TCNs), a convolutional model for time-series processing that has recently emerged as a promising alternative to more complex recurrent architectures. We propose the first NAS tool that explicitly targets the optimization of the most peculiar architectural parameters of TCNs, namely dilation, receptive-field and number of features in each layer. The proposed approach searches for networks that offer good trade-offs between accuracy and number of parameters/operations, enabling an efficient deployment on embedded platforms. Moreover, its fundamental feature is that of being lightweight in terms of search complexity, making it usable even with limited hardware resources. We test the proposed NAS on four real-world, edge-relevant tasks, involving audio and bio-signals: (i) PPG-based Heart-Rate Monitoring, (ii) ECG-based Arrythmia Detection, (iii) sEMG-based Hand-Gesture Recognition, and (iv) Keyword Spotting.Results show that, starting from a single seed network, our method is capable of obtaining a rich collection of Pareto optimal architectures, among which we obtain models with the same accuracy as the seed, and 15.9-152× fewer parameters. Moreover, the NAS finds solutions that Pareto-dominate state-of-the-arthand-tuned models for 3 out of the 4 benchmarks, and are Pareto-optimal on the fourth (sEMG). Compared to three state-of-the-art NAS tools, ProxylessNAS, MorphNet and FBNetV2, our method explores a larger search space for TCNs (up to 1012×) and obtains superior solutions, while requiring low GPU memory and search time. We deploy our NAS outputs on two distinct edge devices, the multicore GreenWaves Technology GAP8 IoT processor and the single-core STMicroelectronics STM32H7 microcontroller. With respect to the state-of-the-art hand-tuned models, we reduce latency and energy of up to 5.5× and 3.8× on the two targets respectively, without any accuracy loss. Matteo Risso, Alessio Burrello, Francesco Conti 0001, Lorenzo Lamberti, Yukai Chen, Luca Benini, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
IEEE Trans. Computers | 9 |
| 2022 | Bioformers: Embedding Transformers for Ultra-Low Power sEMG-based Gesture RecognitionabstractHuman-machine interaction is gaining traction in rehabilitation tasks, such as controlling prosthetic hands or robotic arms. Gesture recognition exploiting surface electromyographic (sEMG) signals is one of the most promising approaches, given that sEMG signal acquisition is non-invasive and is directly related to muscle contraction. However, the analysis of these signals still presents many challenges since similar gestures result in similar muscle contractions. Thus the resulting signal shapes are almost identical, leading to low classification accuracy. To tackle this challenge, complex neural networks are employed, which require large memory footprints, consume relatively high energy and limit the maximum battery life of devices used for classification. This work addresses this problem with the introduction of the Bioformers. This new family of ultra-small attention-based architectures approaches state-of-the-art performance while reducing the number of parameters and operations of 4.9 ×. Additionally, by introducing a new inter-subjects pre-training, we improve the accuracy of our best Bioformer by 3.39 %, matching state-of-the-art accuracy without any additional inference cost. Deploying our best performing Bioformer on a Parallel, Ultra-Low Power (PULP) microcontroller unit (MCU), the GreenWaves GAP8, we achieve an inference latency and energy of 2.72 ms and 0.14 mJ, respectively, 8.0× lower than the previous state-of-the-art neural network, while occupying just 94.2 kB of memory. Alessio Burrello, Francesco Bianco Morghet, Moritz Scherer 0001, Simone Benatti, Luca Benini, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
DATE | 8 |
| 2022 | C-NMT: A Collaborative Inference Framework for Neural Machine TranslationabstractCollaborative Inference (CI) optimizes the latency and energy consumption of deep learning inference through the inter-operation of edge and cloud devices. Albeit beneficial for other tasks, CI has never been applied to the sequence-to-sequence mapping problem at the heart of Neural Machine Translation (NMT). In this work, we address the specific issues of collaborative NMT, such as estimating the latency required to generate the (unknown) output sequence, and show how existing CI methods can be adapted to these applications. Our experiments show that CI can reduce the latency of NMT by up to 44% compared to a non-collaborative approach. Yukai Chen, Roberta Chiaro, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
ISCAS | 5 |
| 2022 | Privacy-preserving Social Distance Monitoring on Microcontrollers with Low-Resolution Infrared Sensors and CNNsabstractLow-resolution infrared (IR) array sensors offer a low-cost, low-power, and privacy-preserving alternative to optical cameras and smartphones/wearables for social distance monitoring in indoor spaces, permitting the recognition of basic shapes, without revealing the personal details of individuals. In this work, we demonstrate that an accurate detection of social distance violations can be achieved processing the raw output of a 8x8 IR array sensor with a small-sized Convolutional Neural Network (CNN). Furthermore, the CNN can be executed directly on a Microcontroller (MCU)-based sensor node.With results on a newly collected open dataset, we show that our best CNN achieves 86.3% balanced accuracy, significantly outperforming the 61% achieved by a state-of-the-art deterministic algorithm. Changing the architectural parameters of the CNN, we obtain a rich Pareto set of models, spanning 70.5-86.3% accuracy and 0.18-75k parameters. Deployed on a STM32L476RGMCU, these models have a latency of 0.73-5.33ms, with an energy consumption per inference of 9.38-68.57$\mu$J. Francesco Daghero, Yukai Chen, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
ISCAS | 9 |
| 2022 | Multi-Complexity-Loss DNAS for Energy-Efficient and Memory-Constrained Deep Neural NetworksabstractNeural Architecture Search (NAS) is increasingly popular to automatically explore the accuracy versus computational complexity trade-off of Deep Learning (DL) architectures. When targeting tiny edge devices, the main challenge for DL deployment is matching the tight memory constraints, hence most NAS algorithms consider model size as the complexity metric. Other methods reduce the energy or latency of DL models by trading off accuracy and number of inference operations. Energy and memory are rarely considered simultaneously, in particular by low-search-cost Differentiable NAS (DNAS) solutions. Matteo Risso, Alessio Burrello, Luca Benini, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
ISLPED | 6 |
| 2022 | Embedding Temporal Convolutional Networks for Energy-efficient PPG-based Heart Rate MonitoringabstractPhotoplethysmography (PPG) sensors allow for non-invasive and comfortable heart rate (HR) monitoring, suitable for compact wrist-worn devices. Unfortunately, motion artifacts (MAs) severely impact the monitoring accuracy, causing high variability in the skin-to-sensor interface. Several data fusion techniques have been introduced to cope with this problem, based on combining PPG signals with inertial sensor data. Until now, both commercial and reasearch solutions are computationally efficient but not very robust, or strongly dependent on hand-tuned parameters, which leads to poor generalization performance. In this work, we tackle these limitations by proposing a computationally lightweight yet robust deep learning-based approach for PPG-based HR estimation. Specifically, we derive a diverse set of Temporal Convolutional Networks for HR estimation, leveraging Neural Architecture Search. Moreover, we also introduce ActPPG, an adaptive algorithm that selects among multiple HR estimators depending on the amount of MAs, to improve energy efficiency. We validate our approaches on two benchmark datasets, achieving as low as 3.84 beats per minute of Mean Absolute Error on PPG-Dalia, which outperforms the previous state of the art. Moreover, we deploy our models on a low-power commercial microcontroller (STM32L4), obtaining a rich set of Pareto optimal solutions in the complexity vs. accuracy space. Alessio Burrello, Daniele Jahier Pagliari, Pierangelo Maria Rapa, Matilde Semilia, Matteo Risso, Tommaso Polonelli, Massimo Poncino, Luca Benini, Simone Benatti |
ACM Trans. Comput. Heal. | 2 |
| 2022 | A deep learning approach for Spatio-Temporal forecasting of new cases and new hospital admissions of COVID-19 spread in Reggio Emilia, Northern Italy
Veronica Sciannameo, Alessia Goffi, Giuseppe Maffeis, Roberta Gianfreda, Daniele Jahier Pagliari, Tommaso Filippini, Pamela Mancuso, Paolo Giorgi-Rossi, Leonardo Alberto Dal Zovo, Angela Corbari, Marco Vinceti, Paola Berchialla |
J. Biomed. Informatics | 5 |
| 2022 | Human Activity Recognition on Microcontrollers with Quantized and Adaptive Deep Neural NetworksabstractHuman Activity Recognition (HAR) based on inertial data is an increasingly diffused task on embedded devices, from smartphones to ultra low-power sensors. Due to the high computational complexity of deep learning models, most embedded HAR systems are based on simple and not-so-accurate classic machine learning algorithms. This work bridges the gap between on-device HAR and deep learning, proposing a set of efficient one-dimensional Convolutional Neural Networks (CNNs) that can be deployed on general purpose microcontrollers (MCUs). Our CNNs are obtained combining hyper-parameters optimization with sub-byte and mixed-precision quantization, to find good trade-offs between classification results and memory occupation. Moreover, we also leverage adaptive inference as an orthogonal optimization to tune the inference complexity at runtime based on the processed input, hence producing a more flexible HAR system. With experiments on four datasets, and targeting an ultra-low-power RISC-V MCU, we show that (i) we are able to obtain a rich set of Pareto-optimal CNNs for HAR, spanning more than 1 order of magnitude in terms of memory, latency, and energy consumption; (ii) thanks to adaptive inference, we can derive >20 runtime operating modes starting from a single CNN, differing by up to 10% in classification scores and by more than 3× in inference complexity, with a limited memory overhead; (iii) on three of the four benchmarks, we outperform all previous deep learning methods, while reducing the memory occupation by more than 100×. The few methods that obtain better performance (both shallow and deep) are not compatible with MCU deployment; (iv) all our CNNs are compatible with real-time on-device HAR, achieving an inference latency that ranges between 9 μs and 16 ms. Their memory occupation varies in 0.05–23.17 kB, and their energy consumption in 0.05 and 61.59 μJ, allowing years of continuous operation on a small battery supply. Francesco Daghero, Alessio Burrello, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
ACM Trans. Embed. Comput. Syst. | 9 |
| 2021 | Ultra-compact binary neural networks for human activity recognition on RISC-V processorsabstractHuman Activity Recognition (HAR) is a relevant inference task in many mobile applications. State-of-the-art HAR at the edge is typically achieved with lightweight machine learning models such as decision trees and Random Forests (RFs), whereas deep learning is less common due to its high computational complexity. In this work, we propose a novel implementation of HAR based on deep neural networks, and precisely on Binary Neural Networks (BNNs), targeting low-power general purpose processors with a RISC-V instruction set. BNNs yield very small memory footprints and low inference complexity, thanks to the replacement of arithmetic operations with bit-wise ones. However, existing BNN implementations on general purpose processors impose constraints tailored to complex computer vision tasks, which result in over-parametrized models for simpler problems like HAR. Therefore, we also introduce a new BNN inference library, which targets ultra-compact models explicitly. With experiments on a single-core RISC-V processor, we show that BNNs trained on two HAR datasets obtain higher classification accuracy compared to a state-of-the-art baseline based on RFs. Furthermore, our BNN reaches the same accuracy of a RF with either less memory (up to 91%) or more energy-efficiency (up to 70%), depending on the complexity of the features extracted by the RF. Francesco Daghero, Daniele Jahier Pagliari, Alessio Burrello, Marco Castellano, Luca Gandolfi, Andrea Calimera, Enrico Macii, Massimo Poncino |
CF | 3 |
| 2021 | Pruning In Time (PIT): A Lightweight Network Architecture Optimizer for Temporal Convolutional NetworksabstractTemporal Convolutional Networks (TCNs) are promising Deep Learning models for time-series processing tasks. One key feature of TCNs is time-dilated convolution, whose optimization requires extensive experimentation. We propose an automatic dilation optimizer, which tackles the problem as a weight pruning on the time-axis, and learns dilation factors together with weights, in a single training. Our method reduces the model size and inference latency on a real SoC hardware target by up to 7.4× and 3×, respectively with no accuracy drop compared to a network without dilation. It also yields a rich set of Pareto-optimal TCNs starting from a single model, outperforming hand-designed solutions in both size and accuracy. Matteo Risso, Alessio Burrello, Daniele Jahier Pagliari, Francesco Conti 0001, Lorenzo Lamberti, Enrico Macii, Luca Benini, Massimo Poncino |
DAC | 3 |
| 2021 | Robust and Energy-Efficient PPG-Based Heart-Rate MonitoringabstractA wrist-worn PPG sensor coupled with a lightweight algorithm can run on a MCU to enable non-invasive and comfortable monitoring, but ensuring robust PPG-based heart-rate monitoring in the presence of motion artifacts is still an open challenge. Recent state-of-the-art algorithms combine PPG and inertial signals to mitigate the effect of motion artifacts. However, these approaches suffer from limited generality. Moreover, their deployment on MCU-based edge nodes has not been investigated. In this work, we tackle both the aforementioned problems by proposing the use of hardware-friendly Temporal Convolutional Networks (TCN) for PPG-based heart estimation. Starting from a single "seed" TCN, we leverage an automatic Neural Architecture Search (NAS) approach to derive a rich family of models. Among them, we obtain a TCN that outperforms the previous state-of-the- art on the largest PPG dataset available (PPGDalia), achieving a Mean Absolute Error (MAE) of just 3.84 Beats Per Minute (BPM). Furthermore, we tested also a set of smaller yet still accurate (MAE of 5.64 - 6.29 BPM) networks that can be deployed on a commercial MCU (STM32L4) which require as few as 5k parameters and reach a latency of 17.1 ms consuming just 0.21 mJ per inference. Matteo Risso, Alessio Burrello, Daniele Jahier Pagliari, Simone Benatti, Enrico Macii, Luca Benini, Massimo Poncino |
ISCAS | 3 |
| 2021 | TCN Mapping Optimization for Ultra-Low Power Time-Series Edge InferenceabstractTemporal Convolutional Networks (TCNs) are emerging lightweight Deep Learning models for Time Series analysis. We introduce an automated exploration approach and a library of optimized kernels to map TCNs on Parallel Ultra-Low Power (PULP) microcontrollers. Our approach minimizes latency and energy by exploiting a layer tiling optimizer to jointly find the tiling dimensions and select among alternative implementations of the causal and dilated 1D-convolution operations at the core of TCNs. We benchmark our approach on a commercial PULP device, achieving up to $103 \times $ lower latency and $20.3 \times $ lower energy than the Cube-AI toolkit executed on the STM32L4 and from $2.9 \times $ to $26.6 \times $ lower energy compared to commercial closed-source and academic open-source approaches on the same hardware target. Alessio Burrello, Alberto Dequino, Daniele Jahier Pagliari, Francesco Conti 0001, Marcello Zanghieri, Enrico Macii, Luca Benini, Massimo Poncino |
ISLPED | 3 |
| 2021 | ACME: An Energy-Efficient Approximate Bus Encoding for I2CabstractIn ultra low power systems with many peripherals, off-chip serial interconnects contribute significantly to the total energy budget. Leveraging the error-resilience characteristics of many embedded applications, the approximate computing paradigm has been applied to serial bus encodings to reduce interconnect consumption. However, the power model considered in previous works was purely capacitive. Accordingly, the objective of these approximate encodings was simply to reduce the transition count. While this works well for most bus standards, one notable exception is represented by I2C, whose open-drain physical connection makes the static energy consumed by logic-0 values on the bus extremely relevant. In this work, we propose ACME, the first approximate serial bus encoding targeting specifically I2C connections. With a simple encoding/decoding scheme, ACME concurrently reduces both the static and dynamic energy on the bus by maximizing the number of logic-1 values in codewords, while simultaneously reducing transitions. Using an accurate bus model and realistic capacitance and resistance values selected according to the I2C standard, we show that our encoding outperforms state-of-the-art solutions and reduces the total energy consumption on the bus by 57% on average, with an error smaller than 0.1%. Daniele Jahier Pagliari, Andrea Calimera, Enrico Macii, Massimo Poncino |
ISLPED | 2 |
| 2021 | Adaptive Random Forests for Energy-Efficient Inference on MicrocontrollersabstractRandom Forests (RFs) are widely used Machine Learning models in low-power embedded devices, due to their hardware friendly operation and high accuracy on practically relevant tasks. The accuracy of a RF often increases with the number of internal weak learners (decision trees), but at the cost of a proportional increase in inference latency and energy consumption. Such costs can be mitigated considering that, in most applications, inputs are not all equally difficult to classify. Therefore, a large RF is often necessary only for (few) hard inputs, and wasteful for easier ones. In this work, we propose an early-stopping mechanism for RFs, which terminates the inference as soon as a high-enough classification confidence is reached, reducing the number of weak learners executed for easy inputs. The early-stopping confidence threshold can be controlled at runtime, in order to favor either energy saving or accuracy. We apply our method to three different embedded classification tasks, on a single-core RISC-V microcontroller, achieving an energy reduction from 38% to more than 90% with a drop of less than 0.5% in accuracy. We also show that our approach outperforms previous adaptive ML methods for RFs. Francesco Daghero, Alessio Burrello, Luca Benini, Andrea Calimera, Enrico Macii, Massimo Poncino, Daniele Jahier Pagliari |
VLSI-SoC | 8 |
| 2021 | Manufacturing as a Data-Driven Practice: Methodologies, Technologies, and ToolsabstractIn recent years, the introduction and exploitation of innovative information technologies in industrial contexts have led to the continuous growth of digital shop floor environments. The new Industry 4.0 model allows smart factories to become very advanced IT industries, generating an ever-increasing amount of valuable data. As a consequence, the necessity of powerful and reliable software architectures is becoming prominent along with data-driven methodologies to extract useful and hidden knowledge supporting the decision-making process. This article discusses the latest software technologies needed to collect, manage, and elaborate all data generated through innovative Internet-of-Things (IoT) architectures deployed over the production line, with the aim of extracting useful knowledge for the orchestration of high-level control services that can generate added business value. This survey covers the entire data life cycle in manufacturing environments, discussing key functional and methodological aspects along with a rich and properly classified set of technologies and tools, useful to add intelligence to data-driven services. Therefore, it serves both as a first guided step toward the rich landscape of the literature for readers approaching this field and as a global yet detailed overview of the current state of the art in the Industry 4.0 domain for experts. As a case study, we discuss, in detail, the deployment of the proposed solutions for two research project demonstrators, showing their ability to mitigate manufacturing line interruptions and reduce the corresponding impacts and costs. Tania Cerquitelli, Daniele Jahier Pagliari, Andrea Calimera, Lorenzo Bottaccioli, Edoardo Patti, Andrea Acquaviva, Massimo Poncino |
Proc. IEEE | 2 |
| 2021 | CRIME: Input-Dependent Collaborative Inference for Recurrent Neural NetworksabstractThe excellent accuracy of Recurrent Neural Networks (RNNs) for time-series and natural language processing comes at the cost of computational complexity. Therefore, the choice between edge and cloud computing for RNN inference, with the goal of minimizing response time or energy consumption, is not trivial. An edge approach must deal with the aforementioned complexity, while a cloud solution pays large time and energy costs for data transmission. Collaborative inference is a technique that tries to obtain the best of both worlds, by splitting the inference task among a network of collaborating devices. While already investigated for other types of neural networks, collaborative inference for RNNs poses completely new challenges, such as the strong influence of input length on processing time and energy, and is greatly unexplored. In this paper, we introduce a Collaborative RNN Inference Mapping Engine(CRIME), which automatically selects the best inference device for each input. CRIME is flexible with respect to the connection topology among collaborating devices, and adapts to changes in the connections statuses and in the devices loads. With experiments on several RNNs and datasets, we show that CRIME can reduce the execution time (or end-node energy) by more than 25% compared to any single-device approach. Daniele Jahier Pagliari, Roberta Chiaro, Enrico Macii, Massimo Poncino |
IEEE Trans. Computers | 1 |
| 2021 | A Microservices-Based Framework for Smart Design and Optimization of PV InstallationsabstractThe design of photovoltaic (PV) installations mostly relies on rule-of-thumb criteria and on gross estimates of the shading patterns, and the few optimized approaches are generally focused on the problem of identifying the most suitable surfaces (e.g., roofs) in a larger geographic area (e.g., city or district). This article proposes a framework to address the design and the optimization of PV installations through a set of microservices focusing on the different variables of the design: identification of the target surfaces, elaboration of weather data, modeling of the PV panel, and floorplanning of the panel on the surface. The microservices architecture ensures extensibility and generality, as the user may execute only a subset of the proposed services or provide novel algorithms to extend the existing ones. Additionally, the framework provides a set of built-in models that allow sensitivity to the distribution of shades and accurate modeling of the power production over time. We show the many benefits of the proposed framework on two different use cases. Sara Vinco, Daniele Jahier Pagliari, Lorenzo Bottaccioli, Edoardo Patti, Enrico Macii, Massimo Poncino |
IEEE Trans. Sustain. Comput. | 2 |
| 2020 | Input-Dependent Edge-Cloud Mapping of Recurrent Neural Networks InferenceabstractGiven the computational complexity of Recurrent Neural Networks (RNNs) inference, IoT and mobile devices typically offload this task to the cloud. However, the execution time and energy consumption of RNN inference strongly depends on the length of the processed input. Therefore, considering also communication costs, it may be more convenient to process short input sequences locally and only offload long ones to the cloud. In this paper, we propose a low-overhead runtime tool that performs this choice automatically. Results based on real edge and cloud devices show that our method is able to simultaneously reduce the total execution time and energy consumption of the system compared to solutions that run RNN inference fully locally or fully in the cloud. Daniele Jahier Pagliari, Roberta Chiaro, Yukai Chen, Sara Vinco, Enrico Macii, Massimo Poncino |
DAC | 1 |
| 2020 | Modeling and Simulation of Cyber-Physical Electrical Energy Systems With SystemC-AMSabstractModern cyber-physical electrical energy systems (CPEES) are characterized by wider adoption of sustainable energy sources and by an increased attention to optimization, with the goal of reducing pollution and wastes. This imposes a need for instruments supporting the design flow, to simulate and validate the behavior of system components and to apply additional optimization and exploration steps. Additionally, each system might be tested with a number of management policies, to evaluate their economic impact. It is thus evident that simulation is a key ingredient in the design flow of CPEES. This paper proposes a framework for CPEES modeling and simulation, that relies on the open-source standard SystemC-AMS. The paper formalizes the information and energy flow in a generic CPEES, by focusing on both AC and DC components, and by including support for mechanical and physical models that represent multiple energy sources and loads. Experimental results, applied to a complex CPEES case study, will prove the effectiveness of the proposed solution, in terms of accuracy, speed up w.r.t. the current state-of-the-art Matlab/Simulink, and support for the design flow. Yukai Chen, Sara Vinco, Daniele Jahier Pagliari, Paolo Montuschi, Enrico Macii, Massimo Poncino |
IEEE Trans. Sustain. Comput. | 3 |
| 2019 | Low-Overhead Power Trace Obfuscation for Smart Meter PrivacyabstractSmart meters communicate to the utility provider fine-grain information about a user's energy consumption, which could be used to infer the user's habits and pose thus a critical privacy risk. State-of-the-art solutions try to obfuscate the readings of a meter either by using a large re-chargeable battery to filter the trace or by adding random noise to alter it. Both solutions, however, have significant drawbacks: large batteries are prohibitively expensive, whereas digitally added noise implies that the user entrusts the utility provider to protect his/her privacy. Daniele Jahier Pagliari, Sara Vinco, Enrico Macii, Massimo Poncino |
DAC | 1 |
| 2019 | Irradiance-Driven Partial Reconfiguration of PV PanelsabstractAdaptive reconfiguration of a photo-voltaic (PV) panel by means of a switch network is a well-known approach to tackle shading issues dynamically and with a reasonable cost. Most of these approaches assume however that the entire panel is reconfigurable, resulting in high installation costs due to the large wiring overhead required by this solution. In this work we propose an architecture in which only a portion of the panel is made reconfigurable, while minimizing the loss in the extracted power with respect to a fully reconfigurable solution. The key feature of our approach is the use of environmental (irradiance and temperature) data to determine the reconfigurable subset at design time. Simulation results show that, by reconfiguring only about 50-70% of a panel, it is possible to achieve up to 45% power increase with respect to a static topology, while losing less than 5% power with respect to full reconfiguration. Daniele Jahier Pagliari, Sara Vinco, Enrico Macii, Massimo Poncino |
DATE | 1 |
| 2019 | Dynamic Beam Width Tuning for Energy-Efficient Recurrent Neural NetworksabstractRecurrent Neural Networks (RNNs) are state-of-the-art models for many machine learning tasks, such as language modeling and machine translation. Executing the inference phase of a RNN directly in edge nodes, rather than in the cloud, would provide benefits in terms of energy consumption, latency and network bandwidth, provided that models can be made efficient enough to run on energy-constrained embedded devices. Daniele Jahier Pagliari, Francesco Panini, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 1 |
| 2019 | CNN-Based Camera-less User Attention Detection for Smartphone Power ManagementabstractThe many sensors hosted by mobile electronic devices are commonly used to recognize user activities and context, in order to provide new functionalities, such as tracking physical activity and sleep cycles. Despite its potential, such context recognition is only employed for power management purposes in very specific scenarios (e.g. in-pocket detection). In this work we present a novel context recognition system able to reliably identify whether a mobile device is not being looked at, and to consequently trigger power management actions such as turning off the display and moving to suspended mode. Our method takes as input the readings from common low-power sensors present in virtually all mobile devices and classifies them using a Convolutional Neural Network. Most importantly, the power-hungry camera sub-system is not used, resulting in an extremely energy-efficient detection strategy. Results show that our system is able to identify scenarios in which a device is not being used with 95.6% accuracy, thus reducing the energy overheads by 91% compared to a standard timeout-based power management and by 58% compared to a system relying on the camera. Daniele Jahier Pagliari, Matteo Ansaldi, Enrico Macii, Massimo Poncino |
ISLPED | 1 |
| 2019 | Fine-Grain Back Biasing for the Design of Energy-Quality Scalable OperatorsabstractEnergy-quality scalable systems are a promising solution to cope with the small energy budgets and high processing demands of mobile and Internet of Things applications. These systems leverage the error resilience of applications to obtain high energy efficiency, at the expense of tolerable reductions in the output quality. Hardware datapath operators able to reconfigure their precision and power consumption at runtime are key components of such systems. However, most implementations of these operators require manual, architecture-specific modifications and tend to have large power overheads compared to standard designs, when working at maximum precision. One promising design-independent alternative is dynamic voltage and accuracy scaling, whose adoption, however, is hindered by incompatibilities with standard design flows. In this paper, we propose a new methodology for the design of energy-quality scalable operators; our solution leverages runtime tuning of transistors threshold voltages to obtain a fine-grain control of the speed and power consumption of standard-cells within an operator. Thanks to the additional flexibility provided by this fine-grain knob, our method overcomes the main limitations of previous solutions, at the cost of a small area overhead. We demonstrate our approach on a 28 nm FDSOI technology; by exploiting the strong effect of back-gate biasing on threshold voltage, we achieve a power consumption reduction of more than 40% compared to the state-of-the-art, for the same precision. Daniele Jahier Pagliari, Yves Durand, David Coriat, Edith Beigné, Enrico Macii, Massimo Poncino |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | A Cross-level Verification Methodology for Digital IPs Augmented with Embedded Timing MonitorsabstractSmart systems are characterized by the integration in a single device of multi-domain subsystems of different technological domains, namely, analog, digital, discrete and power devices, MEMS, and power sources. Such challenges, emerging from the heterogeneous nature of the whole system, combined with the traditional challenges of digital design, directly impact on performance and on propagation delay of digital components. This article proposes a design approach to enhance the RTL model of a given digital component for the integration in smart systems with the automatic insertion of delay sensors, which can detect and correct timing failures. The article then proposes a methodology to verify such added features at system level. The augmented model is abstracted to SystemC TLM, which is automatically injected with mutants (i.e., code mutations) to emulate delays and timing failures. The resulting TLM model is finally simulated to identify timing failures and to verify the correctness of the inserted delay monitors. Experimental results demonstrate the applicability of the proposed design and verification methodology, thanks to an efficient sensor-aware abstraction methodology, by applying the flow to three complex case studies. Sara Vinco, Nicola Bombieri, Daniele Jahier Pagliari, Franco Fummi, Enrico Macii, Massimo Poncino |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2018 | All-digital embedded meters for on-line power estimationabstractModern low power designs use multiple knobs for concurrent dynamic and leakage power optimization; supply voltage and threshold voltage are the most adopted. An efficient control of these knobs needs management policies aware of the power breakdown. This implies the availability of smart on-chip strategies for dynamic and leakage power estimation at runtime. In this paper, we address this issue proposing the implementation of embedded dynamic/static power meters that use an optimized regression model fed with data collected from in-situ activity monitors. The number of sensors, their bitwidth and optimal placement are obtained through an automated design flow. The methodology works for general logic and applies not just to processor cores, but also to application-specific designs. We apply our solution to a representative class of benchmarks, showing that it can achieve an average estimation error smaller than 3%, with limited area and power overheads. Daniele Jahier Pagliari, Valentino Peluso, Yukai Chen, Andrea Calimera, Enrico Macii, Massimo Poncino |
DATE | 1 |
| 2018 | Battery-aware Design Exploration of Scheduling Policies for Multi-sensor DevicesabstractLifetime maximization is a key challenge in battery-powered multi-sensor devices. Battery-aware power management strategies combine task scheduling with dynamic voltage scaling (DVS), accounting for the fact that the power drawn by the device is different from that provided by the battery due to its many non-idealities. However, state-of-the-art techniques in this field do not take into account several important aspects, such as the impact of sensing tasks on the overall power demand, the (operating point dependent) losses due to multiple DC-DC conversions, and the dynamic modifications in battery efficiency caused by different distributions of the currents in the temporal and in the frequency domains. In this work, we propose a novel approach to identify optimal power management solutions, that addresses all these limitations. Specifically, using advanced battery and DC-DC converter models, we propose methods to explore the scheduling space both statically (at design time) and dynamically (at runtime), accounting not only for computation tasks, but also for communication and sensing. With this method, we show that the battery lifetime can be increased by as much as 23.36% if an optimal power management strategy is adopted. Yukai Chen, Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 2 |
| 2018 | Application-Driven Synthesis of Energy-Efficient Reconfigurable-Precision OperatorsabstractThe increasing performance demands in emerging Internet of Things applications clash with the low energy budgets of end-nodes. Therefore, hardware operators able to reconfigure their computational precision at runtime are increasingly employed in these devices, to obtain good-enough results at minimal energy costs. Among the many methods proposed to implement such operators, Dynamic Voltage and Accuracy Scaling (DVAS) is particularly promising, due to its broad applicability and low overheads. However, a straight-forward application of DVAS conflicts with the optimizations performed by classic EDA algorithms, and does not yield the expected results. In this paper, we propose a novel synthesis algorithm for reconfigurable-precision circuits, that allows to integrate DVAS in a standard implementation flow. Moreover, we show how this algorithm can exploit information about the application, namely on the frequency of usage of each precision, to further reduce the total energy consumption. Applying our method to the popular LeNet neural network for digit recognition, we are able to reduce the energy due to Multiply-And-Accumulate (MAC) operations by 25%, compared to a straight-forward application of DVAS. Daniele Jahier Pagliari, Massimo Poncino |
ISCAS | 1 |
| 2018 | Dynamic Bit-width Reconfiguration for Energy-Efficient Deep Learning HardwareabstractDeep learning models have reached state of the art performance in many machine learning tasks. Benefits in terms of energy, bandwidth, latency, etc., can be obtained by evaluating these models directly within Internet of Things end nodes, rather than in the cloud. This calls for implementations of deep learning tasks that can run in resource limited environments with low energy footprints. Research and industry have recently investigated these aspects, coming up with specialized hardware accelerators for low power deep learning. One effective technique adopted in these devices consists in reducing the bit-width of calculations, exploiting the error resilience of deep learning. However, bit-widths are tipically set statically for a given model, regardless of input data. Unless models are retrained, this solution invariably sacrifices accuracy for energy efficiency. Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino |
ISLPED | 1 |
| 2018 | LAPSE: Low-Overhead Adaptive Power Saving and Contrast Enhancement for OLEDsabstractOrganic Light Emitting Diode (OLED) display panels are becoming increasingly popular especially in mobile devices; one of the key characteristics of these panels is that their power consumption strongly depends on the displayed image. In this paper we propose LAPSE, a new methodology to concurrently reduce the energy consumed by an OLED display and enhance the contrast of the displayed image, that relies on image-specific pixel-by-pixel transformations. Unlike previous approaches, LAPSE focuses specifically on reducing the overheads required to implement the transformation at runtime. To this end, we propose a transformation that can be executed in real time, either in software, with low time overhead, or in a hardware accelerator with a small area and low energy budget. Despite the significant reduction in complexity, we obtain comparable results to those achieved with more complex approaches in terms of power saving and image quality. Moreover, our method allows to easily explore the full quality-versus-power tradeoff by acting on a few basic parameters; thus, it enables the runtime selection among multiple display quality settings, according to the status of the system. Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino |
IEEE Trans. Image Process. | 1 |
| 2017 | A methodology for the design of dynamic accuracy operators by runtime back biasabstractMobile and IoT applications must balance increasing processing demands with limited power and cost budgets. Approximate computing achieves this goal leveraging the error tolerance features common in many emerging applications to reduce power consumption. In particular, adequate (i.e., energy/quality-configurable) hardware operators are key components in an error tolerant system. Existing implementations of these operators require significant architectural modifications, hence they are often design-specific and tend to have large overheads compared to accurate units. In this paper, we propose a methodology to design adequate data-path operators in an automatic way, which uses threshold voltage scaling as a knob to dynamically control the power/accuracy tradeoff. The method overcomes the limitations of previous solutions based on supply voltage scaling, in that it introduces lower overheads and it allows fine-grain regulation of this tradeoff. We demonstrate our approach on a state-of-the-art 28nm FDSOI technology, exploiting the strong effect of back biasing on threshold voltage. Results show a power consumption reduction of as much as 39% compared to solutions based only on supply voltage scaling, at iso-accuracy. Daniele Jahier Pagliari, Yves Durand, David Coriat, Anca Mariana Molnos, Edith Beigné, Enrico Macii, Massimo Poncino |
DATE | 1 |
| 2017 | Accelerators for Breast Cancer DetectionabstractAlgorithms used in microwave imaging for breast cancer detection require hardware acceleration to speed up execution time and reduce power consumption. In this article, we present the hardware implementation of two accelerators for two alternative imaging algorithms that we obtain entirely from SystemC specifications via high-level synthesis. The two algorithms present opposite characteristics that stress the design process and the capabilities of commercial HLS tools in different ways: the first is communication bound and requires overlapping and pipelining of communication and computation in order to maximize the application throughput; the second is computation bound and uses complex mathematical functions that HLS tools do not directly support. Despite these difficulties, thanks to HLS, in the span of only 4 months we were able to explore a large design space and derive about 100 implementations with different cost-performance profiles, targeting both a Field-Programmable Gate Array (FPGA) platform and a 32-nm standard-cell Application Specific Integrated Circuit (ASIC) library. In addition, we could obtain results that outperform a previous Register-Transfer Level (RTL) implementation, which confirms the remarkable progress of HLS tools. Daniele Jahier Pagliari, Mario R. Casu, Luca P. Carloni |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2017 | Approximate Energy-Efficient Encoding for Serial InterfacesabstractSerial buses are ubiquitous interconnections in embedded computing systems that are used to interface processing elements with peripherals, such as sensors, actuators, and I/O controllers. Despite their limited wiring, as off-chip connections they can account for a significant amount of the total power consumption of a system-on-chip device. Encoding the information sent on these buses is the most intuitive and affordable way to reduce their power contribution; moreover, the encoding can be made even more effective by exploiting the fact that many embedded applications can tolerate intermediate approximations without a significant impact on the final quality of results, thus trading off accuracy for power consumption. We propose a simple yet very effective approximate encoding for reducing dynamic energy in serial buses. Our approach uses differential encoding as a baseline scheme and extends it with bounded approximations to overcome the intrinsic limitations of differential encoding for data with low temporal correlation. We show that the proposed scheme, in addition to yielding extremely compact codecs, is superior to all state-of-the-art approximate serial encodings over a wide set of traces representing data received or sent from/to sensor or actuators. Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2016 | Serial T0: approximate bus encoding for energy-efficient transmission of sensor signalsabstractOff-chip serial buses are common in embedded systems, and due to the long physical lines, can contribute significantly to their energy consumption. However, these buses are often connected to analog sensors, whose data is inherently affected by noise and A/D errors. Thus, communication can tolerate small approximations, without a significant impact on the system outputs quality. Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino |
DAC | 1 |
| 2016 | Low-overhead adaptive constrast enhancement and power reduction for OLEDs
Daniele Jahier Pagliari, Massimo Poncino, Enrico Macii |
DATE | 1 |
| 2016 | Approximate Differential Encoding for Energy-Efficient Serial CommunicationabstractEmbedded computing systems include several off-chip serial links, that are typically used to interface processing elements with peripherals, such as sensors, actuators and I/O controllers. Because of the long physical lines of these connections, they can contribute significantly to the total energy consumption. On the other hand, many embedded applications are error resilient, i.e. they can tolerate intermediate approximations without a significant impact on the final quality of results. This feature can be exploited in serial buses to explore the trade-off between data approximations and energy consumption. Daniele Jahier Pagliari, Enrico Macii, Massimo Poncino |
ACM Great Lakes Symposium on VLSI | 1 |
| 2015 | Acceleration of microwave imaging algorithms for breast cancer detection via High-Level SynthesisabstractWe present the system-level design of two accelerators for two microwave imaging algorithms for breast cancer detection. The accelerators were designed in SystemC and optimized via High-Level Synthesis (HLS). The two algorithms stress the capabilities of commercial HLS tools in different ways: the first is communication-bound and requires careful pipelining of communication and computation; the second is computation-bound and requires the implementation of mathematical functions that are not properly supported by HLS tools. Still, in the span of four months we were able to design and validate about one hundred alternative implementations, targeting a Zynq SoC platform. Furthermore, we were pleased to obtain results that are superior to a previous RTL implementation, which confirms the remarkable progress of HLS tools. Daniele Jahier Pagliari, Mario R. Casu, Luca P. Carloni |
ICCD | 1 |
| 2015 | An automated design flow for approximate circuits based on reduced precision redundancyabstractReduced Precision Redundancy (RPR) is a popular Approximate Computing technique, in which a circuit operated in Voltage Over-Scaling (VOS) is paired to a reduced-bitwidth and faster replica so that VOS-induced timing errors are partially recovered by the replica, and their impact is mitigated. Previous works have provided various examples of effective implementations of RPR, which however suffer from three limitations: first, these circuits are designed using ad-hoc procedures, and no generalization is provided; second, error impact analysis is carried out statistically, thus neglecting issues like non-elementary data distribution and temporal correlation. Last, only dynamic power was considered in the optimization. In this work we propose a new generalized approach to RPR that allows to overcome all these limitations, leveraging the capabilities of state-of-the-art synthesis and simulation tools. By sacrificing theoretical provability in favor of an empirical input-based analysis, we build a design tool able to automatically add RPR to a preexisting gate-level netlist. Thanks to this method, we are able to confute some of the conclusions drawn in previous works, in particular those related to statistical assumptions on inputs; we show that a given inputs distribution may yield extremely different results depending on their temporal behavior. Daniele Jahier Pagliari, Andrea Calimera, Enrico Macii, Massimo Poncino |
ICCD | 1 |