Giovanni Ansaloni

dblp:18/4692 · DBLP profile ↗
← Back
52ranked-venue papers
4as first author
28since 2021 · last 2026
0000-0002-8940-3775ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 47 · 4 first-author · 27 since 2021Software engineering, systems software and programming languages · 11 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Application-Driven System Technology Co-Optimization for 2.5D Edge AI Platforms
Anna Burdina, David Mallasén, Alexandre Levisse, Pasquale Davide Schiavone, Giovanni Ansaloni, David Atienza 0001
ISLPED5
2025 Systolic Arrays and Structured Pruning Co-design for Efficient Transformers in Edge Systems
abstract
International audience
Pedro Palacios, Rafael Medina 0001, Jean-Luc Rouas, Giovanni Ansaloni, David Atienza 0001
ACM Great Lakes Symposium on VLSI4
2025 Structured pruning for efficient systolic array accelerated cascade Speech-to-Text Translation
Jean-Luc Rouas, Charles Brazier, Leila Ben Letaifa, Rafael Medina 0001, Pedro Palacios, David Atienza 0001, Giovanni Ansaloni
INTERSPEECH7
2025 A Reconfigurable High-Dynamic Range ∆Σ Front-End with Event-Based Decimation for Bandwidth-Efficient Implantable Neural Interfaces
abstract
As the demand for high channel counts and high-resolution recordings of neural activity continues to grow, the increased power and data rate generated impose hard constraints on the telemetry capabilities of wireless implantable neural interfaces. To address this challenge, this work presents a novel system architecture for a reconfigurable readout circuit. It provides per-channel data rate reduction and adaptable bandwidth to match the characteristics and evolution of the neural signals under non-ideal electrode-tissue interactions. The system consists of a 14-bit hybrid continuous-time/discrete-time delta-sigma (CT/DT-∆Σ) analog front-end (AFE) followed by event-based decimation (EBD) which exploits the inherent sparsity in neural signals. The proposed AFE and EBD co-design was simulated using artifact-laden nonhuman primate microwire recordings. Results demonstrate a dynamic range of 76 dB, ensuring artifact robustness, along with up to a two-order-of-magnitude reduction in output data rate and power-area decimation footprint per channel, offering flexibility for high-quality (14 dB NRMSE) and medium-quality (8 dB NRMSE) reconstructions, based on the characteristics of the neural signals recorded at each channel.
Natalia Martínez, Juan Sapriza, Pasquale Davide Schiavone, Giovanni Ansaloni, Luke Bashford, Andrew Jackson 0001, David Atienza 0001, Timothy G. Constandinou
ISCAS4
2025 SideDRAM: Integrating SoftSIMD Datapaths near DRAM Banks for Energy-Efficient Variable Precision Computation
abstract
By interfacing computing logic directly to the DRAM banks, bank-level Compute-near-Memory (CnM) architectures promise to mitigate the bottleneck at the memory interconnect. While this computation paradigm heavily reduces the energy requirements for data movement across the system, current solutions fail to co-optimize hardware and software to further increase efficiency. Instead, in this manuscript, we present SideDRAM , a co-designed bank-level CnM architecture to enable massively parallel and energy-efficient computations near DRAM. In contrast with past solutions, we support flexible data typing and heterogeneous quantization, relying on the robustness of workloads to employ small bitwidths, and enable a row-wide access to the banks to exploit parallelism and spatial locality. As a result, SideDRAM integrates (1) software-defined SIMD (SoftSIMD) datapaths, supporting low-energy computing with flexible precision, (2) an interface to the banks based on very wide registers (VWRs), enabling asymmetric data access to both utilize the full DRAM bank bandwidth and leverage data locality at the datapath, and (3) a low-overhead distributed control plane, allowing the efficient handling of variable data typing. We benchmark SideDRAM as a near-DRAM solution by analyzing the area, performance, and energy consumption of an HBM2 CnM channel executing heterogeneously quantized machine learning models. The results show that, compared to the state-of-the-art FIMDRAM design, energy improvements of up to 67% are achieved when a DeiT-S inference is executed with a batch size of 16 under the same area constraints, resulting in energy-delay-area product (EDAP) savings that reach 83%. When comparing to a massively parallel mixed-signal CnM solution, SideDRAM consistently obtains similar performance and better energy efficiency results (geomean of 15× improvement across workloads) at a lower area overhead.
Rafael Medina 0001, Pengbo Yu, Alexandre Levisse, Dwaipayan Biswas, Marina Zapater, Giovanni Ansaloni, Francky Catthoor, David Atienza 0001
ACM Trans. Embed. Comput. Syst.6
2025 Towards Accurate RISC-V Full System Simulation via Component-Level Calibration
abstract
Full-System (FS) simulation is essential for performance evaluation of complete systems that execute complex applications on a complete software stack consisting of an operating system and user applications. Nevertheless, they require careful fine-tuning against real hardware to obtain reliable performance statistics, which can become tedious, error-prone, and time-consuming with typical trial-and-error approaches. We propose a novel, streamlined, component-level calibration methodology to address these shortcomings to validate FS simulation models. Our methodology greatly accelerates the validation process without sacrificing accuracy. It is Instruction Set Architecture (ISA)-agnostic, and can tackle hardware specifications at different levels of detail. We demonstrate its effectiveness by validating FS models against both open-hardware and IP-protected (closed hardware) RISC-V silicon, achieving a mean error of 19%–23% for the SPEC CPU2017 suite in the two cases. We introduce the first open-source RISC-V-based FS-validated simulation models with a complete and replicable methodology.
Karan Pathak, Joshua Alexander Harrison Klein, Giovanni Ansaloni, Said Hamdioui, Georgi Gaydadjiev, Marina Zapater, David Atienza 0001
ACM Trans. Embed. Comput. Syst.3
2024 FVLLMONTI: The 3D Neural Network Compute Cube $(N^{2}C^{2})$ Concept for Efficient Transformer Architectures Towards Speech-to-Speech Translation
abstract
This multi-partner-project contribution introduces the midway results of the Horizon 2020 FVLLMONTI project. In this project we develop a new and ultra-efficient class of ANN accelerators, the neural network compute cube$(N^{2}C^{2})$, which is specifically designed to execute complex machine learning tasks in a 3D technology, in order to provide the high computing power and ultra-high efficiency needed for future edgeAI applications. We showcase its effectiveness by targeting the challenging class of Transformer ANNs, tailored for Automatic Speech Recognition and Machine Translation, the two fundamental components of speech-to-speech translation. To gain the full benefit of the accelerator design, we develop disruptive vertical transistor technologies and execute design-technology-co-optimization (DTCO) loops from single device, to cell and compute cube level. Further, a hardware-software-co-optimization is executed, e.g. by compressing the executed speech recognition and translation models for energy efficient executing without substantial loss in precision.
Ian O'Connor, Sara Mannaa, Alberto Bosio, Bastien Deveautour, Damien Deleruyelle, Tetiana Obukhova, Cédric Marchand 0002, Jens Trommer, Çigdem Çakirlar, Bruno Neckel Wesling, Thomas Mikolajick, Oskar Baumgartner, Mischa Thesberg, David Pirker, Christoph Lenz, Zlatan Stanojevic, Markus Karner, Guilhem Larrieu, Sylvain Pelloquin, Konstantinous Moustakas, Giovanni Ansaloni, Alireza Amirshahi, David Atienza 0001, Jean-Luc Rouas, Leila Ben Letaifa, Georgeta Bordeall, Charles Brazier, C. Mukherjee 0001, Marina Deng, Marc François, Houssem Rezgui, Reveil Lucas, Cristell Maneux
DATE22
2024 Cross-layer Exploration of 2.5D Energy-Efficient Heterogeneous Chiplets Integration: From System Simulation to Open Hardware
abstract
In the past decade, computing systems have significantly increased in complexity and power consumption. Nowadays, heterogeneous multi-processor systems-on-chip (MPSoCs) integrate many computing cores. Heterogeneous MPSoCs often comprise general-purpose processors and a variety of accelerators, thus supporting specialized functions for the target application domain to minimize overall energy when executing a specific task. The ensuing architectural design space is, therefore, increasingly multi-dimensional, especially in the light of upcoming 2.5D/3D chiplets integration, which, on one side, allows unprecedented system integration possibilities but, on the other, exacerbates data transfer bottlenecks and affects overall power consumption significantly. To traverse such space in search of high-performance/high-efficiency solutions, we introduce a cross-layer approach combining fast explorations with virtual systems with modular open-hardware design frameworks. This paper showcases how these two approaches effectively cross-fertilize: detailed hardware designs are essential in calibrating performance, power and temperature models, and validating simulation outcomes. Conversely, full system simulation is crucial for projecting the impact of design choices towards complex but energy-efficient heterogeneous multi-processor architectures.
Anna Burdina, Gabriel Catel Torres, Pasquale Davide Schiavone, Miguel Peón-Quirós, Giovanni Ansaloni, David Atienza 0001, Marina Zapater
ISLPED5
2024 SAT-Based Exact Modulo Scheduling Mapping for Resource-Constrained CGRAs
abstract
Coarse-Grain Reconfigurable Arrays (CGRAs) represent emerging low-power architectures designed to accelerate Compute-Intensive Loops (CILs). The effectiveness of CGRAs in providing acceleration relies on the quality of mapping: how efficiently the CIL is compiled onto the platform. State-of-the-Art (SoA) compilation techniques utilize modulo scheduling to minimize the Iteration Interval (II) and use graph algorithms like Max-Clique Enumeration to address mapping challenges. Our work approaches the mapping problem through a satisfiability (SAT) formulation. We introduce the Kernel Mobility Schedule (KMS), an ad hoc schedule used with the Data Flow Graph and CGRA architectural information to generate Boolean statements that, when satisfied, yield a valid mapping. Experimental results demonstrate SAT-MapIt outperforming SoA alternatives in almost 50% of explored benchmarks. Additionally, we evaluated the mapping results in a synthesizable CGRA design and emphasized the runtime metrics trends, i.e., energy efficiency and latency, across different CILs and CGRA sizes. We show that a hardware-agnostic analysis performed on compiler-level metrics can optimally prune the architectural design space, while still retaining Pareto-optimal configurations. Moreover, by exploring how implementation details impact cost and performance on real hardware, we highlight the importance of holistic software-to-hardware mapping flows, as the one presented herein.
Cristian Tirelli, Juan Sapriza, Rubén Rodríguez Álvarez, Lorenzo Ferretti, Benoît W. Denkinger, Giovanni Ansaloni, José Miranda 0001, David Atienza 0001, Laura Pozzi 0001
ACM J. Emerg. Technol. Comput. Syst.6
2024 Bank on Compute-Near-Memory: Design Space Exploration of Processing-Near-Bank Architectures
abstract
Near-DRAM computing strategies advocate for providing computational capabilities close to where data is stored. Although this paradigm can effectively address the memory-to-processor communication bottleneck, it also presents new challenges: The strict resource constraints in the memory periphery demand careful tailoring of architectural elements. We herein propose a novel framework and methodology to explore compute-near-memory designs that interface to DRAM memory banks, demonstrating the area, energy, and performance tradeoffs subject to the architectural configuration. We exemplify this methodology by conducting two studies on compute-near-bank designs: 1) analyzing the interaction between control and data resources, and 2) exploring the integration of processing units with different DRAM standards. According to our study, the optimal size ratios between instruction and data capacity vary from$2\times $to$4\times $across benchmarks from representative application domains. The retrieved Pareto-optimal solutions from our framework improve state-of-the-art designs, e.g., achieving a 50% performance increase on matrix operations with 15% energy overhead relative to the FIMDRAM design. In addition, the exploration of DRAM shows the interplay between available internal bandwidth, performance, and area overhead. For example, a threefold increase in bandwidth rises performance by 47% across workloads at a 34% extra area cost.
Rafael Medina 0001, Giovanni Ansaloni, Marina Zapater, Alexandre Levisse, Saeideh Alinezhad Chamazcoti, Timon Evenblij, Dwaipayan Biswas, Francky Catthoor, David Atienza 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 Which Coupled is Best Coupled? An Exploration of AIMC Tile Interfaces and Load Balancing for CNNs
abstract
Due to stringent energy and performance constraints, edge AI computing often employs heterogeneous systems that utilize both general-purpose CPUs and accelerators. Analog in-memory computing (AIMC) is a well-known AI inference solution that overcomes computational bottlenecks by performing matrix-vector multiplication operations (MVMs) in constant time. However, the tiles of AIMC-based accelerators are limited by the number of weights they can hold. State-of-the-art research often sizes neural networks to AIMC tiles (or vice-versa), but does not consider cases where AIMC tiles cannot cover the whole network due to lack of tile resources or the network size. In this work, we study the trade-offs of available AIMC tile resources, neural network coverage, AIMC tile proximity to compute resources, and multi-core load balancing techniques. We first perform a study of single-layer performance and energy scalability of AIMC tiles in the two most typical AIMC acceleration targets: dense/fully-connected layers and convolutional layers. This study guides the methodology with which we approach parameter allocation to AIMC tiles in the context of large edge neural networks, both where AIMC tiles are close to the CPU (tightly-coupled) and cannot share resources across the system, and where AIMC tiles are far from the CPU (loosely-coupled) and can employ workload stealing. We explore the performance and energy trends of six modern CNNs using different methods of load balancing for differently-coupled system configurations with variable AIMC tile resources. We show that, by properly distributing workloads, AIMC acceleration can be made highly effective even on under-provisioned systems. As an example, 5.9x speedup and 5.6x energy gains were measured on an 8-core system, for a 41% coverage of neural network parameters.
Joshua Alexander Harrison Klein, Irem Boybat, Giovanni Ansaloni, Marina Zapater, David Atienza 0001
IEEE Trans. Parallel Distributed Syst.3
2024 An Energy Efficient Soft SIMD Microarchitecture and Its Application on Quantized CNNs
abstract
The ever-increasing computational complexity and energy consumption of today’s applications, such as machine learning (ML) algorithms, not only strain the capabilities of the underlying hardware but also significantly restrict their wide deployment at the edge. Addressing these challenges, novel architecture solutions are required by leveraging opportunities exposed by algorithms, e.g., robustness to small-bitwidth operand quantization and high intrinsic data-level parallelism. However, traditional hardware single instruction multiple data (Hard SIMD) architectures only support a small set of operand bitwidths, limiting performance improvement. To fill the gap, this manuscript introduces a novel pipelined processor microarchitecture for arithmetic computing based on the software-defined SIMD (Soft SIMD) paradigm that can define arbitrary SIMD modes through control instructions at run-time. This microarchitecture is optimized for parallel fine-grained fixed-point arithmetic, such as shift/add. It can also efficiently execute sequential shift-add-based multiplication over SIMD subwords, thanks to zero-skipping and canonical signed digit (CSD) coding. A lightweight repacking unit allows changing subword bitwidth dynamically. These features are implemented within a tight energy and area budget. An energy consumption model is established through post-synthesis for performance assessment. We select heterogeneously quantized (HQ) convolutional neural networks (CNNs) from the ML domain as the benchmark and map it onto our microarchitecture. Experimental results showcase that our approach dramatically outperforms traditional Hard SIMD Multiplier-Adder regarding area and energy requirements. In particular, our microarchitecture occupies up to 59.9% less area than a Hard SIMD that supports fewer SIMD bitwidths, while consuming up to 50.1% less energy on average to execute HQ CNNs.
Pengbo Yu, Flavio Ponzina, Alexandre Levisse, Mohit Gupta 0004, Dwaipayan Biswas, Giovanni Ansaloni, David Atienza 0001, Francky Catthoor
IEEE Trans. Very Large Scale Integr. Syst.6
2023 TiC-SAT: Tightly-Coupled Systolic Accelerator for Transformers
abstract
Transformer models have achieved impressive results in various AI scenarios, ranging from vision to natural language processing. However, their computational complexity and their vast number of parameters hinder their implementations on resource-constrained platforms. Furthermore, while loosely-coupled hardware accelerators have been proposed in the literature, data transfer costs limit their speed-up potential. We address this challenge along two axes. First, we introduce tightly-coupled, small-scale systolic arrays (TiC-SATs), governed by dedicated ISA extensions, as dedicated functional units to speed up execution. Then, thanks to the tightly-coupled architecture, we employ software optimizations to maximize data reuse, thus lowering miss rates across cache hierarchies. Full system simulations across various BERT and Vision-Transformer models are employed to validate our strategy, resulting in substantial application-wide speed-ups (e.g., up to 89.5X for BERT-large). TiC-SAT is available as an open-source framework1.
Alireza Amirshahi, Joshua Alexander Harrison Klein, Giovanni Ansaloni, David Atienza 0001
ASP-DAC3
2023 System-Level Exploration of In-Package Wireless Communication for Multi-Chiplet Platforms
abstract
Multi-Chiplet architectures are being increasingly adopted to support the design of very large systems in a single package, facilitating the integration of heterogeneous components and improving manufacturing yield. However, chiplet-based solutions have to cope with limited inter-chiplet routing resources, which complicate the design of the data interconnect and the power delivery network. Emerging in-package wireless technology is a promising strategy to address these challenges, as it allows to implement flexible chiplet interconnects while freeing package resources for power supply connections. To assess the capabilities of such an approach and its impact from a full-system perspective, herein we present an exploration of the performance of in-package wireless communication, based on dedicated extensions to the gem5-X simulator. We consider different Medium Access Control (MAC) protocols, as well as applications with different runtime profiles, showcasing that current in-package wireless solutions are competitive with wired chiplet interconnects. Our results show how in-package wireless solutions can outperform wired alternatives when running artificial intelligence workloads, achieving up to a 2.64× speed-up when running deep neural networks (DNNs) on a chiplet-based system with 16 cores distributed in four clusters.
Rafael Medina 0001, Joshua Kein, Giovanni Ansaloni, Marina Zapater, Sergi Abadal, Eduard Alarcón, David Atienza 0001
ASP-DAC3
2023 An Open-Hardware Coarse-Grained Reconfigurable Array for Edge Computing
abstract
In this work, we propose an open-hardware low-power coarse-grained reconfigurable array connected to a lightweight microcontroller and enclosed in an application mapping framework. The latter provides complete support to configure kernels in the reconfigurable array, execute applications, and measure performance.
Rubén Rodríguez Álvarez, Benoît W. Denkinger, Juan Sapriza, José Miranda 0001, Giovanni Ansaloni, David Atienza 0001
CF5
2023 Cross Layer Design for the Predictive Assessment of Technology-Enabled Architectures
abstract
There is great interest in “end-to-end” analysis that captures how innovation at the materials, device, and/or archi-tectural levels will impact figures of merit at the application-level. However, there are numerous combinations of devices and architectures to study, and we must establish systematic ways to accurately explore and cull a vast design space. We aim to capture how innovations at the materials/device-level may ultimately impact figures of merit associated with both existing and emerging technologies that may be employed for either logic and/or memory. We will highlight how collaborations with researchers at these levels of the design hierarchy - as well as efforts to help construct well-calibrated device models - can in-turn support architectural design space explorations that will help to identify the most promising ways to use new technologies to support application-level workloads of interest. For given compute workloads, we can then quantitatively assess the potential benefits of technology-driven architectures to identify the most promising paths forward. Because of the large number of potentially interesting device-architecture combinations, it is of the utmost importance to develop well-calibrated analytical modeling tools to more rapidly assess the potential value of a given (likely heterogeneous) solution. We highlight recent efforts and needs in this space.
Michael T. Niemier, Xiaobo Sharon Hu, Liu Liu 0023, Mohammad Mehdi Sharifi, Ian O'Connor, David Atienza 0001, Giovanni Ansaloni, Can Li 0024, Daniel C. Ralph
DATE7
2023 A 16-bit Floating-Point Near-SRAM Architecture for Low-power Sparse Matrix-Vector Multiplication
abstract
State-of-the-art Artificial Intelligence (AI) algorithms, such as graph neural networks and recommendation systems, require floating-point computation of very large matrix multiplications over sparse data. Their execution in resource-constrained scenarios, like edge AI systems, requires a) careful optimization of computing patterns, leveraging sparsity as an opportunity to lower computational requirements, and b) using dedicated hardware. In this paper, we introduce a novel near-memory floating-point computing architecture dedicated to the parallel processing of sparse matrix-vector multiplication (SpMV). This architecture can be integrated at the periphery of memory arrays to exploit the inherent parallelism of memory structures to speed up computation. In addition, it uses its proximity to memory to achieve high computational capability and very low latency. The illustrated implementation, operating at 1GHz, can compute up to 370 MFLOPS (millions of floating-point operations per second) while computing SpMV multiplications, while incurring a modest 17% area overhead when interfaced with a 4KB SRAM array.
Grégoire Eggermann, Marco Rios, Giovanni Ansaloni, Sani R. Nassif, David Atienza 0001
VLSI-SoC3
2023 REMOTE: Re-thinking Task Mapping on Wireless 2.5D Systems-on-Package for Hotspot Removal
abstract
2.5D Systems-on-Package (SoPs) are composed by several chiplets placed on an interposer. They are becoming increasingly popular as they enable easy integration of electronic components in the same package and high fabrication yields. Nevertheless, they introduce a new bottleneck in inter-chiplet communication, which must be routed through the interposer. Such a constraint favors mapping related tasks on computing cores within the same chiplet, leading to thermal hotspots. In-package wireless technology holds promise to reconsider such a position because integrated wireless antennas provide low-latency and high-bandwidth communication paths, thus bypassing the in-terposer bottleneck. Furthermore, in this work, we propose a new task mapping heuristic that leverages in-package wireless technology to improve the thermal behavior of 2.5D SoPs executing complex applications. Combining system simulation and thermal modeling, our results show that we can distribute computation in wireless 2.5D SoPs to reduce peak temperatures by up to 24% through task mapping with a negligible performance impact.
Rafael Medina 0001, Darong Huang 0003, Giovanni Ansaloni, Marina Zapater, David Atienza 0001
VLSI-SoC3
2023 ALPINE: Analog In-Memory Acceleration With Tight Processor Integration for Deep Learning
abstract
Analog in-memory computing (AIMC) cores offers significant performance and energy benefits for neural network inference with respect to digital logic (e.g., CPUs). AIMCs accelerate matrix-vector multiplications, which dominate these applications' run-time. However, AIMC-centric platforms lack the flexibility of general-purpose systems, as they often have hard-coded data flows and can only support a limited set of processing functions. With the goal of bridging this gap in flexibility, we present a novel system architecture that tightly integrates analog in-memory computing accelerators into multi-core CPUs in general-purpose systems. We developed a powerful gem5-based full system-level simulation framework into the gem5-X simulator, ALPINE, which enables an in-depth characterization of the proposed architecture. ALPINE allows the simulation of the entire computer architecture stack from major hardware components to their interactions with the Linux OS. Within ALPINE, we have defined a custom ISA extension and a software library to facilitate the deployment of inference models. We showcase and analyze a variety of mappings of different neural network types, and demonstrate up to 20.5x/20.8x performance/energy gains with respect to a SIMD-enabled ARM CPU implementation for convolutional neural networks, multi-layer perceptrons, and recurrent neural networks.
Joshua Alexander Harrison Klein, Irem Boybat, Yasir Mahmood Qureshi, Martino Dazzi, Alexandre Levisse, Giovanni Ansaloni, Marina Zapater, Abu Sebastian, David Atienza 0001
IEEE Trans. Computers6
2023 Thermal and Voltage-Aware Performance Management of 3-D MPSoCs With Flow Cell Arrays and Integrated SC Converters
abstract
Flow cell arrays (FCAs) concurrently provide efficient on-chip liquid cooling and electrochemical power generation. This technology is especially promising for 3-D multiprocessor systems-on-chip (3-D MPSoCs) realized in deeply scaled technologies, which present very challenging power and thermal requirements. Indeed, FCAs effectively improve power delivery network (PDN) performance, particularly if switched capacitor (SC) converters are employed to decouple the flow cells and the systems-on-chip voltages, allowing each to operate at their optimal point. Nonetheless, the design of FCA-based solutions entails nonobvious considerations and tradeoffs, stemming from their dual role in governing both the thermal and power delivery characteristics of 3-D MPSoCs. Showcasing them in this article, we explore multiple FCA design configurations and demonstrate that this technology can decrease the temperature of a heterogeneous 3-D MPSoC by 78 °C, and its total power consumption by 46%, compared to a high-performance cold-plate-based liquid cooling solution. At the same time, FCAs enable up to 90% voltage drop recovery across dies, using SC converters occupying a small fraction of the chip area. Such outcomes provide an opportunity to boost 3-D MPSoC computing performance by increasing the operating frequency of dies. Leveraging these results, we introduce a novel temperature and voltage-aware model-predictive control (MPC) strategy that optimizes power efficiency during runtime. We achieve application-wide speedups of up to 16% on various machine learning (ML), data mining, and other high-performance benchmarks while keeping the 3-D MPSoC temperature below 83 °C and voltage drops below 5%.
Halima Najibi, Alexandre Levisse, Giovanni Ansaloni, Marina Zapater, Miroslav Vasic, David Atienza 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Overflow-free Compute Memories for Edge AI Acceleration
abstract
Compute memories are memory arrays augmented with dedicated logic to support arithmetic. They support the efficient execution of data-centric computing patterns, such as those characterizing Artificial Intelligence (AI) algorithms. These architectures can provide computing capabilities as part of the memory array structures (In-Memory Computing, IMC) or at their immediate periphery (Near-Memory Computing, NMC). By bringing the processing elements inside (or very close to) storage, compute memories minimize the cost of data access. Moreover, highly parallel (and, hence, high-performance) computations are enabled by exploiting the regular structure of memory arrays. However, the regular layout of memory elements also constrains the data range of inputs and outputs, since the bitwidths of operands and results stored at each address cannot be freely varied. Addressing this challenge, we herein propose a HW/SW co-design methodology combining careful per-layer quantization and inter-layer scaling with lightweight hardware support for overflow-free computation of dot-vector operations. We demonstrate their use to implement the convolutional and fully connected layers of AI models. We embody our strategy in two implementations, based on IMC and NMC, respectively. Experimental results highlight that an area overhead of only 10.5% (for IMC) and 12.9% (for NMC) is required when interfacing with a 2KB subarray. Furthermore, inferences on benchmark CNNs show negligible accuracy degradation due to quantization for equivalent floating-point implementations.
Flavio Ponzina, Marco Rios, Alexandre Levisse, Giovanni Ansaloni, David Atienza 0001
ACM Trans. Embed. Comput. Syst.4
2022 INCLASS: Incremental Classification Strategy for Self-Aware Epileptic Seizure Detection
abstract
Wearable Health Companions allow the unobtrusive monitoring of patients affected by chronic conditions. In particular, by acquiring and interpreting bio-signals, they enable the detection of acute episodes in cardiac and neurological ailments. Nevertheless, the processing of bio-signals is computationally complex, especially when a large number of features are required to obtain reliable detection outcomes. Addressing this challenge, we present a novel methodology, named INCLASS, that iteratively extends employed feature sets at run-time, until a confidence condition is satisfied. INCLASS builds such sets based on code analysis and profiling information. When applied to the challenging scenario of detecting epileptic seizures based on ECG and SpO2 acquisitions, INCLASS obtains savings of up to 54%, while incurring in a negligible loss of detection performance (1.1% degradation of specificity and sensitivity) with respect to always computing and evaluating all features.
Lorenzo Ferretti, Giovanni Ansaloni, Renaud Marquis, Tomás Teijeiro, Philippe Ryvlin, David Atienza 0001, Laura Pozzi 0001
DATE2
2022 Thermal and Power-Aware Run-time Performance Management of 3D MPSoCs with Integrated Flow Cell Arrays
abstract
Flow Cell Arrays (FCA) technology employs microchannels filled with an electrolytic fluid to concurrently provide cooling and power generation to integrated circuits (ICs). This solution is particularly appealing for Three-Dimensional Multi-Processor Systems-on-Chip (3D MPSoCs) realized in deeply scaled technologies, as their extreme power densities result in significant thermal and voltage supply challenges. FCAs provide them with extra power to boost performance. However, the dual effects of FCAs (cooling and power supply) have conflicting trends leading to a complex interplay between temperature, voltage stability, and performance. In this paper, we explore this trade-off by introducing a novel methodology that controls the operating frequency of computing components and the electrolytic coolant flow rate at run-time. Our strategy enables tangible performance gains while abiding by timing, voltage drop, and temperature constraints. We showcase its benefits by targeting a 4-layer 3D MPSoC, achieving up to 24% increase in the operating frequencies and resulting in application speedups of up to 17%, while reducing the costs related to FCA liquid pumping energy.
Halima Najibi, Alexandre Levisse, Giovanni Ansaloni, Marina Zapater, David Atienza 0001
ACM Great Lakes Symposium on VLSI3
2022 Error Resilient In-Memory Computing Architecture for CNN Inference on the Edge
abstract
The growing popularity of edge computing has fostered the development of diverse solutions to support Artificial Intelligence (AI) in energy-constrained devices. Nonetheless, comparatively few efforts have focused on the resiliency exhibited by AI workloads (such as Convolutional Neural Networks, CNNs) as an avenue towards increasing their run-time efficiency, and even fewer have proposed strategies to increase such resiliency. We herein address this challenge in the context of Bit-line Computing architectures, an embodiment of the in-memory computing paradigm tailored towards CNN applications. We show that little additional hardware is required to add highly effective error detection and mitigation in such platforms. In turn, our proposed scheme can cope with high error rates when performing memory accesses with no impact on CNNs accuracy, allowing for very aggressive voltage scaling. Complementary, we also show that CNN resiliency can be increased by algorithmic optimizations in addition to architectural ones, adopting a combined ensembling and pruning strategy that increases robustness while not inflating workload requirements. Experiments on different quantized CNN models reveal that our combined hardware/software approach enables the supply voltage to be reduced to just 650mV, decreasing the energy per inference up to 51.3%, without affecting the baseline CNN classification accuracy.
Marco Rios, Flavio Ponzina, Giovanni Ansaloni, Alexandre Levisse, David Atienza 0001
ACM Great Lakes Symposium on VLSI3
2022 A Formal Framework for Maximum Error Estimation in Approximate Logic Synthesis
abstract
Approximate logic synthesis techniques have become popular in error-resilient systems, where accuracy requirements can be traded for improved energy efficiency. Many of these techniques operate on a circuit by substituting or removing some of its portions under a predefined error constraint; however, the research on systematic methods to determine the error induced by such transformations is still at an early stage. We propose herein a generic framework for modeling maximum error in a circuit, called partition and propagate, which is a fundamental preliminary step for ALS. This framework is based on circuit partitioning and error propagation among the subcircuits. We provide a sound, complete formal description of such framework, and we illustrate how two state-of-the-art algorithms can be subsumed by it. Moreover, we propose a novel gate-level error-modeling algorithm, which is able to identify the whole range of possible errors induced by a given approximate transformation. We compare the three strategies and illustrate the efficiency of the new error-propagation methodology, which is able to identify accurate error bounds and, hence, guide ALS techniques to more valuable solutions.
Ilaria Scarabottolo, Giovanni Ansaloni, George A. Constantinides, Laura Pozzi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Architecting more than Moore: wireless plasticity for massive heterogeneous computer architectures (WiPLASH)
abstract
This paper presents the research directions pursued by the WiPLASH European project, pioneering on-chip wireless communications as a disruptive enabler towards next-generation computing systems for artificial intelligence (AI). We illustrate the holistic approach driving our research efforts, which encompass expertises and abstraction levels ranging from physical design of embedded graphene antennas to system-level evaluation of wirelessly-communicating heterogeneous systems.
Joshua Alexander Harrison Klein, Alexandre Levisse, Giovanni Ansaloni, David Atienza 0001, Marina Zapater, Martino Dazzi, Geethan Karunaratne, Irem Boybat, Abu Sebastian, Davide Rossi 0001, Francesco Conti 0001, Elana Pereira de Santana, Peter Haring Bolívar, Mohamed Saeed, Renato Negra, Kun-Ta Wang, Max Christian Lemme, Akshay Jain 0001, Robert Guirado, Hamidreza Taghvaee, Sergi Abadal
CF3
2021 Exact Neural Networks from Inexact Multipliers via Fibonacci Weight Encoding
abstract
Edge devices must support computationally demanding algorithms, such as neural networks, within tight area/energy budgets. While approximate computing may alleviate these constraints, limiting induced errors remains an open challenge. In this paper, we propose a hardware/software co-design solution via an inexact multiplier, reducing area/power-delay-product requirements by 73/43%, respectively, while still computing exact results when one input is a Fibonacci encoded value. We introduce a retraining strategy to quantize neural network weights to Fibonacci encoded values, ensuring exact computation during inference. We benchmark our strategy on Squeezenet 1.0, DenseNet-121, and ResNet-18, measuring accuracy degradations of only 0.4/1.1/1.7%.
William Andrew Simon, Valérian Ray, Alexandre Levisse, Giovanni Ansaloni, Marina Zapater, David Atienza 0001
DAC4
2021 Running Efficiently CNNs on the Edge Thanks to Hybrid SRAM-RRAM In-Memory Computing
abstract
The increasing size of Convolutional Neural Networks (CNNs) and the high computational workload required for inference pose major challenges for their deployment on resource-constrained edge devices. in this paper, we address them by proposing a novel In-Memory Computing (IMC) architecture. Our IMC strategy allows us to efficiently perform arithmetic operations based on bitline computing, enabling a high degree of parallelism while reducing energy-costly data transfers. Moreover, it features a hybrid memory structure, where a portion of each subarray, dedicated to storing CNN weights, is implemented as high-density, zero-standby-power Resistive RAM. Finally, it exploits an innovative method for storing quantized weights based on their value, named Weight Data Mapping (WDM), which further increases efficiency. Compared to state-of-the-art IMC alternatives, our solution provides up to 93% improvements in energy efficiency and up to 6x less run-time when performing inference on Mobilenet and AlexNet neural networks.
Marco Rios, Flavio Ponzina, Giovanni Ansaloni, Alexandre Levisse, David Atienza 0001
DATE3
2020 Approximate Logic Synthesis: A Survey
abstract
Approximate computing is an emerging paradigm that, by relaxing the requirement for full accuracy, offers benefits in terms of design area and power consumption. This paradigm is particularly attractive in applications where the underlying computation has inherent resilience to small errors. Such applications are abundant in many domains, including machine learning, computer vision, and signal processing. In circuit design, a major challenge is the capability to synthesize the approximate circuits automatically without manually relying on the expertise of designers. In this work, we review methods devised to synthesize approximate circuits, given their exact functionality and an approximability threshold. We summarize strategies for evaluating the error that circuit simplification can induce on the output, which guides synthesis techniques in choosing the circuit transformations that lead to the largest benefit for a given amount of induced error. We then review circuit simplification methods that operate at the gate or Boolean level, including those that leverage classical Boolean synthesis techniques to realize the approximations. We also summarize strategies that take high-level descriptions, such as C or behavioral Verilog, and synthesize approximate circuits from these descriptions.
Ilaria Scarabottolo, Giovanni Ansaloni, George A. Constantinides, Laura Pozzi 0001, Sherief Reda
Proc. IEEE2
2020 Leveraging Prior Knowledge for Effective Design-Space Exploration in High-Level Synthesis
abstract
High-Level Synthesis (HLS) tools allow the generation of a large variety of hardware implementations from the same specification by setting different optimization directives. Each combination of HLS directives returns an implementation of the target application that is based on a particular microarchitecture. Designers are interested only in the subset of implementations that correspond to Pareto-optimal points in the performance versus cost design space. Finding this subset is hard because the relationship between the HLS directives and the Pareto-optimal implementations cannot be foreseen. Hence, designers must default to an exploration of the design space through many time-consuming HLS runs. We present a methodology that infers knowledge from past design explorations to identify high-quality directives for new target applications. To this end, we formulate a novel abstract representation of applications and their associated configuration spaces, introduce a similarity metric to compare quantitatively the configuration spaces of different applications, and a method to infer actionable information from a source space to a target space. The experimental results with the MachSuite benchmarks show that our approach retrieves close approximations of the Pareto frontier of best-performing implementations for the target application, in exchange for a small number of HLS runs.
Lorenzo Ferretti, Jihye Kwon, Giovanni Ansaloni, Giuseppe Di Guglielmo, Luca P. Carloni, Laura Pozzi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 Partition and Propagate: an Error Derivation Algorithm for the Design of Approximate Circuits
abstract
Inexact hardware design techniques have become popular in error-tolerant systems, where energy efficiency is a primary concern. Several techniques aim to identify circuit portions that can be discarded under an error constraint, but research on systematic methods to determine such error is still at an early stage. We herein illustrate a generic, scalable algorithm that determines the influence of each circuit gate on the final output. The algorithm first partitions the graph representing the circuit, then determines the error propagation model of the resulting subgraphs. When applied to existing approximate design frameworks, our solution improves their efficiency and result quality.
Ilaria Scarabottolo, Giovanni Ansaloni, George A. Constantinides, Laura Pozzi 0001
DAC2
2019 Tailoring SVM Inference for Resource-Efficient ECG-Based Epilepsy Monitors
abstract
Event detection and classification algorithms are resilient towards aggressive resource-aware optimisations. In this paper, we leverage this characteristic in the context of smart health monitoring systems. In more detail, we study the attainable benefits resulting from tailoring Support Vector Machine (SVM) inference engines devoted to the detection of epileptic seizures from ECG-derived features. We conceive and explore multiple optimisations, each effectively reducing resource budgets while minimally impacting classification performance. These strategies can be seamlessly combined, which results in 12.5X and 16X gains in energy and area, respectively, with a negligible loss, 3.2% in classification performance.
Lorenzo Ferretti, Giovanni Ansaloni, Laura Pozzi 0001, Amir Aminifar, David Atienza 0001, Leila Cammoun, Philippe Ryvlin
DATE2
2019 Compiler-Assisted Selection of Hardware Acceleration Candidates from Application Source Code
abstract
Hardware design is a difficult task. Beside ensuring functional correctness of an implementation, hardware developers are confronted with multiple and often conflicting constraints, such as performance and area cost targets, that require lengthy explorations. This issue is compounded when considering the acceleration of complex applications, of which some parts are implemented in software, and others are accelerated in hardware. Hardware/Software partitioning must be settled early in the development cycle, and is far from trivial, since at this stage detailed performance measurements are not available, while wrong choices can lead to vastly sub-optimal solutions or to wasted implementation efforts. To address this challenge, we present a framework to automatically identify, from un-modified software code, software segments that are promising candidates for hardware acceleration, to evaluate their potential speedup and resource requirements, and to select a subset of them under resource constraint. Our strategy is based on Intermediate Representation (IR) analysis passes, which we embed in the LLVM compiler toolchain, and does not require any time-consuming synthesis. We explore its effectiveness on the reference software implementation of a complex application, the H.264 Decoder from University of Illinois, and demonstrate that our methodology selects higher-performance sets of accelerators, when compared to strategies only based on profiling information.
Georgios Zacharopoulos 0001, Lorenzo Ferretti, Giovanni Ansaloni, Giuseppe Di Guglielmo, Luca P. Carloni, Laura Pozzi 0001
ICCD3
2019 RegionSeeker: Automatically Identifying and Selecting Accelerators From Application Source Code
abstract
Embedded systems present stringent and often conflicting requirements. On the one side, the need for high performance within a tight energy budget favors inflexible Application Specific Integrated Circuit (ASIC) implementations; on the other side, a short time-to-market demands programmability. Hybrid architectures such as special-purpose customized processors represent an attractive solution, as they are programmable by software, but use dedicated hardware to accelerate parts of the computation. In such a scenario, the capability of automatically identifying the computation parts to be realized in hardware is highly desirable, in order to reduce design time and effort. This paper aims at advancing the state-of-the-art in this field. We recognize that subgraphs of control flow graphs having a single input control point and a single output control point, that we call regions, are good targets for the synthesis of application specific hardware accelerators. We therefore provide a method to identify them and an LLVM-based toolchain (named RegionSeeker) that, analyzing a software application, automatically selects its most profitable regions given an area constraint. Experimental evidence shows that the accelerators identified by RegionSeeker provide a speedup of up to $4.6\boldsymbol {\times }$ and, on average, approximately 30% higher speedup is achieved compared to state-of-the-art identification techniques.
Georgios Zacharopoulos 0001, Lorenzo Ferretti, Emanuele Giaquinta, Giovanni Ansaloni, Laura Pozzi 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2018 A partitioning strategy for exploring error-resilience in circuits: work-in-progress
Ilaria Scarabottolo, Giovanni Ansaloni, Laura Pozzi 0001
CASES2
2018 Circuit carving: A methodology for the design of approximate hardware
abstract
Systems-on-Chip (SoCs) commonly couple low-power processors and dedicated hardware accelerators, which allow the execution of high-workload and/or timing-critical applications while relying on constrained resources. The functions performed by accelerators are often robust with respect to approximations that, when implemented in HW, can lead to circuits with tangibly lower area and power consumption. Research in approximate computing aims at developing effective strategies to explore the ensuing correctness/efficiency trade-off. In this context, we address the challenge of approximate circuit design in an innovative way, called here Circuit Carving, which consists in identifying the maximum portion of an exact circuit that can be discarded from it, or carved out, to derive an inexact version not exceeding an error threshold. We achieve this goal by proposing an algorithm based on binary tree exploration, bounded by conditions extracted from the circuit topology. Our approach can be applied to any combinatorial circuit, without a-priori knowledge of its functionality. The proposed algorithm allows back-tracking in order to never be trapped in local minima, and identifies the exact influence of each circuit gate on the output correctness, resulting in inexact circuits with higher efficiency and accuracy with respect to state-of-the-art greedy strategies.
Ilaria Scarabottolo, Giovanni Ansaloni, Laura Pozzi 0001
DATE2
2018 Lattice-Traversing Design Space Exploration for High Level Synthesis
abstract
This paper describes a design space exploration methodology for High Level Synthesis (HLS) frameworks. Inputs of HLS tools are a description (usually in C/C++) of the functionality of an intended hardware, and a set of optimisation directives that specify its implementation, hence allowing the generation of many design variants with widely varying performance and required resources. The relationship between directives and performance/cost is nonetheless not straightforward, and highly influenced by application-specific characteristics. A major challenge facing designers is then to define effective values for the directives while avoiding time-consuming - and often infeasible - exhaustive explorations. We herein address it by proposing a novel HLS exploration approach which employs a lattice representation of the design space, and a methodology for its navigation. We base our strategy on the observation that Pareto-implementations share a low variance among their configurations. We therefore guide the selection of HLS directives minimising the variance of new candidate solutions, with respect to the best performing ones that have already been visited. By only requiring local searches in the lattice space, our methodology gracefully scales to complex designs. It results in close approximations of the real Pareto frontier, while requiring a lower workload and fewer synthesis runs with respect to existing approaches.
Lorenzo Ferretti, Giovanni Ansaloni, Laura Pozzi 0001
ICCD2
2018 Heterogeneous and Inexact: Maximizing Power Efficiency of Edge Computing Sensors for Health Monitoring Applications
abstract
In the Internet-of-Things (IoT) era, there is an increasing trend to enable intelligent behavior in edge computing sensors. Thus, a new generation of smart wearable devices for health monitoring is being developed, able to perform complex Digital Signal Processing (DSP) routines that extract features of clinical relevance from the acquired data. These new edge computing sensors for personalized healthcare must operate within a tight energy envelope; addressing the ensuing challenge, we herein introduce an inexact and heterogeneous edge computing architecture, specifically tailored to the bio-DSP domain. We observe that bio-signal analysis applications present task-level parallelism, intensive computational hotspots and a high degree of resilience towards errors. These characteristics drive our new bio-DSP edge node architecture design composed of multiple processing cores, a Coarse-Grained Reconfigurable Array (CGRA) accelerator, and hardware-software co-design support to become resilient to a non-zero probability of bit-flips at runtime. All these characteristics enable our new bio-DSP architecture to operate with an ultra-low voltage operating point. Indeed our results indicate that the energy benefits attained from the inclusion of all these characteristics in bio-DSP architectures are more than additive: task parallelism is harnessed both at the processor and the accelerator level, and the high tolerance of the CGRA towards voltage down-scaling is exploited to further decrease the IoT edge bio-DSP system energy envelope.
Soumya Basu 0002, Loris Duch, Miguel Peón-Quirós, David Atienza 0001, Giovanni Ansaloni, Laura Pozzi 0001
ISCAS5
2017 A Synchronization-Based Hybrid-Memory Multi-Core Architecture for Energy-Efficient Biomedical Signal Processing
abstract
In the last decade, improvements on technology scaling have enabled the design of a novel generation of wearable biosensing monitors. These smart Wireless Body Sensor Nodes (WBSNs) are able to acquire and process biological signals, such as electrocardiograms, for periods of time extending from hours to days. The energy required for the on-node digital signal processing (DSP) is a crucial limiting factor in the conception of these devices. To address this design challenge, we introduce a domain-specific ultra-low power (ULP) architecture dedicated to bio-signal processing. The platform features a light-weight strategy to support different operating modes and synchronization among cores. Our approach effectively reduces the power consumption, harnessing the intrinsic parallelism and the workload requirements characterizing the target domain. Operations at low voltage levels are supported by a heterogeneous memory subsystem comprising a standard-cell based ultra-low voltage reliable partition. Experimental results show that, when executing real-world bio-signal DSP applications, a state-of-the-art multi-core architecture can improve its energy efficiency in up to 50 percent by utilizing our proposed approach, outperforming traditional single-core alternatives.
Rubén Braojos, Daniele Bortolotti, Andrea Bartolini, Giovanni Ansaloni, Luca Benini, David Atienza 0001
IEEE Trans. Computers4
2017 An Inexact Ultra-low Power Bio-signal Processing Architecture With Lightweight Error Recovery
abstract
The energy efficiency of digital architectures is tightly linked to the voltage level (Vdd) at which they operate. Aggressive voltage scaling is therefore mandatory when ultra-low power processing is required. Nonetheless, the lowest admissible Vdd is often bounded by reliability concerns, especially since static and dynamic non-idealities are exacerbated in the near-threshold region, imposing costly guard-bands to guarantee correctness under worst-case conditions. A striking alternative, explored in this paper, waives the requirement for unconditional correctness, undergoing more relaxed constraints. First, after a run-time failure, processing correctly resumes at a later point in time. Second, failures induce a limited Quality-of-Service (QoS) degradation. We focus our investigation on the practical scenario of embedded bio-signal analysis, a domain in which energy efficiency is key, while applications are inherently error-tolerant to a certain degree. Targeting a domain-specific multi-core platform, we present a study of the impact of inexactness on application-visible errors. Then, we introduce a novel methodology to manage them, which requires minimal hardware resources and a negligible energy overhead. Experimental evidence show that, by tolerating 900 errors/hour, the resulting inexact platform can achieve an efficiency increase of up to 24%, with a QoS degradation of less than 3%.
Soumya Basu 0002, Loris Duch, Rubén Braojos, Giovanni Ansaloni, Laura Pozzi 0001, David Atienza 0001
ACM Trans. Embed. Comput. Syst.4
2014 Ultra-Low Power Design of Wearable Cardiac Monitoring Systems
abstract
This paper presents the system-level architecture of novel ultra-low power wireless body sensor nodes (WBSNs) for real-time cardiac monitoring and analysis, and discusses the main design challenges of this new generation of medical devices. In particular, it highlights first the unsustainable energy cost incurred by the straightforward wireless streaming of raw data to external analysis servers. Then, it introduces the need for new cross-layered design methods (beyond hardware and software boundaries) to enhance the autonomy of WBSNs for ambulatory monitoring. In fact, by embedding more onboard intelligence and exploiting electrocardiogram (ECG) specific knowledge, it is possible to perform real-time compressive sensing, filtering, delineation and classification of heartbeats, while dramatically extending the battery lifetime of cardiac monitoring systems. The paper concludes by showing the results of this new approach to design ultra-low power wearable WBSNs in a real-life platform commercialized by SmartCardia. This wearable system allows a wide range of applications, including multi-lead ECG arrhythmia detection and autonomous sleep monitoring for critical scenarios, such as monitoring of the sleep state of airline pilots.
Rubén Braojos, Hossein Mamaghanian, Alair Dias Junior, Giovanni Ansaloni, David Atienza 0001, Francisco J. Rincón, Srinivasan Murali
DAC4
2014 Hardware/software approach for code synchronization in low-power multi-core sensor nodes
abstract
Latest embedded bio-signal analysis applications, targeting low-power Wireless Body Sensor Nodes (WBSNs), present conflicting requirements. On one hand, bio-signal analysis applications are continuously increasing their demand for high computing capabilities. On the other hand, long-term signal processing in WBSNs must be provided within their highly constrained energy budget. In this context, parallel processing effectively increases the power efficiency of WBSNs, but only if the execution can be properly synchronized among computing elements. To address this challenge, in this work we propose a hardware/software approach to synchronize the execution of bio-signal processing applications in multi-core WBSNs. This new approach requires little hardware resources and very few adaptations in the source code. Moreover, it provides the necessary flexibility to execute applications with an arbitrarily large degree of complexity and parallelism, enabling considerable reductions in power consumption for all multi-core WBSN execution conditions. Experimental results show that a multi-core WBSN architecture using the illustrated approach can obtain energy savings of up to 40%, with respect to an equivalent single-core architecture, when performing advanced bio-signal analysis.
Rubén Braojos, Ahmed Yasir Dogan, Ivan Beretta, Giovanni Ansaloni, David Atienza 0001
DATE4
2014 Power-efficient joint compressed sensing of multi-lead ECG signals
abstract
Compressed Sensing (CS) is a new acquisition-compression paradigm for low-complexity energy-aware sensing and compression. By merging both sampling and compression, CS is very promising to develop practical ultra-low power readout systems for wireless bio-signal monitoring devices, where large amounts of sensor data need to be transferred through power-hungry wireless links. Lately CS has been successfully applied for real-time energy-aware single-lead ECG compression on resource-constrained Wireless Body Sensor Network (WBSN) motes [1]. Building on our previous work, in this paper we propose a new and promising approach for joint compression of multi-lead ECG signals, where strong correlations exist between them. This situation that exhibit strong correlations, can be exploited to reduce even further amount of data to be transmitted wirelessly, thus addressing the important challenge of ultra-low-power embedded monitoring of multi-lead ECG signals.
Hossein Mamaghanian, Giovanni Ansaloni, David Atienza 0001, Pierre Vandergheynst
ICASSP2
2013 A methodology for embedded classification of heartbeats using random projections
abstract
Smart Wireless Body Sensor Nodes (WBSNs) are a novel class of unobtrusive, battery-powered devices allowing the continuous monitoring and real-time interpretation of a subject's bio-signals. One of its most relevant applications is the acquisition and analysis of Electrocardiograms (ECGs). These low-power WBSN designs, while able to perform advanced signal processing to extract information on hearth conditions of subjects, are usually constrained in terms of computational power and transmission bandwidth. It is therefore beneficial to identify in the early stages of analysis which parts of an ECG acquisition are critical and activate only in these cases detailed (and computationally intensive) diagnosis algorithms. In this paper, we introduce and study the performance of a real-time optimized neuro-fuzzy classifier based on random projections, which is able to discern normal and pathological heartbeats on an embedded WBSN. Moreover, it exposes high confidence and low computational and memory requirements. Indeed, by focusing on abnormal heartbeats morphologies, we proved that a WBSN system can effectively enhance its efficiency, obtaining energy savings of as much as 63% in the signal processing stage and 68% in the subsequent wireless transmission when the proposed classifier is employed.
Rubén Braojos, Giovanni Ansaloni, David Atienza 0001
DATE2
2013 Synchronizing code execution on ultra-low-power embedded multi-channel signal analysis platforms
abstract
Embedded biosignal analysis involves a considerable amount of parallel computations, which can be exploited by employing low-voltage and ultra-low-power (ULP) parallel computing architectures. By allowing data and instruction broadcasting, single instruction multiple data (SIMD) processing paradigm enables considerable power savings and application speedup, in turn allowing for a lower voltage supply for a given workload. The state-of-the-art multi-core architectures for biosignal analysis however lack a bare, yet smart, synchronization technique among the cores, allowing lockstep execution of algorithm parts that can be performed using the SIMD, even in the presence of data-dependent execution flows. In this paper, we propose a lightweight synchronization technique to enhance an ULP multi-core processor, resulting in improved energy efficiency through lockstep SIMD execution. Our results show that the proposed improvements accomplish tangible power savings, up to 64% for an 8-core system operating at a workload of 89 MOps/s while exploiting voltage scaling.
Ahmed Yasir Dogan, Rubén Braojos, Jeremy Constantin, Giovanni Ansaloni, Andreas Peter Burg, David Atienza 0001
DATE4
2012 Embedded real-time ECG delineation methods: A comparative evaluation
abstract
Wireless sensor nodes (WSNs) have recently evolved to include a fair amount of computational power, so that advanced signal processing algorithms can now be embedded even in these extremely low-power platforms. An increasingly successful field of application of WSNs is tele-healthcare, which enables continuous monitoring of subjects, even outside a medical environment. In particular, the design of solutions for automated and remote electrocardiogram (ECG) analysis has attracted considerable research interest in recent years, and different algorithms for delineation of normal and pathological heart rhythms have been proposed. In this paper, some of the most promising techniques for filtering and delineation of ECG signals are explored and comparatively evaluated, describing their implementation on the state-of-the-art IcyHeart WSN. The goal of this paper is to explore the trade-offs implied in the different settings and the impact of design choices for implementing “smart” WSNs dedicated to monitoring ECG bio-signals.
Rubén Braojos, Giovanni Ansaloni, David Atienza 0001, Francisco J. Rincón
BIBE2
2012 IcyHeart: Highly integrated ultra-low-power SoC solution for unobtrusive and energy efficient wireless cardiac monitoring: Research project for the benefit of specific groups (FP7, Capacities)
abstract
The objective of the IcyHeart project is to investigate and demonstrate a highly integrated and power-efficient microelectronic solution for remote monitoring of a subject's electrocardiogram (ECG) signals. A complete System-on-a-Chip (SoC) is being developed that embarks on a single chip an ultra-low-power signal acquisition front-end with analogue-to-digital converter (ADC) for ECG, a low-power digital signal processor (DSP) and a low-energy radio frequency (RF) transceiver. These features, for the first time, coexist on a single die. Energy efficient signal processing algorithms targeting ECG, and expandable to other bio-signals, are embedded and run on the on-chip DSP. The final IcyHeart product will consist of a tiny PCB embarking IcyHeart SoC and all the necessary discrete components and powering circuit. The outcome of the project is expected to generate high market value for the European SMEs developing novel cardio-monitoring products in home and professional environments, and to create high societal impact for several categories of European citizens requiring miniature, comfortable and easy-to-use wireless tele-healthcare solutions.
Marios Milis, Kyriacos Michaelides, Anastasis Kounoudes, Giovanni Ansaloni, David Atienza 0001, Frédéric Giroud, Pierre-François Ruedi, Frederic Masson
BIBE4
2012 Integrated Kernel Partitioning and Scheduling for Coarse-Grained Reconfigurable Arrays
abstract
Coarse-grained reconfigurable arrays (CGRAs) are a promising class of architectures conjugating flexibility and efficiency. Devising effective methodologies to map applications onto CGRAs is a challenging task, due to their parallel execution paradigm and constrained hardware resources. In order to handle complex applications, it is important to devise efficient strategies to partition a kernel into pieces that obey resource constraint and methodologies to schedule them on the underlying hardware. In this paper, we tackle these problems by proposing algorithms to address partitioning based on recursive searches over abstract trees. A novel scheduling strategy is also described that, leveraging differences in delays of various operations, is able to efficiently map operations on CGRA architectures. Experimental evidence on kernels derived from a diverse set of data flow graphs and EEMBC benchmarks demonstrate the efficacy of the described methods, which, when combined, achieve a higher runtime performance on a given mesh size than state-of-the-art approaches (as much as 38% for the benchmark applications considered).
Giovanni Ansaloni, Kazuyuki Tanimura, Laura Pozzi 0001, Nikil Dutt
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2011 Slack-aware scheduling on Coarse Grained Reconfigurable Arrays
abstract
Coarse Grained Reconfigurable Arrays (CGRAs) are a promising class of architectures conjugating flexibility and efficiency. Devising effective methodologies to map applications onto CGRAs is a challenging task, due to their parallel execution paradigm and sparse interconnection topology. In this paper we present a scheduling framework that is able to efficiently map operations on CGRA architectures. It leverages differences in delays of various operations, which a reconfigurable architecture always exhibits at run-time, to effectively route data. We call this ability “slack-awareness”. Experimental evidence showcases the benefit of slack-aware scheduling in a coarse-grained re-configurable environment, as more complex applications can be mapped for a given mesh size and more efficient schedules can be achieved, compared to the state of the art methods.
Giovanni Ansaloni, Laura Pozzi 0001, Kazuyuki Tanimura, Nikil Dutt
DATE1
2011 EGRA: A Coarse Grained Reconfigurable Architectural Template
abstract
Reconfigurable arrays combine the benefit of spatial execution, typical of hardware solutions, with that of programmability, present in microprocessors. When mapping software applications (or parts of them) onto hardware, however, fine-grain arrays, such as field-programmable gate arrays (FPGAs), often provide more flexibility than is needed, and do not implement coarser-level operations efficiently. Therefore, coarse grained reconfigurable arrays (CGRAs) have been proposed to this aim. Most CGRA design emerged in research present ad-hoc solutions in many aspects; in this paper we propose an architectural template to enable design space exploration of different possible CGRA designs. We called the template expression-grained reconfigurable array (EGRA), as its ability to generate complex computational cells, executing expressions as opposed to single operations, is a defining feature. Other notable EGRA characteristics include the ability to support heterogeneous cells and different storage requirements through various memory interfaces. The performed design explorations, as shown trough the experimental data provided, can effectively drive designers to further close the performance gap between reconfigurable and hardwired logic by providing guidelines on architectural design choices. Performance results on a number of embedded applications show that EGRA instances can be used as a reconfigurable fabric for customizable processors, outperforming more traditional CGRA designs.
Giovanni Ansaloni, Paolo Bonzini, Laura Pozzi 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2009 Heterogeneous coarse-grained processing elements: A template architecture for embedded processing acceleration
abstract
Reconfigurable Architectures are good candidates for application accelerators that cannot be set in stone at production time. FPGAs however, often suffer from the area and performance penalty intrinsic in gate-level reconfigurability. To reduce this overhead, coarse-grained reconfigurable arrays (CGRAs) are reconfigurable at the ALU level, but a successful design needs more than computational power-the main bottleneck usually being memory transfers. Just like the integration of hardwired multiplier and memory blocks enabled FPGAs to efficiently implement digital signal processing applications, in this paper we study a customizable architecture template based on heterogeneous processing elements (multipliers, ALU clusters and memories) that provides enough flexibility to realize fast pipelined implementations of various loop kernels on a CGRA.
Giovanni Ansaloni, Paolo Bonzini, Laura Pozzi 0001
DATE1
2008 Compiling custom instructions onto expression-grained reconfigurable architectures
abstract
While customizable processors aim at combining the flexibility of general purpose processors with the speed and power advantages of custom circuits, commercially available processors are often limited by the inability to reconfigure the application-specific features after manufacturing. Even though reconfigurable array-based accelerators are available, their performance is often unacceptable, and comes with other disadvantages such as the size of the configuration bitstream. Additionally, compilation support is limited for existing Coarse Grain Reconfigurable Arrays (CGRAs).
Paolo Bonzini, Giovanni Ansaloni, Laura Pozzi 0001
CASES2