Benoît W. Denkinger

dblp:264/9303 · also Benoît Walter Denkinger · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0002-1959-2013ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2024 SAT-Based Exact Modulo Scheduling Mapping for Resource-Constrained CGRAs
abstract
Coarse-Grain Reconfigurable Arrays (CGRAs) represent emerging low-power architectures designed to accelerate Compute-Intensive Loops (CILs). The effectiveness of CGRAs in providing acceleration relies on the quality of mapping: how efficiently the CIL is compiled onto the platform. State-of-the-Art (SoA) compilation techniques utilize modulo scheduling to minimize the Iteration Interval (II) and use graph algorithms like Max-Clique Enumeration to address mapping challenges. Our work approaches the mapping problem through a satisfiability (SAT) formulation. We introduce the Kernel Mobility Schedule (KMS), an ad hoc schedule used with the Data Flow Graph and CGRA architectural information to generate Boolean statements that, when satisfied, yield a valid mapping. Experimental results demonstrate SAT-MapIt outperforming SoA alternatives in almost 50% of explored benchmarks. Additionally, we evaluated the mapping results in a synthesizable CGRA design and emphasized the runtime metrics trends, i.e., energy efficiency and latency, across different CILs and CGRA sizes. We show that a hardware-agnostic analysis performed on compiler-level metrics can optimally prune the architectural design space, while still retaining Pareto-optimal configurations. Moreover, by exploring how implementation details impact cost and performance on real hardware, we highlight the importance of holistic software-to-hardware mapping flows, as the one presented herein.
Cristian Tirelli, Juan Sapriza, Rubén Rodríguez Álvarez, Lorenzo Ferretti, Benoît W. Denkinger, Giovanni Ansaloni, José Miranda 0001, David Atienza 0001, Laura Pozzi 0001
ACM J. Emerg. Technol. Comput. Syst.5
2023 An Open-Hardware Coarse-Grained Reconfigurable Array for Edge Computing
abstract
In this work, we propose an open-hardware low-power coarse-grained reconfigurable array connected to a lightweight microcontroller and enclosed in an application mapping framework. The latter provides complete support to configure kernels in the reconfigurable array, execute applications, and measure performance.
Rubén Rodríguez Álvarez, Benoît W. Denkinger, Juan Sapriza, José Miranda 0001, Giovanni Ansaloni, David Atienza 0001
CF2
2023 X-HEEP: An Open-Source, Configurable and Extendible RISC-V Microcontroller
abstract
X-HEEP (eXtendable Heterogeneous Energy-Efficient Platform) is an open-source1, configurable, and extensible single-core RISC-V microcontroller developed at the Embedded Systems Laboratory (ESL) of EPFL for edge-computing platforms. X-HEEP can be used standalone as a low-cost microcontroller, or it can be integrated into existing platforms to act like a peripheral subsystem, or it can be extended and customized with external peripherals and accelerators nimbly. The latter is particularly appealing for novel accelerators, memories, or peripherals designers who desire a simple controller to drive their IP and communicate with the external world using software functions. X-HEEP is built on top of existing, mature open-source IPs such as CPUs, peripherals, and many other building blocks from the OpenHW Group, the PULP team from ETH Zurich and the University of Bologna, and lowRISC. Its contribution includes its expandability, configurability, and agile use, targetting a large number of users to take one step further towards the democratization of open-source hardware.
Pasquale Davide Schiavone, Simone Machetti, Miguel Peón-Quirós, José Miranda 0001, Benoît W. Denkinger, Thomas Christoph Müller, Rubén Rodríguez Álvarez, Saverio Nasturzio, David Atienza 0001
CF5
2023 Acceleration of Control Intensive Applications on Coarse-Grained Reconfigurable Arrays for Embedded Systems
abstract
Embedded systems confront two opposite goals: low-power operation and high performance. The current trend to reach these goals is toward heterogeneous platforms, including multi-core architectures with heterogeneous cores and hardware accelerators. The latter can be divided into custom accelerators (e.g., ASICs) and programmable domain-specific cores (e.g., DSIPs). VWR2A Denkinger et al. 2022 is a programmable architecture that integrates high computational density and low power memory structures. The flexibility of VWR2A allows a large portion of applications to be covered, resulting in better performance and energy efficiency than ASICs and general-purpose processors. However, while this has been well studied for data-intensive kernels, this is not the case for control-intensive kernels —code with complex if-else and nested loop structures. Traditionally, control-intensive code is left to be executed by the host processor. This situation unnecessarily restricts the potential impact of energy-efficient acceleration, especially at the application level. In this paper, we evaluate the performance and energy consumption of VWR2A for control-intensive code and compare it with an ARM Cortex-M4 processor and a RISC-V Ibex processor. The performance and energy consumption are evaluated at the kernel and application levels. Our results confirm that VWR2A is faster and more energy-efficient than the two considered general-purpose processors also for control-intensive code.
Benoît W. Denkinger, Miguel Peón-Quirós, Mario Konijnenburg, David Atienza 0001, Francky Catthoor
IEEE Trans. Computers1
2022 VWR2A: a very-wide-register reconfigurable-array architecture for low-power embedded devices
abstract
Edge-computing requires high-performance energy-efficient embedded systems. Fixed-function or custom accelerators, such as FFT or FIR filter engines, are very efficient at implementing a particular functionality for a given set of constraints. However, they are inflexible when facing application-wide optimizations or functionality upgrades. Conversely, programmable cores offer higher flexibility, but often with a penalty in area, performance, and, above all, energy consumption. In this paper, we propose VWR2A, an architecture that integrates high computational density and low power memory structures (i.e., very-wide registers and scratchpad memories). VWR2A narrows the energy gap with similar or better performance on FFT kernels with respect to an FFT accelerator. Moreover, VWR2A flexibility allows to accelerate multiple kernels, resulting in significant energy savings at the application level.
Benoît W. Denkinger, Miguel Peón-Quirós, Mario Konijnenburg, David Atienza 0001, Francky Catthoor
DAC1
2020 Modular Design and Optimization of Biomedical Applications for Ultralow Power Heterogeneous Platforms
abstract
In the last years, remote health monitoring is becoming an essential branch of health care with the rapid development of wearable sensors technology. To meet the demand of new more complex applications and ensuring adequate battery lifetime, wearable sensors have evolved into multicore systems with advanced power-saving capabilities and additional heterogeneous components. In this article, we present an approach that applies optimization and parallelization techniques uncovered by modern ultralow power (ULP) platforms in the SW layers with the goal of improving the mapping and reducing the energy consumption of biomedical applications. Additionally, we investigate the benefit of integrating domain-specific accelerators to further reduce the energy consumption of the most computationally expensive kernels. Using 30-s excerpts of signals from two public databases, we apply the proposed optimization techniques on well-known modules of biomedical benchmarks from the state-of-the-art and two complete applications. We observe speed-ups of 5.17× and energy savings of 41.6% for the multicore implementation using a cluster of 8 cores with respect to single-core wearable sensor designs when processing a standard 12-lead electrocardiogram (ECG) signal analysis. Additionally, we conclude that the minimum workload required to take advantage of parallelization for a heartbeat classifier corresponds to the processing of 3-lead ECG signals, with a speed-up of 2.96× and energy savings of 19.3%. Moreover, we observe additional energy savings of up to 7.75% and 16.8% by applying power management and memory scaling to the multicore implementation of the 3-lead beat classifier and 12-lead ECG analysis, respectively. Finally,by integrating hardware (HW) acceleration we observe overall energy savings of up to 51.3% for the 12-lead ECG analysis.
Elisabetta De Giovanni, Fabio Montagna, Benoît W. Denkinger, Simone Machetti, Miguel Peón-Quirós, Simone Benatti, Davide Rossi 0001, Luca Benini, David Atienza 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3