EDBT 2026 Demo / reviewers in the wild / expert
Germain Haugou
dblp:05/2743
· DBLP profile ↗
8ranked-venue papers
0as first author
2since 2021 · last 2021
0009-0008-3211-1254ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 since 2021Software engineering, systems software and programming languages · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 38% Processor architecture and microarchitecture · 30% Parallel and multicore computing · 11% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › clustered architecture
cluster-based multiprocessor |
0.5 | 1 | 2021 | Energy-Efficient Hardware-Accelerated Synchronization for Shared-L1-Memory Multiprocessor Clusters · IEEE Trans. Parallel Distributed Syst. 2021 |
Energy-efficient computing › voltage scaling
near-threshold computing |
0.1 | 1 | 2021 | Energy-Efficient Hardware-Accelerated Synchronization for Shared-L1-Memory Multiprocessor Clusters · IEEE Trans. Parallel Distributed Syst. 2021 |
Hardware accelerators and domain-specific architectures
many-core accelerator |
0.1 | 1 | 2012 | Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012 |
Parallel and multicore computing › many-core systems
many-core computing |
0.1 | 1 | 2012 | Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012 |
GPUs and heterogeneous computing › heterogeneous programming models
OpenCL |
0.0 | 1 | 2012 | Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2012 | Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012 |
Methods — techniques the papers use, named apart from their topics
hardware-accelerated synchronization · 0.5fine-grain power management · 0.5OpenCV · 0.1OpenCL · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | GVSoC: A Highly Configurable, Fast and Accurate Full-Platform Simulator for RISC-V based IoT ProcessorsabstractThe last few years have seen the emergence of IoT processors: ultra-low power systems-on-chips (SoCs) combining lightweight and flexible micro-controller units (MCUs), often based on open-ISA RISC-V cores, with application-specific accelerators to maximize performance and energy efficiency. Overall, this heterogeneity level requires complex hardware and a full-fledged software stack to orchestrate the execution and exploit platform features. For this reason, enabling agile design space exploration becomes a crucial asset for this new class of low-power SoCs. In this scenario, high-level simulators play an essential role in breaking the speed and design effort bottlenecks of cycle-accurate simulators and FPGA prototypes, respectively, while preserving functional and timing accuracy. We present GVSoC, a highly configurable and timing-accurate event-driven simulator that combines the efficiency of C++ models with the flexibility of Python configuration scripts. GVSoC is fully open-sourced, with the intent to drive future research in the area of highly parallel and heterogeneous RISC-V based IoT processors, leveraging three foundational features: Python-based modular configuration of the hardware description, easy calibration of platform parameters for accurate performance estimation, and high-speed simulation. Experimental results show that GVSoC enables practical functional and performance analysis and design exploration at the full-platform level (processors, memory, peripherals and IOs) with a speed-up of 2500× with respect to cycle-accurate simulation with errors typically below 10% for performance analysis. Nazareno Bruschi, Germain Haugou, Giuseppe Tagliavini, Francesco Conti 0001, Luca Benini, Davide Rossi 0001 |
ICCD | 2 |
| 2021 | Energy-Efficient Hardware-Accelerated Synchronization for Shared-L1-Memory Multiprocessor ClustersabstractThe steeply growing performance demands for highly power- and energy-constrained processing systems such as end-nodes of the Internet-of-Things (IoT) have led to parallel near-threshold computing (NTC), joining the energy-efficiency benefits of low-voltage operation with the performance typical of parallel systems. Shared-L1-memory multiprocessor clusters are a promising architecture, delivering performance in the order of GOPS and over 100 GOPS/W of energy-efficiency. However, this level of computational efficiency can only be reached by maximizing the effective utilization of the processing elements (PEs) available in the clusters. Along with this effort, the optimization of PE-to-PE synchronization and communication is a critical factor for performance. In this article, we describe a light-weight hardware-accelerated synchronization and communication unit (SCU) for tightly-coupled clusters of processors. We detail the architecture, which enables fine-grain per-PE power management, and its integration into an eight-core cluster of RISC-V processors. To validate the effectiveness of the proposed solution, we implemented the eight-core cluster in advanced 22 nm FDX technology and evaluated performance and energy-efficiency with tunable microbenchmarks and a set of real-life applications and kernels. The proposed solution allows synchronization-free regions as small as 42 cycles, over 41× smaller than the baseline implementation based on fast test-and-set access to L1 memory when constraining the microbenchmarks to 10 percent synchronization overhead. When evaluated on the real-life DSP-applications, the proposed SCU improves performance by up to 92 and 23 percent on average and energy efficiency by up to 98 and 39 percent on average. Florian Glaser, Giuseppe Tagliavini, Davide Rossi 0001, Germain Haugou, Qiuting Huang, Luca Benini |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2019 | Hardware-Accelerated Energy-Efficient Synchronization and Communication for Ultra-Low-Power Tightly Coupled ClustersabstractParallel ultra low power computing is emerging as an enabler to meet the growing performance and energy efficiency demands in deeply embedded systems such as the end-nodes of the internet-of-things (IoT). The parallel nature of these systems however adds a significant degree of complexity as processing elements (PEs) need to communicate in various ways to organize and synchronize execution. Naive implementations of these central and non-trivial mechanisms can quickly jeopardize overall system performance and limit the achievable speedup and energy efficiency. To avoid this bottleneck, we present an event-based solution centered around a technology-independent, light-weight and scalable (up to 16 cores) synchronization and communication unit (SCU) and its integration into a shared-memory multicore cluster. Careful design and tight coupling of the SCU to the data interfaces of the cores allows to execute common synchronization procedures with a single instruction. Furthermore, we present hardware support for the common barrier and lock synchronization primitives with a barrier latency of only eleven cycles, independent of the number of involved cores. We demonstrate the efficiency of the solution based on experiments with a post-layout implementation of the multicore cluster in a 22 nm CMOS process where the SCU constitutes less than 2 % of area overhead. Our solution supports parallel sections as small as 100 or 72 cycles with a synchronization overhead of just 10 %, an improvement of up to 14× or 30× with respect to cycle count or energy, respectively, compared to a test-and-set based implementation. Florian Glaser, Germain Haugou, Davide Rossi 0001, Qiuting Huang, Luca Benini |
DATE | 2 |
| 2018 | An Open-Source Verification Framework for Open-Source Cores: A RISC-V Case StudyabstractThe complexity and heterogeneity of digital devices used in embedded systems is increasing everyday and delivering a bug-free design is still a very complex task. The interest for open-source hardware in real products is demanding for tools and advanced methodologies for verification to provide high reliability to open and free IPs. In this work, an open-source evolutionary optimizer has been used to create functional test programs that improve the verification test set for an open-source microprocessor, enhancing in this way, the verification level of the device. The verification programs are generated to optimize code coverage metrics and are tested against a high-level model to find device incorrectnesses during the generation time. A perturbation mechanism has been included in the verification framework to cover parts of the device under verification not reachable with only software stimuli such as interrupts or memory stalls. The proposed methodology uncovered 10 bugs still present in the RTL description of the analyzed device and demonstrated the effectiveness of open-source verification tools for the next generation of open-source RISC-V microprocessors. Pasquale Davide Schiavone, Ernesto Sánchez 0001, Annachiara Ruospo, Francesco Minervini, Florian Zaruba, Germain Haugou, Luca Benini |
VLSI-SoC | 6 |
| 2016 | Enabling OpenVX support in mW-scale parallel acceleratorsabstractmW-scale parallel accelerators are a promising target for application domains such as the Internet of Thing (IoT), which require a strong compliance with a limited power budget combined with high performance capabilities. An important use case is given by smart sensing devices featuring increasingly sophisticated vision capabilities, at the cost of an increasing amount of near-sensor computation power. OpenVX is an emerging standard for the embedded vision, and provides a C-based application programming interface and a runtime environment. OpenVX is designed to maximize functional and performance portability across diverse hardware platforms. However, state-of-the-art implementations rely on memory-hungry data structures, which cannot be supported in constrained devices. In this paper we propose an alternative and novel approach to provide OpenVX support in mW-scale parallel accelerators. Our main contributions are: (i) an extension to the original OpenVX model to support static management of application graphs in the form of binary files; (ii) the definition of a companion runtime environment providing a lightweight support to execute binary graphs in a resource-constrained environment. Our approach achieves 68% memory footprint reduction and 3× execution speed-up compared to a baseline implementation. At the same time, data memory bandwidth is reduced by 10% and energy efficiency is improved by 2×. Giuseppe Tagliavini, Germain Haugou, Andrea Marongiu, Luca Benini |
CASES | 2 |
| 2015 | A framework for optimizing OpenVX applications performance on embedded manycore acceleratorsabstractNowadays Computer Vision application are ubiquitous, and their presence on embedded devices is more and more widespread. Heterogeneous embedded systems featuring a clustered manycore accelerator are a very promising target to execute embedded vision algorithms, but the code optimization for these platforms is a challenging task. Moreover, designers really need support tools that are both fast and accurate. In this work we introduce ADRENALINE, an environment for development and optimization of OpenVX applications targeting manycore accelerators. ADRENALINE consists of a custom OpenVX run-time and a virtual platform, and overall it is intended to provide support to enhance performance of embedded vision applications. Giuseppe Tagliavini, Germain Haugou, Andrea Marongiu, Luca Benini |
SCOPES | 2 |
| 2013 | Improving simulation speed and accuracy for many-core embedded platforms with ensemble modelsabstractIn this paper, we introduce a novel modeling technique to reduce the time associated with cycle-accurate simulation of parallel applications deployed on many-core embedded platforms. We introduce an ensemble model based on artificial neural networks that exploits (in the training phase) multiple levels of simulation abstraction, from cycle-accurate to cycle-approximate, to predict the cycle-accurate results for unknown application configurations. We show that high-level modeling can be used to significantly reduce the number of low-level model evaluations provided that a suitable artificial neural network is used to aggregate the results. We propose a methodology for the design and optimization of such an ensemble model and we assess the proposed approach for an industrial simulation framework based on STMicroelectronics STHORM (P2012) many-core computing fabric. Edoardo Paone, Nazanin Vahabi, Vittorio Zaccaria, Cristina Silvano, Diego Melpignano, Germain Haugou, Thierry Lepley |
DATE | 6 |
| 2012 | Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applicationsabstractP2012 is an area- and power-efficient many-core computing accelerator based on multiple globally asynchronous, locally synchronous processor clusters. Each cluster features up to 16 processors with independent instruction streams sharing a multi-banked one-cycle access L1 data memory, a multi-channel DMA engine and specialized hardware for synchronization and aggressive power management. P2012 is 3D stacking ready and can be customized to achieve extreme area and energy efficiency by adding domain-specific HW IPs to the cluster. The first P2012 SoC prototype in 28nm CMOS will sample in Q3, featuring four 16-processor clusters, a 1MB L2 memory and delivering 80GOPS (with 32 bit single precision floating point support) in 18mm2 with 2W power consumption (worst-case). P2012 can run standard OpenCL™ and proprietary Native Programming Model SW components to achieve the highest level of control on application-to-resource mapping. A dedicated version of the OpenCV vision library is provided in the P2012 SW Development Kit to enable visual analytics acceleration. This paper will discuss preliminary performance measurements of common feature extraction and tracking algorithms, parallelized on P2012, versus sequential execution on ARM CPUs. Diego Melpignano, Luca Benini, Eric Flamand, Bruno Jego, Thierry Lepley, Germain Haugou, Fabien Clermidy, Denis Dutoit |
DAC | 6 |