EDBT 2026 Demo / reviewers in the wild / expert
Denis Dutoit
dblp:85/9944
· DBLP profile ↗
11ranked-venue papers
1as first author
1since 2021 · last 2025
0000-0001-7786-0947ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 50% Hardware accelerators and domain-specific architectures · 38% GPUs and heterogeneous computing · 12% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
many-core accelerator |
0.1 | 1 | 2012 | Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012 |
Parallel and multicore computing › many-core systems
many-core computing |
0.1 | 1 | 2012 | Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012 |
GPUs and heterogeneous computing › heterogeneous programming models
OpenCL |
0.0 | 1 | 2012 | Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2012 | Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applications · DAC 2012 |
Methods — techniques the papers use, named apart from their topics
OpenCV · 0.1OpenCL · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NET4EXA: Pioneering the Future of Interconnects for Supercomputing and AIabstractNET4EXA aims to develop a next-generation high-performance interconnect for HPC and AI systems, addressing the increasing demands of large-scale infrastructures, such as those required for training Large Language Models. Building upon the proven BXI (Bull eXascale Interconnect) European technology used in TOP15 supercomputers, NET4EXA will deliver the new BXI release, BXIv3, a complete hardware and software interconnect solution, including switch and network interface components. The project will integrate a fully functional pilot system at TRL 8, ready for deployment into upcoming exascale and post-exascale systems from 2025 onward. Leveraging prior research from European initiatives like RED-SEA, the previous achievements of consortium partners and over 20 years of expertise from BULL, NET4EXA also lays the groundwork for the future generation of BXI, BXIv4, providing analysis and preliminary design. The project will use a hybrid development and co-design approach, combining commercial switch technology with custom IP and FPGA-based NICs. Performances of NET4EXA BXIv3 interconnect will be evaluated using a broad portfolio of benchmarks, scientific scalable applications, and AI workloads. Michele Martinelli, Roberto Ammendola, Andrea Biagioni, Carlotta Chiarini, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Pier Stanislao Paolucci, Elena Pastorelli, Pierpaolo Perticaroli, Luca Pontisso, Cristian Rossi, Francesco Simula, Piero Vicini, David Colin, Gregoire Pichon, Alexandre Louvet, John Gliksberg, Matteo Turisini, Andrea Monterubbiano, Jean-Philippe Nomine, Denis Dutoit, Hugo Taboada, Lilia Zaourar, Mohamed Benazouz, Angelos Bilas, Fabien Chaix, Manolis Katevenis, Nikolaos Chrysos, Evangelos Mageiropoulos, Christos Kozanitis, Thomas Moen, Steffen Persvold, Einar Rustad, Sandro Fiore, Fabrizio Granelli, Simone Pezzuto, Raffaello Potestio, Luca Tubiana, Philippe Velha, Flavio Vella, Daniele De Sensi, Salvatore Pontarelli |
DSD | 23 |
| 2020 | POPSTAR: a Robust Modular Optical NoC Architecture for Chiplet-based 3D Integrated SystemsabstractSilicon photonics technology is now gaining maturity with increasing levels of design complexity from devices to large photonic integrated circuits. Close integration of control electronics with 3D assembly of photonics and CMOS opens the way to high-performance computing architectures partitioned in chiplets connected by optical NoC on silicon photonic interposers. In this paper, we give an overview of our works on optical links and NoC for manycore systems, from low-level control of photonic devices to high-level system optimization of the optical communications. We detail the POPSTAR optical NoC topology and architecture (Processors On Photonic Silicon interposer Terascale ARchitecture) with electro-optical interface chiplets, the corresponding nested spiral topology for single-writer multiple- reader links and the associated control electronics, in charge of high-speed drivers, thermal stabilization and handling of the protocol stack, from data integrity to flow-control, routing and arbitration of the optical communications. The strengths and opportunities for this architecture will be discussed, with a shift in system & implementation constraints with respect to previous optical NoC proposals, and new challenges to be addressed. Yvain Thonnart, Stéphane Bernabé, Jean Charbonnier, Christian Bernard, David Coriat, César Fuguet Tortolero, Pierre Tissier, Benoît Charbonnier, Stephane Malhouitre, Damien Saint-Patrice, Myriam Assous, Aditya Narayan, Ayse K. Coskun, Denis Dutoit, Pascal Vivet |
DATE | 14 |
| 2017 | Paving the Way Towards a Highly Energy-Efficient and Highly Integrated Compute Node for the Exascale Revolution: The ExaNoDe ApproachabstractPower consumption and high compute density are the key factors to be considered when building a compute node for the upcoming Exascale revolution. Current architectural design and manufacturing technologies are not able to provide the requested level of density and power efficiency to realise an operational Exascale machine. A disruptive change in the hardware design and integration process is needed in order to cope with the requirements of this forthcoming computing target. This paper presents the ExaNoDe H2020 research project aiming to design a highly energy efficient and highly integrated heterogeneous compute node targeting Exascale level computing, mixing low-power processors, heterogeneous co-processors and using advanced hardware integration technologies with the novel UNIMEM Global Address Space memory system. Alvise Rigo, Christian Pinto, Kevin Pouget, Daniel Raho, Denis Dutoit, Pierre-Yves Martinez, Chris Doran, Luca Benini, Iakovos Mavroidis, Manolis Marazakis, Valeria Bartsch, Guy Lonsdale, Antoniu Pop, John Goodacre, Annaik Colliot, Paul M. Carpenter, Petar Radojkovic, Dirk Pleiter, Dominique Drouin, Benoît Dupont de Dinechin |
DSD | 5 |
| 2014 | Thermal analysis and model identification techniques for a logic + WIDEIO stacked DRAM test chipabstractHigh temperature is one of the limiting factors and major concerns in 3D-chip integration. In this paper we use a 3D test chip (WIDEIO DRAM on top of a logic die) equipped with temperature sensors and heaters to explore thermal effects. We correlated real temperature measurements with the power dissipated by the heaters using model learning techniques. The resulting compact thermal model is able to predict temperatures at chip locations far from the temperature sensors and to infer the power dissipation at any location of the chip. Results are verified by mean of an off-sample validation technique and show a high accuracy of the compact thermal model when compared with silicon measurements. Francesco Beneventi, Andrea Bartolini, Pascal Vivet, Denis Dutoit, Luca Benini |
DATE | 4 |
| 2014 | EUROSERVER: Energy Efficient Node for European Micro-ServersabstractEUROSERVER is a collaborative project that aims to dramatically improve data centre energy-efficiency, cost, and software efficiency. It is addressing these important challenges through the coordinated application of several key recent innovations: 64-bit ARM cores, 3D heterogeneous silicon-on-silicon integration, and fully-depleted silicon-on-insulator (FD SOI) process technology, together with new software techniques for efficient resource management, including resource sharing and workload isolation. We are pioneering a system architecture approach that allows specialized silicon devices to be built even for low-volume markets where NRE costs are currently prohibitive. The EUROSERVER device will embed multiple silicon "chiplets" on an active silicon interposer. Its system architecture is being driven by requirements from three use cases: data centres and cloud computing, telecom infrastructures, and high-end embedded systems. We will build two fully integrated full-system prototypes, based on a common micro-server board, and targeting embedded servers and enterprise servers. Yves Durand, Paul M. Carpenter, Stefano Adami, Angelos Bilas, Denis Dutoit, Alexis Farcy, Georgi Gaydadjiev, John Goodacre, Manolis Katevenis, Manolis Marazakis, Emil Matús, Iakovos Mavroidis, John Thomson |
DSD | 5 |
| 2013 | 3D integration for power-efficient computingabstract3D stacking is currently seen as a breakthrough technology for improving bandwidth and energy efficiency in multi-core architectures. The expectation is to solve major issues such as external memory pressure and latency while maintaining reasonable power consumption. In this paper, we show some advances in this field of research, starting with memory interface solutions as WIDEIO experience on a real chip for solving DRAM accesses issue. We explain the integration of a 512-bit memory interface in a Network-on-Chip multi-core framework and we show the performance we can achieve, these results being based on a 65nm prototype integrating 10µm diameter Through Silicon Vias. We then present the potentiality of new fine grain 3D stacking technology for power-efficient memory hierarchy. We expose an innovative 3D stacked multi-cache strategy aimed at lowering memory latency and external memory bandwidth requirements and thus demonstrating the efficiency of 3D stacking to rethink architectures for obtaining unequalled performances in power efficiency. Denis Dutoit, Eric Guthmuller, Ivan Miro Panades |
DATE | 1 |
| 2013 | 3D stacking for multi-core architectures: From WIDEIO to distributed cachesabstract3D stacking has been viewed as a breakthrough solution for increasing performance in multi-core architectures. The hope is to solve some of the main issues in current multi-core architectures: external memory pressure and latency; I/O bottleneck; communication power consumption. In this paper, some advances of this field of research are shown, starting with a WIDEIO experience on a real chip for solving DRAM accesses issue. The integration of a 512 bit-width bus is demonstrated in a Network-on-Chip (NoC) multi-core framework and the resulting performance based on a 65nm prototype with 10μm diameter Through Silicon Vias (TSV). The potentiality of 3D scaling thanks to 3D asynchronous Network-on-Chip implementation is then shown. Finally, an innovative 3D stacked distributed cache strategy aimed at lowering memory latency and external memory bandwidth requirements is presented. This new memory partitioning demonstrates the efficiency of 3D stacking to rethink architectures for addressing multi-core scaling challenges. Fabien Clermidy, Denis Dutoit, Eric Guthmuller, Ivan Miro Panades, Pascal Vivet |
ISCAS | 2 |
| 2012 | Platform 2012, a many-core computing accelerator for embedded SoCs: performance evaluation of visual analytics applicationsabstractP2012 is an area- and power-efficient many-core computing accelerator based on multiple globally asynchronous, locally synchronous processor clusters. Each cluster features up to 16 processors with independent instruction streams sharing a multi-banked one-cycle access L1 data memory, a multi-channel DMA engine and specialized hardware for synchronization and aggressive power management. P2012 is 3D stacking ready and can be customized to achieve extreme area and energy efficiency by adding domain-specific HW IPs to the cluster. The first P2012 SoC prototype in 28nm CMOS will sample in Q3, featuring four 16-processor clusters, a 1MB L2 memory and delivering 80GOPS (with 32 bit single precision floating point support) in 18mm2 with 2W power consumption (worst-case). P2012 can run standard OpenCL™ and proprietary Native Programming Model SW components to achieve the highest level of control on application-to-resource mapping. A dedicated version of the OpenCV vision library is provided in the P2012 SW Development Kit to enable visual analytics acceleration. This paper will discuss preliminary performance measurements of common feature extraction and tracking algorithms, parallelized on P2012, versus sequential execution on ARM CPUs. Diego Melpignano, Luca Benini, Eric Flamand, Bruno Jego, Thierry Lepley, Germain Haugou, Fabien Clermidy, Denis Dutoit |
DAC | 8 |
| 2011 | 3D Embedded multi-core: Some perspectivesabstract3D technologies using Through Silicon Vias (TSV) have not yet proved their viability for being deployed in large-range products. In this paper, we investigate three promising perspectives for short to medium terms adoption of such technology in high-end System-on-Chip built around multi-core architectures: the wide bus concept will help solving high bandwidth requirements with external memory. 3D Network-on-Chip is a promising solution for increased modularity and scalability. We show that an efficient implementation provides an available bandwidth outperforming classical interfaces. Finally, we put in perspective the active interposer concept which aims at simplifying and improving power, test and debug management. Fabien Clermidy, Florian Darve, Denis Dutoit, Walid Lafi, Pascal Vivet |
DATE | 3 |
| 2011 | Reconfiguration of a 3GPP-LTE telecommunication application on a 22-core NoC-based system-on-chipabstractThe MAGALI chip is a 65nm digital baseband dedicated to advanced telecommunication applications. Based on a 15-router mesh Network-on-Chip - NoC, it embeds 22 Processing Elements - PE performing the different functions of a complex baseband in a programmable manner. The NoC framework is based on asynchronous routers which provide a complete Globally Asynchronous Locally Synchronous -- GALS framework. Thanks to this structure, each PE is a synchronous island with a programmable frequency. On the architectural side, the NoC supports fast reconfiguration thanks to a distributed scheme. In this demonstration, we propose to show the reconfiguration capabilities of the MAGALI chip on three modes of a 3GPP-LTE receiver part, by switching on-the-fly between these three modes. The resulting transmission quality with different levels of upcoming signal noises is shown. Fabien Clermidy, Nicolas Cassiau, N. Coste, Denis Dutoit, M. Fantini, Dimitri Ktenas, Romain Lemaire, L. Stefanizzi |
NOCS | 4 |
| 2011 | 3D NoC using through silicon Via: An asynchronous implementationabstract3D stacking is seen as one of the most interesting technologies for System-on-Chip (SoC) developments. However, 3D technologies using Through Silicon Vias (TSV) have not yet proved their viability for being deployed in large-range of products. In this paper, we are investigating 3D Network-on-Chip has a promising solution for increased modularity and scalability. We show that an efficient implementation based on asynchronous logic provides an available bandwidth of 64GB/s for only 700 TSV, outperforming classical interfaces while simplifying the assembly process. We also point out the benefit in terms of power consumption for these new interfaces with a gain of 5 times compared to classical LPDDR2 interfaces. Pascal Vivet, Denis Dutoit, Yvain Thonnart, Fabien Clermidy |
VLSI-SoC | 2 |