EDBT 2026 Demo / reviewers in the wild / expert
Jean-Marc Philippe
dblp:28/6008
· DBLP profile ↗
13ranked-venue papers
2as first author
3since 2021 · last 2023
0000-0002-6062-8145ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | eProcessor: European, Extendable, Energy-Efficient, Extreme-Scale, Extensible, Processor EcosystemabstractThe eProcessor project aims at creating a RISC-V full stack ecosystem. The eProcessor architecture combines a high-performance out-of-order core with energy-efficient accelerators for vector processing and artificial intelligence with reduced-precision functional units. The design of this architecture follows a hardware/software co-design approach with relevant application use cases from the high-performance computing, bioinformatics and artificial intelligence domains. Two eProcessor prototypes will be developed based on two fabricated eProcessor ASICs integrated into a computer-on-module. Lluc Alvarez, Abraham Ruiz, Arnau Bigas-Soldevilla, Pavel Kuroedov, Alberto González 0004, Hamsika Mahale, Noe Bustamante, Albert Aguilera, Francesco Minervini, Javier Salamero, Oscar Palomar, Vassilis Papaefstathiou, Antonis Psathakis, Nikolaos Dimou, Michalis Giaourtas, Iasonas Mastorakis, Giorgos Ieronymakis, Georgios-Michail Matzouranis, Vassilis Flouris, Nikolaos Kossifidis, Manolis Marazakis, Bhavishya Goel, Madhavan Manivannan, Ahsen Ejaz, Panagiotis Strikos, Mateo Vázquez, Ioannis Sourdis, Pedro Trancoso, Per Stenström, Jens Hagemeyer, Lennart Tigges, Nils Kucza, Jean-Marc Philippe, Ioannis Papaefstathiou |
CF | 33 |
| 2023 | A-DECA: An Automated Design Space Exploration Approach for Computing Architectures to Develop Efficient High-Performance Many-Core ProcessorsabstractHigh-performance many-core processors have complex computing architectures with many design parameters related to different levels (CPU macro/micro-architecture, interconnect, memory, specific accelerators, etc.). Design Space Exploration (DSE) is key to tackle the challenges related to the design of such processors, especially in the early stages. This work introduces A-DECA, a highly modular DSE approach for automating the exploration of design parameters. A-DECA combines simulators, models, and exploration strategies to derive relevant objective estimations while preserving a reasonable execution time. Thus, it provides a full methodology enabling the exploration of the design space in an easy-to-use, automatic, and effective way. A-DECA is evaluated in the context of next-generation HPC processors with various applications. We combine simulation tools and analytical formulations to assess PPA (Performance, Power, and Area). Based on an efficient implementation of a multi-objective genetic algorithm for the exploration strategy, current results show a great reduction of design space optimization by around 30% compared to the initial population. A-DECA optimizes the objectives and automatically returns a set of configurations with different characteristics allowing the architect to choose the best design according to the application context. Lilia Zaourar, Alice Chillet, Jean-Marc Philippe |
DSD | 3 |
| 2021 | Analysis of on-chip communication properties in accelerator architectures for deep neural networksabstractDeep neural networks (DNNs) algorithms are expected to be core components of next-generation applications. These high performance sensing and recognition algorithms are key enabling technologies of smarter systems that make appropriate decisions about their environment. The integration of these compute-intensive and memory-hungry algorithms into embedded systems will require the use of specific energy-efficient hardware accelerators. The intrinsic parallelism of DNNs algorithms allows for the use of a large number of small processing elements, and the tight exploitation of data reuse can significantly reduce power consumption. To meet these features, many dataflow models and on-chip communication proposals have been studied in recent years. This paper proposes a comprehensive study of on-chip communication properties based on the analysis of application-specific features, such as data reuse and communication models, as well as the results of mapping these applications to architectures of different sizes. In addition, the influence of mechanisms such as broadcast and multicast on performance and energy efficiency is analyzed. This study leads to the definition of overarching features to be integrated into next-generation on-chip communication infrastructures for CNN accelerators. Hana Krichene, Jean-Marc Philippe |
NOCS | 2 |
| 2018 | PNeuro: A scalable energy-efficient programmable hardware accelerator for neural networksabstractArtificial intelligence and especially Machine Learning recently gained a lot of interest from the industry. Indeed, new generation of neural networks built with a large number of successive computing layers enables a large amount of new applications and services implemented from smart sensors to data centers. These Deep Neural Networks (DNN) can interpret signals to recognize objects or situations to drive decision processes. However, their integration into embedded systems remains challenging due to their high computing needs. This paper presents PNeuro, a scalable energy-efficient hardware accelerator for the inference phase of DNN processing chains. Simple programmable processing elements architectured in SIMD clusters perform all the operations needed by DNN (convolutions, pooling, non-linear functions, etc.). An FDSOI 28 nm prototype shows an energy efficiency of 700 GMACS/s/W at 800 MHz. These results open important perspectives regarding the development of smart energy-efficient solutions based on Deep Neural Networks. Alexandre Carbon, Jean-Marc Philippe, Olivier Bichler, Renaud Schmit, Benoît Tain, David Briand, Nicolas Ventroux, Michel Paindavoine, Olivier Brousse |
DATE | 2 |
| 2018 | Task management on fully heterogeneous micro-server system: Modeling and resolution strategiesabstractSummary Many of today's important applications of our everyday lives, eg, weather forecast, design of plane and car shapes, medical analysis, or search engine queries depend on massively parallel computer programs executed in data centers. A large amount of energy is used to power them, and it is of primary importance to compute more efficiently to sustain the increasing demand of computing power while keeping energy consumption reasonable. One promising research path in this domain is heterogeneous systems since specific computing resources (processors, accelerators, etc) are more adapted to efficiently execute parts of applications. Nevertheless, the exploitation of these platforms raises new challenges in terms of application management optimization. The aim of our work is to determine effective algorithms to exploit these heterogeneous platforms by finding appropriate application mapping and scheduling to optimize the execution time and energy consumption with respect to various constraints. To achieve this goal, there is a need of a detailed modeling of the applications and the underlying hardware to be able to find realistic solutions. In this paper, we propose such a model, provide two implementations with state‐of‐the‐art tools, and propose a fast greedy online resolution algorithm and preliminary mapping and scheduling numerical results. Lilia Zaourar, Massinissa Ait Aba, David Briand, Jean-Marc Philippe |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | The M2DC Project: Modular Microserver DataCentreabstractThe Modular Microserver DataCentre (M2DC) project will investigate, develop and demonstrate a modular, highly-efficient, cost-optimized server architecture composed of heterogeneous microserver computing resources, being able to be tailored to meet requirements from various application domains such as image processing, cloud computing or HPC. M2DC will be built on three main pillars: a flexible server architecture that can be easily customised, maintained and updated, advanced management strategies and system efficiency enhancements (SEE), well-defined interfaces to surrounding software data centre ecosystem. Mariano Cecowski, Giovanni Agosta, Ariel Oleksiak, Michal Kierzynka, Micha vor dem Berge, Wolfgang Christmann, Stefan Krupop, Mario Porrmann, Jens Hagemeyer, René Griessl, Meysam Peykanu, Lennart Tigges, Sven Rosinger, Daniel Schlitt, Christian Pieper, Carlo Brandolese, William Fornaciari, Gerardo Pelosi, Robert Plestenjak, Justin Cinkelj, Loïc Cudennec, Thierry Goubier, Jean-Marc Philippe, Udo Janssen, Chris Adeniyi-Jones |
DSD | 23 |
| 2015 | Exploration and design of embedded systems including neural algorithms
Jean-Marc Philippe, Alexandre Carbon, Olivier Brousse, Michel Paindavoine |
DATE | 1 |
| 2015 | NeuroDSP Accelerator for Face Detection ApplicationabstractNeuro-Inspired Vision approach, based on models from biology, allows to reduce the computational complexity. One of these models - The Hmax model - shows that the recognition of an object in the visual cortex mobilizes V1, V2 and V4 areas. From the computational point of view, V1 corresponds to the area of the directional filters (for example Gabor filters or wavelet filters). This information is then processed in the area V2 in order to obtain local maxima. This new information is then sent to an artificial neural network. This neural processing module corresponds to area V4 of the visual cortex and is intended to categorize objects present in the scene. In order to realize autonomous vision systems (low-power consumption) with such treatments inside, we studied a new architeure of a Neural Processor named NeuroDSP. We describe in this paper an optimized Hmax model implementation on this Neural Processor for a face detection application. Michel Paindavoine, Olivier Boisard, Alexandre Carbon, Jean-Marc Philippe, Olivier Brousse |
ACM Great Lakes Symposium on VLSI | 4 |
| 2014 | Parallel Architecture Benchmarking: From Embedded Computing to HPC, a FiPS Project PerspectiveabstractWith the growing numbers of both parallel architectures and related programming models, the benchmarking tasks become very tricky since parallel programming requires architecture-dependent compilers and languages as well as high programming expertise. More than just comparing architectures with synthetic benchmarks, benchmarking is also more and more used to design specialized systems composed of heterogeneous computing resources to optimize the performance or performance/watt ratio (e.g. embedded systems designers build System-on-Chip (SoC) out of dedicated and well-chosen components). In the High-Performance-Computing (HPC) domain, systems are designed with symmetric and scalable computing nodes built to deliver the highest performance on a wide variety of applications. However, HPC is now facing cost and power consumption issues which motivate the design of heterogeneous systems. This is one of the rationales of the European FiPS project, which proposes to develop hardware architecture and software methodology easing the design of such systems. Thus, having a fair comparison between architectures while considering an application is of growing importance. Unfortunately, porting it on all available architectures using the related programming models is impossible. To tackle this challenge, we introduced a novel methodology to evaluate and to compare parallel architectures in order to ease the work of the programmer. Based on the usage of micro benchmarks, code profiling and characterization tools, this methodology introduces a semi-automatic prediction of sequential applications performances on a set of parallel architectures. In addition, performance estimation is correlated with the cost of other criteria such as power or portability effort. Introduced for targeting vision-based embedded applications, our methodology is currently being extended to target more complex applications from HPC world. This paper extends our work with new experiments and early results on a real HPC application of DNA sequencing. Yves Lhuillier, Jean-Marc Philippe, Alexandre Guerre, Michal Kierzynka, Ariel Oleksiak |
EUC | 2 |
| 2014 | HARS: A hardware-assisted runtime software for embedded many-core architecturesabstractThe current trend in embedded computing consists in increasing the number of processing resources on a chip. Following this paradigm, cluster-based many-core accelerators with a shared hierarchical memory have emerged. Handling synchronizations on these architectures is critical since parallel implementations speed-ups of embedded applications strongly depend on the ability to exploit the largest possible number of cores while limiting task management overhead. This article presents the combination of a low-overhead complete runtime software and a flexible hardware accelerator for synchronizations called HARS (Hardware-Assisted Runtime Software). Experiments on a multicore test chip showed that the hardware accelerator for synchronizations has less than 1% area overhead compared to a cluster of the chip while reducing synchronization latencies (up to 2.8 times compared to a test-and-set implementation) and contentions. The runtime software part offers basic features like memory management but also optimized execution engines to allow the easy and efficient extraction of the parallelism in applications with multiple programming models. By using the hardware acceleration as well as a very low overhead task scheduling software technique, we show that HARS outperforms an optimized state-of-the-art task scheduler by 13% for the execution of a parallel application. Yves Lhuillier, Maroun Ojail, Alexandre Guerre, Jean-Marc Philippe, Karim Ben Chehida, Farhat Thabet, Caaliph Andriamisaina, Chafic Jaber, Raphaël David |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2013 | An efficient and flexible hardware support for accelerating synchronization operations on the STHORM many-core architectureabstractThe current trend in embedded computing consists in increasing the number of processing resources on a chip. Following this paradigm, the STMicroelectronics/CEA Platform 2012 (P2012) project designed an area- and power-efficient many-core accelerator as an answer to the needs of computing power of next-generation data-intensive embedded applications. Synchronization handling on this architecture was critical since speed-ups of parallel implementations of embedded applications strongly depend on the ability to exploit the largest possible number of cores while limiting task management overhead. This paper presents the HardWare Synchronizer (HWS), a flexible hardware accelerator for synchronization operations in the P2012 architecture. Experiments on a multi-core test chip showed that the HWS has less than 1% area overhead while reducing synchronization latencies (up to 2.8 times) and contentions. Farhat Thabet, Yves Lhuillier, Caaliph Andriamisaina, Jean-Marc Philippe, Raphaël David |
DATE | 4 |
| 2007 | On-line Routing of Reconfigurable Functions for Future Self-Adaptive Systems - Investigations within the ÆTHER ProjectabstractThe progress in hardware technologies for implementing portable, low power and low cost electronic systems for consumer products has been major the last years. The complexity of embedded systems will further increase at a rate which is not met by the development of advanced CAD tools for managing the large design space. This will likely lead to increased design problems regarding system implementation, test and verification. In the next 15-20 years, it is likely that the consumer products are based on computing devices which are grouped together in networks including thousands or even millions of nodes. The ÆTHER project deals with managing the complexity of such systems based on emerging technologies for future applications. This paper presents how the design complexity can be managed at the hardware level by integrating self-adaptive characteristics, and how the trade-off in performance and flexibility can be optimized to fulfill all application requirements while reducing the design complexity. Katarina Paulsson, Michael Hübner 0001, Jürgen Becker 0001, Jean-Marc Philippe, Christian Gamrat |
FPL | 4 |
| 2006 | An energy-efficient ternary interconnection link for asynchronous systemsabstractWe introduce a new ternary link including a binary-to-ternary encoder and a ternary-to-binary decoder in voltage-mode multiple-valued logic (MVL). This link improves the transistor count compared to existing designs and it has no DC current path. The complete link was simulated with SPICE and a 0.13mum CMOS technology. It additionally shows interesting advantages on power consumption for global interconnects compared to full-swing signaling binary systems (up to 56.4% less energy consumption). Its low propagation delay is also an advantage in the design of high-speed on-chip links for asynchronous systems Jean-Marc Philippe, E. Kinvi-Boh, Sébastien Pillement, Olivier Sentieys |
ISCAS | 1 |