Romain Lemaire

dblp:54/671 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Generative binary memory: Pseudo-Replay class-Incremental learning on binarized embeddings
abstract
In dynamic environments where new concepts continuously emerge, Deep Neural Networks (DNNs) must adapt by learning new classes while retaining previously acquired ones. This challenge is addressed by Class-Incremental Learning (CIL). This paper introduces Generative Binary Memory (GBM), a novel CIL pseudo-replay approach which generates synthetic binary pseudo-exemplars. Relying on Bernoulli Mixture Models (BMMs), GBM effectively models the multi-modal characteristics of class distributions, in a latent, binary space. With a specifically-designed feature binarizer, our approach applies to any conventional DNN. GBM also natively supports Binary Neural Networks (BNNs) for highly-constrained model sizes in embedded systems. The experimental results demonstrate that GBM achieves higher than state-of-the-art average accuracy on CIFAR100 ( + 2.9 % ) and TinyImageNet ( + 1.5 % ) for a ResNet-18 equipped with our binarizer. GBM also outperforms emerging CIL methods for BNNs, with + 3.1 % in final accuracy and × 4.7 memory reduction, on CORE50.
Yanis Basso-Bert, William Guicquero, Anca Mariana Molnos, Romain Lemaire, Antoine Dupret
Neural Networks4
2026 Towards Experience Replay for Class-Incremental Learning in Fully-Binary Networks
abstract
Binary Neural Networks (BNNs) are a promising approach to enable Artificial Neural Network (ANN) implementation on ultra-low power edge devices. Such devices may compute data in highly dynamic environments, in which the classes targeted for inference can evolve or even novel classes may appear, requiring continual learning. Class Incremental Learning (CIL) is an important type of continual learning for classification problems, yet it has been scarcely addressed in the context of BNNs. Furthermore, most of existing BNNs models are not fully binary, as they require several real-valued network layers, at the input, the output, and for batch normalization. This article goes a step further, enabling class incremental learning in Fully-Binarized NNs (FBNNs) through four main contributions. We firstly revisit the FBNN design and its training procedure that is suitable to CIL. Secondly, we explore loss balancing, a method to tradeoff the performance of past and current classes. Thirdly, we propose a semi-supervised method to pre-train the feature extractor of the FBNN for transferable representations. Fourthly, two conventional CIL methods, i . e ., Latent and Native replay, are thoroughly compared. These contributions are exemplified first on the CIFAR100 dataset, before being scaled up to the CORE50 continual learning benchmark. The final results based on our 3Mb FBNN on CORE50 , exhibit performance that is at par with, or better than conventional, larger, real-valued NN models.
Yanis Basso-Bert, Anca Mariana Molnos, Romain Lemaire, William Guicquero, Antoine Dupret
ACM Trans. Embed. Comput. Syst.3
2025 J3DAI: A tiny DNN-Based Edge AI Accelerator for 3D-Stacked CMOS Image Sensor
abstract
This paper presents J3DAI, a tiny deep neural network-based hardware accelerator for a 3-layer 3D-stacked CMOS image sensor featuring an artificial intelligence (AI) chip integrating a Deep Neural Network (DNN)-based accelerator. The DNN accelerator is designed to efficiently perform neural network tasks such as image classification and segmentation. This paper focuses on the digital system of J3DAI, highlighting its Performance-Power-Area (PPA) characteristics and showcasing advanced edge AI capabilities on a CMOS image sensor.To support hardware, we utilized the Aidge comprehensive software framework, which enables the programming of both the host processor and the DNN accelerator. Aidge supports post-training quantization, significantly reducing memory footprint and computational complexity, making it crucial for deploying models on resource-constrained hardware like J3DAI.Our experimental results demonstrate the versatility and efficiency of this innovative design in the field of edge AI, showcasing its potential to handle both simple and computationally intensive tasks.
Benoît Tain, Raphael Millet, Romain Lemaire, Michal Szczepanski, Laurent Alacoque, Emmanuel Pluchart, Sylvain Choisnet, Rohit Prasad, Jérôme Chossat, Pascal Pierunek, Pascal Vivet, Sébastien Thuries
ISLPED3
2024 On Class-Incremental Learning for Fully Binarized Convolutional Neural Networks
abstract
Recent advances in Binary Neural Networks (BNNs) are opening up new possibilities for disruptive hardware accelerators. This paper extends prior work on incremental learning to BNNs, by proposing a specifically-designed fully-binarized net-work and evaluating it on two learning variants, i.e., native and latent replay. The proposed BNN achieves a 53.3% test accuracy on the CIFAR-100 benchmark while relying on a binary-only arithmetic, for a 4.1Mb model size. Given a class-incremental learning experimental setup, we evaluate the influence of replay buffer size on the strategy, highlighting a turning point where latent replay offers a better classification performance than Native replay. In addition, our approach exhibits robustness against a large number of successive retrainings with an accuracy always 10% higher than a full-precision counterpart.
Yanis Basso-Bert, William Guicquero, Anca Mariana Molnos, Romain Lemaire, Antoine Dupret
ISCAS4
2021 Mont-Blanc 2020: Towards Scalable and Power Efficient European HPC Processors
abstract
The Mont-Blanc 2020 (MB2020) project has triggered the development of the next generation industrial processor for Big Data and High Performance Computing (HPC). MB2020 is paving the way to the future low-power European processor for exascale, defining the System-on-Chip (SoC) architecture and implementing new critical building blocks to be integrated in such an SoC. In this paper, we first present an overview of the MB2020 project, then we describe our experimental infrastructure, the requirements of relevant applications, and the IP blocks developed in the project. Finally, we present our emulation-based final demonstrator and explain how it integrates within our first generation of HPC processors.
Adrià Armejach, Bine Brank, Jordi Cortina, François Dolique, Timothy Hayes 0001, Nam Ho, Pierre-Axel Lagadec, Romain Lemaire, Guillem López-Paradís, Laurent Marliac, Miquel Moretó, Pedro Marcuello, Dirk Pleiter, Xubin Tan, Said Derradji
DATE8
2020 M3D-ADTCO: Monolithic 3D Architecture, Design and Technology Co-Optimization for High Energy Efficient 3D IC
abstract
Monolithic 3D (M3D) stands now as the ultimate technology to side step Moore’s Law stagnation. Due to its nanoscale Monolithic Inter-tier Via (MIV), M3D enables an ultrahigh density interconnect between Logic and Memory that is required in the field of highly energy efficient 3D integrated circuits (3D-ICs) designed for new abundant data computing systems. At design level, M3D still suffers from a lack of commercial tools, especially for Place and Route, precluding the capability to provide signoff M3D GDS. In this paper, we introduce M3D-ADTCO, an architecture, design and technology co-optimization platform aimed at providing signoff M3D GDS. It relies on a M3D Process Design Kit and the use of a commercial Place and Route tool. We demonstrate an area reduction of 23.61 % at iso performance and power compared to a 2D RISC-V micro-controller based System on Chip (SoC) while creating space to increase (2x) the RISC-V instruction memory.
Sébastien Thuries, Olivier Billoint, Sylvain Choisnet, Romain Lemaire, Pascal Vivet, Perrine Batude, Didier Lattard
DATE4
2017 A Programmable Inbound Transfer Processor for Active Messages in Embedded Multicore Systems
abstract
The "Internet of Things" requires new multicore computing devices with very high energy-efficiency. We propose an improved architecture of these embedded devices with emphasis on the efficiency of data transfers. By performing data re-organization at transport layer within the NoC infrastructure, we avoid the need for intermediate buffers for data distribution and organization. To complement "smart DMA” that structure the traffic at source side, we use a simple programmable processor to reorganize incoming data at target side. By doing this, only useful data is transported on the network, and unpacking at destination restores their structure in the most suitable way for the application, without the need of duplication. We have prototyped an inbound data processor in a MIPS-based multicore architecture. Applied on an image compression application, we save up to 33% memory footprint and divide the processing latency by a factor of 2.
Yves Durand, Christian Bernard, Romain Lemaire, César Fuguet Tortolero, Emilie Garat
DSD3
2016 MCAPI-compliant Hardware Buffer Manager mechanism to support communication in multi-core architectures
Thiago R. da Rosa, Thomas Mesquida, Romain Lemaire, Fabien Clermidy
DATE3
2013 A dynamic stream link for efficient data flow control in NoC based heterogeneous MPSoC
abstract
As Systems-on-Chip size increase, the communication costs become critical and Networks-on-Chip (NoC) bring innovative solutions. Efficient stream-based protocols over NoC have been widely studied to address dataflow communications. They are usually controlled by a set of static parameters. However, new applications, such as high-resolution video decoders, present more data-dependent behaviors forcing communication protocols to support higher dynamicity. For this purpose, we present in this paper dynamic stream links for stream-based end-to-end NoC communications by introducing two link protocols, both independent of the transfer size, allowing to improve the hardware/software control flexibility. The proposed protocols have been modeled in a MPSoC virtual platform and the hardware cost evaluated. Based on simulations, we provide guidelines to exploit these protocols according to application needs.
Claude Helmstetter, Sylvain Basset, Romain Lemaire, Fabien Clermidy, Pascal Vivet, Michel Langevin, Chuck Pilkington, Pierre G. Paulin, Didier Fuin
ASP-DAC3
2013 HW-SW integration for energy-efficient/variability-aware computing
abstract
Recent trends in embedded system architectures brought a rapid shift towards multicore, heterogeneous and reconfigurable platforms. This imposes a large effort for programmers to develop their applications to efficiently exploit the underlying architecture. In addition, process variability issues lead to performance and power uncertainties, impacting expected quality of service and energy efficiency of the running software. In particular, variability may lead to sub-optimal runtime task allocation.
Gasser Ayad, Andrea Acquaviva, Enrico Macii, Brahim Sahbi, Romain Lemaire
DATE5
2011 Reconfiguration of a 3GPP-LTE telecommunication application on a 22-core NoC-based system-on-chip
abstract
The MAGALI chip is a 65nm digital baseband dedicated to advanced telecommunication applications. Based on a 15-router mesh Network-on-Chip - NoC, it embeds 22 Processing Elements - PE performing the different functions of a complex baseband in a programmable manner. The NoC framework is based on asynchronous routers which provide a complete Globally Asynchronous Locally Synchronous -- GALS framework. Thanks to this structure, each PE is a synchronous island with a programmable frequency. On the architectural side, the NoC supports fast reconfiguration thanks to a distributed scheme. In this demonstration, we propose to show the reconfiguration capabilities of the MAGALI chip on three modes of a 3GPP-LTE receiver part, by switching on-the-fly between these three modes. The resulting transmission quality with different levels of upcoming signal noises is shown.
Fabien Clermidy, Nicolas Cassiau, N. Coste, Denis Dutoit, M. Fantini, Dimitri Ktenas, Romain Lemaire, L. Stefanizzi
NOCS7
2010 Distributed Sequencing for Resource Sharing in Multi-applicative Heterogeneous NoC Platforms
abstract
In the context of heterogeneous NoC architectures for embedded systems, it is today mandatory to support multiple applications given the plurality of standards and usages. While static reconfiguration between applications has already been extensively studied, we propose a potential increase in hardware resource usage by enabling concurrent or overlapping applications on the top of a heterogeneous NoC platform. In this paper, we describe a distributed sequencing protocol allowing hardware resource sharing between several applications. This protocol ensures correct synchronization of the processing between hardware resources without the need of a global fine-grain scheduler on the system, thus alleviating the pressure on the run-time system. The proposed protocol has been integrated and validated in a NoC-based digital baseband for 4G SDR telecom applications, and was integrated on a manufactured chip on a STMicroelectronics CMOS 65 nm LP technology.
Yvain Thonnart, Romain Lemaire, Fabien Clermidy
NOCS2
2009 An Open and Reconfigurable Platform for 4G Telecommunication: Concepts and Application
abstract
Advanced telecommunication applications require more and more flexibility. Static and run-time configuration mechanisms, as well as plug-in computing units are two solutions to solve this issue. In this paper, we propose an open architecture, dedicated to complex data-flow applications, fulfilling these requirements thanks to a distributed configuration scheme and a standard interface for data flow computing units. Based on this architecture, an open platform demonstrator has been designed. Simulation results on a 3GPP/LTE application are presented, showing a reconfiguration time overhead of only 2.6%.
Fabien Clermidy, Romain Lemaire, Xavier Popon, Dimitri Ktenas, Yvain Thonnart
DSD2
2009 Abstract Description of System Application and Hardware Architecture for Hardware/Software Code Generation
abstract
The deployment of a system application over a hardware architecture is a costly phase in the design process. This cost increases when dealing with complex applications in terms of computation requirements and exchange of data and for advanced architectures with complex and configurable communication infrastructures. The usage of abstract models for application, architecture and mapping is a key element for automatic hardware/software code generation and for the final deployment. In this paper, we present languages for abstract modeling of application, architecture, meta-mapping and mapping and we introduce a code generation flow. The use of those models allows the extraction and exploitation of architectural and application information for specific code generation to a target platform. A case study of modeling and deploying a complex 4G telecommunication application on a heterogeneous and multi core platform is presented.
Amin El Mrabti, Hamed Sheibanyrad, Frédéric Rousseau 0001, Frédéric Pétrot, Romain Lemaire, Jérôme Martin
DSD5
2009 A Communication and configuration controller for NoC based reconfigurable data flow architecture
abstract
While network-on-chip aspects such as topologies, routing strategies or quality-of-service have been largely studied, the mapping of real applications on distributed NoC-based architecture is still an open issue. In this paper, we address this issue for complex reconfigurable data-flow applications. We introduce the concept of communication and configuration controller (CCC) which interacts both with the usual network interface and the IP core structure. The proposed CCC is a programmable template-based architecture, which provides solutions to manage reconfiguration flows, data synchronizations and global control signaling. An implementation of the CCC is presented, and its performances in a 65 nm technology are discussed through a concrete telecommunication application.
Fabien Clermidy, Romain Lemaire, Yvain Thonnart, Pascal Vivet
NOCS2