Mathieu Léonardon

dblp:204/5714 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-9973-843XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Event Classification of Accelerometer Data for Industrial Package Monitoring with Embedded Deep Learning
abstract
Package monitoring is an important topic in industrial applications, with significant implications for operational efficiency and ecological sustainability. In this study, we propose an approach that employs an embedded system, placed on reusable packages, to detect their state (on a Forklift, in a Truck, or in an undetermined location). We aim to design a system with a lifespan of several years, corresponding to the lifespan of reusable packages. Our analysis demonstrates that maximizing device lifespan requires minimizing wake time. We propose a pipeline that includes data processing, training, and evaluation of the deep learning model designed for imbalanced, multi-class time series data collected from an embedded sensor. The method uses a one-dimensional Convolutional Neural Network architecture to classify accelerometer data from the IoT device. Before training, two data augmentation techniques are tested to solve the imbalance problem of the dataset: the Synthetic Minority Oversampling TEchnique and the ADAptive SYNthetic sampling approach. After training, compression techniques are implemented to have a small model size. On the considered two-class problem, the methodology yields a precision of 94.54% for the first class and 95.83% for the second class, while compression techniques reduce the model size by a factor of four. The trained model is deployed on the IoT device, where it operates with a power consumption of 316 mW during inference.
Manon Renault, Hamoud Younes, Hugo Tessier, Ronan Le Roy, Bastien Pasdeloup, Mathieu Léonardon
COMPSAC6
2025 FPGA-Oriented Design Space Exploration of a Real-Time Road Scene Semantic Segmentation Deep Neural Network
abstract
The growing interest in autonomous driving technologies requires the creation of efficient real-time systems to understand road scenes. Semantic segmentation, an essential task in computer vision, is crucial in this scenario and has become a viable solution for real-time applications largely due to deep learning models. In terms of hardware for systems with real-time constraints, embedded GPUs are a straightforward solution and provide an easy deployment. At the same time, FPGAs have proven to be more efficient for embedded machine vision tasks, especially in terms of power consumption. However, implementing large and complex models on resource-constrained devices such as FPGA and reaching a high frame rate and a low latency while preserving the semantic segmentation performance is a major challenge, due to the required trade-off between resources and accuracy. Consequently, the choice of model is not straightforward and requires joint consideration of implementation complexity and accuracy on the task. This work aims to demonstrate that, with appropriate training and FPGA-oriented redesign, a low-complexity neural network can match the performance of more complex state-of-the-art networks.
Hugo Le Blevec, Mathieu Léonardon, Stefan Weithoffer, Matthieu Arzel
FPGA2
2025 Input Resolution Downsizing as a Compression Technique for Vision Deep Learning Systems
abstract
Model compression is a critical area of research in deep learning, particularly in vision, driven by the need to lighten models memory or computational footprints. While numerous methods for model compression have been proposed, most focus on pruning, quantization, or knowledge distillation. In this work, we delve into an under-explored avenue: reducing the resolution of the input image as a complementary approach to other types of compression. By systematically investigating the impact of input resolution reduction on both classification and semantic segmentation tasks and on convnets and transformer-based architectures, we demonstrate that this strategy provides an interesting alternative for model compression. Our experimental results on standard benchmarks highlight the potential of this method, achieving competitive performance while significantly reducing computational and memory requirements. This study establishes input resolution reduction as a viable and promising direction in the broader landscape of model compression techniques for vision applications.
Jérémy Morlier, Mathieu Léonardon, Vincent Gripon
IJCNN2
2025 Bit-Width-Aware Design Environment for Few-Shot Learning on Edge AI Hardware
abstract
In this study, we propose an implementation methodology of real-time few-shot learning on tiny FPGA SoCs such as the PYNQ-Z1 board with arbitrary fixed-point bit-widths. Tensil-based conventional design environments limited hardware implementations to fixed-point bit-widths of 16 or 32 bits. To address this, we adopt the FINN framework, enabling implementations with arbitrary bit-widths. Several customizations and minor adjustments are made, including: 1.Optimization of Transpose nodes to resolve data format mismatches, 2.Addition of handling for converting the final "reduce mean" operation to Global Average Pooling (GAP). These adjustments allow us to reduce the bit-width while maintaining the same accuracy as the conventional realization, and achieve approximately twice the throughput in evaluations using CIFAR-10 dataset.
R. Kanda, Hugo Le Blevec, Naoya Onizawa, Mathieu Léonardon, Vincent Gripon, Takahiro Hanyu
ISCAS4
2025 A Prototyping Framework for P4-Programmable Traffic Managers
abstract
International audience
Karl La Grassa, André Béliveau, Mathieu Léonardon, Jean-Pierre David, Matthieu Arzel, Yvon Savaria
RSP3
2024 PEFSL: A deployment Pipeline for Embedded Few-Shot Learning on a FPGA SoC
abstract
This paper tackles the challenges of implementing few-shot learning on embedded systems, specifically FPGA SoCs, a vital approach for adapting to diverse classification tasks, especially when the costs of data acquisition or labeling prove to be prohibitively high. Our contributions encompass the development of an end-to-end open-source pipeline for a few-shot learning platform for object classification on a FPGA SoCs. The pipeline is built on top of the Tensil open-source framework, facilitating the design, training, evaluation, and deployment of DNN backbones tailored for few-shot learning. Additionally, we showcase our work's potential by building and deploying a low-power, low-latency demonstrator trained on the MiniImageNet dataset with a dataflow architecture. The proposed system has a latency of 30 ms while consuming 6.2 W on the PYNQ-Z1 board.
Lucas Grativol Ribeiro, Lubin Gauthier, Mathieu Léonardon, Jérémy Morlier, Antoine Lavrard-Meyer, Guillaume Muller 0001, Virginie Fresse, Matthieu Arzel
ISCAS3
2022 MOL-Based In-Memory Computing of Binary Neural Networks
abstract
Convolutional neural networks (CNNs) have proven very effective in a variety of practical applications involving artificial intelligence (AI). However, the layer depth of CNN deepens as user applications become more sophisticated, resulting in a huge number of operations and increased memory size. The massive amount of the produced intermediate data leads to intensive data movement between memory and computing cores causing a real bottleneck. In-memory computing (IMC) aims to address this bottleneck by directly computing inside memory, eliminating energy-intensive and time-consuming data movement. On the other hand, the emerging binary neural networks (BNNs), which is a special case of CNN, show a number of hardware-friendly properties, including memory saving. In BNN, the costly floating-point multiply-and-accumulate is replaced with lightweight bitwise XNOR and popcount operations. In this article, we propose an IMC programmable architecture targeting efficient implementation of BNN. Computational memories based on the recently introduced memristor overwrite logic (MOL) design style are employed. The architecture, which is presented in semiparallel and parallel models, efficiently executes the advanced quantization algorithm of XNOR-Net BNN. Performance evaluation based on the CIFAR-10 dataset demonstrates between$1.24\times $and$3\times $speedup and 49% and 99% energy saving compared to state-of-the-art implementations and up to 273-image/s/W throughput efficiency.
Khaled Alhaj Ali, Amer Baghdadi, Elsa Dupraz, Mathieu Léonardon, Mostafa Rizk, Jean-Philippe Diguet
IEEE Trans. Very Large Scale Integr. Syst.4
2018 Custom Low Power Processor for Polar Decoding
abstract
Cloud Radio Access Network is foreseen as one of the key features of the future 5G mobile communication standard. In this context, all the baseband processing is intended to be performed on CPUs in order to keep a high level of flexibility. The challenge is then to propose efficient software implementations of baseband processing algorithms that guarantee a sufficient throughput, while limiting the energy consumption. In this paper, as an alternative to general purpose processors, we propose an implementation of an Application Specific Instruction set Processor customized for the Successive Cancellation decoding of polar codes. The resulting software decoder achieves throughputs similar to state-of-the-art ARM processor implementations, while reducing the energy consumption by a factor 10.
Mathieu Léonardon, Camille Leroux, David Binet, J. M. Pierre Langlois, Christophe Jégo, Yvon Savaria
ISCAS1