EDBT 2026 Demo / reviewers in the wild / expert
Matthieu Arzel
dblp:16/1399
· DBLP profile ↗
17ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-8774-7662ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LEASARD: Low-Energy Deep Neural Networks for Autonomous Search-and-Rescue DronesabstractInternational audience Panagiotis Papadakis, Isabelle Fantoni, Jean-Philippe Diguet, Matthieu Arzel |
COMPSAC | 4 |
| 2025 | FPGA-Oriented Design Space Exploration of a Real-Time Road Scene Semantic Segmentation Deep Neural NetworkabstractThe growing interest in autonomous driving technologies requires the creation of efficient real-time systems to understand road scenes. Semantic segmentation, an essential task in computer vision, is crucial in this scenario and has become a viable solution for real-time applications largely due to deep learning models. In terms of hardware for systems with real-time constraints, embedded GPUs are a straightforward solution and provide an easy deployment. At the same time, FPGAs have proven to be more efficient for embedded machine vision tasks, especially in terms of power consumption. However, implementing large and complex models on resource-constrained devices such as FPGA and reaching a high frame rate and a low latency while preserving the semantic segmentation performance is a major challenge, due to the required trade-off between resources and accuracy. Consequently, the choice of model is not straightforward and requires joint consideration of implementation complexity and accuracy on the task. This work aims to demonstrate that, with appropriate training and FPGA-oriented redesign, a low-complexity neural network can match the performance of more complex state-of-the-art networks. Hugo Le Blevec, Mathieu Léonardon, Stefan Weithoffer, Matthieu Arzel |
FPGA | 4 |
| 2025 | A Prototyping Framework for P4-Programmable Traffic ManagersabstractInternational audience Karl La Grassa, André Béliveau, Mathieu Léonardon, Jean-Pierre David, Matthieu Arzel, Yvon Savaria |
RSP | 5 |
| 2024 | PEFSL: A deployment Pipeline for Embedded Few-Shot Learning on a FPGA SoCabstractThis paper tackles the challenges of implementing few-shot learning on embedded systems, specifically FPGA SoCs, a vital approach for adapting to diverse classification tasks, especially when the costs of data acquisition or labeling prove to be prohibitively high. Our contributions encompass the development of an end-to-end open-source pipeline for a few-shot learning platform for object classification on a FPGA SoCs. The pipeline is built on top of the Tensil open-source framework, facilitating the design, training, evaluation, and deployment of DNN backbones tailored for few-shot learning. Additionally, we showcase our work's potential by building and deploying a low-power, low-latency demonstrator trained on the MiniImageNet dataset with a dataflow architecture. The proposed system has a latency of 30 ms while consuming 6.2 W on the PYNQ-Z1 board. Lucas Grativol Ribeiro, Lubin Gauthier, Mathieu Léonardon, Jérémy Morlier, Antoine Lavrard-Meyer, Guillaume Muller 0001, Virginie Fresse, Matthieu Arzel |
ISCAS | 8 |
| 2022 | DWT Collusion Resistant Video Watermarking Using Tardos Family CodesabstractA fingerprinting process is an efficient means of protecting multimedia content and preventing illegal distribution. The goal is to find individuals who were engaged in the production and illicit distribution of a multimedia product. We investigated discrete wavelet transform (DWT) based blind video watermarking strategy tied with probabilistic fingerprinting codes to avoid collusion among higher-resolution videos. We used FFmpeg to run a variety of collusion attacks (e.g., averaging, darkening, and lighten) on high resolution video and compared the most often suggested code generator and decoders in the literature to find at least one colluder within the necessary code length. The Laarhoven codes generator and nearest neighbor search (NNS) decoder outperforms all other suggested generators and decoders in the literature in terms of computational time, colluder detection and resources. Gaëtan Le Guelvouit, Jean Dion, Frédéric Guilloud, Matthieu Arzel |
IPAS | 5 |
| 2018 | A fully flexible circuit implementation of clique-based neural networks in 65-nm CMOSabstractClique-based neural networks implement low-complexity functions working with a reduced connectivity between neurons. Thus, they address very specific applications operating with a very low energy budget. This paper proposes a flexible and iterative neural architecture able to implement multiple types of clique-based neural networks of up to 3968 neurons. The circuit has been integrated in a ST 65-nm CMOS ASIC and validated in the context of ECG classification. The network core reacts in 83ns to a stimulation and occupies a 0.21mm2silicon area. Benoit Larras, Paul Chollet, Cyril Lahuec, Fabrice Seguin, Matthieu Arzel |
ISCAS | 5 |
| 2018 | 40 Gop/s/mm2 fixed-point operators for Brain Computer Interface in 65 nm CMOSabstractThe performance of non-invasive Brain-Computer Interface (BCI) depends on the computing performance of the system which solves the inverse problem. So the number of basic operations computed per second determines the BCI's resolution. An architecture with pipelined and parallelized flow is then required, and each operator in this architecture must be optimised to reach the highest possible computing performance. This paper presents the implementation of a fixed-point reciprocal and an inverse square root operators for the STMicroelectronics 65 nm CMOS technology. This paper follows previous works that optimise these operators on FPGA target. Each operator reaches a computing performance of about 40 Gop/s/mm2, which improves the literature results by a factor of 5. Thus, this works fits well for portable and high performance BCI applications. Erwan Libessart, Matthieu Arzel, Cyril Lahuec, Francesco Paolo Andriulli |
ISCAS | 2 |
| 2017 | A 65-nm CMOS 7fJ per synaptic event clique-based neural network in scalable architectureabstractTo operate under severe energy constraints, clique-based neural networks are good candidates. They benefit from a reduced exchange of information between low-complexity processing units with no performance degradation. This paper proposes a modular, flexible and scalable architecture validated by an ST 65-nm CMOS ASIC implementation for a 30-neuron clique-based neural network circuit. With 0.8V power supply, 150nA unitary current and a low performance degradation, the neuron energy consumption is reduced to only 7fJ per synaptic event. The network occupies a 41,820μm2 silicon area. Benoit Larras, Paul Chollet, Cyril Lahuec, Fabrice Seguin, Matthieu Arzel |
ISCAS | 5 |
| 2017 | A Scaling-Less Newton-Raphson Pipelined Implementation for a Fixed-Point Reciprocal OperatorabstractThe reciprocal is a widespread operation in digital signal processing architectures. A usual method consists in using the Newton-Raphson algorithm or its derivatives, either in floating or in fixed-point formats. With the former format, the standardized format of the mantissa makes the implementation easier, but for the fixed-point format there are many possibilities. This forces a design with scaling of the input in order to respect a predetermined work range. Having the input in a known range makes it possible to compute a first approximation with coefficients stored in memory blocks. With this method, it is hard to propose a “ready to use” IP for all the fixed-point formats. In this letter, a novel architecture, which does not require scaling, is proposed. This design is totally pipelined, ROM-less and can be directly used in any architecture. The implementation was optimized to reach a maximum clock frequency of 740 MHz on a Virtex-7 Field-Programmable Gate Array (FPGA). Erwan Libessart, Matthieu Arzel, Cyril Lahuec, Francesco Paolo Andriulli |
IEEE Signal Process. Lett. | 2 |
| 2016 | Ouessant: Flexible integration of dedicated coprocessors in Systems on Chip
Pierre-Henri Horrein, Philip-Dylan Gleonec, Erwan Libessart, André Lalevee, Matthieu Arzel |
DATE | 5 |
| 2016 | AutoReloc: Automated Design Flow for Bitstream Relocation on Xilinx FPGAsabstractDynamic and partial reconfiguration of Field Programmable Gate Arrays (FPGA) enable to reuse logic resources for several applications which are scheduled in a sequential order or which are loaded on demand. A fraction of the design on the FPGA is then substituted by another logic function while the rest of the system on the chip stays unaffected. If a design provides several partial reconfigurable areas, the configuration bitstream representing the logic function to be configured in this region has to be adapted to the physical requirements of this chip area. This can be achieved by deploying a repository with all possible configuration bitstreams for all possible regions. It is obvious that storage space can quickly become a limiting parameter in reconfigurable designs. For this purpose, bitstream relocation provides a less storage greedy approach. Only one representation as bitstream of an application needs to be stored. During the configuration process, a relocation algorithm manipulates the bitstream in order to suit it to the respective reconfigurable area. However, reconfigurable regions have to fulfill strong constraints for a relocation to be possible, which makes the selection and placement of reconfigurable regions a complex process. Unfortunately this is not automated by tools so far. In this paper, an approach to automate the development of such relocatable bitstreams is presented along with new algorithms related to relocation specific steps. This approach results in functional designs with minimal intervention from the designer. André Lalevee, Pierre-Henri Horrein, Matthieu Arzel, Michael Hübner 0001, Sandrine Vaton |
DSD | 3 |
| 2015 | Design of analog subthreshold Encoded Neural Network circuit in sub-100nm CMOSabstractEncoded Neural Networks (ENN) associate low-complexity algorithm with a storage capacity much larger than Hopfield Neural Networks' (HNN) for the same number of nodes. They are thus promising for implementing large scale neural networks mimicking the functioning of the human brain. The implementation of such a network on chip requires reducing the power consumption of the nodes to the femtojoule range to compare to human brain figures. Moreover, the circuit area must be reduced as much as possible. To address these challenges, this paper proposes a subthreshold analog ENN designed for the ST 65nm CMOS process. The designed circuit accepts power supply between 0.3V and 0.86V with currents below 300nA. In a network of 30 computation nodes, it yields a 32fJ energy consumption per decoding per node. The ENN converges only 21ns after being stimulated. Finally, the node core, i.e. without synapse, has a surface area of only 9.5µm2, and each synapse 3.6µm2. Benoit Larras, Cyril Lahuec, Fabrice Seguin, Matthieu Arzel |
IJCNN | 4 |
| 2014 | Low-complexity layered BP-based detection and decoding for a NB-LDPC coded MIMO systemabstractIn this paper, the combination of a low-complexity Multiple-input Multiple-output based on belief propagation (MIMO-BP) detector with a Non-Binary Low-Density Parity-Check (NB-LDPC) decoder is investigated. Such detection and decoding algorithms can enable an equivalent representation based on a larger Joint Factor Graph (JFG). Shuffle schedule can therefore be used jointly and simultaneously on the detector and the decoder. Actions are undertaken at the detector, decoder and the iterative receiver levels in order to reduce overall complexity. Indeed, applying the proposed low-complexity BP-based detection greatly reduces the number of operations per iteration (divided by ten), with a negligible performance penalty. EXtrinsic Information Transfer (EXIT) charts enable to analyze the convergence behaviour of the proposed iterative receiver. This analysis is used to find the best set of parameters enabling a detection-decoding process with a performance closest to the full-complexity system. Ali Haroun, Charbel Abdel Nour, Matthieu Arzel, Christophe Jégo |
ICC | 3 |
| 2014 | Symbol-based BP detection for MIMO systems associated with non-binary LDPC codesabstractIn this paper, an efficient iterative receiver for digital communication systems is investigated. It combines a Non-Binary Low-Density Parity-Check (NB-LDPC) decoder with a high-order constellation demapper and a Multiple-Input Multiple-Output (MIMO) detector. A suboptimal MIMO detector based on the Belief Propagation (BP) algorithm is investigated as an alternative to a Maximum Likelihood (ML) detector. Extrinsic information is exchanged between the detector and the decoder thanks to an iterative process. EXtrinsic Information Transfer (EXIT) charts enable to analyze the convergence behaviour of the proposed MIMO-BP detector. A suitable schedule for the different types of iterations is finally proposed to reduce the receiver latency and improve its error correction performance. Ali Haroun, Charbel Abdel Nour, Matthieu Arzel, Christophe Jégo |
WCNC | 3 |
| 2013 | Analog implementation of encoded neural networksabstractEncoded neural networks mix the principles of associative memories and error-correcting decoders. Their storage capacity has been shown to be much larger than Hopfield Neural Networks'. This paper introduces an analog implementation of this new type of network. The proposed circuit has been designed for the 1V supply ST CMOS 65nm process. It consumes 1165 times less energy than a digital equivalent circuit while being 2.7 times more efficient in terms of combined speed and surface. Benoit Larras, Cyril Lahuec, Matthieu Arzel, Fabrice Seguin |
ISCAS | 3 |
| 2012 | Hardware acceleration of SVM-based traffic classification on FPGAabstractUnderstanding the composition of the Internet traffic has many applications nowadays, mainly tracking bandwidth consuming applications, QoS-based traffic engineering and lawful interception of illegal traffic. Although many classification methods such as Support Vector Machines (SVM) have demonstrated their accuracy, not enough attention has been paid to the practical implementation of lightweight classifiers. In this paper, we consider the design of a real-time SVM classifier at many Gbps to allow online detection of categories of applications. Our solution is based on the design of a hardware accelerated SVM classifier on a FPGA board. Tristan Groleat, Matthieu Arzel, Sandrine Vaton |
IWCMC | 2 |
| 2006 | Semi-iterative analog turbo decodingabstractThis paper presents a novel analog turbo decoding architecture allowing analog decoders for long frame lengths to be implemented on a single chip. This is made possible by suitably using slicing techniques which allow hardware reuse and reconfigurability. The architecture is applied to a DVB-RCS-like code. It shows a reduction of occupied chip area by a factor of ten when compared to a conventional slice design with no significant performance degradation. A single 27mm20.25mum BiCMOS decoder can then decode any frame length from 40 up to 1824 bits Matthieu Arzel, Fabrice Seguin, Cyril Lahuec, Michel Jézéquel |
ISCAS | 1 |