EDBT 2026 Demo / reviewers in the wild / expert
Bertrand Le Gal
dblp:03/4695
· DBLP profile ↗
27ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0003-2269-8756ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 6 first-author · 6 since 2021Computer networks · 6 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-Probability-Frequency Fast Successive Cancellation (DPF-FSC) Decoding of Non-Binary Polar CodeabstractIn low-power wide-area network (LPWAN) protocols (such as LoRaWAN), orthogonalq-ary modulation is often employed because of its high capacity. It is well established thatq-ary coded modulation (CM) operates close to the Shannon limit at very low SNR. In this scheme, symbols from a non-binary (NB) error correction code defined over GF(q) are directly mapped to a set ofqorthogonal modulation symbols. However, the high decoding complexity for large alphabet sizes (q∈ [64 1024]) has so far prevented practical adoption. This study presents the first high-throughput software decoder for non-binary polar (Non-Binary polar (NB-p)) codes. We introduce the Dual-Probability-Frequency Fast Successive Cancellation (DPF-FSC) decoding algorithm, which operates jointly in probability and frequency domains to reduce the computational complexity. Combined with graph pruning techniques and optimized for parallel execution on multicore SIMD architectures, our implementation achieves real-time performance suitable for software-defined radio (SDR) and cloud-RAN systems. In particular, we demonstrate that our proposed DPF-FSC decoder achieves throughputs ranging from several Mbps on low-power embedded processors to over 500 Mbps on modern multicore CPUs. The most optimized implementation (D2b) reaches up to 900 Mbps on high-end processors for codes over GF(16) with 256 symbols and a 20% coding rate in single-core configuration. The throughput is reduced for larger Galois fields (GF(64) and GF(256)), but can still reach throughputs 40 to 400 Mbps, respectively, depending on the coding rate. Abdallah Abdallah, Bertrand Le Gal, Camille Monière, Emmanuel Boutillon |
IEEE Internet Things J. | 2 |
| 2025 | Differential Delay Mitigation in Multi-Antenna Aircraft TransmissionsabstractIn multi-antenna transmit diversity systems, accurate Channel State Information estimation is essential for reliable data decoding. In aeronautical telemetry, the spatial separation of antennas onboard aircraft, which can span several meters, introduces non negligible differential delays between received signals which degrades channel estimation, thereby impacting receiver performance. To address this issue, we propose a novel approach where both pilot sequences and transmitted data are Space-Time Block Coded to enhance CSI estimation by mitigating the degradation caused by differential delay. Thus leading to improved receiver robustness in multi-antenna aircraft telemetry. Simulation results demonstrate significantly better cross-correlation properties for the proposed pilot sequences over IRIG-106 ones. Complete communication system simulation using AWGN channel confirms that the new sequences provide gains in CSI estimation without computational complexity increase. Oussama Ait Sidi Ali, Romain Tajan, Bertrand Le Gal, Alain Thomas |
VTC2025-Spring | 3 |
| 2025 | DynHaMo: Dynamic Hardware-Based Monitoring Dedicated to Attacks DetectionabstractNumerous attacks compromising processor security have been developed over decades, including some targeting the microarchitecture, such as side-channel or transient attacks, or control-flow hijacking attacks. As these attacks target processor microarchitectural features and bypass software-level mitigation techniques, they are considered a serious threat. In order to mitigate these attacks while limiting the impact on performance, various detection methods have been proposed. Indeed, detection techniques offer solutions to limit the execution of costly countermeasures, only after attacks detection, limiting the induced performance overhead. However, detection techniques in the literature suffer from several drawbacks, including non-real-time detection, significant increase in execution time, or make the hypothesis of a trusted Operating System (OS). In this work, we introduce DynHaMo that addresses these issues by detecting attacks targeting the microarchitecture, such as Cache-based Side-Channel Attacks (CSCAs) and Return-Oriented Programming (ROP) attacks, at run-time by taking advantage of dynamic instruction insertion at the hardware level. DynHaMo, is a light-weight hardware micro-decoding unit capable of monitoring microarchitectural events on the fly. For evaluation purposes, DynHaMo has been integrated into a RISC-V core, assessed through multiple benchmarks and attack codes, and implemented on an FPGA platform. We evaluated our solution under high workloads to demonstrate the efficiency of the approach and its robustness to noise. The evaluation results show a detection accuracy of 99.3% on average, with 0.7% false negative and 1.2% false positive on average. Juliette Pottier, Maria Mendez Real, Bertrand Le Gal, Sébastien Pillement |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2023 | Securing a RISC-V architecture: A dynamic approachabstractThe SecureV (also known as SecV) project offers an innovative, open-source hardware, secure, and high-performance processor core based on the RISC-V ISA. The originality of the approach lies in the integration of a complete solution to increase security based on dynamic code transformation, covering 4 of the 5 NIST11National Institute of Standards and Technology functions of cybersecurity via monitoring (identify, detect), obfuscation (protect), and dynamic adaptation (react). Sébastien Pillement, Maria Mendez Real, J. Pottier, T. Nieddu, Bertrand Le Gal, Sébastien Faucou, Jean-Luc Béchennec, Mikaël Briday, Sylvain Girbal, Jimmy Le Rhun, Olivier Gilles, Daniel Gracia Pérez, André Sintzoff, Jean-Roch Coulon |
DATE | 5 |
| 2023 | Model Based Design of FMCW Radar Processing Systems on FPGA PlatformsabstractThe use of high level synthesis (HLS) tools is progressing in the industrial and academic worlds as they have gained in maturity. In the last few years, they are commonly applied to design hardware accelerators for ASIC and FPGA targets. Indeed, they enable signal processing algorithm integration in a shorter amount of time than methodologies based on handmade RTL design. Moreover, suitable behavioral description facilitate design space exploration to fulfil system constraints while minimizing other parameters. In this article, HLS tool abilities to design complete signal processing systems are evaluated. To this end, a commonly used FMCW radar signal processing system is deployed. The article presents demonstrate that a large set of trade-off solutions can be produced thanks to the described models. In a second time, the interest of the proposed approach is compared to optimized multicore SIMD implementations that are also relevant to give flexibility to FMCW radar embedded systems. Hugues Almorin, Bertrand Le Gal, Christophe Jégo, Vincent Kissel |
DSD | 2 |
| 2023 | Implementation of an Assignment Algorithm for Object Tracking on a FPGA MPSoCabstractAn assignment algorithm is a key step for multi-object tracking. In order to guarantee real-time execution on embedded systems, it is important to choose an efficient algorithm that matches the application context and the implementation constraints. In this paper, we present a comparative study of different implementations of an auction-based assignment algorithm for asymmetric problems on a System On Module (SOM) target based on an ARM processor and an FPGA circuit. This implementation differs with previous works [1], [2] by the context of multi-object tracking on an embedded system, which implies limited assignment problems, as well as by the choice of a more efficient assignment algorithm. Parallelization strategies were applied to achieve high-efficiency levels for both hardware and software implementations. The experimental results obtained show that, first, software implementation is, in most cases, sufficient to achieve the real-time performances and secondly, that proposed implementations perform better than related work ones (~ 130x for software and ~ 1.7x for hardware). Denis Shemonaev, Bertrand Le Gal, Christophe Jégo, Anthony Besseau |
DSD | 2 |
| 2023 | The Smart Kalman Filter: A Deep Learning-Based Approach for Time-Varying Channel EstimationabstractIn digital wireless communications, the received signal can be strongly altered by the environment and may contain Inter-Symbol Interference (ISI). To remove or reduce the ISI i.e. equalize, the impulse response of the propagation channel can be estimated. The Kalman Filter (KF) is an inescapable estimation algorithm in linear systems because of its optimality in terms of Minimum Mean Square Error (MMSE) under certain assumptions. However, in real conditions, implementations of KF are often difficult because of the necessity of hand-tuning parameters. In this paper, we present the Smart Kalman Filter (SKF), a hybrid architecture that combines a KF and the power of neural networks to extract relevant features from data to benefit from an adaptive KF that is automatically well tuned over time. We demonstrate in this paper that the proposed SKF is up to 5dB better at low Signal-to-Noise Ratio (SNR) and 3dB better at high SNR than Least Square (LS) algorithm in a time-varying channel estimation context with abrupt Doppler frequency variations. Antoine Siebert, Guillaume Ferré, Bertrand Le Gal, Aurélien Fourny |
PIMRC | 3 |
| 2023 | High-performance hard-input LDPC decoding on multi-core devices for optical space links
Bertrand Le Gal, Christophe Jégo, Vincent Pignoly |
J. Syst. Archit. | 1 |
| 2023 | Real-time energy-efficient software and hardware implementations of a QCSP communication system
Camille Monière, Bertrand Le Gal, Emmanuel Boutillon |
J. Syst. Archit. | 2 |
| 2020 | Model-based Design of Hardware SC Polar Decoders for FPGAsabstractPolar codes are a new error correction code family that should be benchmarked and evaluated in comparison to LDPC and turbo-codes. Indeed, recent advances in the 5G digital communication standard recommended the use of polar codes in EMBB control channels. However, in many cases, the implementation of efficient FEC hardware decoders is challenging. Specialised knowledge is required to enable and facilitate testing, rapid design iterations, and fast prototyping. In this article, a model-based design methodology to generate efficient hardware SC polar code decoders is presented. With HLS design process and tools, we demonstrate how FPGA system designers can quickly develop complex hardware systems with good performances. The favourable impact of design space exploration is underlined on achievable performances when a relevant computation model is used. The flexibility of the abstraction layers is evaluated. Hardware decoder generation efficiency is assessed and compared to competing approaches. It is shown that the fine-tuning of computation parallelism, bit length, pruning level, and working frequency help to design high-throughput decoders with moderate hardware complexities. Decoding throughputs higher than 300 Mbps are achieved on an Xilinx Virtex-7 device and on an Altera Stratix IV device. Yann Delomier, Bertrand Le Gal, Jérémie Crenne, Christophe Jégo |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2018 | ADMM hardware decoder for regular LDPC codes using a NISC-based architectureabstractThe Alternate Direction Method of Multipliers (ADMM) approach is an original method for LDPC decoding based on the linear programming (LP) technique. It introduces a novelty at the error correction performance level. Nevertheless, this method can be toughly implemented due to its high computational complexity. In this paper, an implementation of the ADMM LP decoding algorithm on an FPGA target is presented. Its hardware resource cost is evaluated and compared with the state of the art LDPC decoders using the belief propagation (BP) decoding approach. The preliminary logic synthesis results show that an LP based hardware decoder for LDPC codes should be viable for applications with tough error correction requirements. However, additional research works are required to reach equivalent hardware complexity and throughput performances that are similar to traditional BP based LDPC decoders. Imen Debbabi, Bertrand Le Gal, Nadia Khouja, Fethi Tlili, Christophe Jégo |
WCNC | 2 |
| 2018 | Implementation aspects of a pipeline ADMM-based LP decoding of LDPC convolutional codesabstractThe enforcement of the linear programming (LP) decoding was recently extended to LDPC convolutional codes (LDPC-CC). It was demonstrated that their convolutional structure suites the message-passing based representation of the LP problem, thanks to the application of the alternating directions method of multipliers (ADMM). In this paper, a modified formulation of the ADMM-LP problem is introduced for pipeline decoding of LDPC-CCs based on the layered schedule. Moreover, an assessment of the fixed-point format required for hardware implementation is provided to evaluate the complexity of the underlying decoding algorithm. Simulations show that an 8-bit quantization scheme yields negligible error correction performances loss of about 0.05 dB regarding the floating-precision ADMM based LDPC-CC decoder. Further, an algorithmic optimization of the Euclidean projection skipping technique, described in [1] [2], is proposed. This improved version enables to reduce the computational complexity of the decoding process without memory penalty contrary to the original formulation. Hayfa Ben Thameur, Bertrand Le Gal, Nadia Khouja, Fethi Tlili, Christophe Jégo |
WCNC | 2 |
| 2017 | A survey on decoding schedules of LDPC convolutional codes and associated hardware architecturesabstractLow-density parity-check convolutional codes (LDPC-CC) have interesting error correction features. They have a great potential to become a key error-correcting codes for enhancing reliability of modern digital communication systems, optical systems and storage devices. On the implementation side, however, the design of low-cost low-power and high-throughput LDPC-CC decoders remains challenging. This survey paper provides an overview of the state-of-the-art of different algorithmic optimizations proposed to ease and improve the LDPC-CC decoder implementations. To this end, a summary of the available decoding scheduling approaches is provided. Besides, a complexity analysis, performance comparison and an architectural analysis of LDPC-CC decoders based on different scheduling techniques is detailed. Hayfa Ben Thameur, Bertrand Le Gal, Nadia Khouja, Fethi Tlili, Christophe Jégo |
ISCC | 2 |
| 2016 | Multicore implementation of LDPC decoders based on ADMM algorithmabstractAlternate direction method of multipliers (ADMM) technique has recently been proposed for LDPC decoding. Even though it improves the error rate performance compared with traditional message passing (MP) techniques, it shows a higher computation complexity. In this article, the ADMM decoding algorithm is first described. Then, its computation complexity is analyzed. Finally, an optimized version which benefits from the multi-core processors architecture as well as the ADMM algorithm s parallelism is presented. The optimized version of the ADMM decoder can achieve up to 30 Mbps for standardized LDPC codes on a laptop x86 processor. Therefore, it could guide an efficient GPU implementation for real-time and high-throughput decoding systems requiring correction performances beyond MP-Sum Product Algorithm (SPA) capabilities. Imen Debbabi, Nadia Khouja, Fethi Tlili, Bertrand Le Gal, Christophe Jégo |
ICASSP | 4 |
| 2016 | Memory reduction techniques for successive cancellation decoding of polar codesabstractPolar coding is a new coding scheme that asymptotically achieves the capacity of several communication channels. Polar codes can be decoded with a successive cancellation (SC) decoder. In terms of hardware implementation, architectural performance of SC decoders is limited by the memory complexity. In this paper, two complementary methods are proposed to reduce the memory footprint of current state-of-the-art SC decoders. These methods must also applicable to SC-List decoders. The impacts the decoding performance in a rather negligible manner (<0.02dB), as shown by perormed simulations. The association of both methods allows a reduction of 16 ∼35% of the memory complexity for SC decoders depending on their quantization format. Bertrand Le Gal, Camille Leroux, Christophe Jégo |
ICASSP | 1 |
| 2016 | A scalable 3-phase polar decoderabstractIn this paper, we propose a 3-phase polar codes Successive Cancellation (SC) decoder. Benefiting from the local properties of the decoding tree, 3 zones are defined and associated to 3 distinct sub-decoders. This approach reduces the memory footprint while guaranteeing a better scalability in comparison with state of the art SC decoders. Several 3-phase SC decoders are implemented on an FPGA circuit and compare favorably to state-of-the-art implementations in terms of distributed hardware resources (LUT, D-FF) and throughput. Moreover, the memory of the decoder is reduced by 37% for a N = 221polar codes. Bertrand Le Gal, Camille Leroux, Christophe Jégo |
ISCAS | 1 |
| 2016 | Evaluation of the hardware complexity of the ADMM approach for LDPC decodingabstractLinear Programming (LP) is a novel technique for LDPC decoding. With the advance of the Alternate Direction Method of Multipliers (ADMM) approach, a significant step towards LP LDPC decoding scalability and optimization is made possible. Yet, this innovative decoding technique has not been implemented in hardware. Its hardware complexity has neither been estimated nor compared with traditional techniques. In this paper, an overview of the ADMM approach and its error correction performances for LDPC decoding is provided. Then, its computation complexity is evaluated to show the hardware feasibility of ADMM-based LDPC decoders. Our analysis is mainly carried at two levels. First, a quantitative complexity analysis is reported and a comparison with traditional LDPC decoders is given. Second, a proposal of a partially parallel architecture is described and its hardware complexity is evaluated then compared with state-of-the-art LDPC decoders. Imen Debbabi, Nadia Khouja, Fethi Tlili, Bertrand Le Gal, Christophe Jégo |
WCNC | 4 |
| 2016 | High-Throughput Multi-Core LDPC Decoders Based on x86 ProcessorabstractLow-Density Parity-Check (LDPC) codes are an efficient way to correct transmission errors in digital communication systems. Although initially targeting strictly to ASICs due to computation complexity, LDPC decoders have been recently ported to multicore and many-core systems. Most works focused on taking advantage of GPU devices. In this paper, we propose an alternative solution based on a layered OMS/NMS LDPC decoding algorithm that can be efficiently implemented on a multi-core device using Single Instruction Multiple Data (SIMD) and Single Program Multiple Data (SPMD) programming models. Several experimentations were performed on a x86 processor target. Throughputs up to 170 Mbps were achieved on a single core of an INTEL Core i7 processor when executing 20 layered-based decoding iterations. Throughputs reaches up to 560 Mbps on four INTEL Core-i7 cores. Experimentation results show that the proposed implementations achieved similar BER correction performance than previous works. Moreover, much higher throughputs have been achieved by comparison with all previous GPU and CPU works. They range from x1.4 to x8 by comparison with recent GPU works. Bertrand Le Gal, Christophe Jégo |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | A Flexible SoC and Its Methodology for Parser-Based ApplicationsabstractEmbedded systems are being increasingly network interconnected. They are required to interact with their environment through text-based protocol messages. Parsing such messages is control dominated. The work presented in this article attempts to accelerate message parsers using a codesign-based approach. We propose a generic architecture associated with an automated design methodology that enables SoC/SoPC system generation from high-level specifications of message protocols. Experimental results obtained on a Xilinx ML605 board show acceleration factors ranging from four to 11. Both static and dynamic reconfigurations of coprocessors are discussed and then evaluated so as to reduce the system hardware complexity. Bertrand Le Gal, Yérom-David Bromberg, Laurent Réveillère, Jigar Solanki |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2014 | GPU-like on-chip system for decoding LDPC codesabstractRapid prototyping is an important step in the development and the verification of computationally demanding tasks of digital communication systems, such as Forward Error Correction (FEC) decoding. The goal is to replace time-consuming simulations based on abstract models of the system with real-time experiments under real-world conditions. GPU-like architecture is a promising approach to fully exploit the potential of FPGA-based acceleration platforms. In this article, an application-specific GPU-like architecture and a complete compilation framework for decoding LDPC codes are proposed. The interest in an application-specific GPU in comparison with current GPUs is detailed. Finally, real-time experimentations demonstrate the potential of the GPU-like decoder to investigate both algorithmic and architectural issues. Bertrand Le Gal, Christophe Jégo |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2013 | Design space exploration for partially reconfigurable architectures in real-time systems
François Duhem, Fabrice Muller, Willy Aubry, Bertrand Le Gal, Daniel Négru, Philippe Lorenzini |
J. Syst. Archit. | 4 |
| 2012 | Design of multi-mode application-specific cores based on high-level synthesis
Emmanuel Casseau, Bertrand Le Gal |
Integr. | 2 |
| 2011 | Assertion support in high-level synthesis design flow
Aurélien Ribon, Bertrand Le Gal, Christophe Jégo, Dominique Dallet |
FDL | 2 |
| 2009 | Design and implementation of a reconfigurable decimation and channel selection filter for GSM and UMTS radio standardsabstractThis work presents a low-power multistandard decimation and channel selection filter architecture. The filter is suitable after an over-sampling sigma-delta converter and performs decimation in two stages. The first stage is a modified structure of the cascade of integrators-combs (CIC) filter and allows reducing sampling rate downto only the double of the Nyquist frequency. The second stage composed of classical FIR filter, has relaxed specifications and performs channel selection. Implementation of the proposed filter for UMTS and GSM standards shows good filtering performances. The signal to ratio measured for UMTS is 14,65 dB and for GSM 26,96 dB which satisfy largely the standards requirements. Implementation on ASIC 65-nm process technology shows power consumption gain of 14% in comparison to previously proposed low-power architecture. Nadia Khouja, Khaled Grati, Adel Ghazel, Bertrand Le Gal |
WCNC | 4 |
| 2008 | A new orthogonal online digital calibration for time-interleaved analog-to-digital convertersabstractModern communication technologies need faster analog-to-digital converters (ADC). To significantly increase the sampling rate of an ADC, time-interleaved ADC (TIADC) is an efficient solution. A M-channels TIADC is composed of M ADCs which operate at interleaved sampling times. Due to the manufacturing process, the main drawback of a TIADC system is that the M ADCs are not exactly the same. This means that offset, gain and time mismatch errors are introduced. As a result, these errors cause distortions in the output sampled signal and introduce unwanted tones and noise, and hence, reduce the spurious free dynamic range (SFDR) as well as the signal to noise ratio (SNR). In this paper, we propose a new orthogonal online digital calibration, for timing skew, offset and gain mismatches, based on Code Division Multiple Access (CDMA) technique already used in communications. Our calibration is online, this means that errors can be estimated while the ADC is running. Since most of the calibration processes are carried out on the digital outputs, very little change is needed on the analog part of the ADC. Simulations results showed the efficiency of our proposed calibration architecture. Guillaume Ferré, Maher Jridi, Lilian Bossuet, Bertrand Le Gal, Dominique Dallet |
ISCAS | 4 |
| 2008 | Dynamic Memory Access Management for High-Performance DSP Applications Using High-Level SynthesisabstractMultimedia applications such as video and image processing are often characterized by a huge number of data accesses. In many digital signal processing applications, array access patterns are regular and periodic. In these cases, optimized architectures using pipelined memory access controllers can be generated. In this paper, we focus on implementing memory interfacing modules that can be automatically generated from a high-level synthesis tool and which can efficiently handle predictable address patterns as well as random ones (i.e., dynamic address computations). The benefits of balancing dynamic address computations from datapath to dedicated computation units in the memory controller is also analyzed as well as operator bitwidth optimization and data locality to save power consumption and reduce latency. Bertrand Le Gal, Emmanuel Casseau, Sylvain Huet |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2005 | Hardware Virtual Components Compliant with Communication System StandardsabstractIn this paper, we focus on the design of a communication system based on reusing IP cores. Traditional methods for designing hardware cores for this kind of applications use a RTL specification. However, they suffer from heavy limitations that prevent them from efficiently addressing the algorithmic complexity and the high flexibility required by the various application profiles. For this reason, we propose to raise the abstraction level of the specification and introduce the notion of architectural flexibility by benefiting from the emerging high-level synthesis tools. From a single behavioral-level VHDL specification, we are able to generate a variety of architectures, compliant with the most important communication standards. This technique has been successfully applied to the most important IP cores (synchronization IP, Viterbi IP and Reed-Solomon decoder IP cores) of the DVB-DSNG digital video-broadcasting standard. Nabil Abdelli, Pierre Bomel, Emmanuel Casseau, Anne-Marie Fouilliart, Christophe Jégo, Philippe Kajfasz, Bertrand Le Gal, Nathalie Le Heno |
DSD | 7 |