EDBT 2026 Demo / reviewers in the wild / expert
François Leduc-Primeau
dblp:47/7909
· DBLP profile ↗
21ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-5528-8510ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 12 · 6 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer networks
4 papers |
Physical-layer communications · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware reliability and fault tolerance · 60% Integrated circuit design · 40% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Physical-layer communications
channel coding |
1.1 | 4 | 2021 | Noisy Density Evolution With Asymmetric Deviation Models · IEEE Trans. Commun. 2021 Modeling and Energy Optimization of LDPC Decoder Circuits With Timing Violations · IEEE Trans. Commun. 2018 Relaxed Half-Stochastic Belief Propagation · IEEE Trans. Commun. 2013 |
Physical-layer communications › channel coding › error control coding › block codes
LDPC codes |
1.1 | 4 | 2021 | Noisy Density Evolution With Asymmetric Deviation Models · IEEE Trans. Commun. 2021 Modeling and Energy Optimization of LDPC Decoder Circuits With Timing Violations · IEEE Trans. Commun. 2018 Relaxed Half-Stochastic Belief Propagation · IEEE Trans. Commun. 2013 |
Physical-layer communications › channel coding › decoding algorithms › iterative decoding
density evolution |
0.5 | 1 | 2021 | Noisy Density Evolution With Asymmetric Deviation Models · IEEE Trans. Commun. 2021 |
Hardware reliability and fault tolerance
soft errors |
0.5 | 1 | 2021 | Noisy Density Evolution With Asymmetric Deviation Models · IEEE Trans. Commun. 2021 |
Physical-layer communications › channel coding › decoding algorithms
belief propagation |
0.5 | 3 | 2021 | Relaxed Half-Stochastic Belief Propagation · IEEE Trans. Commun. 2013 Noisy Density Evolution With Asymmetric Deviation Models · IEEE Trans. Commun. 2021 Dithered Belief Propagation Decoding · IEEE Trans. Commun. 2012 |
Integrated circuit design
low-power circuit design |
0.3 | 1 | 2018 | Modeling and Energy Optimization of LDPC Decoder Circuits With Timing Violations · IEEE Trans. Commun. 2018 |
Methods — techniques the papers use, named apart from their topics
density evolution · 1.8asymmetric deviation models · 1.0offset min-sum algorithm · 0.7sum-product algorithm · 0.2dithered belief propagation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Foundation Model for Massive MIMO Precoding with an Adaptive Per-User Rate-Power TradeoffabstractDeep learning (DL) has emerged as a solution for precoding in massive multiple-input multiple-output (mMIMO) systems due to its capacity to learn the characteristics of the propagation environment. However, training such a model requires high-quality, local datasets at the deployment site, which are often difficult to collect. We propose a transformer-based foundation model for mMIMO precoding that seeks to minimize the energy consumption of the transmitter while dynamically adapting to per-user rate requirements. At equal energy consumption, zero-shot deployment of the proposed foundation model significantly outperforms zero forcing, and approaches weighted minimum mean squared error performance with 8× less complexity. To address model adaptation in data-scarce settings, we introduce a data augmentation method that finds training samples similar to the target distribution by computing the cosine similarity between the outputs of the pre-trained feature extractor. Our work enables the implementation of DL-based solutions in practice by addressing challenges of data availability and training complexity. Moreover, the ability to dynamically configure per-user rate requirements can be leveraged by higher level resource allocation and scheduling algorithms for greater control over energy efficiency, spectral efficiency and fairness. Jérôme Emery, Ali Hasanzadeh Karkan, Jean-François Frigon, François Leduc-Primeau |
PIMRC | 4 |
| 2024 | Online Energy-Efficient Beam Bandwidth Partitioning in mmWave Mobile NetworksabstractThis paper studies beam bandwidth partitioning problem in mobile millimeter-wave (mmWave) and multiple antennas networks. The main novelty is to flexibly optimize the beamforming bandwidth with the aim to minimize the energy consumption of the system while guaranteeing the data requirements of all mobile users. We formulate the problem as an integer nonlinear programming problem. To efficiently solve the problem, we design a deep reinforcement learning using the proximal policy optimization approach and train a deep neural network in an on-policy manner. Then, for comparison purposes, we develop low-complexity online iterative accurate solutions. We show that our approach achieves better performance compared to the iterative solutions and is able to achieve at least 4% less energy consumption and more than 12% energy efficiency gains. Zoubeir Mlika, Tri Nhu Do, Adel Larabi, Jennie Diem Vo, Jean-François Frigon, François Leduc-Primeau |
VTC Fall | 6 |
| 2023 | Quark: An Integer RISC-V Vector Processor for Sub-Byte Quantized DNN InferenceabstractIn this paper, we present Quark, an integer RISC-V vector processor specifically tailored for sub-byte DNN inference. Quark is implemented in GlobalFoundries' 22FDX FD-SOI technology. It is designed on top of Ara, an open-source 64-bit RISC-V vector processor. To accommodate sub-byte DNN inference, Quark extends Ara by adding specialized vector instructions to perform sub-byte quantized operations. We also remove the floating-point unit from Quarks' lanes and use the CVA6 RISC-V scalar core for the re-scaling operations that are required in quantized neural network inference. This makes each lane of Quark 2 times smaller and 1.9 times more power efficient compared to the ones of Ara. In this paper we show that Quark can run quantized models at sub-byte precision. Notably we show that for 1-bit and 2-bit quantized models, Quark can accelerate computation of Conv2d over various ranges of inputs and kernel sizes. MohammadHossein AskariHemmat, Théo Dupuis, Yoan Fournier, Nizar El Zarif, Matheus A. Cavalcante, Matteo Perotti, Frank K. Gürkaynak, Luca Benini, François Leduc-Primeau, Yvon Savaria, Jean-Pierre David |
ISCAS | 9 |
| 2022 | Flexible Unsupervised Learning for Massive MIMO Subarray Hybrid BeamformingabstractHybrid beamforming is a promising technology to improve the energy efficiency of massive MIMO systems. In particular, subarray hybrid beamforming can further decrease power consumption by reducing the number of phase-shifters. However, designing the hybrid beamforming vectors is a complex task due to the discrete nature of the subarray connections and the phase-shift amounts. Finding the optimal connections between RF chains and antennas requires solving a non-convex problem in a large search space. In addition, conventional solutions assume that perfect channel state information (CSI) is available, which is not the case in practical systems. Therefore, we propose a novel unsupervised learning approach to design the hybrid beamforming for any subarray structure while supporting quantized phase-shifters and noisy CSI. One major feature of the proposed architecture is that no beamforming codebook is required, and the neural network is trained to take into account the phase-shifter quantization. Simulation results show that the proposed deep learning solutions can achieve higher sum-rates than existing methods. Hamed Hojatian, Jérémy Nadal, Jean-François Frigon, François Leduc-Primeau |
GLOBECOM | 4 |
| 2021 | Improving the Energy-Efficiency of a Kalman Filter Using Unreliable MemoriesabstractKalman filters are widely used for real-time estimation of dynamic systems, and they sometimes need to be implemented on energy-constrained devices. A Kalman filter implementation from unreliable memories is considered, where the flipping probability of a bit in a memory cell directly depends on its energy consumption. The degradation in estimation performance caused by the noise in the memory is theoretically investigated. Updated equations are then developed for the Kalman filter, taking into account the new source of noise from the unreliable memory. Finally, a method is proposed to optimize the bit energy allocation in the memory, and it is shown from numerical simulations that this method allows for important energy gains. Jonathan Kern, Elsa Dupraz, Abdeldjalil Aïssa-El-Bey, François Leduc-Primeau |
ICASSP | 4 |
| 2021 | Power-Efficient Deep Neural Networks with Noisy Memristor ImplementationabstractThis paper considers Deep Neural Network (DNN) linear-nonlinear computations implemented on memristor cross-bar substrates. To address the case where true memristor conductance values may differ from their target values, it introduces a theoretical framework that characterizes the effect of conductance value variations on the final inference computation. With only second-order moment assumptions, theoretical results on tracking the mean, variance, and covariance of the layer-by-layer noisy computations are given. By allowing the possibility of amplifying certain signals within the DNN, power consumption is characterized and then optimized via KKT conditions. Simulation results verify the accuracy of the proposed analysis and demonstrate the significant power efficiency gains that are possible via optimization for a target mean squared error. Elsa Dupraz, Lav R. Varshney, François Leduc-Primeau |
ITW | 3 |
| 2021 | Noisy Density Evolution With Asymmetric Deviation ModelsabstractThis paper considers low-density parity-check (LDPC) decoders affected by deviations introduced by the electronic device on which the decoder is implemented. Noisy density evolution (DE) that allows to theoretically study the performance of these LDPC decoders can only consider symmetric deviation models due to the all-zero codeword assumption. A novel DE method is proposed that admits the use of asymmetric deviation models, thus widening the range of faulty implementations that can be analyzed. DE equations are provided for three noisy decoders: belief propagation, Gallager B, and quantized min-sum (MS). Simulation results confirm that the proposed DE accurately predicts the performance of LDPC decoders with asymmetric deviations. Furthermore, asymmetric versions of the Gallager B and MS decoders are proposed to compensate the effect of asymmetric deviations. The parameters of these decoders are then optimized using the proposed DE, leading to better ensemble thresholds and improved finite-length performance in the presence of asymmetric deviations. Elsa Dupraz, François Leduc-Primeau |
IEEE Trans. Commun. | 2 |
| 2021 | Unsupervised Deep Learning for Massive MIMO Hybrid BeamformingabstractHybrid beamforming is a promising technique to reduce the complexity and cost of massive multiple-input multiple-output (MIMO) systems while providing high data rate. However, the hybrid precoder design is a challenging task requiring channel state information (CSI) feedback and solving a complex optimization problem. This paper proposes a novel RSSI-based unsupervised deep learning method to design the hybrid beamforming in massive MIMO systems. Furthermore, we propose i) a method to design the synchronization signal (SS) in initial access (IA); and ii) a method to design the codebook for the analog precoder. We also evaluate the system performance through a realistic channel model in various scenarios. We show that the proposed method not only greatly increases the spectral efficiency especially in frequency-division duplex (FDD) communication by using partial CSI feedback, but also has near-optimal sum-rate and outperforms other state-of-the-art full-CSI solutions. Hamed Hojatian, Jérémy Nadal, Jean-François Frigon, François Leduc-Primeau |
IEEE Trans. Wirel. Commun. | 4 |
| 2020 | RSSI-Based Hybrid Beamforming Design with Deep LearningabstractHybrid beamforming is a promising technology for 5G millimetre-wave communications. However, its implementation is challenging in practical multiple-input multiple-output (MIMO) systems because non-convex optimization problems have to be solved, introducing additional latency and energy consumption. In addition, the channel-state information (CSI) must be either estimated from pilot signals or fed back through dedicated channels, introducing a large signaling overhead. In this paper, a hybrid precoder is designed based only on received signal strength indicator (RSSI) feedback from each user. A deep learning method is proposed to perform the associated optimization with reasonable complexity. Results demonstrate that the obtained sum-rates are very close to the ones obtained with full-CSI optimal but complex solutions. Finally, the proposed solution allows to greatly increase the spectral efficiency of the system when compared to existing techniques, as minimal CSI feedback is required. Hamed Hojatian, Vu Nguyen Ha, Jérémy Nadal, Jean-François Frigon, François Leduc-Primeau |
ICC | 5 |
| 2020 | Overlap-Save FBMC ReceiversabstractFuture communication systems are foreseen to support several services with different requirements. Waveform designs based on filter-bank multi-carrier with offset quadrature amplitude modulation (FBMC/OQAM) can offer interesting advantages in this context, such as low out-of-band power leakage and high spectral efficiency due to the lack of guard intervals. However, downsides of FBMC/OQAM with respect to a typical orthogonal frequency-division multiplexing (OFDM) solution include higher latency, higher complexity and difficulties in adapting some existing OFDM techniques such as MIMO Alamouti. To address these issues, novel FBMC receivers suitable for short prototype filters are proposed. Based on the Overlap-Save algorithm, the proposed receivers improve error-rate performance on multipath channels and support asynchronous communication. We show that complexity can be further reduced by efficiently processing blocks of FBMC symbols jointly, and that user mobility support can be traded off for additional complexity reductions in a flexible way through polynomial decomposition of the equalizer stage. Finally, we show that a block-Alamouti scheme can be applied, and we propose a MIMO equalizer with improved error-rate performance on time-varying channels, compared to the typical FBMC block-Alamouti equalizer. Jérémy Nadal, François Leduc-Primeau, Charbel Abdel Nour, Amer Baghdadi |
IEEE Trans. Wirel. Commun. | 2 |
| 2019 | Training Modern Deep Neural Networks for Memory-Fault RobustnessabstractBecause deep neural networks (DNNs) rely on a large number of parameters and computations, their implementation in energy-constrained systems is challenging. In this paper, we investigate the solution of reducing the supply voltage of the memories used in the system, which results in bit-cell faults. We explore the robustness of state-of-the-art DNN architectures towards such defects and propose a regularizer meant to mitigate their effects on accuracy. Our experiments clearly demonstrate the interest of operating the system in a faulty regime to save energy without reducing accuracy. Ghouthi Boukli Hacene, François Leduc-Primeau, Amal Ben Soussia, Vincent Gripon, François Gagnon |
ISCAS | 2 |
| 2018 | A Block FBMC Receiver Designed for Short FiltersabstractIn this paper, a new filter-bank multi-carrier (FBMC) receiver targeted at short prototype filters (PFs) is presented. In addition to the typical advantages of FBMC modulation such as improved frequency containment and support of relaxed synchronization, the proposed receiver enables accurate one-tap equalization of the signal despite the use of a short PF, while significantly reducing the complexity by merging the equalization and the standard FBMC receiver. An adapted frame structure is proposed that occupies the same radio resources as a 4G/LTE orthogonal frequency-division multiplexing (OFDM) frame while providing a similar data rate. Moreover, this frame structure is shown to readily support block-based Alamouticoded multiple-input multiple-output transmissions. Simulation results show that the proposed FBMC receiver can outperform an OFDM system on 4G/LTE channel models in both single antenna and Alamouti configurations, while having a better robustness to synchronization errors than a frequency-spread FBMC receiver. Jérémy Nadal, François Leduc-Primeau, Charbel Abdel Nour, Amer Baghdadi |
ICC | 2 |
| 2018 | Modeling and Energy Optimization of LDPC Decoder Circuits With Timing ViolationsabstractThis paper proposes a “quasi-synchronous” design approach for signal processing circuits, in which timing violations are permitted, but without the need for a hardware compensation mechanism. The case of a low-density parity-check (LDPC) decoder is studied, and a method for accurately modeling the effect of timing violations at a high level of abstraction is presented. The error-correction performance of code ensembles is then evaluated using density evolution, while taking into account the effect of timing faults. Following this, several quasi-synchronous LDPC decoder circuits based on the offset min-sum algorithm are optimized, providing a 23%-40% reduction in energy consumption or energy-delay product, while achieving the same performance and occupying the same area as conventional synchronous circuits. François Leduc-Primeau, Frank R. Kschischang, Warren J. Gross |
IEEE Trans. Commun. | 1 |
| 2017 | VLSI Implementation of Deep Neural Network Using Integral Stochastic ComputingabstractThe hardware implementation of deep neural networks (DNNs) has recently received tremendous attention: many applications in fact require high-speed operations that suit a hardware implementation. However, numerous elements and complex interconnections are usually required, leading to a large area occupation and copious power consumption. Stochastic computing (SC) has shown promising results for low-power area-efficient hardware implementations, even though existing stochastic algorithms require long streams that cause long latencies. In this paper, we propose an integer form of stochastic computation and introduce some elementary circuits. We then propose an efficient implementation of a DNN based on integral SC. The proposed architecture has been implemented on a Virtex7 field-programmable gate array, resulting in 45% and 62% average reductions in area and latency compared with the best reported architecture in the literature. We also synthesize the circuits in a 65-nm CMOS technology, and we show that the proposed integral stochastic architecture results in up to 21% reduction in energy consumption compared with the binary radix implementation at the same misclassification rate. Due to fault-tolerant nature of stochastic architectures, we also consider a quasi-synchronous implementation that yields 33% reduction in energy consumption with respect to the binary radix implementation without any compromise on performance. Arash Ardakani, François Leduc-Primeau, Naoya Onizawa, Takahiro Hanyu, Warren J. Gross |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Hardware implementation of FIR/IIR digital filters using integral stochastic computationabstractStochastic computing (SC) has received much recent attention due to its inherent fault-tolerance and low implementation cost compared to binary radix representations. SC has been proposed for various signal processing applications such as digital filters. The prior art in stochastic FIR filters can accurately implement the desired filtering function for low-order filters, however, their accuracy degrades as the filter order increases. Moreover, stochastic IIR filters demonstrate high hardware complexity and degraded accuracy. In this paper, we propose an architecture for high-order FIR filters with negligible accuracy loss compared to fixed-point implementation. The proposed architecture requires fewer random number generators. We also describe a novel cascaded second-order direct-form II structure for IIR filters. The implementation results of the proposed design show an improvement in latency and hardware complexity compared to the stochastic architectures reported to date. Arash Ardakani, François Leduc-Primeau, Warren J. Gross |
ICASSP | 2 |
| 2015 | Energy optimization of LDPC decoder circuits with timing violationsabstractThis paper presents a quasi-synchronous design approach for signal processing circuits, in which timing violations are permitted, but without the need for a hardware compensation mechanism. A quasi-synchronous low-density parity-check decoder processing circuit based on the offset min-sum algorithm is designed, achieving the same performance and occupying the same area as a conventional synchronous circuit, but using up to 28% less energy. François Leduc-Primeau, Frank R. Kschischang, Warren J. Gross |
ICC | 1 |
| 2014 | Cluster-based associative memories built from unreliable storageabstractWe consider associative memories based on clustered graphs that were recently introduced. These memories are almost optimal in terms of the amount of storage they require (efficiency), and allow retrieving messages with low complexity. We study an unreliable implementation of the memory and compare its error rate and storage efficiency with that of a reliable implementation. We present analytical and simulation results that indicate that the proposed memory structure can tolerate a large number of faults at a reasonable cost, thereby making it a good candidate for achieving highly efficient circuit implementations of associative memories. François Leduc-Primeau, Vincent Gripon, Michael G. Rabbat, Warren J. Gross |
ICASSP | 1 |
| 2013 | Relaxed Half-Stochastic Belief PropagationabstractLow-density parity-check codes are attractive for high throughput applications because of their low decoding complexity per bit, but also because all the codeword bits can be decoded in parallel. However, achieving this in a circuit implementation is complicated by the number of wires required to exchange messages between processing nodes. Decoding algorithms that exchange binary messages are interesting for fully-parallel implementations because they can reduce the number and the length of the wires, and increase logic density. This paper introduces the Relaxed Half-Stochastic (RHS) decoding algorithm, a binary message belief propagation (BP) algorithm that achieves a coding gain comparable to the best known BP algorithms that use real-valued messages. We derive the RHS algorithm by starting from the well-known Sum-Product algorithm, and then derive a low-complexity version suitable for circuit implementation. We present extensive simulation results on two standardized codes having different rates and constructions, including low bit error rate results. These simulations show that RHS can converge faster on average than existing state-of-the-art decoding algorithms, leading to improvements in throughput and energy efficiency. François Leduc-Primeau, Saied Hemati, Shie Mannor, Warren J. Gross |
IEEE Trans. Commun. | 1 |
| 2012 | Dithered Belief Propagation DecodingabstractWe introduce two dithered belief propagation decoding algorithms to lower the error floor with a minimal hardware overhead. One of the algorithms can additionally improve the decoding performance in the waterfall region using a large iteration limit but with a negligible increase in the average time complexity. François Leduc-Primeau, Saied Hemati, Shie Mannor, Warren J. Gross |
IEEE Trans. Commun. | 1 |
| 2010 | Lowering Error Floors Using Dithered Belief PropagationabstractWe propose dithered belief propagation decoding algorithms to reduce the number of decoding failures of a belief propagation decoder and lower the error floor. The random nature of the algorithms enables a low hardware complexity compared to previously reported techniques. We introduce two dithering methods that target check node operations and channel input values, respectively. We present simulation results that confirm the error rate gains in the floor region, and that relate those gains with the maximum number of decoding iterations. The results show that the first algorithm can achieve good error rate gains with a low iteration limit. For the second algorithm, results show that with a large iteration limit, high FER gains are possible. Furthermore the average time complexity remains the same as that of a standard belief propagation algorithm. François Leduc-Primeau, Saied Hemati, Shie Mannor, Warren J. Gross |
GLOBECOM | 1 |
| 2009 | A Relaxed Half-Stochastic Iterative Decoder for LDPC CodesabstractThis paper presents a Relaxed Half-Stochastic (RHS) low-density parity-check (LDPC) decoding algorithm that uses some elements of the sum-product algorithm (SPA) in its variable nodes, but maintains the low-complexity interleaver and check node structures characteristic of stochastic decoders. The algorithm relies on the principle of successive relaxation to convert binary stochastic streams to a log-likelihood ratio (LLR) representation. Simulations of a (2048, 1723) RS-LDPC code show that the RHS algorithm can outperform 100-iterations floating-point SPA decoding. We describe approaches for low-complexity implementation of the RHS algorithm. Furthermore, we show how the stochastic nature of the belief representation can be exploited to lower the error floor. François Leduc-Primeau, Saied Hemati, Warren J. Gross, Shie Mannor |
GLOBECOM | 1 |