EDBT 2026 Demo / reviewers in the wild / expert
David Bol
dblp:88/677
· DBLP profile ↗
38ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-2678-1613ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimization and Integration of Low-Area Low-Energy and High-SQNR FFT Accelerators in ULP MCUsabstractPower-constrained embedded applications often execute Fast Fourier Transforms (FFT) on ultra-low-power microcontrollers (ULP MCUs). They can integrate memory-based hardware (HW) FFT accelerators, with additional silicon area being traded for energy and latency improvements. This trade-off is modulated by the radix of the FFT accelerator. Moreover, a fixed-point representation with a fixed wordlength is generally used, which limits the signal-to-quantization-noise ratio (SQNR) and requires a data scaling method to avoid overflows. Higher accelerator radices reduce the scaling method SQNR and increase its area overhead. In this work, we explore the trade-offs between energy, area, latency and SQNR in FFT accelerators integrated within ULP MCUs. An in-depth analysis is conducted based on design implementations in a 65-nm LP CMOS technology with a configurable radix from 2 to$2^{4}$. Energy and latency improvements of radix-$2^{4}$reach a factor$2.83\times $and$3.98\times $, respectively. However, the logic area increases by$2.14\times $. We also propose a new scaling method, called index-encoded hierarchical block floating-point (IH-BFP). It provides the high SQNR of existing methods and relies on the deterministic access pattern of the FFT to achieve energy savings of 9.7 % and logic area savings of 32.7 %. Pol Maistriaux, Jérôme Louveaux, David Bol |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | A Narrowband RF Front End in 22-nm FD-SOI Featuring a Programmable Low-Noise Amplifier with a Configurable Noise-Power Trade-OffabstractWireless receivers are usually designed to guarantee a certain performance level in the worst-case operating conditions. However, as these conditions are rarely encountered, the receivers consume more power than necessary to correctly demodulate the incoming signal. This power overhead could be avoided by programmable receivers, which adapt their performance to the conditions of the wireless channel. In this paper, we propose a highly programmable RF front end for narrowband low-power wide-area networks. It trades off a noise figure between 7 and 9dB for a power consumption between 3.65 and 1.81mW thanks to a dynamically programmable low-noise amplifier. Integrated into a LoRa receiver, the proposed programmable RF front end can bring up to 46% of power savings. Marco Gonzalez, Pol Maistriaux, David Bol |
ISCAS | 3 |
| 2024 | A Parametric Power Model of Multi-Band Sub-6 GHz Cellular Base Stations Using On-Site MeasurementsabstractThe increasing energy consumption of mobile networks has emerged as a critical concern for mobile telecommunication operators, requiring measures to curb or reverse the historical upward trajectory in order to align with the sector’s decarbonization target. Meanwhile, 5G-NR is massively deployed in the networks to improve quality of service, with the hope of simultaneously improving energy efficiency thanks to enhanced power-saving features. However, up-to-date 5G-enabled base stations, that support higher bandwidths with more transceivers at higher frequencies, raise concerns about their absolute power consumption, especially at low traffic loads. Proper power models are therefore needed to identify the key levers for energy savings in mobile networks under real traffic loads. This paper addresses this challenge by first providing a parametric power consumption model applicable to commercial sub- 6 GHz cellular base stations. Then, numerical model parameters are estimated by combining on-site measurements from operators with radio equipment documentation from manufacturers. The uncertainty of model predictions is assessed to be in the range of $\mathbf{1 0 - 2 0 \%}$. Moreover, we show that the proposed average power model aligns well with measurements of equipment that use existing power-saving features. Estimates of power consumption are also provided for typical 3-sector single-band macro base stations in active mode, e.g., 1-4 kW when equipped with traditional radio units, and 2-4 kW when using active antenna units. Louis Golard, Youssef Agram, François Rottenberg, François Quitin, David Bol, Jérôme Louveaux |
PIMRC | 5 |
| 2024 | Cross-Domain Optimization of Low-Power Mixed-Signal Sensor Systems Under Classification Accuracy ConstraintsabstractOptimizing mixed-signal systems-on-chips (SoCs) is a challenging task, especially when they involve both analog building blocks and machine-learning (ML) algorithms which make the overall system performance hard to predict. This paper proposes a methodology for optimizing complex mixed-signal sensor systems, which consists of three steps: (1) modeling Pareto-optimal local performance of individual circuits, (2) building a macromodel of the mixed-signal system performance including local non-idealities, and (3) finding the optimal set of system-level parameters using numerical optimization. The methodology is applied to the use case of an electrocardiogram (ECG) sensor system with embedded arrhythmia classification, with the objective of minimizing the system power consumption while imposing constraints on inference accuracy. Different numerical optimization methods are compared to evaluate their respective performance on the minimization task: gradient-based, genetic, or bayesian optimization. The proposed methodology achieves the optimization in a reasonable execution time of 10 hours, which had not been previously demonstrated on such complex mixed-signal system. Among the different methods tested, bayesian optimization provides the best success rate of 97% and fastest convergence speed. Remi Dekimpe, David Bol |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | A Combined Analytical and Simulation-Based Methodology for Quantifying the Noise-Power-Area Trade-Offs in Biomedical AmplifiersabstractLow-noise operation is one of the most important performance criteria for low-power amplifiers targeting biopotential acquisition. While advanced circuit architectures exist to minimize the intrinsic noise, an analytical formalism is still lacking to estimate the lowest achievable noise level without performing extensive circuit-level optimizations. This work proposes a hybrid methodology mixing theoretical analyses and a limited number of simulations to estimate and minimize the input-referred noise and the noise efficiency factor of various biomedical amplifiers topologies. Compared to previous works, accurate bias-dependent noise models are obtained thanks to simulations of single devices and allow this methodology to successfully take into account the thermal and flicker noise sources from MOS transistors analytically. The optimal noise-current-area trade-off is then derived, showing the fundamental limits of the architecture. In this paper, the proposed methodology is applied to a current-reuse amplifier topology designed for two applications. The specifications for each application are obtained from a system-level perspective, including an input high-pass filter whose noise is considered analytically. Simulation-based optimization results show a good agreement with the analytical approach, proving that the methodology can be used for noise estimation, comparison between architectures, and to extract meaningful design guidelines. Sylvain Favresse, David Bol, Denis Flandre |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | Post-Silicon Optimization of a Highly Programmable 64-MHz PLL Achieving 2.7-5.7 μWabstractHierarchical optimization methods used in the design of complex mixed-signal systems require accurate behavioral models to avoid the long simulation times of transistor-level SPICE simulations of the whole system. However, robust behavioral models that accurately model circuit non-idealities and their complex interactions must be very complex themselves and are hardly achievable. Post-silicon tuning, which is already widely used for the calibration of analog building blocks, is an interesting alternative to speed up the optimization of these complex systems. However, post-silicon tuning usually focuses on single-objective problems in blocks with a limited number of degrees of freedom. In this paper, we propose a post-silicon “hardware-in-the-loop” optimization method to solve multi-objective problems in mixed-signal systems with numerous degrees of freedom. We use this method to optimize the noise-power trade-off of a 64-MHz phase-locked loop (PLL) based on a back-bias-controlled ring oscillator. A genetic algorithm was run based on measurements of the 22-nm fully-depleted silicon-on-insulator prototype to find the Pareto-optimal configurations in terms of power and long-term jitter. The obtained Pareto front gives a range of power consumption between 2.7 and 5.7 μW, corresponding to an RMS long-term jitter between 88 and 45 ns. Whereas the simulation-based optimization would require more than a year using the genetic algorithm based on SPICE simulations, we conducted the post-silicon optimization in only 17 h. Marco Gonzalez, David Bol |
DATE | 2 |
| 2023 | Technical and Ecological Limits of 2.45-GHz Wireless Power Transfer for Battery-Less SensorsabstractWireless power transfer (WPT) is a possible alternative to batteries for supplying smart sensors. It is often described as “greener" because it avoids the use of batteries. However, this claim usually lacks the proper assessment of the environmental impacts of WPT and is particularly questionable given the poor efficiency of wirelessly transmitting power over a few meters of distance. In this paper, we design and thoroughly optimize a complete state-of-the-art WPT system and compare it to an equivalent battery-powered solution to assess whether it effectively represents an eco-friendly alternative. The proposed WPT system is a 2.45-GHz simultaneous wireless information and power transfer (SWIPT) system with beamforming and a high peak-to-average power ratio waveform for supplying a battery-less passive infrared sensor for room occupancy tracking. The final prototype of the smart sensor consumes 4and can operate in steady state at 5from the remote power head with an overall power transfer efficiency of 17/and can cold-start at 3.5. We carry out a life-cycle assessment (LCA) of the environmental impacts through four indicators, i.e., primary energy demand, global warming potential, terrestrial ecotoxicity, and freshwater consumption. Results show that for a 10-year lifetime, the WPT alternative has at least 5.5 to 10.3× higher environmental impacts than its battery-powered equivalent. This demonstrates the importance of LCA during the design of IoT smart sensors and shows that the use of meter-range 2.45¯/GHz SWIPT should be limited to applications where conventional battery-powered systems cannot be used. Marco Gonzalez, Pengcheng Xu 0002, Remi Dekimpe, Maxime Schramme, Ivan Stupia, Thibault Pirson, David Bol |
IEEE Internet Things J. | 7 |
| 2023 | Bottom-Up and Top-Down Approaches for the Design of Neuromorphic Processing Systems: Tradeoffs and Synergies Between Natural and Artificial IntelligenceabstractWhile Moore’s law has driven exponential computing power expectations, its nearing end calls for new avenues for improving the overall system performance. One of these avenues is the exploration of alternative brain-inspired computing architectures that aim at achieving the flexibility and computational efficiency of biological neural processing systems. Within this context, neuromorphic engineering represents a paradigm shift in computing based on the implementation of spiking neural network architectures in which processing and memory are tightly colocated. In this article, we provide a comprehensive overview of the field, highlighting the different levels of granularity at which this paradigm shift is realized and comparing design approaches that focus on replicating natural intelligence (bottom-up) versus those that aim at solving practical artificial intelligence applications (top-down). First, we present the analog, mixed-signal, and digital circuit design styles, identifying the boundary between processing and memory through time multiplexing, in-memory computation, and novel devices. Then, we highlight the key tradeoffs for each of the bottom-up and top-down design approaches, survey their silicon implementations, and carry out detailed comparative analyses to extract design guidelines. Finally, we identify necessary synergies and missing elements required to achieve a competitive advantage for neuromorphic systems over conventional machine-learning accelerators in edge computing applications and outline the key ingredients for a framework toward neuromorphic intelligence. Charlotte Frenkel, David Bol, Giacomo Indiveri |
Proc. IEEE | 2 |
| 2023 | A 7T-NDR Dual-Supply 28-nm FD-SOI Ultra-Low Power SRAM With 0.23-nW/kB Sleep Retention and 0.8 pJ/32b Access at 64 MHz With Forward Back BiasabstractThis work presents a 16kB ultra-low power (ULP) SRAM macro in 28nm FD-SOI with high energy efficiency in active mode and ultra-low leakage (ULL) in sleep mode, embedded in the SleepRider micro-controller unit (MCU) intended for IoT edge applications. The proposed SRAM integrates custom 7T ULL bitcells based on negative differential resistance (NDR) structures and a pMOS-only write port, achieving$2.1\times $lower area than previous NDR-based bitcells. A dual-supply strategy combined with negative-wordline write-assist concurrently provides worst-case data retention and correct write operations, up to the 64-MHz MCU target frequency. The SRAM macro periphery combines several low-power techniques to extract the full potential of the novel 7T bitcells, reaching an unprecedented speed-energy-leakage optimum with only 2.5% area overhead. Adaptive forward body biasing (FBB) further improves active mode performance while ensuring robustness against PVT variations. Measurement results showcase a minimum energy point of 0.78pJ per 32b access (assuming 50% read/write) at 0.5V and 64MHz. Moreover, leakage power drops from 296nW/kB at 0.5V in idle conditions to 0.23nW/kB in sleep at the 0.46V data retention voltage (DRV), yielding more than$1000\times $leakage reduction. As such, the proposed SRAM achieves an excellent trade-off between area, leakage and energy in the 10-to-100MHz frequency range. Adrian Kneip, David Bol |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | UFBBR: A Unified Frequency and Back-Bias Regulation Unit for Ultralow-Power Microcontrollers in 28-nm FDSOIabstractSensitivity to process, voltage, and temperature (PVT) variations constitutes a serious obstacle in ultralow- voltage/ultralow-power (ULV/ULP) circuits and systems. To address this challenge, we propose a unified frequency/back-bias regulation (UFBBR) macro embedded in a custom ULP ARM Cortex-M4 microcontroller unit (MCU) manufactured in 28-nm FDSOI technology. The UFBBR technique combines the generation of a 32-to-80 MHz system clock and asymmetric adaptive back biasing for PVT compensation. Relying on a novel dual-output frequency-locked loop, it senses both the logic speed and the N/PMOS process imbalance using back-bias-controlled oscillators, and generates adequate forward back-bias voltages with digitally-controlled oscillators followed by switched-capacitor charge-pumps for fast current actuation. Compared to a situation with zero back biasing and appropriate frequency/voltage margins, the UFBBR provides$15\times $of frequency boosting at 0.4 V or 180 mV of voltage reduction at 64 MHz. It leverages software-programmable configuration knobs to achieve a fast wake-up of$8 ~\mu \text{s}$and an in-lock power of$22 ~\mu \text{W}$, with an area overhead below 0.032 mm2. In sleep, it drives the back-bias voltages towards 0 V and disables the clock for minimum power consumption. These features help the MCU system achieve a minimum energy point of$5.5 ~\mu \text{W}$/MHz and a sleep power of$7.7 ~\mu \text{W}$. Maxime Schramme, David Bol |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | A Low-Complexity LoRa Synchronization Algorithm Robust to Sampling Time OffsetsabstractLoRaWAN is nowadays one of the most popular protocols for low-power Internet of Things communications. Although its physical layer, namely LoRa, has been thoroughly studied in the literature, aspects related to the synchronization of LoRa receivers have received little attention so far. The estimation and correction of carrier frequency and sampling time offsets (STOs) are, however, crucial to attain the low sensitivity levels offered by the LoRa spread-spectrum modulation. The goal of this article is to build a low-complexity, yet efficient synchronization algorithm capable of correcting both offsets. To this end, a complete analytical model of a LoRa signal corrupted by these offsets is first derived. Using this model, we propose a new estimator for the STO. We also show that the estimations of the carrier frequency and the STOs cannot be performed independently. Therefore, to avoid a complex joint estimation of both offsets, an iterative low-complexity synchronization algorithm is proposed. To reach a packet error rate of 10−3, performance evaluations show that the proposed receiver requires only 1 or 2 dB higher signal-to-noise ratio than a theoretical perfectly synchronized receiver, while incurring a very low computational overhead. Mathieu Xhonneux, Orion Afisiadis, David Bol, Jérôme Louveaux |
IEEE Internet Things J. | 3 |
| 2022 | Accurate and Insightful Closed-Form Prediction of Subthreshold SRAM Hold Failure RateabstractThe failure probabilities of industrial SRAM cells fall below the ppm (10−6) range, disqualifying the computational-intensive Monte-Carlo simulations for efficient robustness assessment. Starting from a novel two-dimensional threshold voltage imbalance representation, we propose a new methodology for fast and accurate prediction of the subthreshold SRAM hold stability failure rate. The probability is derived in a closed form which involves the transistor threshold-voltage standard deviations and only requires the two quick DC extractions of the worst- and best-case static noise margins. We validate our approach on a Six-Transistor (6T) bitcell in 28 nm Fully Depleted Silicon-On-Insulator (FD-SOI) CMOS technology. Our method turns out to be especially insightful for comparative and sensitivity analyses, for instance to study the effect of supply voltage downscaling or temperature variations. Finally, we show that the achieved accuracy and the capability of estimating extremely low failure probabilities (down to 10−9), combined with the important gain in simulation and post-processing cost, makes our methodology attractive compared to other recent modelling works. Léopold Van Brandt, Roghayeh Saeidi, David Bol, Denis Flandre |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | A Family of Current References Based on 2T Voltage References: Demonstration in 0.18-μm With 0.1-nA PTAT and 1.1-μA CWT 38-ppm/°C DesignsabstractThe robustness of current and voltage references to process, voltage and temperature (PVT) variations is paramount to the operation of integrated circuits in real-world conditions. However, while recent voltage references can meet most of these requirements with a handful of transistors, current references remain rather complex, requiring significant design time and silicon area. In this paper, we present a family of simple current references consisting of a two-transistor (2T) ultra-low-power voltage reference, buffered onto a voltage-to-current converter by a single transistor. Two topologies are fabricated in a 0.18-$\mu \text{m}$partially-depleted silicon-on-insulator (SOI) technology and measured over 10 dies. First, a 7T nA-range proportional-to-absolute-temperature (PTAT) reference intended for constant-$g_{m}$biasing of subthreshold operational amplifiers demonstrates a 0.096-nA current with a line sensitivity (LS) of 1.48 %/V, a temperature coefficient (TC) of 0.75 %/°C, and a variability$(\sigma /\mu)$of 1.66 %. Then, two 4T+1R$\mu \text{A}$-range constant-with-temperature (CWT) references with (resp. without) TC calibration exhibit a 1.09-$\mu \text{A}$(resp. 0.99-$\mu \text{A}$) current with a 0.21-%/V (resp. 0.20-%/V) LS, a 38-ppm/°C (resp. 290-ppm/°C) TC, and a 0.87-% (resp. 0.65-%)$(\sigma /\mu)$. In addition, portability to common scaled CMOS technologies, such as 65-nm bulk and 28-nm fully-depleted SOI, is discussed and validated through post-layout simulations. Martin Lefebvre 0002, David Bol |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Comprehensive Analytical Comparison of Ring Oscillators in FDSOI Technology: Current Starving Versus Back-Bias ControlabstractBack-bias control is a new degree of freedom brought by fully-depleted silicon-on-insulator (FDSOI) CMOS technologies, which can be used to control the oscillation frequency of voltage-controlled ring oscillators (VCROs). The resulting VCRO architecture is called a back-bias-controlled oscillator (BBCO). This paper compares it with the conventional current-starved ring oscillator (CSRO) topology in terms of power consumption and phase noise figure-of-merit (FoM), while taking practical design constraints of process-voltage-temperature (PVT) robustness and frequency tuning range into account. The proposed comprehensive analysis takes advantage of relevant and compact analytical models, as well as extensive pre-layout simulation results. The comparison is made at four different target oscillation frequencies, which are representative of frequency synthesis for WiFi/Bluetooth/LPWAN wireless communications and of clock generation for smartphone/Internet-of-Things processors: 300 MHz, 868 MHz, 2.45 GHz, and 5.18 GHz. In 28-nm FDSOI technology, the results demonstrate that BBCOs can intrinsically reach 1.69 to$4.63\times $lower minimum power consumption and slightly better FoM values than CSROs. Maxime Schramme, Léopold Van Brandt, Denis Flandre, David Bol |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | Moore's Law and ICT Innovation in the AnthropoceneabstractIn information and communication technologies (ICTs), innovation is intrinsically linked to empirical laws of exponential efficiency improvement such as Moore's law. By following these laws, the industry achieved an amazing relative decoupling between the improvement of key performance indicators (KPIs), such as the number of transistors, from physical resource usage such as silicon wafers. Concurrently, digital ICTs came from almost zero greenhouse gas emission (GHG) in the middle of the twentieth century to direct annual carbon footprint of approximately 1400 MT CO2e today. Given the fact that we have to strongly reduce global GHG emissions to limit global warming below 2°C, it is not clear if the simple follow-up of these trends can decrease the direct GHG emissions of the ICT sector on a trajectory compatible with Paris agreement. In this paper, we analyze the recent evolution of energy and carbon footprints from three ICT activity sub-sectors: semiconductor manufacturing, wireless Internet access and datacenter usage. By adopting a Kaya-like decomposition in technology affluence and efficiency factors, we find out that the KPI increase failed to reach an absolute decoupling with respect to total energy consumption because the technology affluence increases more than the efficiency. The same conclusion holds for GHG emissions except for datacenters, where recent investment in renewable energy sources lead to an absolute GHG reduction over the last years, despite a moderate energy increase. We formulate hypotheses for this absence of absolute decoupling from three scientific fields: ecological economics, economics of technology and sociology of technology. We argue that aligning direct GHG emissions of the ICT sector on a trajectory compatible with Paris agreement requires an ecological transition in innovation by adopting sobriety in addition to efficiency. David Bol, Thibault Pirson, Remi Dekimpe |
DATE | 1 |
| 2021 | Impact of Analog Non-Idealities on the Design Space of 6T-SRAM Current-Domain Dot-Product Operators for In-Memory ComputingabstractIn-memory computing provides unprecedented power and area efficiency for the execution of convolutional neural networks by using memory bitcells to perform dot-product (DP) operations in the analog domain. Yet, these operators suffer from analog non-idealities (ANIs) that degrade the inference accuracy. This paper proposes design guidelines inferred from a holistic simulation-based analysis of the impact of ANIs on the accuracy-efficiency trade-off that affects current-domain DP operators based on conventional 6T-SRAM bitcell arrays. We define a custom SNR metric aware of the DP operand distribution to quantify decision errors associated with various ANIs, over ranges of input/output resolution and hardware design parameters. We find out that non-linearity and local mismatch are the dominant ANIs limiting the design space, while IR drops turn out to be critical only when targeting high parallelism. We then quantify the accuracy-efficiency trade-off related to these dominant ANIs across the design space and propose optimal design choices. We notably identify that using larger operators can either improve or worsen the SNR depending on the target output resolution. Furthermore, we show that hardware calibration techniques which mitigate mismatch help to recover a fraction of the lost SNR, with greater effectiveness when scaling down the supply voltage. Adrian Kneip, David Bol |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2020 | A 28-nm Convolutional Neuromorphic Processor Enabling Online Learning with Spike-Based RetinasabstractIn an attempt to follow biological information representation and organization principles, the field of neuromorphic engineering is usually approached bottom-up, from the biophysical models to large-scale integration in silico. While ideal as experimentation platforms for cognitive computing and neuroscience, bottom-up neuromorphic processors have yet to demonstrate an efficiency advantage compared to specialized neural network accelerators for real-world problems. Top-down approaches aim at answering this difficulty by (i) starting from the applicative problem and (ii) investigating how to make the associated algorithms hardware-efficient and biologically-plausible. In order to leverage the data sparsity of spike-based neuromorphic retinas for adaptive edge computing and vision applications, we follow a top-down approach and propose SPOON, a 28-nm event-driven CNN (eCNN). It embeds online learning with only 16.8-% power and 11.8-% area overheads with the biologically-plausible direct random target projection (DRTP) algorithm. With an energy per classification of 313nJ at 0.6V and a 0.32-mm2area for accuracies of 95.3% (on-chip training) and 97.5% (off-chip training) on MNIST, we demonstrate that SPOON reaches the efficiency of conventional machine learning accelerators while embedding on-chip learning and being compatible with event-based sensors, a point that we further emphasize with N-MNIST benchmarking. Charlotte Frenkel, Jean-Didier Legat, David Bol |
ISCAS | 3 |
| 2019 | A 65-nm 738k-Synapse/mm2 Quad-Core Binary-Weight Digital Neuromorphic Processor with Stochastic Spike-Driven Online LearningabstractRecent trends in the field of artificial neural networks (ANNs) and convolutional neural networks (CNNs) investigate weight binarization for full on-chip weight storage to minimize circuit resources and to avoid the high energy cost of off-chip memory accesses. In parallel, spiking neural network (SNN) architectures are explored to further reduce power when processing sparse event-based data streams, while on-chip spike-based online learning targets applications constrained in power and resources during the training phase. However, leveraging high-density on-chip online learning in binary-weight SNNs is still an open challenge. In this work, we demonstrate MorphIC, a quad-core binary-weight digital neuromorphic processor embedding a stochastic version of the spike-driven synaptic plasticity (S-SDSP) learning rule and a hierarchical routing fabric for large-scale chip interconnection. The MorphIC SNN processor embeds a total of 2k leaky integrate-and-fire (LIF) neurons and more than two million plastic synapses for an active silicon area of 2.86mm2in 65nm CMOS, achieving a high density of 738k synapses/mm2. Charlotte Frenkel, Jean-Didier Legat, David Bol |
ISCAS | 3 |
| 2019 | A 0.4V 0.5fJ/cycle TSPC Flip-Flop in 65nm LP CMOS with Retention Mode Controlled by Clock-Gating CellsabstractIn this paper, we propose a low-overhead solution to ensure contention-free data retention in clock-gated true single-phase-clock (TSPC) flip-flops (FF) at ultra-low voltage (ULV). It relies on a retention feedback loop added to the TSPC FF and controlled by the clock-gating module. When the clock is gated, the retention is enabled, which drives the FF in retention mode. This limits the energy overhead induced by the added feedback loop and makes the FF contention-free. Moreover, as several FFs typically share the same clock-gating module, the control signal generation overhead is also kept low. The proposed 19T TSPC FF with retention mode was implemented as a standard cell in 65nm LP CMOS. The FF energy is 0.5fJ/cycle at 0.4V, from post-layout simulations and for a typical 25% activity factor, which is 62% reduction compared to the conventional 24T master-slave FF. Experimental validation of a prototyped Cortex-M0 testchip including the integration of the proposed FF into synthesis and place/route flow validates its robust operation at ULV. Ludovic Moreau, Remi Dekimpe, David Bol |
ISCAS | 3 |
| 2019 | A battery-less BLE smart sensor for room occupancy tracking supplied by 2.45-GHz wireless power transfer
Remi Dekimpe, Pengcheng Xu 0002, Maxime Schramme, Pierre Gérard, Denis Flandre, David Bol |
Integr. | 6 |
| 2018 | Gradient importance sampling: An efficient statistical extraction methodology of high-sigma SRAM dynamic characteristicsabstractThe impact of within-die transistor variability has increased with CMOS technology scaling up to the point where it has emerged as a systematic problem for the designer. Estimating extremely low failure rate, i.e. “high-sigma” probabilities, by the conventional Monte Carlo (MC) approach requires millions of simulation runs, making it an impractical approach for circuit designers. To overcome this problem, alternative failure estimation methodologies, which require a smaller number of runs have been proposed. In this paper, we propose a novel methodology called “gradient importance sampling” (GIS) for fast statistical extraction of high-sigma circuit characteristics. It is based on conventional Importance Sampling combined with a gradient-based approach to find the most probable failure point (MPFP). By applying GIS to extract SRAM dynamic characteristics in 28nm FDSOI CMOS, we show that the proposed methodology is straightfor-ward, computationally efficient and the results are in line with those obtained via standard MC. To the best of our knowledge, the GIS results are the best in their class for low failure rate estimation. Thomas Haine, Johan Segers, Denis Flandre, David Bol |
DATE | 4 |
| 2018 | Multilevel Half-Rate Phase Detector for Clock and Data Recovery Circuits
Cecilia Gimeno, David Bol, Denis Flandre |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | A fully-synthesized 20-gate digital spike-based synapse with embedded online learningabstractNeuromorphic engineering aims at building cognitive systems made of electronic neuron and synapse circuits. These emerging computing architectures have a high potential for real-world problems that are difficult to formalize and program, such as vision or sensorimotor control. In order to leverage the potential of neuromorphic engineering and study cognition principles in physical systems, the development of autonomous online learning is a key feature. However, to develop scalable systems that can be used in realistic applications, it is crucial to design compact and low-power hardware platforms. Here we analyze a spike-driven synaptic plasticity (SDSP) learning rule and show that it is particularly well suited for highly compact digital synapse implementations, especially if compared to conventional spike-timing-dependent plasticity (STDP) rules. Furthermore, we designed an asynchronous fully-synthesizable digital synapse circuit with embedded SDSP-based online learning features, and with programmability options for versatile computing. The proposed synapse implementation requires only 20 gates for a compact area of 25μm2 in a 28nm FDSOI CMOS process. Charlotte Frenkel, Giacomo Indiveri, Jean-Didier Legat, David Bol |
ISCAS | 4 |
| 2017 | Integration of level shifting in a TSPC flip-flop for low-power robust timing closure in dual-Vdd ULV circuitsabstractUltra-low-voltage (ULV) operation of logic circuits is an interesting solution to reduce power consumption in digital circuits. However the always-on blocks at nominal Vdd are necessary for functionality or I/O communications, which induces complex timing closure. In this paper, we propose a 5.4fJ/cycle 0.4V to 1.2V level-shifting flip-flop in 28nm FDSOI, which simplifies the clock tree constraints between the power domains. François Stas, David Bol |
ISCAS | 2 |
| 2017 | A 0.4V 0.08fJ/cycle retentive True-Single-Phase-Clock 18T Flip-Flop in 28nm FDSOI CMOSabstractIn this paper, we propose an 18-transistor (18T) True-Single-Phase-Clock (TSPC) Flip-Flop (FF) with static data retention based on two forward-conditional feedback loops, without increasing the clock load, in comparison to the baseline TSPC architecture. The proposed FF was implemented for ultra-low-voltage (ULV) operation in 28nm FDSOI CMOS. The performances of the proposed FF extracted from measurements of clock dividers are compared to reference designs including the conventional M-S FF, the baseline TSPC FF and a recently-proposed retentive TSPC FF. Compared to the conventional MS FF, the proposed FF shows respectively 5%, 60% and 30% improvements at 0.4V in maximum frequency, energy/cycle and leakage power. François Stas, David Bol |
ISCAS | 2 |
| 2016 | Comparative analysis of redundancy schemes for soft-error detection in low-cost space applicationsabstractSingle-Event Effects are an increasingly important issue in electronic circuits due to technology scaling, efficient error detection schemes are thus required for circuits dedicated to radiative environments, such as in space applications. This work shows that the widespread spatial and temporal redundancy schemes exhibit widely different performances depending on technology, environment and circuit architecture parameters. Following these results, three new redundancy schemes are proposed and compared: one of them, the Forward Temporal Redundancy, stands out as it achieves full error detection with limited timing penalty at only 100% sequential and 45% combinational overheads for a benchmark pipelined MIPS microprocessor. Charlotte Frenkel, Jean-Didier Legat, David Bol |
VLSI-SoC | 3 |
| 2015 | Analysis and optimization for dynamic read stability in 28nm SRAM bitcellsabstractThe importance of the dynamic analysis for SRAM operation increases as a result of shrinking access cycle time, voltage scaling and increased process variations. In this paper, quantitative study of the dynamic read noise margin (DNM) is introduced showing the evolution from the static read noise margin (SNM) to DNM through cumulative dynamic effects in 28nm FDSOI. The impact of parasitic capacitances on the DNM is further analyzed. Finally, we show that by sizing for a 150-mV DNM instead of a 150-mV SNM and by inserting two 0.5fF extra caps in the bitcell allows reducing the pull-down NMOS width by a factor 3.5×. Ahmed T. Elthakeb, Thomas Haine, Denis Flandre, Yehea I. Ismail, Hamdy Abd Elhamid, David Bol |
ISCAS | 6 |
| 2014 | Bellevue: A 50MHz variable-width SIMD 32bit microcontroller at 0.37V for processing-intensive wireless sensor nodesabstractIn the context of wireless sensor nodes for the Internet-of-Things, there is a need for low-power high-performance computing cores for video monitoring applications. In this paper we present a custom 50MHz 32-bit microcontroller running at 0.37V built on a 65nm LP/GP CMOS process. Part of an energy-harvesting SoC with on-chip CMOS imager, it features adaptive voltage scaling, low-power 1.55μW sleep mode, and a variable-width SIMD pipeline and multiply/divide unit, achieving 7.7μW/MHz overall. François Botman, Julien De Vos, Sebastien Bernard, François Stas, Jean-Didier Legat, David Bol |
ISCAS | 6 |
| 2014 | Data-Dependent Operation Speed-Up Through Automatically Inserted Signal Transition Detectors for Ultralow Voltage Logic CircuitsabstractWith the advent of mobile electronics requiring ever more computing power from a limited energy supply, there is a need for efficient systems capable of maximizing this ratio. Architectural enhancements must therefore be designed to enable high performance, all the while maintaining the power advantage. The technique proposed in this paper allows the acceleration of combinatorial circuits beyond the performance generally achievable by conventional synthesis and timing closure, by exploiting the data-dependent delay variations inherent in such circuits. Through the automatic insertion of transition detectors within the target circuit, the progress of operations underway can be monitored and prematurely completed, thereby increasing the operation speed from the worst toward the average case. In addition, a synthesis flow is proposed to increase the proportion of fast paths, thereby increasing the technique's impact. The proposed technique was applied automatically to a series of benchmark circuits, and the synthesis results show it to achieve good performance, with an average increase of 29% over conventional synthesis, for an average energy increase of${<}{21\%}$overall. François Botman, David Bol, Jean-Didier Legat, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Impact of back gate biasing schemes on energy and robustness of ULV logic in 28nm UTBB FDSOI technologyabstractMinimum energy per operation is typically achieved in the subthreshold region where low speed and low robustness are two challenging problems. This paper studies the impact of Back Biasing (BB) schemes on these features for FDSOI technology. We show that Forward BB can help cover a wider design space in term of optimal frequency of operation while keeping minimum energy. Asymmetric BB between NMOS and PMOS can mitigate the effect of systematic mismatch on Minimum Energy Point (MEP) and robustness. With optimal asymmetric BB, we achieve either a MEP reduction up to 18% or a 36× speedup at the MEP. Guerric de Streel, David Bol |
ISLPED | 2 |
| 2012 | Towards Green Cryptography: A Comparison of Lightweight Ciphers from the Energy Viewpoint
Stéphanie Kerckhof, François Durvaux, Cédric Hocquet, David Bol, François-Xavier Standaert |
CHES | 4 |
| 2010 | Robustness-aware sleep transistor engineering for power-gated nanometer subthreshold circuitsabstractIn ultra-low-power applications with long standby periods, power-gating technique can be combined with sub-threshold operation to minimize energy. However, in nanometer technologies, we show in this paper that the introduction of the sleep transistor threatens subthreshold circuit robustness because of noise margin degradation. An increase in Vddto maintain robustness limits the achievable sleep-mode leakage power reduction to 100× with up to 60% active-mode energy penalty. We therefore propose a framework to engineer the sleep transistor under robustness constraint, which shows that a std-Vtlong-channel MOSFET is the optimum sleep transistor with 170× leakage reduction at only 20% energy penalty. David Bol, Cédric Hocquet, Denis Flandre, Jean-Didier Legat |
ISCAS | 1 |
| 2010 | Nanometer MOSFET Effects on the Minimum-Energy Point of Sub-45nm Subthreshold Logic - Mitigation at Technology and Circuit LevelsabstractSubthreshold operation of digital circuits enables minimum energy consumption. In this article, we observe that minimum energy E min of subthreshold logic dramatically increases when reaching 45nm CMOS node. We demonstrate by circuit simulation and analytical modeling that this increase comes from the combined effects of variability, gate leakage, and Drain-Induced Barrier Lowering (DIBL) effect. We then investigate the new impact of individual MOSFET parameters L g , V t , and T ox on E min in sub-45nm technologies. We further propose an optimum MOSFET selection, which favors low-V t mid-L g devices in 45nm CMOS technology. The use of such optimum MOSFETs yields 35% E min reduction for a benchmark multiplier with good speed performances and negligible area overhead. This optimum MOSFET selection can easily be integrated into a standard EDA tool flow by appropriate selection of the standard cell library. We finally demonstrate that undoped-channel fully-depleted Silicon-On-Insulator (SOI) technology brings 60% E min reduction with baseline MOSFETs thanks to strong mitigation of variability and short-channel effects. This study reveals a new (à priori counterintuitive) paradigm in device optimization for subthreshold logic: relaxing gate leakage constraints to improve robustness against short-channel effects and variability. Additionally, we propose pre-Silicon BSIM4 MOSFET model cards for realistic subthreshold circuit simulations including variability in bulk and fully depleted SOI technologies, which are made available online. David Bol, Denis Flandre, Jean-Didier Legat |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2009 | Technology flavor selection and adaptive techniques for timing-constrained 45nm subthreshold circuitsabstractWe investigate techniques to design 45nm minimum-energy subthreshold CMOS circuits under timing constraints, considering the practical case of an 8-bit multiplier. We first show that technology flavor and Vt selections shift minimum-energy point to different operating frequencies, thereby enabling minimum energy in either low- or mid-performance applications. However, we demonstrate that independent dual-Vt assignment to save leakage in non-critical paths is not feasible. We then show that reverse adaptive body biasing (ABB) is potentially more efficient to compensate for global process/temperature variations than adaptive voltage scaling and forward ABB. Nevertheless, its practical efficiency is limited by the affordable VBB range and the value of the body-effect coefficient. David Bol, Denis Flandre, Jean-Didier Legat |
ISLPED | 1 |
| 2009 | Nanometer MOSFET effects on the minimum-energy point of 45nm subthreshold logicabstractIn this paper, we observe that minimum energy Emin of subthreshold logic dramatically increases when reaching 45nm node. We demonstrate by circuit simulation and analytical modeling that this increase comes from the combined effects of variability, gate leakage and DIBL. We then investigate the new impact of MOSFET parameters on Emin in nanometer technologies. We finally propose an optimum MOSFET selection intended for subthreshold circuit designers, which favors low-Vt mid-Lg devices in standard 45nm GP technology. The use of such optimum MOSFETs yields 35% Emin reduction for a benchmark multiplier with good speed performances and negligible area overhead. David Bol, Dina Kamel, Denis Flandre, Jean-Didier Legat |
ISLPED | 1 |
| 2009 | Interests and Limitations of Technology Scaling for Subthreshold LogicabstractSubthreshold logic is an efficient technique to achieve ultralow energy per operation for low-to-medium throughput applications. In this paper, the interests and limitations of technology scaling for subthreshold logic are investigated from 0.25 mum to 32 nm nodes. Scaling to 90/65 nm nodes is shown to be highly desirable for medium-throughput applications (1-10 MHz) due to great dynamic energy reduction. However, this interest is limited at 45/32 nm nodes by high static energy due to degraded subthreshold swing and delay variability. Moreover, for low-throughput applications (10-100 kHz), this limitation is worsened by the increase of minimum supply voltage to achieve sufficient functional yield, which results in bad energy efficiency starting at 0.13 mum node. Upsizing the channel length is proposed as a straightforward circuit-level technique to efficiently mitigate these effects. At 32 nm node, this technique reduces energy per operation by 60% at medium throughput and by two orders of magnitude at low throughput. David Bol, Renaud Ambroise, Denis Flandre, Jean-Didier Legat |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2008 | Analysis and minimization of practical energy in 45nm subthreshold logic circuitsabstractOver the last decade, the design of ultra-low-power digital circuits in subthreshold regime has been driven by the quest for minimum energy per operation. In this contribution, we observe that operating at minimum-energy point is not straightforward as design constraints from real-life applications have an important impact on energy. Therefore, we introduce the alternative concept of practical energy, taking functional-yield and throughput constraints on minimum Vddinto account. In this context, we demonstrate for the first time the detrimental impact of DIBL on minimum Vdd. Practical energy gives a useful analysis framework of circuit optimization to reach minimum-energy point, while considering the throughput as an input variable dictated by the application. From simulation of a benchmark multiplier in 45 nm technology, we find out that practical energy can be far higher than minimum energy point, in the case of low-throughput applications (ap 10-100 kOp/s) because of static leakage energy and robustness-limited minimum Vdd. With the proposed framework, we investigate the capability of conventional optimization techniques to make practical energy meet minimum energy point. Amongst these techniques, channel length upsize is shown to be more efficient than MTCMOS power gating, body biasing, Vt selection or device width upsize, as it increases robustness while simultaneously reducing static leakage energy. A small length upsize with low area overhead is shown to reduce practical energy at low throughput to less than 2.1 times the minimum energy level. At medium throughput, it even brings practical energy 30% lower than minimum energy level without optimization techniques. David Bol, Renaud Ambroise, Denis Flandre, Jean-Didier Legat |
ICCD | 1 |
| 2006 | Low-Cost Elliptic Curve Digital Signature Coprocessor for Smart CardsabstractThis paper proposes different low-cost coprocessors for public key authentication on 8-bit smart cards. Elliptic curve cryptography is used for its efficiency per bit of key and the Elliptic Curve Digital Signature Algorithm is chosen. For this functionality, an area constrained coprocessor is probably the best approach to perform the most computer-intensive operations at an acceptable speed considering the limited memory and power of the selected platform. For that purpose, the scalar point multiplication in GF(2m) in both affine and projective coordinates was implemented in order to compare their performances with the same level of optimization and the same technology. A hardware/ software co-design strategy was also used to avoid the need of a dedicated register file. Guerric Meurice de Dormale, Renaud Ambroise, David Bol, Jean-Jacques Quisquater, Jean-Didier Legat |
ASAP | 3 |