Peter Debacker

dblp:118/8568 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
9since 2021 · last 2023
0000-0003-3825-5554ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2023 Evaluating the Effects of FeFET Device Variability on Charge Sharing Based AiMC Accelerator
abstract
With shrinking device geometries, the device-to-device (D2D) variability of emerging memory devices, affect the performance of deep neural networks (DNN) and restrict the use of standalone devices to be used as Analog in-Memory Computing (AiMC) accelerators. For FeFET based AiMC designs, such effects can be controlled by using resistors or capacitors connected in series with FeFETs. However, in both cases, the effect of variability will again start dominating when the programming window shrinks further due to scaling towards smaller nanometer dimensions. Additionally, while resistance based solutions will add to the existing IR drop issues within the macro, capacitor based approach avoids such a case. In this paper, we evaluate a FeFET and capacitor based AiMC accelerator and study how finetuning cell operating parameters along with design modifications of the AiMC macro, can help in achieving a variability aware AiMC design. Using a 22-nm CMOS PDK augmented with FeFET model based on measured data, evaluation on a DNN with CIFAR-10 benchmark, demonstrated that following our approach, the implemented DNN achieves an accuracy of ∼85% which is an improvement of ∼13% over an untuned AiMC macro.
Swatilekha Majumdar, Stefan Cosemans, Arindam Mallik, Peter Debacker, Francky Catthoor, Jan Van Houdt
ISCAS4
2022 Learn to Learn on Chip: Hardware-aware Meta-learning for Quantized Few-shot Learning at the Edge
abstract
Recent years have seen a growing trend of deploying deep neural network-based applications on edge devices. Many of these applications, such as biometric identification, activity tracking, user preference learning, etc., require fine-tuning of the trained networks for user personalization. One way to prepare these models to handle new, unseen tasks, is to pre-train them on a distribution of known tasks. This observation has led to increasing research into meta-learning based few-shot learning techniques. However, basic meta-learning approaches do not account for the limited memory and computational resources during on-chip training. We propose a modified meta-learning algorithm that enables quantized fine-tuning to optimally condition the models for on-chip few shot learning. The modification involves the inclusion of target hardware constraints upfront in the meta-learning process. Block floating point datatypes with low precision mantissa bits are utilized in the forward and backward passes, to allow hardware-friendly adaptation. Experiments show that our algorithm provides better initializations than conventional algorithms, more suitable for efficient quantized fine-tuning. This allows the few-shot learner to achieve better convergence, in terms of accuracy and speed. Extensive experiments are also performed to analyze the impact of initialization on quantized fine-tuning and further corroborate the benefits of our method.
Nitish Satya Murthy, Peter Vrancx, Nathan Laubeuf, Peter Debacker, Francky Catthoor, Marian Verhelst
SEC4
2022 Tiny ci-SAR A/D Converter for Deep Neural Networks in Analog in-Memory Computation
abstract
This paper presents a tiny charge injection-Successive Approximation (ci-SAR) A/D converter (ADC) to be integrated at the periphery of analog Matrix Vector Multiplication (MVM) accelerators for Deep Neural Network (DNN) inference. Derived from the ci-SAR ADC, this converter exploits a single charge injecting cell to minimize area and energy consumption. The ADC exhibits a signal-to-noise and distortion ratio of 30.5 dB, at 5 bits of nominal resolution. The energy per conversion is 86 fJ, running at 34 MS/s, with a silicon area of $75 \mu m ^{2}$, in 22 nm technology node. From the results of our analytical framework, an SRAM-based Analog in-Memory Compute (AiMC) array, including the proposed ADC at 5 bits of resolution, can achieve an energy efficiency of 1650 TOPs/W.
Michele Caselli, Debjyoti Bhattacharjee, Arindam Mallik, Peter Debacker, Diederik Verkest
ISCAS4
2022 Write-Verify Scheme for IGZO DRAM in Analog in-Memory Computing
abstract
Large weight variations cause significant degradation in Deep Neural Networks (DNNs) accuracy in machine learning (ML) context. Indium-Gallium-Zinc-Oxide (IGZO) DRAM compute cell is a promising option for Analog in-Memory Computing (AiMC) accelerators, but its applicability requires the compensation of large variations affecting the voltage threshold of the readout device. This paper proposes a write-verify scheme for IGZO-based AiMC accelerator, designed in 22-nm technology, based on a compensation loop operating on the analog weight value stored in the IGZO DRAM cell. After the compensation routine, the IGZO IONnormalized variation drops from $\pm \mathbf{2 7} \%$ to $\pm 3 \%$ in simulation. With sufficiently large weight reuse, the additional energy spent for the compensation of the entire arrays becomes negligible, recovering the 2000 TOPS/W baseline performance of an ideal IGZO array without the write-verify.
Michele Caselli, Subhali Subhechha, Peter Debacker, Arindam Mallik, Diederik Verkest
ISCAS3
2022 Dynamic Quantization Range Control for Analog-in-Memory Neural Networks Acceleration
abstract
Analog in Memory Computing (AiMC) based neural network acceleration is a promising solution to increase the energy efficiency of deep neural networks deployment. However, the quantization requirements of these analog systems are not compatible with state-of-the-art neural network quantization techniques. Indeed, while the quantization of the weights and activations is considered by modern deep neural network quantization techniques, AiMC accelerators also impose the quantization of each Matrix Vector Multiplication (MVM) result. In most demonstrated AiMC implementations, the quantization range of MVM results is considered a fixed parameter of the accelerator. This work demonstrates that dynamic control over this quantization range is possible but also desirable for analog neural networks acceleration. An AiMC compatible quantization flow coupled with a hardware aware quantization range driving technique is introduced to fully exploit these dynamic ranges. Using CIFAR-10 and ImageNet as benchmarks, the proposed solution results in networks that are both more accurate and more robust to the inherent vulnerability of analog circuits than fixed quantization range based approaches.
Nathan Laubeuf, Jonas Doevenspeck, Ioannis A. Papistas, Michele Caselli, Stefan Cosemans, Peter Vrancx, Debjyoti Bhattacharjee, Arindam Mallik, Peter Debacker, Diederik Verkest, Francky Catthoor, Rudy Lauwereins
ACM Trans. Design Autom. Electr. Syst.9
2022 Carbon Nanotube SRAM in 5-nm Technology Node Design, Optimization, and Performance Evaluation - Part I: CNFET Transistor Optimization
abstract
In this article, we propose a carbon nanotube (CNT) field-effect transistor (CNFET)-based static random access memory (SRAM) design at the 5-nm technology node that is optimized based on the tradeoff between performance, stability, and power efficiency. In addition to size optimization, physical model parameters including CNT density, CNT diameter, and CNFET flat band voltage are evaluated and optimized for CNFET SRAM performance improvement. Optimized CNFET SRAM is compared with state-of-the-art 7-nm FinFET SRAM cell based on Arizona State University [ASAP 7-nm FinFET predictive technology models (PTM)] library. We find that the read, write EDPs, and static power of the proposed CNFET SRAM cell are improved by 67.6%, 71.5%, and 43.6%, respectively, compared with the FinFET SRAM cell, with slightly better stability. CNT interconnects both inside and in-between CNFET SRAM cells are considered to compose an all-carbon-based SRAM (ACS) array which will be discussed in the Part II of this article. A 7-nm FinFET SRAM cell with copper interconnects is implemented and used for comparison.
Rongmei Chen, Yuanqing Cheng, Souhir Elloumi, Kangwei Xu, Vihar P. Georgiev, Kai Ni 0004, Peter Debacker, A. Asenov, Aida Todri
IEEE Trans. Very Large Scale Integr. Syst.10
2022 Carbon Nanotube SRAM in 5-nm Technology Node Design, Optimization, and Performance Evaluation - Part II: CNT Interconnect Optimization
abstract
The size and parameter optimization for the 5-nm carbon nanotube field effect transistor (CNFET) static random access memory (SRAM) cell was presented in Part I of this article. Based on that work, we propose a carbon nanotube (CNT) SRAM array composed of the schematically optimized CNFET SRAM and CNT interconnects. We consider the interconnects inside the CNFET SRAM cell composed of metallic single-wall CNT (M-SWCNT) bundles to represent the metal layers 0 and 1 (M0 and M1). We investigate the layout structure of CNFET SRAM cell considering CNFET devices, M-SWCNT interconnects, and metal electrode Palladium with CNT (Pd-CNT) contacts. Two versions of cell layout designs are explored and compared in terms of performance, stability, and power efficiency. Furthermore, we implement a 16 Kbit SRAM array composed of the proposed CNFET SRAM cells, multiwall CNT (MWCNTs) inter-cell interconnects and Pd-CNT contacts. Such an array shows significant advantages, with the read and write overall energy-delay product (EDP), static power consumption, and core area of$0.28\times $,$0.52\times $, and$0.76\times $respectively to 7-nm FinFET-SRAM array with copper interconnects, whereas the read and write static noise margins are 6% and 12% respectively larger than the FinFET counterpart.
Rongmei Chen, Yuanqing Cheng, Souhir Elloumi, Kangwei Xu, Vihar P. Georgiev, Kai Ni 0004, Peter Debacker, A. Asenov, Aida Todri
IEEE Trans. Very Large Scale Integr. Syst.10
2021 Noise tolerant ternary weight deep neural networks for analog in-memory inference
abstract
Analog in memory computing (AiMC) is a promising hardware solution to efficiently perform inference with deep neural networks (DNNs). Similar to digital DNN accelerators, AiMC systems benefit from aggressively quantized DNNs. In addition, AiMC systems also suffer from noise on activations and weights. Training strategies to condition DNNs against weight noise can increase the efficiency of AiMC systems by enabling the use of more compact but more noisy weight memory devices. In this work, we utilize noise-aware training and introduce gradual noise training and network width scaling to increase the tolerance of DNNs against weight noise. Our results show that noise-aware training and gradual noise training drastically lowers the impact of weight noise without changing the network size. By utilizing network width scaling, the weight noise tolerance is increased even more with the penalty of more network parameters.
Jonas Doevenspeck, Peter Vrancx, Nathan Laubeuf, Arindam Mallik, Peter Debacker, Diederik Verkest, Rudy Lauwereins, Wim Dehaene
IJCNN5
2021 Design-Technology Space Exploration for Energy Efficient AiMC-Based Inference Acceleration
abstract
Extremely energy-efficient convolutional neural network inference (CNN) is recently enabled by analog in-memory compute (AiMC). The integration of AiMC in a primarily digital inference system brings new challenges ranging from device specifications to defining novel system architectures. A novel framework to evaluate the impact of AiMC array at the system level is presented. The framework is used to model a SRAM-based 1024 × 512 prototype AiMC array, capable of energy efficiency of upto 675 TMACs/W. The proposed framework allows modelling different compute cells, array dimensions, operating voltage, and activation buffer energy models which can be used to determine the overall energy efficiency for various CNN workloads.
Debjyoti Bhattacharjee, Nathan Laubeuf, Stefan Cosemans, Ioannis A. Papistas, Arindam Mallik, Peter Debacker, Myung Hee Na, Diederik Verkest
ISCAS6
2019 A Comparative Analysis on the Impact of Bank Contention in STT-MRAM and SRAM Based LLCs
abstract
Spin Transfer Torque Magnetic RAM (STT-MRAM) is being extensively considered as a promising replacement for Last Level Caches (LLC), due to its high density, low leakage and non-volatility. However, writes to STT-MRAM are energy intensive and have a high latency. While the high dynamic energy consumption during writes can be compensated by the low static energy consumption, the high latency results in performance degradation. This work shows that in contrast to SRAM-based LLCs, the performance degradation for STT-MRAM is primarily due to bank contention, when trying to satisfy a read request while the bank is being written. We holistically explore the effects of cache banking and cache contention on energy and performance in the LLC of mobile multicore systems, with in-order cores or with out-of-order cores. The detail of the analysis is enabled by highly accurate cache models, based on a 28nm SRAM industry compiler, and an in-house developed STT-MRAM compiler, which generates full STT-MRAM macro designs with silicon-validated MTJ stack and complete parasitic extraction at the 28nm node. Our results show that there is a clear difference in the energy-performance optimal banking configuration between STT-MRAM caches and SRAM caches. These low contention STT-MRAM cache designs with the optimal number of banks save at least 60% cache energy while losing at most single digit percentages in system performance compared to SRAM cache designs. This show an increased potential of using STT-MRAM as a replacement for SRAM in an LLC.
Timon Evenblij, Christian Tenllado, Manu Perumkunnil Komalan, Francky Catthoor, Sushil Sakhare, Peter Debacker, Gouri Sankar Kar, Arnaud Furnémont, Nicolas Bueno, José Ignacio Gómez
ICCD6
2017 Vertical M1 Routing-Aware Detailed Placement for Congestion and Wirelength Reduction in Sub-10nm Nodes
abstract
Aggressive pitch scaling in sub-10nm nodes has introduced complex design rules which make routing extremely challenging. Cell architectures have also been changed to meet the design rules. For example, metal layers below M1 are used to gain additional routing resources. New cell architectures wherein inter-row M1 routing is allowed force consideration of vertical alignment of cells. In this work, we propose a mixed-integer linear programming (MILP)-based, detailed placement optimization to maximize direct vertical M1 routing utilization for congestion and wirelength reduction.
Peter Debacker, Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001, Praveen Raghavan, Lutong Wang
DAC1
2017 IR-drop aware Design & technology co-optimization for N5 node with different device and cell height options
abstract
In this paper we propose a novel Design-Technology Co-Optimization (DTCO) framework that enables PDK generation and design implementation of sub-10nm technology nodes. The framework allows to study the impact of different technology options at design level and use effective design Power, Performance and Area (PPA) to decide on right technology option. Design implementation flow is IR-drop aware, allowing integration of optimized Power Delivery Network (PDN) for different device/cell options. Using N5-like technology node assumptions (contacted poly and metallization pitch of 42 and 32nm), we generate digital PDKs for different device (finFET, 2 & 3 nanowires) and standard cell options (3, 2 or 1 fins & 7.5 or 6-Tracks cell height). Different PDKs have been used to implement and characterize a wire dominated circuit. Our study shows that the design PDN/IR-drop awareness is fundamental to complete DTCO approach for sub-10nm nodes. Using our dedicated design methodology we reach the IR-drop target of 2.5% VDD (on the lowest metal layers), while minimizing the area degradation induced by the PDN. Further, we demonstrate that such optimized PDN is mandatory to enable the 20% area gain when moving from 7.5 to 6-Tracks cell height. Finally, we show that the impact of different device options is in range of 15% Power, 2X Performance and 20% Area, further validating the need of a fully integrated DTCO.
Luca Mattii, Dragomir Milojevic, Peter Debacker, Yasser Sherazi, Mladen Berekovic, Praveen Raghavan
ICCAD3
2017 Patterning Aware Design Optimization of Selective Etching in N5 and Beyond
abstract
The aggressive scaling towards N5 causes the difficulty in controlling process variations for line and block (cut) printing, while the margin for edge placement errors (EPE) is decreasing. Selective etching, implemented by self-aligned multiple patterning with alternating materials, introduces better EPE margins with self-aligned blocks (SAB). In this paper, we propose a systematic study on design space for SAB with various options of design rules. We also develop a post SAB optimization framework to evaluate the impacts of different rules. Based on the experiments on 6-track industrial benchmarks, we explore possible options for closure of current physical design flow.
Yibo Lin, Peter Debacker, Darko Trivkovic, Ryoung-Han Kim, Praveen Raghavan, David Z. Pan
ICCD2
2017 MILP-Based Optimization of 2-D Block Masks for Timing-Aware Dummy Segment Removal in Self-Aligned Multiple Patterning Layouts
abstract
Self-aligned multiple patterning, due to its low overlay error, has emerged as the leading option for 1-D gridded back-end-of-line (BEOL) in sub-14-nm nodes. To form actual routing patterns from a uniform “sea of wires,” cut masks are needed for line-end cutting or realization of space between routing segments. The line-end cutting results in nonfunctional (i.e., dummy fill) patterns that change wire capacitance, and hence design timing and power. Therefore, to remove such dummy fill patterns, extra 2-D block masks are used. However, 2-D block masks cannot remove arbitrary dummy fill patterns, due to design rule constraints on the block mask shapes. In this paper, we address the timing-aware optimization of 2-D block mask layouts under various sets of mask rules that are derived from mask patterning technology options (e.g., 193i and 193d) for foundry 7-/5-nm (N7/N5) BEOL. Our central contribution is a mixed integer linear programming (MILP) optimization that minimizes timing impact due to dummy metal segments while satisfying block mask rules and metal density constraints. We also propose a distributed optimization flow to improve the scalability. With our optimizer, we recover up to 84% of the worst negative slack impact from dummy segments, with up to 64% dummy removal rate. We further extend our MILP to a co-optimization of cut and block masks. This paper gives new insights into fundamental limits of benefit from emerging cut and block mask technology options.
Peter Debacker, Kwangsoo Han, Andrew B. Kahng, Hyein Lee 0001, Praveen Raghavan, Lutong Wang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2016 Capturing True Workload Dependency of BTI-induced Degradation in CPU Components
abstract
Atomistic-based approaches accurately model Bias Temperature Instability phenomena, but they suffer from prolonged execution times, preventing their seamless integration in system-level analysis flows. In this paper we present a comprehensive flow that combines the accuracy of Capture Emission Time (CET) maps with the efficiency of the Compact Digital Waveform (CDW) representation. That way, we capture the true workload-dependent BTI-induced degradation of selected CPU components. First, we show that existing works that assume constant stress patterns fail to account for workload dependency leading to fundamental estimation errors. Second, we evaluate the impact of different real workloads on selected CPU sub-blocks from a commercial processor design. To the best of our knowledge, this is the first work that combines atomistic property and true workload-dependency for variability analysis.
Dimitrios Stamoulis, Simone Corbetta, Dimitrios Rodopoulos, Pieter Weckx, Peter Debacker, Brett H. Meyer, Ben Kaczer, Praveen Raghavan, Dimitrios Soudris, Francky Catthoor, Zeljko Zilic
ACM Great Lakes Symposium on VLSI5
2016 How much cost reduction justifies the adoption of monolithic 3D ICs at 7nm node?
abstract
In this paper we study power, performance, and cost (PPC) tradeoffs for 2-tier, gate-level, full-chip GDS monolithic 3D ICs (M3D) built using a foundry-grade 7nm bulk FinFET technology. We first develop highly-accurate wafer and die cost models for 2D and M3D to study PPC tradeoffs. In our study, both 2D and M3D designs are optimized in terms of the number of BEOL metal layers used for routing to obtain the best possible PPC values. We develop a new CAD methodology for 2-tier gate-level M3D, named Projected 2D Flow, that allows us to accurately compare RC parasitics of equivalent nets in both 2D and M3D designs. Our experiments based on two different circuit types (BEOL-dominant vs. FEOL-dominant) confirm that M3D designs indeed offer a significant footprint saving. However, to our surprise, the PPC quality of M3D turns out to be worse than that of 2D by 34% due to the high wafer cost of M3D. Our study also reveals that M3D wafer yield should be as high as 90% of 2D wafer yield, and the M3D device manufacturing cost should be less than 33% of that of 2D to justify the adoption of M3D technology at the 7nm era. Lastly, and counter-intuitively, our study shows that FEOL-dominant circuit shows more PPC benefits from M3D technology than BEOL-dominant circuit.
Bon Woong Ku, Peter Debacker, Dragomir Milojevic, Praveen Raghavan, Sung Kyu Lim
ICCAD2
2016 Physical Design Solutions to Tackle FEOL/BEOL Degradation in Gate-level Monolithic 3D ICs
abstract
In this paper, we develop physical design tools and methodologies to tackle the inter-tier performance variations caused by low temperature manufacturing in 2-tier gate-level monolithic 3D ICs (M3D). First, we model the top tier front-end-of-line (FEOL) device mobility degradation and its impact on cell delay/power values. Next, we quantify the impact of tungsten interconnect and cost-driven metal layer saving in the back-end-of-line (BEOL) of the bottom tier. These device and interconnect degradation models are used in our new full-chip M3D physical design flow named Derated 2D. This flow overcomes the well-known drawback of the state-of-the-art Shrunk 2D that requires shrinking of layout objects and RC parasitics. Also, Derated 2D performs low-temperature process-aware tier partitioning to effectively keep timing-critical components in the bottom tier. Moreover, Derated 2D conducts timing-driven monolithic inter-tier via (MIV) planning to cope with the resistivity increase in tungsten BEOL. Lastly, Derated 2D offers an effective timing closure solution through a post-route optimization. Experiments based on a foundry-grade 7nm FinFET process design kit (PDK) show that Derated 2D achieves up to 36% performance improvement and 10% energy saving compared with Shrunk 2D. Using a post-route optimization, Derated 2D further improves timing under various FEOL/BEOL degradation settings at a minimum energy overhead.
Bon Woong Ku, Peter Debacker, Dragomir Milojevic, Praveen Raghavan, Diederik Verkest, Aaron Thean, Sung Kyu Lim
ISLPED2