Doo Seok Jeong

dblp:190/7424 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0001-7954-2213ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 5 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 IterL2Norm: Fast Iterative L2-Normalization
abstract
Transformer-based large language models are a memory-bound model whose operation is based on a large amount of data that are marginally reused. Thus, the data movement between a host and accelerator likely dictates the total wall-clock time. Layer normalization is one of the key workloads in the transformer model, following each of multi-head attention and feed-forward network blocks. To reduce data movement, layer normalization needs to be performed on the same chip as the matrix-matrix multiplication engine. To this end, we introduce an iterative L2-normalization method for 1D input (IterL2Norm), ensuring fast convergence to the steady-state solution within five iteration steps and high precision, outperforming the fast inverse square root algorithm in six out of nine cases for FP32 and five out of nine for BFloat16 across the embedding lengths used in the OPT models. Implemented in 32/28nm CMOS, the IterL2Norm macro normalizes d-dimensional vectors, where 64 ≤$d$≤ 1024, with a latency of 116–227 cycles at 100MHz/l.05V.
Changmin Ye, Yonguk Sim, Youngchae Kim, SeongMin Jin, Doo Seok Jeong
DATE5
2025 Computing-in-Memory Dataflow for Minimal Buffer Traffic
abstract
Computing-In-Memory (CIM) offers a potential solution to the memory wall issue and can achieve high energy efficiency by minimizing data movement, making it a promising architecture for edge AI devices. Lightweight models like MobileNet and EfficientNet, which utilize depthwise convolution for feature extraction, have been developed for these devices. However, CIM macros often face challenges in accelerating depthwise convolution, including underutilization of CIM memory and heavy buffer traffic. The latter, in particular, has been overlooked despite its significant impact on latency and energy consumption. To address this, we introduce a novel CIM dataflow that significantly reduces buffer traffic by maximizing data reuse and improving memory utilization during depthwise convolution. The proposed dataflow is grounded in solid theoretical principles, fully demonstrated in this paper. When applied to MobileNet and EfficientNet models, our dataflow reduces buffer traffic by 77.4-87.0%, leading to a total reduction in data traffic energy and latency by 10.1-17.9% and 15.6-27.8%, respectively, compared to the baseline (conventional weight-stationary dataflow).
Choongseok Song, Doo Seok Jeong
ICCD2
2025 Optimal strategy for mapping spiking neural networks onto manycore neuromorphic processors
abstract
Manycore digital neuromorphic event processors execute ad hoc event routing between spiking neurons distributed across multiple cores. Due to limited hardware resources, such as on-chip memory capacity, only a limited number of neurons and their fan-in weights can be accommodated per core. This challenge is particularly significant for convolutional layers, where placing neurons from the same layer in different cores hinders weight reuse, as the same weight must be duplicated across cores. To address this, we propose an optimal mapping method for spiking units across multiple cores, considering hardware resource constraints. This method is based on the discrete Lagrange Multiplier Method, which uses a total memory usage as an objective function alongside constraint functions (memory usage per core). Our results show that this method achieves optimal spiking unit distributions with high core memory utilization (> 70%) for the reduced ResNet models.
Changmin Ye, Doo Seok Jeong
ISCAS2
2024 Optimal data distribution in FeFET-based computing-in-memory macros
abstract
Computing-in-memory (CIM) may offer a power-efficient solution to the acceleration of major workloads for memory-bound deep neural networks given memory and processing units on the same die, particularly, when incorporating the processing units into the memory domains. Further, CIM macros utilizing nonvolatile memory with multibit data significantly boost their data density and realize zero standby power by gating power when idle. Ferroelectric Field-Effect-Transistor (FeFET) is a leading contender for this type of CIM. In this work, we designed mixed signal CIM macros based on FeFETs and identified their optimal performance with the size of a sub-array (nM× nw) addressed at one cycle, where nwis the number of FeFETs representing a single w-bit weight. The simulations performed identified the optimal sub-array $n_M^{\ast} \times w/2$ for w-bit weights with different $n_M^{\ast}$ (i.e., parallelism) for different weight resolution w, which highlights a ∼29× improvement in figure of merit for 8-bit weights compared with the case of no weight-splitting (nw= 1).
Yonguk Sim, Choongseok Song, Jongwook Jeon, Daewoong Kwon, Doo Seok Jeong
ISCAS6
2023 LaCERA: Layer-centric event-routing architecture
Changmin Ye, Vladimir Kornijcuk, Donghyung Yoo, Jeeson Kim, Doo Seok Jeong
Neurocomputing5
2023 Training Spiking Neural Networks Using Lessons From Deep Learning
abstract
The brain is the perfect place to look for inspiration to develop more efficient neural networks. The inner workings of our synapses and neurons provide a glimpse at what the future of deep learning might look like. This article serves as a tutorial and perspective showing how to apply the lessons learned from several decades of research in deep learning, gradient descent, backpropagation, and neuroscience to biologically plausible spiking neural networks (SNNs). We also explore the delicate interplay between encoding data as spikes and the learning process; the challenges and solutions of applying gradient-based learning to SNNs; the subtle link between temporal backpropagation and spike timing-dependent plasticity; and how deep learning might move toward biologically plausible online learning. Some ideas are well accepted and commonly used among the neuromorphic engineering community, while others are presented or justified for the first time here. A series of companion interactive tutorials complementary to this article using our Python package,snnTorch, are also made available: https://snntorch.readthedocs.io/en/latest/tutorials/index.html.
Jason Kamran Eshraghian, Max Ward 0001, Emre Neftci, Xinxin Wang 0002, Gregor Lenz, Girish Dwivedi, Mohammed Bennamoun, Doo Seok Jeong, Wei Lu 0003
Proc. IEEE8
2021 CBP: backpropagation with constraint on weight precision using a pseudo-Lagrange multiplier method
abstract
Backward propagation of errors (backpropagation) is a method to minimize objective functions (e.g., loss functions) of deep neural networks by identifying optimal sets of weights and biases. Imposing constraints on weight precision is often required to alleviate prohibitive workloads on hardware. Despite the remarkable success of backpropagation, the algorithm itself is not capable of considering such constraints unless additional algorithms are applied simultaneously. To address this issue, we propose the constrained backpropagation (CBP) algorithm based on the pseudo-Lagrange multiplier method to obtain the optimal set of weights that satisfy a given set of constraints. The defining characteristic of the proposed CBP algorithm is the utilization of a Lagrangian function (loss function plus constraint function) as its objective function. We considered various types of constraints — binary, ternary, one-bit shift, and two-bit shift weight constraints. As a post-training method, CBP applied to AlexNet, ResNet-18, ResNet-50, and GoogLeNet on ImageNet, which were pre-trained using the conventional backpropagation. For most cases, the proposed algorithm outperforms the state-of-the-art methods on ImageNet, e.g., 66.6\%, 74.4\%, and 64.0\% top-1 accuracy for ResNet-18, ResNet-50, and GoogLeNet with binary weights, respectively. This highlights CBP as a learning algorithm to address diverse constraints with the minimal performance loss by employing appropriate constraint functions. The code for CBP is publicly available at \url{https://github.com/dooseokjeong/CBP}.
Guhyun Kim, Doo Seok Jeong
NeurIPS2
2021 Hardware-Efficient Emulation of Leaky Integrate-and-Fire Model Using Template-Scaling-Based Exponential Function Approximation
abstract
We present a method to emulate a leaky integrate-and-fire (LIF) model in a field-programmable gate array (FPGA) in a hardware-efficient manner. The simplified spike-response model (SRM0) is chosen as an LIF model. For the hardware-efficient implementation of SRM0, we adopt the template-scaling-based exponential function approximation (TS-EFA). This method allows high precision and low latency exponential function approximations with the efficient use of hardware resources. We subsequently propose an algorithm for SRM0, which leverages the advantage of TS-EFA. An implementation of 512 neurons conforming to SRM0in an FPGA highlights (i) high precision of SRM0emulation (mean squared error of membrane potential approximation: 4×10-12- 1×10-10), (ii) low latency (eight clock cycles), and (iii) high efficiency in hardware usage (only 125b memory per neuron).
Jeeson Kim, Vladimir Kornijcuk, Changmin Ye, Doo Seok Jeong
IEEE Trans. Circuits Syst. I Regul. Pap.4
2020 Simplified calcium signaling cascade for synaptic plasticity
Vladimir Kornijcuk, Guhyun Kim, Doo Seok Jeong
Neural Networks4
2019 Stochastic Learning with Back Propagation
abstract
Despite of remarkable progress on deep learning, its hardware implementation beyond deep learning acceleration is still behind the software deep learning due in part to lack of hardware-compatible learning algorithm. In this paper, a learning method called the stochastic learning with backpropagation (SLBP) algorithm was proposed. The network of concern consists of ternary synaptic weight, favorable to be implemented in a resistance-based crossbar array. Every training epoch, the SLBP algorithm evaluates weight update probability at which the corresponding weight is updated in a stochastic manner. The algorithm was used to train a denoising autoencoder, which identified the successful reduction in noise (increase in peak signal-to-noise ratio by approximately 68%). Notably, the SLBP algorithm achieves an 86% reduction in memory usage compared with a real-valued autoencoder trained using a backpropagation algorithm.
Guhyun Kim, Cheol Seong Hwang, Doo Seok Jeong
ISCAS3
2018 Pointer Based Routing Scheme for On-chip Learning in Neuromorphic Systems
abstract
A look-up table (LUT)-based spike-routing approach is often used in inference-only neuromorphic systems due to its excellent reconfigurability. The challenge is to apply this approach also to on-chip learning that requires a search of a lengthy LUT for all relevant synapses to a firing neuron. To solve this issue, we propose a pointer-based routing scheme that remarkably accelerates spike-routing at the cost of an additional LUT (pointer LUT). Our theoretical estimations suggest that the proposed routing scheme at 1 GHz clock speed supports a spiking neural network of up to 107synapses and more than 105neurons firing at 50 Hz without spike traffic congestion. The scheme needs approximately 32 MB memory. To verify experimentally, the proposed routing scheme was implemented on a Xilinx Virtex 7 FPGA board deploying an array of leaky integrate-and-fire neurons.
Vladimir Kornijcuk, Doo Seok Jeong
IJCNN2
2018 A Physical Unclonable Function With Redox-Based Nanoionic Resistive Memory
abstract
Emerging non-volatile reduction-oxidation (redox)-based resistive switching memories (ReRAMs) exhibit a unique set of characteristics that make them promising candidates for the next generation of low-cost, low-power, tiny, and secure physical unclonable functions (PUFs). Their underlying stochastic ionic conduction behavior, intrinsic nonlinear current-voltage characteristics, and their well-known nano-fabrication process variability might normally be considered disadvantageous ReRAM features. However, using a combination of a novel architecture and special peripheral circuitry, this paper exploits these non-idealities in a physical one-way function, nonlinear resistive PUF, potentially applicable to a variety of cyber-physical security applications. We experimentally verify the performance of valency change mechanism (VCM)-based ReRAM in nano-fabricated crossbar arrays across multiple dies and runs. In addition to supporting a massive pool of challenge-response pairs (CRPs), using a combination of experiment and simulation our proposed PUF exhibits a reliability of 98.67%, a uniqueness of 49.85%, a diffuseness of 49.86%, a uniformity of 47.28%, and a bit-aliasing of 47.48%.
Jeeson Kim, Taimur Ahmed, Hussein Nili, Doo Seok Jeong, Paul Beckett, Sharath Sriram, Damith Chinthana Ranasinghe, Omid Kavehei
IEEE Trans. Inf. Forensics Secur.5