EDBT 2026 Demo / reviewers in the wild / expert
Sugil Lee
dblp:221/2869
· DBLP profile ↗
16ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0003-3092-6501ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mitigating the Impact of ReRAM I-V Nonlinearity and IR Drop via Fast Offline Network TrainingabstractReRAM crossbar arrays (RCAs) have the potential to provide extremely high efficiency for accelerating deep neural networks (DNNs). However, one crucial challenge for RCA-based DNN accelerators is functional inaccuracy due to nonidealities present in RCA hardware. While nonideality-aware training (NAT) could be used to mitigate the effect of nonidealities, with currently available methods it would take months to train even a medium size convolutional neural network (CNN). In this article we propose a nonideality prediction method that enables very fast training of RCA-based neural networks, and show its feasibility through NAT of DNNs. Our key ideas include 1) weight-centric nonideality modeling and 2) data-dependence elimination by tailored input randomization. Our experimental results using a multilayer perceptron and CNNs demonstrate that our method is very fast ($100\sim 15$$000\times $faster training speed) while achieving much better-crossbar-level accuracy ($2 \sim 90\times $lower-RMS error) and post-retraining validated accuracy than previous methods. Sugil Lee, Mohamed E. Fouda, Chenghao Quan, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | FlexInt: A New Number Format for Robust Sub-8-Bit Neural Network InferenceabstractWhile previous work has demonstrated that even large DNNs can be quantized to very low precision (sub-8-bit integers), concerns over robustness across different types of networks and datasets have led to a more serious consideration of floating-point (FP) formats in the industry. However, at 8 bits and below, there is no universally accepted FP format or one that provides robust performance on diverse data distributions. Thus in this paper, based on our analysis of integer (INT) and FP formats, we propose a novel number format called FlexInt, with a high dynamic range similar to FP, yet low max rounding error, targeting efficient representation of DNNs for inference at 8 bits and below. We also propose a novel FlexInt MAC (Multiply-Accumulate) hardware architecture. Our experimental results using large networks on image classification and natural language processing demonstrate that our FlexInt can deliver more robust performance and far superior worst-case accuracy, compared to both INT and FP across various data distributions; has a hardware overhead similar to that of FP; and can consistently make near-Pareto-optimal area-accuracy trade-offs across diverse networks. Minuk Hong, Hyeon Uk Sim, Sugil Lee, Jongeun Lee |
ICCAD | 3 |
| 2023 | NTT-PIM: Row-Centric Architecture and Mapping for Efficient Number-Theoretic Transform on PIMabstractRecently DRAM-based PIMs (processing-in-memories) with unmodified cell arrays have demonstrated impressive performance for accelerating AI applications. However, due to the very restrictive hardware constraints, PIM remains an accelerator for simple functions only. In this paper we propose NTT-PIM, which is based on the same principles such as no modification of cell arrays and very restrictive area budget, but shows state-of-the-art performance for a very complex application such as NTT, thanks to features optimized for the application’s characteristics, such as in-place update and pipelining via multiple buffers. Our experimental results demonstrate that our NTT-PIM can outperform previous best PIM-based NTT accelerators in terms of runtime by 1.7 ∼ 17× while having negligible area and power overhead. Jaewoo Park 0006, Sugil Lee, Jongeun Lee |
DAC | 2 |
| 2023 | Offline Training-Based Mitigation of IR Drop for ReRAM-Based Deep Neural Network AcceleratorsabstractRecently, resistive RAM (ReRAM)-based hardware accelerators showed unprecedented performance compared the digital accelerators. Technology scaling causes an inevitable increase in interconnect wire resistance, which leads to IR drops that could limit the performance of ReRAM-based accelerators. These IR drops deteriorate the signal integrity and quality, especially in the crossbar structures which are used to build high-density ReRAMs. Hence, finding a software solution, which can predict the effect of IR drop without involving expensive hardware or SPICE simulations, is very desirable. In this article, we propose two neural networks models to predict the impact of the IR drop problem. These models are used to evaluate the performance of the different deep neural network (DNN) models including binary and quantized neural networks showing similar performance (i.e., recognition accuracy) to the golden validation (i.e., SPICE-based DNN validation). In addition, these predication models are incorporated into the DNN training framework to efficiently retrain the DNN models and bridge the accuracy drop. To further enhance the validation accuracy, we propose incremental training methods. The DNN validation results, done through SPICE simulations, show very high improvement in performance close to the baseline performance, which demonstrates the efficacy of the proposed method even with challenging datasets, such as CIFAR10 and SVHN. Sugil Lee, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Training-Free Stuck-At Fault Mitigation for ReRAM-Based Deep Learning AcceleratorsabstractAlthough Resistive RAMs can support highly efficient matrix–vector multiplication, which is very useful for machine learning and other applications, the nonideal behavior of hardware, such as stuck-at fault (SAF) and IR drop is an important concern in making ReRAM crossbar array-based deep learning accelerators. Previous work has addressed the nonideality problem through either redundancy in hardware, which requires a permanent increase of hardware cost, or software retraining, which may be even more costly or unacceptable due to its need for a training dataset as well as high computation overhead. In this article, we propose a very lightweight method that can be applied on top of existing hardware or software solutions. Our method, called forward-parameter tuning (FPT), takes advantage of a certain statistical property existing in the activation data of neural network layers, and can mitigate the impact of mild nonidealities in ReRAM crossbar arrays (RCAs) for deep learning applications without using any hardware, a dataset, or gradient-based training. Our experimental results using MNIST, CIFAR-10, and CIFAR-100, and ImageNet datasets in binary and multibit networks demonstrate that our technique is very effective, both alone and together with previous methods, up to 20% fault rate, which is higher than even some of the previous remapping methods. We also evaluate our method in the presence of other nonidealities, such as variability and IR drop. Furthermore, we provide an analysis based on the concept of the effective fault rate (EFR), which not only demonstrates that EFR can be a useful tool to predict the accuracy of faulty RCA-based neural networks but also explains why mitigating the SAF problem is more difficult with multibit neural networks. Chenghao Quan, Mohamed E. Fouda, Sugil Lee, Giju Jung, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Accurate Prediction of ReRAM Crossbar Performance Under I-V Nonlinearity and IR DropabstractDespite the promise of extremely efficient matrix-vector multiplication (MVM) by ReRAM crossbar arrays (RCAs), maintaining high accuracy has been challenging due to nonidealities such as wire resistance (also known as IR drop) and I-V nonlinearity (i.e., voltage-dependent conductance). For system architects, a fast method to accurately predict the MVM output of an RCA under nonidealities is highly desirable. While IR drop alone without I-V nonlinearity can be efficiently predicted, the existence of I-V nonlinearity makes the problem much harder. In this paper we propose a novel algorithm based on iterative refinement, which can predict with high accuracy the outcome of an MVM operation on an RCA in the presence of both I-V nonlinearity and IR drop. Our experiments using binary RCAs of different sizes demonstrate that our proposed method is order-of-magnitude more accurate than previous methods in terms of RMS error. We also present case studies predicting hardware-realistic accuracy of binarized neural networks on RCAs as well as nonideality-aware retraining, demonstrating the efficacy of our method for early design space exploration of ReRAM-based accelerators. Sugil Lee, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
ICCD | 1 |
| 2022 | MLogNet: A Logarithmic Quantization-Based Accelerator for Depthwise Separable ConvolutionabstractIn this article, we propose a novel logarithmic quantization-based deep neural network (DNN) architecture for depthwise separable convolution (DSC) networks. Our architecture is based on selective two-word logarithmic quantization (STLQ), which improves accuracy greatly over logarithmic-scale quantization while retaining the speed and area advantage of logarithmic quantization. On the other hand, it also comes with the synchronization problem due to variable-latency processing elements (PEs), which we address through a novel architecture and a compile-time optimization technique. Our architecture is dynamically reconfigurable to support various combinations of depthwise versus pointwise convolution layers efficiently. Our experimental results using layers from MobileNetV2 and ShuffleNetV2 demonstrate that our architecture is significantly faster and more area efficient than previous DSC accelerator architectures as well as previous accelerators utilizing logarithmic quantization. Jooyeon Choi, Hyeon Uk Sim, Sangyun Oh, Sugil Lee, Jongeun Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Automated Log-Scale Quantization for Low-Cost Deep Neural NetworksabstractQuantization plays an important role in deep neural network (DNN) hardware. In particular, logarithmic quantization has multiple advantages for DNN hardware implementations, and its weakness in terms of lower performance at high precision compared with linear quantization has been recently remedied by what we call selective two-word logarithmic quantization (STLQ). However, there is a lack of training methods designed for STLQ or even logarithmic quantization in general. In this paper we propose a novel STLQ-aware training method, which significantly out-performs the previous state-of-the-art training method for STLQ. Moreover, our training results demonstrate that with our new training method, STLQ applied to weight parameters of ResNet-18 can achieve the same level of performance as state-of-the-art quantization method, APoT, at 3-bit precision. We also apply our method to various DNNs in image enhancement and semantic segmentation, showing competitive results. Sangyun Oh, Hyeon Uk Sim, Sugil Lee, Jongeun Lee |
CVPR | 3 |
| 2021 | Cost- and Dataset-free Stuck-at Fault Mitigation for ReRAM-based Deep Learning AcceleratorsabstractResistive RAMs can implement extremely efficient matrix vector multiplication, drawing much attention for deep learning accelerator research. However, high fault rate is one of the fundamental challenges of ReRAM crossbar array-based deep learning accelerators. In this paper we propose a dataset-free, cost-free method to mitigate the impact of stuck-at faults in ReRAM crossbar arrays for deep learning applications. Our technique exploits the statistical properties of deep learning applications, hence complementary to previous hardware or algorithmic methods. Our experimental results using MNIST and CIFAR-10 datasets in binary networks demonstrate that our technique is very effective, both alone and together with previous methods, up to 20 % fault rate, which is higher than the previous remapping methods. We also evaluate our method in the presence of other non-idealities such as variability and IR drop. Giju Jung, Mohamed E. Fouda, Sugil Lee, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
DATE | 3 |
| 2021 | Fast and Low-Cost Mitigation of ReRAM Variability for Deep Learning ApplicationsabstractTo overcome the programming variability (PV) of ReRAM crossbar arrays (RCAs), the most common method is program-verify, which, however, has high energy and latency overhead. In this paper we propose a very fast and low-cost method to mitigate the effect of PV and other variability for RCA-based DNN (Deep Neural Network) accelerators. Leveraging the statistical properties of DNN output, our method called Online Batch-Norm Correction (OBNC) can compensate for the effect of programming and other variability on RCA output without using on-chip training or an iterative procedure, and is thus very fast. Also our method does not require a nonideality model or a training dataset, hence very easy to apply. Our experimental results using ternary neural networks with binary and 4-bit activations demonstrate that our OBNC can recover the baseline performance in many variability settings and that our method outperforms a previously known method (VCAM) by large margins when input distribution is asymmetric or activation is multi-bit. Sugil Lee, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
ICCD | 1 |
| 2020 | Learning to Predict IR Drop with Effective Training for ReRAM-based Neural Network HardwareabstractDue to the inevitability of the IR drop problem in passive ReRAM crossbar arrays, finding a software solution that can predict the effect of IR drop without the need of expensive SPICE simulations, is very desirable. In this paper, two simple neural networks are proposed as software solution to predict the effect of IR drop. These networks can be easily integrated in any deep neural network framework to incorporate the IR drop problem during training. As an example, the proposed solution is integrated in BinaryNet framework and the test validation results, done through SPICE simulations, show very high improvement in performance close to the baseline performance, which demonstrates the efficacy of the proposed method. In addition, the proposed solution outperforms the prior work on challenging datasets such as CIFAR10 and SVHN. Sugil Lee, Giju Jung, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
DAC | 1 |
| 2020 | Architecture-Accuracy Co-optimization of ReRAM-based Low-cost Neural Network ProcessorabstractResistive RAM (ReRAM) is a promising technology with such advantages as small device size and in-memory-computing capability. However, designing optimal AI processors based on ReRAMs is challenging due to the limited precision, and the complex interplay between quality of result and hardware efficiency. In this paper we present a study targeting a low-power low-cost image classification application. We discover that the trade-off between accuracy and hardware efficiency in ReRAM-based hardware is not obvious and even surprising, and our solution developed for a recently fabricated ReRAM device achieves both the state-of-the-art efficiency and empirical assurance on the high quality of result. Segi Lee, Sugil Lee, Jongeun Lee, Jong-Moon Choi, Do-Wan Kwon, Seung-Kwang Hong, Kee-Won Kwon |
ACM Great Lakes Symposium on VLSI | 2 |
| 2019 | On-chip memory optimization for high-level synthesis of multi-dimensional data on FPGAabstractIt is very challenging to design an on-chip memory architecture for high-performance kernels with large amount of computation and data. The on-chip memory architecture must support efficient data access from both the computation part and the external memory part, which often have very different expectations about how data should be accessed and stored. Previous work provides only a limited set of optimizations. In this paper we show how to fundamentally restructure on-chip buffers, by decoupling logical array view from the physical buffer view, and providing general mapping schemes for the two. Our framework considers the entire data flow from the external memory to the computation part in order to minimize resource usage without creating performance bottleneck. Our experimental results demonstrate that our proposed technique can generate solutions that reduce memory usage significantly (2X over the conventional method), and successfully generate optimized on-chip buffer architectures without costly design iterations for highly optimized computation kernels. Daewoo Kim, Sugil Lee, Jongeun Lee |
ASP-DAC | 2 |
| 2019 | Successive Log Quantization for Cost-Efficient Neural Networks Using Stochastic ComputingabstractDespite the multifaceted benefits of stochastic computing (SC) such as low cost, low power, and flexible precision, SC-based deep neural networks (DNNs) still suffer from the long-latency problem, especially for those with high precision requirements. While log quantization can be of help, it has its own accuracy-saturation problem due to uneven precision distribution. In this paper we propose successive log quantization (SLQ), which extends log quantization with significant improvements in precision and accuracy, and apply it to state-of-the-art SC-DNNs. SLQ reuses the existing datapath of log quantization, and thus retains its advantages such as simple multiplier hardware. Our experimental results demonstrate that our SLQ can significantly extend both the accuracy and efficiency of SC-DNNs over the state-of-the-art solutions, including linear-quantized and log-quantized SC-DNNs, achieving less than 1~1.5%p accuracy drop for AlexNet, SqueezeNet, and VGG-S at mere 4~5-bit weight resolution. Sugil Lee, Hyeon Uk Sim, Jooyeon Choi, Jongeun Lee |
DAC | 1 |
| 2019 | Double MAC on a DSP: Boosting the Performance of Convolutional Neural Networks on FPGAsabstractDeep learning workloads, such as convolutional neural networks (CNNs) are important due to increasingly demanding high-performance hardware acceleration. One distinguishing feature of a deep learning workload is that it is inherently resilient to small numerical errors and thus works very well with low precision hardware. We propose a novel method called double multiply-and-accumulate (MAC) to theoretically double the computation rate of CNN accelerators by packing two MAC operations into one digital signal processing block of off-the-shelf field-programmable gate arrays (FPGAs). We overcame several technical challenges by exploiting the mode of operation in the CNN accelerator. We have validated our method through FPGA synthesis and Verilog simulation, and evaluated our method by applying it to the state-of-the-art CNN accelerator. The double MAC approach used can double the computation throughput of a CNN layer. On the network level (all convolution layers combined), the performance improvement varies depending on the CNN application and FPGA size, from 14% to more than 80% over a highly optimized state-of-the-art accelerator solution, without sacrificing the output quality significantly. Sugil Lee, Daewoo Kim, Dong Nguyen 0001, Jongeun Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | Sign-magnitude SC: getting 10X accuracy for free in stochastic computing for deep neural networksabstractStochastic computing (SC) is a promising computing paradigm for applications with low precision requirement, stringent cost and power restriction. One known problem with SC, however, is the low accuracy especially with multiplication. In this paper we propose a simple, yet very effective solution to the low-accuracy SC-multiplication problem, which is critical in many applications such as deep neural networks (DNNs). Our solution is based on an old concept of sign-magnitude, which, when applied to SC, has unique advantages. Our experimental results using multiple DNN applications demonstrate that our technique can improve the efficiency of SC-based DNNs by about 32X in terms of latency over using bipolar SC, with very little area overhead (about 1%). Aidyn Zhakatayev, Sugil Lee, Hyeon Uk Sim, Jongeun Lee |
DAC | 2 |