VLDB 2026 Research / reviewers in the wild / expert
Jongeun Lee
dblp:64/5375 · also Jong-eun Lee
· DBLP profile ↗
93ranked-venue papers
15as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 77 · 13 first-author · 21 since 2021Software engineering, systems software and programming languages · 18 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlowQ: Fixed-point Low-precision Post-Training Quantization Framework for Efficient and Accurate SNN InferenceabstractWe propose FlowQ, a post-training quantization (PTQ) framework for spiking neural networks (SNNs) that balances accuracy and hardware-efficiency through quantizer design and calibration-based scale optimization. While using different scales for weights and membrane potentials preserves accuracy, it typically incurs high hardware cost. In contrast, shared scale factors reduce hardware complexity but lead to significant accuracy degradation. FlowQ bridges this gap by using a hardware-friendly quantizer with different scales that differ by a power-of-two, allowing multiplications to be replaced with simple bit-shift operations for negligible overhead. To further improve accuracy, we present FlowTune, a calibration algorithm that iteratively optimizes FlowQ’s scale factors by minimizing mean-squared error, outperforming the commonly used absolute max-based scaling in SNN PTQ. Extensive experiments on CIFAR-10, CIFAR-100, DVS-Gestures, and ImageNet demonstrate the effectiveness of our approach. For example, on CIFAR-10 with VGG-16 at 4-bit precision, FlowQ achieves only a 1.48% accuracy drop. Compared to shared scale factors, FlowQ improves accuracy by 77.4% with just 1% energy and 0.25% area overhead. Faaiz Asim, Sanhtet Aung, Jongeun Lee |
ASP-DAC | 3 |
| 2026 | EP-HDC: Hyperdimensional Computing with Encrypted Parameters for High-Throughput Privacy-Preserving InferenceabstractWhile homomorphic encryption (HE) provides strong privacy protection, its high computational cost has restricted its application to simple tasks. Recently, hyperdimensional computing (HDC) applied to HE has shown promising performance for privacy-preserving machine learning (PPML). However, when applied to more realistic scenarios such as batch inference, the HDC-based HE has still very high compute time as well as high encryption and data transmission overheads. To address this problem, we propose HDC with encrypted parameters (EP-HDC), which is a novel PPML approach featuring client-side HE, i.e., inference is performed on a client using a homomorphically encrypted model. Our EP-HDC can effectively mitigate the encryption and data transmission overhead, as well as providing high scalability with many clients while providing strong protection for user data and model parameters. In addition to application examples for our client-side PPML, we also present design space exploration involving quantization, architecture, and HE-related parameters. Our experimental results using the BFV scheme and the Face/Emotion datasets demonstrate that our method can improve throughput and latency of batch inference by orders of magnitude over previous PPML methods ($36.52 \sim 1068 \times$ and $6.45 \sim 733 \times$, respectively) with <1% accuracy degradation. Jaewoo Park 0006, Chenghao Quan, Jongeun Lee |
ASP-DAC | 3 |
| 2026 | AccelOrb: FPGA Acceleration of Orb v2 for Fast Molecular DynamicsabstractMachine learning force fields (MLFFs) offer high accuracy at reduced computational cost, but their repeated evaluation at every molecular dynamics (MD) timestep leads to prohibitive runtime, limiting simulations to short physical timescales. In this work, we accelerate Orb v2, a Graph Neural Network–based MLFF, by targeting its dominant Attention Interaction Network (AIN), which accounts for approximately 97% of inference time. We identify on-chip memory constraints and a latency-memory trade-off in block-based computation as the key challenges to efficient acceleration. To address these challenges, we propose AccelOrb, a memory-aware FPGA acceleration of AIN based on a single streaming kernel that employs a non-uniform tiling strategy to balance latency and memory footprint, and selectively maps parameters beyond BRAM. Implemented on a resource-constrained AMD Alveo U50 FPGA, our approach achieves 9.19× and 8.18× speedup over CPU and GPU baselines while operating within tight power and resource budgets. Sunjae Kim, Gwanhong Park, Jeawoo Lim, Faaiz Asim, Jongeun Lee |
FCCM | 5 |
| 2026 | COMET: Co-Optimization of CNN Models Using Efficient-Hardware OBC TechniquesabstractConvolutional Neural Networks (CNNs) achieve remarkable accuracy in vision tasks, yet their computational complexity challenges low-power edge deployment. In this work, we present COMET, a framework of CNN models that employ efficient hardware offset-binary coding (OBC) techniques to enable co-optimization of performance and resource utilization. The approach formulates CNN inference using OBC representations applied separately to inputs (Scheme A) and weights (Scheme B), enabling exploitation of bit-width asymmetry. The shift–accumulate operation is modified by incorporating offset-term with the pre-scaled bias. Leveraging symmetries in Schemes A and B, we introduce four look-up table (LUT) techniques—parallel, shared, split, and hybrid—and evaluate their efficiency. Building on this foundation, we develop a general matrix multiplication core using theim2coltransformation for efficient CNN acceleration. We consider LeNet-5 and All-CNN-C to demonstrate that the OBC-GEMM core efficiently supports modern workloads. Evaluation shows that COMET enables efficient FPGA deployment compared to state-of-the-art designs, with negligible accuracy loss, demonstrating its efficiency and scalability across diverse network architectures. Mohd. Tasleem Khan, George Goussetis, Mathini Sellathurai, Yuan Ding 0001, João F. C. Mota, Jongeun Lee |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | SPIMA: Scalable and Cost-Efficient Sparse Matrix Multiplication via Processing in DRAM ArrayabstractSparse matrix multiplication (SpMM) is a critical kernel used in a wide range of applications, but irregular memory access patterns and memory bandwidth bottleneck as well as load imbalance make the efficient and scalable processing on parallel architectures a significant challenge. Motivated by the memory-bound nature of SpMM computation, we propose a cost-effective SpMM accelerator based on a DRAM processing-in-memory (PIM) approach. Our design introduces a novel dataflow to exploit high bank-level parallelism and reuse both input and output data even for highly sparse matrices. Our proposed architecture, SPIMA, features multiple input buffers for scheduling DRAM access, a output buffer and vector register files working holistically, co-designed to maximize the performance of our novel dataflow. Our experimental results using various sparse matrices demonstrate that our proposed dataflow and architecture are robust in terms of matrix size and sparsity. Compared with the state-of-the-art accelerators implemented on PIM, ASIC, and FPGA, we estimate that our PIM architecture can yield competitive performance with highly sparse matrices. Tairali Assylbekov, Minsang Yu, Jaewoo Park 0006, Mingon Kim, Seungsu Kim, Jongeun Lee |
ICCAD | 6 |
| 2025 | Integrating Meta-analysis in Multi-modal Brain Studies with Graph-Based Attention Transformer
Hyoungshin Choi, Jongeun Lee, Bo-yong Park, Hyunjin Park |
MICCAI (12) | 3 |
| 2025 | Mitigating the Impact of ReRAM I-V Nonlinearity and IR Drop via Fast Offline Network TrainingabstractReRAM crossbar arrays (RCAs) have the potential to provide extremely high efficiency for accelerating deep neural networks (DNNs). However, one crucial challenge for RCA-based DNN accelerators is functional inaccuracy due to nonidealities present in RCA hardware. While nonideality-aware training (NAT) could be used to mitigate the effect of nonidealities, with currently available methods it would take months to train even a medium size convolutional neural network (CNN). In this article we propose a nonideality prediction method that enables very fast training of RCA-based neural networks, and show its feasibility through NAT of DNNs. Our key ideas include 1) weight-centric nonideality modeling and 2) data-dependence elimination by tailored input randomization. Our experimental results using a multilayer perceptron and CNNs demonstrate that our method is very fast ($100\sim 15$$000\times $faster training speed) while achieving much better-crossbar-level accuracy ($2 \sim 90\times $lower-RMS error) and post-retraining validated accuracy than previous methods. Sugil Lee, Mohamed E. Fouda, Chenghao Quan, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | Extending Neural Processing Unit and Compiler for Advanced Binarized Neural NetworksabstractBinarized neural networks (BNNs) are one of the most promising approaches to deploy deep neural network models on resource-constrained devices. However, there is very little support on compilers and programmable accelerators for BNNs especially with the modern BNNs that use scale factors and skip connections to maximize network performance. In this paper we present a set of methods to extend a neural processing unit (NPU) and a compiler to support modern BNNs. Our novel ideas include (i) batch-norm folding for binarized layers with scale factors and skip connections, (ii) efficient handling of convolutions with few input channels, and (iii) bit-packing pipelining. Our evaluation using BiRealNet-18 on an FPGA board demonstrates that our compiler-architecture hybrid approach can yield significant speedups for binary convolution layers over the baseline NPU. Also our approach gives 3.6~5.5 $\times$ better end-to-end performance on BiRealNet-18 compared with previous BNN compiler approaches. Minjoon Song, Faaiz Asim, Jongeun Lee |
ASPDAC | 3 |
| 2024 | FlexInt: A New Number Format for Robust Sub-8-Bit Neural Network InferenceabstractWhile previous work has demonstrated that even large DNNs can be quantized to very low precision (sub-8-bit integers), concerns over robustness across different types of networks and datasets have led to a more serious consideration of floating-point (FP) formats in the industry. However, at 8 bits and below, there is no universally accepted FP format or one that provides robust performance on diverse data distributions. Thus in this paper, based on our analysis of integer (INT) and FP formats, we propose a novel number format called FlexInt, with a high dynamic range similar to FP, yet low max rounding error, targeting efficient representation of DNNs for inference at 8 bits and below. We also propose a novel FlexInt MAC (Multiply-Accumulate) hardware architecture. Our experimental results using large networks on image classification and natural language processing demonstrate that our FlexInt can deliver more robust performance and far superior worst-case accuracy, compared to both INT and FP across various data distributions; has a hardware overhead similar to that of FP; and can consistently make near-Pareto-optimal area-accuracy trade-offs across diverse networks. Minuk Hong, Hyeon Uk Sim, Sugil Lee, Jongeun Lee |
ICCAD | 4 |
| 2024 | Cold-start Bundle Recommendation via Popularity-based Coalescence and Curriculum HeatingabstractHow can we recommend cold-start bundles to users? The cold-start problem in bundle recommendation is crucial because new bundles are continuously created on the Web for various marketing purposes. Despite its importance, existing methods for cold-start item recommendation are not readily applicable to bundles. They depend overly on historical information, even for less popular bundles, failing to address the primary challenge of the highly skewed distribution of bundle interactions. In this work, we propose CoHeat (Popularity-based Coalescence and Curriculum Heating), an accurate approach for cold-start bundle recommendation. CoHeat first represents users and bundles through graph-based views, capturing collaborative information effectively. To estimate the user-bundle relationship more accurately, CoHeat addresses the highly skewed distribution of bundle interactions through a popularity-based coalescence approach, which incorporates historical and affiliation information based on the bundle's popularity. Furthermore, it effectively learns latent representations by exploiting curriculum learning and contrastive learning. CoHeat demonstrates superior performance in cold-start bundle recommendation, achieving up to 193% higher nDCG@20 compared to the best competitor. Hyunsik Jeon, Jongeun Lee, Jeongin Yun, U Kang |
WWW | 2 |
| 2024 | Perching and Grasping Using a Passive Dynamic Bioinspired GripperabstractThe ability to grasp objects broadens the application range of unmanned aerial vehicles (UAVs) by allowing interactions with the environment. The difficulty in performing a midair grasp is the high probability of impact between the UAV's foot and the target. For a successful grasp, the foot must smoothly absorb the energy of impact and simultaneously engage with the target in a short period of time. We present a bioinspired passive dynamic foot in which the claws are actuated solely by the impact energy. Our gripper simultaneously resolves the issue of smooth absorption of the impact energy and fast closure of the claws by linking the motion of an ankle linkage and the claws through soft tendons. We study the dynamics of impact and use the stiffness of the tendon as our design/control parameter to adjust the mechanics of the gripper for smooth recycling of the impact energy. Our gripper closes within 45 ms after initial contact with the impacting object without requiring any controller or actuation energy. An electroadhesive locking mechanism attached to the tendon locks the claws within 20 ms after reaching closed configuration. We demonstrated the effectiveness of our gripper by integrating it in an UAV and performing a variety of passive dynamic perching and grasping tasks. Amir Firouzeh, Jongeun Lee, Hyunsoo Yang, Kyu-Jin Cho |
IEEE Trans. Robotics | 2 |
| 2023 | NTT-PIM: Row-Centric Architecture and Mapping for Efficient Number-Theoretic Transform on PIMabstractRecently DRAM-based PIMs (processing-in-memories) with unmodified cell arrays have demonstrated impressive performance for accelerating AI applications. However, due to the very restrictive hardware constraints, PIM remains an accelerator for simple functions only. In this paper we propose NTT-PIM, which is based on the same principles such as no modification of cell arrays and very restrictive area budget, but shows state-of-the-art performance for a very complex application such as NTT, thanks to features optimized for the application’s characteristics, such as in-place update and pipelining via multiple buffers. Our experimental results demonstrate that our NTT-PIM can outperform previous best PIM-based NTT accelerators in terms of runtime by 1.7 ∼ 17× while having negligible area and power overhead. Jaewoo Park 0006, Sugil Lee, Jongeun Lee |
DAC | 3 |
| 2023 | Hyperdimensional Computing as a Rescue for Efficient Privacy-Preserving Machine Learning-as-a-ServiceabstractMachine learning models are often provisioned as a cloud-based service where the clients send their data to the service provider to obtain the result. This setting is commonplace due to the high value of the models, but it requires the clients to forfeit the privacy that the query data may contain. Homomorphic encryption (HE) is a promising technique to address this adversity. With HE, the service provider can take encrypted data as a query and run the model without decrypting it. The result remains encrypted, and only the client can decrypt it. All these benefits come at the cost of computational cost because HE turns simple floating-point arithmetic into the computation between long (of degree ≥ 1024) polynomials. Previous work has proposed to tailor deep neural networks for efficient computation over encrypted data, but already high computational cost is again amplified by HE, hindering performance improvement. In this paper we show hyperdimensional computing can be a rescue for privacy-preserving machine learning over encrypted data. We find that the advantage of hyperdimensional computing in performance is amplified when working with HE. This observation led us to design HE-HDC, a machine-learning inference system that uses hyperdimensional computing with HE. We carefully structure the machine learning service so that the server will perform only the HE-friendly computation. Moreover, we adapt the computation and HE parameters to expedite computation while preserving accuracy and security. Our experimental result based on real measurements shows that HE-HDC outperforms existing systems by 26 ~ 3000 x times with comparable classification accuracy. Jaewoo Park 0006, Chenghao Quan, Hyungon Moon, Jongeun Lee |
ICCAD | 4 |
| 2023 | Aggregately Diversified Bundle Recommendation via Popularity Debiasing and Configuration-Aware Reranking
Hyunsik Jeon, Jongjin Kim 0001, Jaeri Lee, Jongeun Lee, U Kang |
PAKDD (3) | 4 |
| 2023 | Partial Sum Quantization for Reducing ADC Size in ReRAM-Based Neural Network AcceleratorsabstractWhile resistive random-access memory (ReRAM) crossbar arrays have the potential to significantly accelerate deep neural network (DNN) training through fast and low-cost matrix–vector multiplication, peripheral circuits like analog-to-digital converters (ADCs) create a high overhead. These ADCs consume over half of the chip power and a considerable portion of the chip cost. To address this challenge, we propose advanced quantization techniques that can significantly reduce the ADC overhead of ReRAM crossbar arrays (RCAs). Our methodology interprets ADC as a quantization mechanism, allowing us to scale the range of ADC input optimally along with the weight parameters of a DNN, resulting in multiple-bit reduction in ADC precision. This approach reduces ADC size and power consumption by several times, and it is applicable to any DNN type (binarized or multibit) and any RCA size. Additionally, we propose ways to minimize the overhead of the digital scaler, which is an essential part of our scheme and sometimes required. Our experimental results using ResNet-18 on the ImageNet dataset demonstrate that our method can reduce the size of the ADC by 32 times compared to ISAAC with only a minimal accuracy loss degradation of 0.24%. We also present evaluation results in the presence of ReRAM nonideality (such as stuck-at fault). Azat Azamat, Faaiz Asim, Jongeun Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | Offline Training-Based Mitigation of IR Drop for ReRAM-Based Deep Neural Network AcceleratorsabstractRecently, resistive RAM (ReRAM)-based hardware accelerators showed unprecedented performance compared the digital accelerators. Technology scaling causes an inevitable increase in interconnect wire resistance, which leads to IR drops that could limit the performance of ReRAM-based accelerators. These IR drops deteriorate the signal integrity and quality, especially in the crossbar structures which are used to build high-density ReRAMs. Hence, finding a software solution, which can predict the effect of IR drop without involving expensive hardware or SPICE simulations, is very desirable. In this article, we propose two neural networks models to predict the impact of the IR drop problem. These models are used to evaluate the performance of the different deep neural network (DNN) models including binary and quantized neural networks showing similar performance (i.e., recognition accuracy) to the golden validation (i.e., SPICE-based DNN validation). In addition, these predication models are incorporated into the DNN training framework to efficiently retrain the DNN models and bridge the accuracy drop. To further enhance the validation accuracy, we propose incremental training methods. The DNN validation results, done through SPICE simulations, show very high improvement in performance close to the baseline performance, which demonstrates the efficacy of the proposed method even with challenging datasets, such as CIFAR10 and SVHN. Sugil Lee, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Training-Free Stuck-At Fault Mitigation for ReRAM-Based Deep Learning AcceleratorsabstractAlthough Resistive RAMs can support highly efficient matrix–vector multiplication, which is very useful for machine learning and other applications, the nonideal behavior of hardware, such as stuck-at fault (SAF) and IR drop is an important concern in making ReRAM crossbar array-based deep learning accelerators. Previous work has addressed the nonideality problem through either redundancy in hardware, which requires a permanent increase of hardware cost, or software retraining, which may be even more costly or unacceptable due to its need for a training dataset as well as high computation overhead. In this article, we propose a very lightweight method that can be applied on top of existing hardware or software solutions. Our method, called forward-parameter tuning (FPT), takes advantage of a certain statistical property existing in the activation data of neural network layers, and can mitigate the impact of mild nonidealities in ReRAM crossbar arrays (RCAs) for deep learning applications without using any hardware, a dataset, or gradient-based training. Our experimental results using MNIST, CIFAR-10, and CIFAR-100, and ImageNet datasets in binary and multibit networks demonstrate that our technique is very effective, both alone and together with previous methods, up to 20% fault rate, which is higher than even some of the previous remapping methods. We also evaluate our method in the presence of other nonidealities, such as variability and IR drop. Furthermore, we provide an analysis based on the concept of the effective fault rate (EFR), which not only demonstrates that EFR can be a useful tool to predict the accuracy of faulty RCA-based neural networks but also explains why mitigating the SAF problem is more difficult with multibit neural networks. Chenghao Quan, Mohamed E. Fouda, Sugil Lee, Giju Jung, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Centered Symmetric Quantization for Hardware-Efficient Low-Bit Neural Networks
Faaiz Asim, Jaewoo Park 0006, Azat Azamat, Jongeun Lee |
BMVC | 4 |
| 2022 | An Empirical Study on How People Perceive AI-generated MusicabstractMusic creation is difficult because one must express one's creativity while following strict rules. The advancement of deep learning technologies has diversified the methods to automate complex processes and express creativity in music composition. However, prior research has not paid much attention to exploring the audiences' subjective satisfaction to improve music generation models. In this paper, we evaluate human satisfaction with the state-of-the-art automatic symbolic music generation models using deep learning. In doing so, we define a taxonomy for music generation models and suggest nine subjective evaluation metrics. Through an evaluation study, we obtained more than 700 evaluations from 100 participants, using the suggested metrics. Our evaluation study reveals that the token representation method and models' characteristics affect subjective satisfaction. Through our qualitative analysis, we deepen our understanding of AI-generated music and suggested evaluation metrics. Lastly, we present lessons learned and discuss future research directions of deep learning models for music creation. Hyeshin Chu, Joohee Kim, Seongouk Kim, Hongkyu Lim, Hyunwook Lee, Seungmin Jin, Jongeun Lee, Taehwan Kim 0013, Sungahn Ko |
CIKM | 7 |
| 2022 | Non-uniform Step Size Quantization for Accurate Post-training Quantization
Sangyun Oh, Hyeon Uk Sim, Jounghyun Kim, Jongeun Lee |
ECCV (11) | 4 |
| 2022 | Squeezing Accumulators in Binary Neural Networks for Extremely Resource-Constrained ApplicationsabstractThe cost and power consumption of BNN (Binarized Neural Network) hardware is dominated by additions. In particular, accumulators account for a large fraction of hardware overhead, which could be effectively reduced by using reduced-width accumulators. However, it is not straightforward to find the optimal accumulator width due to the complex interplay between width, scale, and the effect of training. In this paper we present algorithmic and hardware-level methods to find the optimal accumulator size for BNN hardware with minimal impact on the quality of result. First, we present partial sum scaling, a top-down approach to minimize the BNN accumulator size based on advanced quantization techniques. We also present an efficient, zero-overhead hardware design for partial sum scaling. Second, we evaluate a bottom-up approach that is to use saturating accumulator, which is more robust against overflows. Our experimental results using CIFAR-10 dataset demonstrate that our partial sum scaling along with our optimized accumulator architecture can reduce the area and power consumption of datapath by 15.50% and 27.03%, respectively, with little impact on inference performance (less than 2%), compared to using 16-bit accumulator. Azat Azamat, Jaewoo Park 0006, Jongeun Lee |
ICCAD | 3 |
| 2022 | Accurate Prediction of ReRAM Crossbar Performance Under I-V Nonlinearity and IR DropabstractDespite the promise of extremely efficient matrix-vector multiplication (MVM) by ReRAM crossbar arrays (RCAs), maintaining high accuracy has been challenging due to nonidealities such as wire resistance (also known as IR drop) and I-V nonlinearity (i.e., voltage-dependent conductance). For system architects, a fast method to accurately predict the MVM output of an RCA under nonidealities is highly desirable. While IR drop alone without I-V nonlinearity can be efficiently predicted, the existence of I-V nonlinearity makes the problem much harder. In this paper we propose a novel algorithm based on iterative refinement, which can predict with high accuracy the outcome of an MVM operation on an RCA in the presence of both I-V nonlinearity and IR drop. Our experiments using binary RCAs of different sizes demonstrate that our proposed method is order-of-magnitude more accurate than previous methods in terms of RMS error. We also present case studies predicting hardware-realistic accuracy of binarized neural networks on RCAs as well as nonideality-aware retraining, demonstrating the efficacy of our method for early design space exploration of ReRAM-based accelerators. Sugil Lee, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
ICCD | 3 |
| 2022 | MLogNet: A Logarithmic Quantization-Based Accelerator for Depthwise Separable ConvolutionabstractIn this article, we propose a novel logarithmic quantization-based deep neural network (DNN) architecture for depthwise separable convolution (DSC) networks. Our architecture is based on selective two-word logarithmic quantization (STLQ), which improves accuracy greatly over logarithmic-scale quantization while retaining the speed and area advantage of logarithmic quantization. On the other hand, it also comes with the synchronization problem due to variable-latency processing elements (PEs), which we address through a novel architecture and a compile-time optimization technique. Our architecture is dynamically reconfigurable to support various combinations of depthwise versus pointwise convolution layers efficiently. Our experimental results using layers from MobileNetV2 and ShuffleNetV2 demonstrate that our architecture is significantly faster and more area efficient than previous DSC accelerator architectures as well as previous accelerators utilizing logarithmic quantization. Jooyeon Choi, Hyeon Uk Sim, Sangyun Oh, Sugil Lee, Jongeun Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Specializing CGRAs for Light-Weight Convolutional Neural NetworksabstractDeep neural network (DNN) processing units, or DPUs, are one of the most energy-efficient platforms for DNN applications. However, designing new DPUs for every DNN model is very costly and time consuming. In this article, we propose an alternative approach: to specialize coarse-grained reconfigurable architectures (CGRAs), which are already quite capable of delivering high performance and high energy efficiency for compute-intensive kernels. We identify a small set of architectural features on a baseline CGRA to enable high-performance mapping of depthwise convolution (DWC) and pointwise convolution (PWC) kernels, which are the most important building block in recent light-weight DNN models. Our experimental results using MobileNets demonstrate that our proposed CGRA enhancement can deliver$8\sim 18\times $improvement in area-delay product (ADP) depending on layer type, over a baseline CGRA with a state-of-the-art CGRA compiler. Moreover, our proposed CGRA architecture can also speed up 3-D convolution with similar efficiency as previous work, demonstrating the effectiveness of our architectural features beyond depthwise separable convolution (DSC) layers. Jungi Lee 0001, Jongeun Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Automated Log-Scale Quantization for Low-Cost Deep Neural NetworksabstractQuantization plays an important role in deep neural network (DNN) hardware. In particular, logarithmic quantization has multiple advantages for DNN hardware implementations, and its weakness in terms of lower performance at high precision compared with linear quantization has been recently remedied by what we call selective two-word logarithmic quantization (STLQ). However, there is a lack of training methods designed for STLQ or even logarithmic quantization in general. In this paper we propose a novel STLQ-aware training method, which significantly out-performs the previous state-of-the-art training method for STLQ. Moreover, our training results demonstrate that with our new training method, STLQ applied to weight parameters of ResNet-18 can achieve the same level of performance as state-of-the-art quantization method, APoT, at 3-bit precision. We also apply our method to various DNNs in image enhancement and semantic segmentation, showing competitive results. Sangyun Oh, Hyeon Uk Sim, Sugil Lee, Jongeun Lee |
CVPR | 4 |
| 2021 | Cost- and Dataset-free Stuck-at Fault Mitigation for ReRAM-based Deep Learning AcceleratorsabstractResistive RAMs can implement extremely efficient matrix vector multiplication, drawing much attention for deep learning accelerator research. However, high fault rate is one of the fundamental challenges of ReRAM crossbar array-based deep learning accelerators. In this paper we propose a dataset-free, cost-free method to mitigate the impact of stuck-at faults in ReRAM crossbar arrays for deep learning applications. Our technique exploits the statistical properties of deep learning applications, hence complementary to previous hardware or algorithmic methods. Our experimental results using MNIST and CIFAR-10 datasets in binary networks demonstrate that our technique is very effective, both alone and together with previous methods, up to 20 % fault rate, which is higher than the previous remapping methods. We also evaluate our method in the presence of other non-idealities such as variability and IR drop. Giju Jung, Mohamed E. Fouda, Sugil Lee, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
DATE | 4 |
| 2021 | NP-CGRA: Extending CGRAs for Efficient Processing of Light-weight Deep Neural NetworksabstractCoarse-grained reconfigurable architectures (CGRAs) can provide both high energy efficiency and flexibility, making them well-suited for machine learning applications. However previous work on CGRAs has a very limited support for deep neural networks (DNNs), especially for recent lightweight models such as depthwise separable convolution (DSC), which are an important workload for mobile environment. In this paper, we propose a set of architecture extensions and a mapping scheme to greatly enhance CGRA's performance for DSC kernels. Our experimental results using MobileNets demonstrate that our proposed CGRA enhancement can deliver 8~ 18x improvement in area-delay product depending on layer type, over a baseline CGRA with a state-of-the-art CGRA compiler. Moreover, our proposed CGRA architecture can also speed up 3D convolution with similar efficiency as previous work, demonstrating the effectiveness of our architectural features beyond DSC layers. Jungi Lee 0001, Jongeun Lee |
DATE | 2 |
| 2021 | Quarry: Quantization-based ADC Reduction for ReRAM-based Deep Neural Network AcceleratorsabstractReRAM (Resistive Random-Access Memory) crossbar arrays have the potential to provide extremely fast and low-cost DNN (Deep Neural Network) acceleration. However, peripheral circuits, in particular ADCs (Analog-Digital Converters), can be a large overhead and/or slow down the operation considerably. In this paper we propose to use advanced quantization techniques to reduce the ADC overhead of ReRAM crossbar arrays. Our method does not require any hardware change but can reduce the overhead of ADC greatly. Our methodology is also general, having no restriction in terms of DNN type (binarized or multi-bit) or ReRAM crossbar array size. Our experimental results using ResNet on ImageNet dataset demonstrate that our method can reduce the size of ADC by 32× compared with ISAAC at very little accuracy loss of 0.24%p. Azat Azamat, Faaiz Asim, Jongeun Lee |
ICCAD | 3 |
| 2021 | Fast and Low-Cost Mitigation of ReRAM Variability for Deep Learning ApplicationsabstractTo overcome the programming variability (PV) of ReRAM crossbar arrays (RCAs), the most common method is program-verify, which, however, has high energy and latency overhead. In this paper we propose a very fast and low-cost method to mitigate the effect of PV and other variability for RCA-based DNN (Deep Neural Network) accelerators. Leveraging the statistical properties of DNN output, our method called Online Batch-Norm Correction (OBNC) can compensate for the effect of programming and other variability on RCA output without using on-chip training or an iterative procedure, and is thus very fast. Also our method does not require a nonideality model or a training dataset, hence very easy to apply. Our experimental results using ternary neural networks with binary and 4-bit activations demonstrate that our OBNC can recover the baseline performance in many variability settings and that our method outperforms a previously known method (VCAM) by large margins when input distribution is asymmetric or activation is multi-bit. Sugil Lee, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
ICCD | 3 |
| 2020 | Learning to Predict IR Drop with Effective Training for ReRAM-based Neural Network HardwareabstractDue to the inevitability of the IR drop problem in passive ReRAM crossbar arrays, finding a software solution that can predict the effect of IR drop without the need of expensive SPICE simulations, is very desirable. In this paper, two simple neural networks are proposed as software solution to predict the effect of IR drop. These networks can be easily integrated in any deep neural network framework to incorporate the IR drop problem during training. As an example, the proposed solution is integrated in BinaryNet framework and the test validation results, done through SPICE simulations, show very high improvement in performance close to the baseline performance, which demonstrates the efficacy of the proposed method. In addition, the proposed solution outperforms the prior work on challenging datasets such as CIFAR10 and SVHN. Sugil Lee, Giju Jung, Mohamed E. Fouda, Jongeun Lee, Ahmed M. Eltawil, Fadi J. Kurdahi |
DAC | 4 |
| 2020 | Architecture-Accuracy Co-optimization of ReRAM-based Low-cost Neural Network ProcessorabstractResistive RAM (ReRAM) is a promising technology with such advantages as small device size and in-memory-computing capability. However, designing optimal AI processors based on ReRAMs is challenging due to the limited precision, and the complex interplay between quality of result and hardware efficiency. In this paper we present a study targeting a low-power low-cost image classification application. We discover that the trade-off between accuracy and hardware efficiency in ReRAM-based hardware is not obvious and even surprising, and our solution developed for a recently fabricated ReRAM device achieves both the state-of-the-art efficiency and empirical assurance on the high quality of result. Segi Lee, Sugil Lee, Jongeun Lee, Jong-Moon Choi, Do-Wan Kwon, Seung-Kwang Hong, Kee-Won Kwon |
ACM Great Lakes Symposium on VLSI | 3 |
| 2020 | SparTANN: sparse training accelerator for neural networks with threshold-based sparsificationabstractWhile sparsity has been exploited in many inference accelerators, not much work is done for training accelerators. Exploiting sparsity in training accelerators involves multiple issues, including where to find sparsity, how to exploit sparsity, and how to create more sparsity. In this paper we present a novel sparse training architecture that can exploit sparsity in gradient tensors in both back propagation and weight update computation. We also propose a single-pass sparsification algorithm, which is a hardware-friendly version of a recently proposed sparse training algorithm, that can create additional sparsity aggressively during training. Our experimental results using large networks such as AlexNet and GoogleNet demonstrate that our sparse training architecture can accelerate convolution layer training time by 4.20~8.88× over baseline dense training without accuracy loss, and further increase the training speed by 7.30~11.87× over the baseline with minimal accuracy loss. Hyeon Uk Sim, Jooyeon Choi, Jongeun Lee |
ISLPED | 3 |
| 2019 | On-chip memory optimization for high-level synthesis of multi-dimensional data on FPGAabstractIt is very challenging to design an on-chip memory architecture for high-performance kernels with large amount of computation and data. The on-chip memory architecture must support efficient data access from both the computation part and the external memory part, which often have very different expectations about how data should be accessed and stored. Previous work provides only a limited set of optimizations. In this paper we show how to fundamentally restructure on-chip buffers, by decoupling logical array view from the physical buffer view, and providing general mapping schemes for the two. Our framework considers the entire data flow from the external memory to the computation part in order to minimize resource usage without creating performance bottleneck. Our experimental results demonstrate that our proposed technique can generate solutions that reduce memory usage significantly (2X over the conventional method), and successfully generate optimized on-chip buffer architectures without costly design iterations for highly optimized computation kernels. Daewoo Kim, Sugil Lee, Jongeun Lee |
ASP-DAC | 3 |
| 2019 | XOMA: exclusive on-chip memory architecture for energy-efficient deep learning accelerationabstractState-of-the-art deep neural networks (DNNs) require hundreds of millions of multiply-accumulate (MAC) computations to perform inference, e.g. in image-recognition tasks. To improve the performance and energy efficiency, deep learning accelerators have been proposed, realized both on FPGAs and as custom ASICs. Generally, such accelerators comprise many parallel processing elements, capable of executing large numbers of concurrent MAC operations. From the energy perspective, however, most consumption arises due to memory accesses, both to off-chip external memory, and on-chip buffers. In this paper, we propose an on-chip DNN co-processor architecture where minimizing memory accesses is the primary design objective. To the maximum possible extent, off-chip memory accesses are eliminated, providing lowest-possible energy consumption for inference. Compared to a state-of-the-art ASIC, our architecture requires 36% fewer external memory accesses and 53% less energy consumption for low-latency image classification. Hyeon Uk Sim, Jason Helge Anderson, Jongeun Lee |
ASP-DAC | 3 |
| 2019 | Log-quantized stochastic computing for memory and computation efficient DNNsabstractFor energy efficiency, many low-bit quantization methods for deep neural networks (DNNs) have been proposed. Among them, logarithmic quantization is being highlighted showing acceptable deep learning performance. It also simplifies high-cost multipliers as well as reducing memory footprint drastically. Meanwhile, stochastic computing (SC) was proposed for low-cost DNN acceleration and the recently proposed SC multiplier improved the accuracy and latency significantly which are main drawbacks of SC. However, in their binary-interfaced system which yet costs much less than storing all stochastic stream, quantization is basically linear as same as conventional fixed-point binary. We applied logarithmically quantized DNNs to the state-of-the-art SC multiplier and studied how it can benefit. We found that SC multiplication on logarithmically quantized input is more accurate and it can help fine-tuning process. Furthermore, we designed the much low-cost SC-DNN accelerator utilizing the reduced complexity of inputs. Finally, while logarithmic quantization benefits data flow, proposed architecture achieves 40% and 24% less area and power consumption than the previous SC-DNN accelerator. Its area X latency product is smaller even than the shifter based accelerator. Hyeon Uk Sim, Jongeun Lee |
ASP-DAC | 2 |
| 2019 | Efficient FPGA implementation of local binary convolutional neural networkabstractBinarized Neural Networks (BNN) has shown a capability of performing various classification tasks while taking advantage of computational simplicity and memory saving. The problem with BNN, however, is a low accuracy on large convolutional neural networks (CNN). Local Binary Convolutional Neural Network (LBCNN) compensates accuracy loss of BNN by using standard convolutional layer together with binary convolutional layer and can achieve as high accuracy as standard AlexNet CNN. For the first time we propose FPGA hardware design architecture of LBCNN and address its unique challenges. We present performance and resource usage predictor along with design space exploration framework. Our architecture on LBCNN AlexNet shows 76.6% higher performance in terms of GOPS, 2.6X and 2.7X higher performance density in terms of GOPS/Slice, and GOPS/DSP compared to previous FPGA implementation of standard AlexNet CNN. Aidyn Zhakatayev, Jongeun Lee |
ASP-DAC | 2 |
| 2019 | Successive Log Quantization for Cost-Efficient Neural Networks Using Stochastic ComputingabstractDespite the multifaceted benefits of stochastic computing (SC) such as low cost, low power, and flexible precision, SC-based deep neural networks (DNNs) still suffer from the long-latency problem, especially for those with high precision requirements. While log quantization can be of help, it has its own accuracy-saturation problem due to uneven precision distribution. In this paper we propose successive log quantization (SLQ), which extends log quantization with significant improvements in precision and accuracy, and apply it to state-of-the-art SC-DNNs. SLQ reuses the existing datapath of log quantization, and thus retains its advantages such as simple multiplier hardware. Our experimental results demonstrate that our SLQ can significantly extend both the accuracy and efficiency of SC-DNNs over the state-of-the-art solutions, including linear-quantized and log-quantized SC-DNNs, achieving less than 1~1.5%p accuracy drop for AlexNet, SqueezeNet, and VGG-S at mere 4~5-bit weight resolution. Sugil Lee, Hyeon Uk Sim, Jooyeon Choi, Jongeun Lee |
DAC | 4 |
| 2019 | Cost-effective stochastic MAC circuits for deep neural networks
Hyeon Uk Sim, Jongeun Lee |
Neural Networks | 2 |
| 2019 | Double MAC on a DSP: Boosting the Performance of Convolutional Neural Networks on FPGAsabstractDeep learning workloads, such as convolutional neural networks (CNNs) are important due to increasingly demanding high-performance hardware acceleration. One distinguishing feature of a deep learning workload is that it is inherently resilient to small numerical errors and thus works very well with low precision hardware. We propose a novel method called double multiply-and-accumulate (MAC) to theoretically double the computation rate of CNN accelerators by packing two MAC operations into one digital signal processing block of off-the-shelf field-programmable gate arrays (FPGAs). We overcame several technical challenges by exploiting the mode of operation in the CNN accelerator. We have validated our method through FPGA synthesis and Verilog simulation, and evaluated our method by applying it to the state-of-the-art CNN accelerator. The double MAC approach used can double the computation throughput of a CNN layer. On the network level (all convolution layers combined), the performance improvement varies depending on the CNN application and FPGA size, from 14% to more than 80% over a highly optimized state-of-the-art accelerator solution, without sacrificing the output quality significantly. Sugil Lee, Daewoo Kim, Dong Nguyen 0001, Jongeun Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | DPS: dynamic precision scaling for stochastic computing-based deep neural networksabstractStochastic computing (SC) is a promising technique with advantages such as low-cost, low-power, and error-resilience. However so far SC-based CNN (convolutional neural network) accelerators have been kept to relatively small CNNs only, primarily due to the inherent precision disadvantage of SC. At the same time, previous SC architectures do not exploit the dynamic precision capability, which can be crucial in providing efficiency as well as flexibility in SC-CNN implementations. In this paper we present a DPS (dynamic precision scaling) SC-CNN that is able to exploit dynamic precision with very low overhead, along with the design methodology for it. Our experimental results demonstrate that our DPS SC-CNN is highly efficient and accurate up to ImageNet-targeting CNNs, and show efficiency improvements over conventional digital designs ranging in 50~100% in operations-per-area depending on the DNN and the application scenario, while losing less than 1% in recognition accuracy. Hyeon Uk Sim, Saken Kenzhegulov, Jongeun Lee |
DAC | 3 |
| 2018 | Sign-magnitude SC: getting 10X accuracy for free in stochastic computing for deep neural networksabstractStochastic computing (SC) is a promising computing paradigm for applications with low precision requirement, stringent cost and power restriction. One known problem with SC, however, is the low accuracy especially with multiplication. In this paper we propose a simple, yet very effective solution to the low-accuracy SC-multiplication problem, which is critical in many applications such as deep neural networks (DNNs). Our solution is based on an old concept of sign-magnitude, which, when applied to SC, has unique advantages. Our experimental results using multiple DNN applications demonstrate that our technique can improve the efficiency of SC-based DNNs by about 32X in terms of latency over using bipolar SC, with very little area overhead (about 1%). Aidyn Zhakatayev, Sugil Lee, Hyeon Uk Sim, Jongeun Lee |
DAC | 4 |
| 2018 | FPGA Architecture Enhancements for Efficient BNN ImplementationabstractBinarized neural networks (BNNs) are ultra-reduced precision neural networks, having weights and activations restricted to single-bit values. BNN computations operate on bitwise data, making them particularly amenable to hardware implementation. In this paper, we first analyze BNN implementations on contemporary commercial 20nm FPGAs. We then propose two lightweight architectural changes that significantly improve the logic density of FPGA BNN implementations. The changes involve incorporating additional carry-chain circuitry into logic elements, where the additional circuitry is connected in a specific way to benefit BNN computations. The architectural changes are evaluated in the context of state-of-the-art Intel and Xilinx FPGAs and shown to provide over 2x area reduction in the key BNN computational task (the XNOR-popcount sub-circuit), at a modest performance cost of less than 2%. Jin Hee Kim, Jongeun Lee, Jason Helge Anderson |
FPT | 2 |
| 2018 | Architecture Exploration of Standard-Cell and FPGA-Overlay CGRAs Using the Open-Source CGRA-ME FrameworkabstractWe describe an open-source software framework,CGRA-ME, for the modeling and exploration of coarse-grained reconfigurable architectures (CGRAs). CGRAs are programmable hardware devices having large ALU-like logic blocks, and datapath bus-style inter-connect. CGRAs are positioned between fine-grained FPGAs and standard-cell ASICs on the spectrum of programmability - they are less flexible than FPGAs, yet are more flexible than ASICs. With CGRA-ME, an architect can describe a CGRA architecture in an XML-based language. The framework also allows the architect to map benchmarks onto the architecture and provides automatic generation of Verilog RTL for the modeled architecture. This allows the architect to simulate for verification purposes, and perform synthesis to either an ASIC or FPGA-overlay implementation of the CGRA, assessing performance, area, and power consumption. In an experimental study, we use CGRA-ME to model, map benchmarks onto, and evaluate several variants of a widely known CGRA, considering both standard-cell and FPGA-overlay physical realizations of the CGRA. S. Alexander Chin, Kuang Ping Niu, Matthew J. P. Walker, Shizhang Yin, Alexander Mertens, Jongeun Lee, Jason Helge Anderson |
ISPD | 6 |
| 2018 | An Efficient and Accurate Stochastic Number Generator Using Even-Distribution CodingabstractStochastic computing (SC) is a promising approach for low-power and low-cost applications with the added benefit of high error tolerance. However, the high overhead of generating stochastic bitstreams can offset the advantages of SC especially when a large number of bitstreams are needed. In this paper, we propose a new stochastic number generator (SNG) that significantly reduces area and energy while improving accuracy. Experimental results show that the proposed SNG can reduce energy by more than 72% compared with the state-of-the-art designs. Aidyn Zhakatayev, Kyounghoon Kim, Kiyoung Choi, Jongeun Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | Scalable stochastic-computing accelerator for convolutional neural networksabstractStochastic Computing (SC) is an alternative design paradigm particularly useful for applications where cost is critical. SC has been applied to neural networks, as neural networks are known for their high computational complexity. However previous work in this area has critical limitations such as the fully-parallel architecture assumption, which prevent them from being applicable to recent ones such as convolutional neural networks, or ConvNets. This paper presents the first SC architecture for ConvNets, shows its feasibility, with detailed analyses of implementation overheads. Our SC-ConvNet is a hybrid between SC and conventional binary design, which is a marked difference from earlier SC-based neural networks. Though this might seem like a compromise, it is a novel feature driven by the need to support modern ConvNets at scale, which commonly have many, large layers. Our proposed architecture also features hybrid layer composition, which helps achieve very high recognition accuracy. Our detailed evaluation results involving functional simulation and RTL synthesis suggest that SC-ConvNets are indeed competitive with conventional binary designs, even without considering inherent error resilience of SC. Hyeon Uk Sim, Dong Nguyen 0001, Jongeun Lee, Kiyoung Choi |
ASP-DAC | 3 |
| 2017 | A New Stochastic Computing Multiplier with Application to Deep Convolutional Neural NetworksabstractStochastic computing (SC) allows for extremely low cost and low power implementations of common arithmetic operations. However inherent random fluctuation error and long latency of SC lead to the degradation of accuracy and energy efficiency when applied to convolutional neural networks (CNNs). In this paper we address the two critical problems of SC-based CNNs, by proposing a novel SC multiply algorithm and its vector extension, SC-MVM (Matrix-Vector Multiplier), under which one SC multiply takes just a few cycles, generates much more accurate results, and can be realized with significantly less cost, as compared to the conventional SC method. Our experimental results using CNNs designed for MNIST and CIFAR-10 datasets demonstrate that not only is our SC-based CNN more accurate and 40X~490X more energy-efficient in computation than the conventional SC-based ones, but ours can also achieve lower area-delay product and lower energy compared with bitwidth-optimized fixed-point implementations of the same accuracy. Hyeon Uk Sim, Jongeun Lee |
DAC | 2 |
| 2017 | Double MAC: Doubling the performance of convolutional neural networks on modern FPGAsabstractThis paper presents a novel method to double the computation rate of convolutional neural network (CNN) accelerators by packing two multiply-and-accumulate (MAC) operations into one DSP block of off-the-shelf FPGAs (called Double MAC). While a general SIMD MAC using a single DSP block seems impossible, our solution is tailored for the kind of MAC operations required for a convolution layer. Our preliminary evaluation shows that not only can our Double MAC approach increase the computation throughput of a CNN layer by twice with essentially the same resource utilization, the network level performance can also be improved by 14~84% over a highly optimized state-of-the-art accelerator solution depending on the CNN hyper-parameters. Dong Nguyen 0001, Daewoo Kim, Jongeun Lee |
DATE | 3 |
| 2017 | Design space exploration of FPGA accelerators for convolutional neural networksabstractThe increasing use of machine learning algorithms, such as Convolutional Neural Networks (CNNs), makes the hardware accelerator approach very compelling. However the question of how to best design an accelerator for a given CNN has not been answered yet, even on a very fundamental level. This paper addresses that challenge, by providing a novel framework that can universally and accurately evaluate and explore various architectural choices for CNN accelerators on FPGAs. Our exploration framework is more extensive than that of any previous work in terms of the design space, and takes into account various FPGA resources to maximize performance including DSP resources, on-chip memory, and off-chip memory bandwidth. Our experimental results using some of the largest CNN models including one that has 16 convolutional layers demonstrate the efficacy of our framework, as well as the need for such a high-level architecture exploration approach to find the best architecture for a CNN model. Atul Rahman, Sangyun Oh, Jongeun Lee, Kiyoung Choi |
DATE | 3 |
| 2017 | FPGA implementation of convolutional neural network based on stochastic computingabstractThere has been a body of research to use stochastic computing (SC) for the implementation of neural networks, in the hope that it will reduce the area cost and energy consumption. However, no working neural network system based on stochastic computing has been demonstrated to support the viability of SC-based deep neural networks in terms of both recognition accuracy and cost/energy efficiency. In this demonstration we present an SC-based deep nenural network system that is highly accurate and efficient. Our system takes an input image and processes it with a convolutional neural network implemented on an FPGA using stochastic computing to recognize the input image, with nearly the same accuracy as conventional binary implementations. Daewoo Kim, Mansureh S. Moghaddam, Hossein Moradian, Hyeon Uk Sim, Jongeun Lee, Kiyoung Choi |
FPT | 5 |
| 2017 | Accurate and Efficient Stochastic Computing Hardware for Convolutional Neural NetworksabstractThis paper presents an efficient unipolar stochastic computing hardware for convolutional neural networks (CNNs). It includes stochastic ReLU and optimized max function, which are key components in a CNN. To avoid the range limitation problem of stochastic numbers and increase the signal-to-noise ratio, we perform weight normalization and upscaling. In addition, to reduce the overhead of binary-to-stochastic conversion, we propose a scheme for sharing stochastic number generators among the neurons in a CNN. Experimental results show that our approach outperforms the previous ones based on stochastic computing in terms of accuracy, area, and energy consumption. Joonsang Yu, Kyounghoon Kim, Jongeun Lee, Kiyoung Choi |
ICCD | 3 |
| 2017 | Efficient Execution of Stream Graphs on Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-grained reconfigurable architectures (CGRAs) can provide extremely energy-efficient acceleration for applications that are rich in arithmetic operations such as digital signal processing and multimedia applications. Since those applications are often naturally represented by stream graphs, it is very compelling to develop optimization strategies for stream graphs on CGRAs. One unique property of stream graphs is that they contain many kernels or loops, which creates both advantages and challenges when it comes to mapping them to CGRAs. This paper addresses two main problems with it, namely, many-buffer problem and control overhead problem, and presents our results of optimizing the execution of stream graphs for CGRAs including our low-cost architecture extensions. Our evaluation results demonstrate that our software and hardware optimizations can help generate highly efficient mapping of stream applications to CGRAs, with 3.4× speedup on average at the application level over CPU-only execution, which is significant. Sangyun Oh, Hongsik Lee, Jongeun Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | An energy-efficient random number generator for stochastic circuitsabstractStochastic circuits provide very high efficiency in terms of gate area and power consumption compared with conventional binary logic. However, they require random bit streams generated by stochastic number generators (SNGs), which account for a significant portion of area and energy offsetting their merits. In this paper, we propose a new SNG that significantly reduces area and energy while improving accuracy in progressive precision. Experimental results show that the proposed SNG reduces energy by more than 72% compared to the state-of-the-art designs. Kyounghoon Kim, Jongeun Lee, Kiyoung Choi |
ASP-DAC | 2 |
| 2016 | Communication-aware mapping of stream graphs for multi-GPU platformsabstractStream graphs can provide a natural way to represent many applications in multimedia and DSP domains. Though the exposed parallelism of stream graphs makes it relatively easy to map them to GP (General Purpose)-GPUs, very large stream graphs as well as how to best exploit multi-GPU platforms to achieve scalable performance poses great challenges for stream graph mapping. Previous work considers either a single GPU only or is based on a crude heuristic that achieves a very low degree of workload balancing, and thus shows only limited scalability. In this paper we present a highly scalable GP-GPU mapping technique for large stream graphs with the following highlights: (1) an accurate GPU performance estimation model for subsets of stream graphs, (2) a novel partitioning heuristic exploiting stream graph's structural properties, and (3) ILP (Integer Linear Programming) formulation of the mapping problem. Our experimental results on a real GPU platform demonstrate that our technique can generate scalable performance for up to 4 GPUs with large stream graphs, and can generate highly optimized multi-GPU code especially for compute-bound ones. Dong Nguyen 0001, Jongeun Lee |
CGO | 2 |
| 2016 | Dynamic energy-accuracy trade-off using stochastic computing in deep neural networksabstractThis paper presents an efficient DNN design with stochastic computing. Observing that directly adopting stochastic computing to DNN has some challenges including random error fluctuation, range limitation, and overhead in accumulation, we address these problems by removing near-zero weights, applying weight-scaling, and integrating the activation function with the accumulator. The approach allows an easy implementation of early decision termination with a fixed hardware design by exploiting the progressive precision characteristics of stochastic computing, which was not easy with existing approaches. Experimental results show that our approach outperforms the conventional binary logic in terms of gate area, latency, and power consumption. Kyounghoon Kim, Jungki Kim, Joonsang Yu, Jungwoo Seo, Jongeun Lee, Kiyoung Choi |
DAC | 5 |
| 2016 | Efficient FPGA acceleration of Convolutional Neural Networks using logical-3D compute array
Atul Rahman, Jongeun Lee, Kiyoung Choi |
DATE | 2 |
| 2016 | Mapping Imperfect Loops to Coarse-Grained Reconfigurable ArchitecturesabstractNested loops represent a significant portion of application runtime in multimedia and DSP applications, an important domain of applications for coarse-grained reconfigurable architectures (CGRAs). While conventional approaches to mapping nested loops utilize only a single-dimensional pipelining, which is either along the innermost loop or along an outer loop, in this paper, we explore an orthogonal approach of pipelining along multiple loop dimensions by first flattening the loop nest. To remedy the inevitable problem of repetitive outer-loop computation in flattened loops, we present a small set of special operations that can effectively reduce the number and frequency of micro-operations in the pipelined loop. We also present a loop transformation technique that can make our special operations applicable to a broader range of loops, including those with triangular iteration spaces. Our experimental results using imperfect loops from StreamIt benchmarks demonstrate that our special operations can cover a large portion of operations in flattened loops, improve performance of nested loops by nearly 30% over using loop flattening only, and achieve near-ideal executions on CGRAs for imperfect loops. Hyeon Uk Sim, Hongsik Lee, Seongseok Seo, Jongeun Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | Efficient High-Level Synthesis for Nested Loops of Nonrectangular Iteration SpacesabstractMost existing solutions to pipelining nested loops are developed for general purpose processors, and may not work efficiently for field-programmable gate arrays due to loop control overhead. This is especially true when the nested loops have nonrectangular iteration spaces (IS). Thus we propose a novel method that can transform triangular IS-the most frequently found type of nonrectangular IS-into rectangular ones, so that other loop transformations can be effectively applied and the overall performance of nested loops can be maximized. Our evaluation results using the state-of-the-art Vivado high-level synthesis tool demonstrate that our technique can improve the performance of nested loops with nonrectangular IS significantly. Hyeon Uk Sim, Atul Rahman, Jongeun Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Optimizing stream program performance on CGRA-based systemsabstractCoarse-Grained Reconfigurable Architectures (CGRAs), often used as coprocessors for DSP and multimedia kernels, can deliver highly energy-efficient execution for compute-intensive kernels. Simultaneously, stream applications, which consist of many actors and channels connecting them, can provide natural representations for DSP applications, and therefore be a good match for CGRAs. We present our results of mapping DSP applications written in StreamIt language to CGRAs, along with our mapping flow. One important challenge in mapping is how to manage the multitude of kernels in the application for the limited local memory of a CGRA, for which we present a novel integer linear programming-based solution. Our evaluation results demonstrate that our software and hardware optimizations can help generate highly efficient mapping of stream applications to CGRAs, enabling far more energy-efficient executions (7× worse to 50× better) compared to using state-of-the-art GP-GPUs. Hongsik Lee, Dong Nguyen 0001, Jongeun Lee |
DAC | 3 |
| 2014 | Improving performance of loops on DIAM-based VLIW architecturesabstractRecent studies show that very long instruction word (VLIW) architectures, which inherently have wide datapath (e.g. 128 or 256 bits for one VLIW instruction word), can benefit from dynamic implied addressing mode (DIAM) and can achieve lower power consumption and smaller code size with a small performance overhead. Such overhead, which is claimed to be small, is mainly caused by the execution of additionally generated special instructions for conveying information that cannot be encoded in reduced instruction bit-width. In this paper, however, we show that the performance impact of applying DIAM on VLIW architecture cannot be overlooked expecially when applications possess high level of instruction level parallelism (ILP), which is mostly the case for loops because of the result of aggressive code scheduling. We also propose a way to relieve the performance degradation especially focusing on loops since loops spend almost 90% of total execution time in programs and tend to have high ILP. We first implement the original DIAM compilation technique in a compiler, and augment it with the proposed loop optimization scheme to show that ours can clearly alleviate the performance loss caused by the excessive number of additional instructions, with the help of slightly modified hardware. Moreover, the well-known loop unrolling scheme, which would produce denser code in loops at the cost of substantial code size bloating, is integrated into our compiler. The experiment result shows that the loop unrolling technique, combined with our augmented DIAM scheme, produces far better code in terms of performance with quite an acceptable amount of code increase. Jinyong Lee, Jongeun Lee, Yunheung Paek |
LCTES | 3 |
| 2014 | Design and optimization for embedded and real-time computing systems and applications
Jongeun Lee, Steve Goddard, Chin-Fu Kuo |
J. Syst. Archit. | 1 |
| 2014 | Configurable range memory for effective data reuse on programmable acceleratorsabstractWhile programmable accelerators such as application-specific processors and reconfigurable architectures can dramatically speed up compute-intensive kernels of an application, application performance can still be severely limited by the communication between processors. To minimize the communication overhead, a shared memory such as a scratchpad memory may be employed between the main processor and the accelerator coprocessor. However, this setup poses a significant challenge to the main processor, which now must manage data on the scratchpad explicitly, resulting in superfluous data copying due to the inflexibility of a scratchpad. In this article, we present an enhancement of a scratchpad, Configurable Range Memory (CRM), whose address range can be reprogrammed to minimize unnecessary data copying between processors and therefore promote data reuse on the accelerator, and also present a software management algorithm for the CRM. Our experimental results involving detailed simulation of full multimedia applications demonstrate that our CRM architecture can reduce the communication overhead quite effectively, reducing the kernel execution time by up to 28% and the application runtime by up to 12.8%, in addition to considerable system energy reduction, compared to the conventional architecture based on a scratchpad. Jongeun Lee, Seongseok Seo, Jong Kyung Paek, Kiyoung Choi |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2013 | Compiling control-intensive loops for CGRAs with state-based full predicationabstractPredication is an essential technique to accelerate kernels with control flow on CGRAs. While state-based full predication (SFP) can remove wasteful power consumption on issuing/decoding instructions from conventional full predication, generating code for SFP is challenging for general CGRAs, especially when there are multiple conditionals to be handled due to exploiting data level parallelism. In this paper, we present a novel compiler framework addressing central issues such as how to express the parallelism between multiple conditionals, and how to allocate resources to them to maximize the parallelism. In particular, by separating the handling of control flow and data flow, our framework can be integrated with conventional mapping algorithms for mapping data flow. Experimental results demonstrate that our framework can find and exploit parallelism between multiple conditionals, thereby leading to 2.21 times higher performance on average than a naive approach. Kyuseung Han, Kiyoung Choi, Jongeun Lee |
DATE | 3 |
| 2013 | Fast shared on-chip memory architecture for efficient hybrid computing with CGRAsabstractWhile Coarse-Grained Reconfigurable Architectures (CGRAs) are very efficient at handling regular, compute-intensive loops, their weakness at control-intensive processing and the need for frequent reconfiguration require another processor, for which usually a main processor is used. To minimize the overhead arising in such collaborative execution, we integrate a dedicated sequential processor (SP) with a reconfigurable array (RA), where the crucial problem is how to share the memory between SP and RA while keeping the SP's memory access latency very short. We present a detailed architecture, control, and program example of our approach, focusing on our optimized on-chip shared memory organization between SP and RA. Our preliminary results demonstrate that our optimized memory architecture is very effective in reducing kernel execution times (23.5% compared to a more straightforward alternative), and our approach can reduce the RA control overhead and other sequential code execution time in kernels significantly, resulting in up to 23.1% reduction in kernel execution time, compared to the conventional system using the main processor for sequential code execution. Jongeun Lee, Yeonghun Jeong, Sungsok Seo |
DATE | 1 |
| 2013 | Evaluator-executor transformation for efficient pipelining of loops with conditionalsabstractControl divergence poses many problems in parallelizing loops. While predicated execution is commonly used to convert control dependence into data dependence, it often incurs high overhead because it allocates resources equally for both branches of a conditional statement regardless of their execution frequencies. For those loops with unbalanced conditionals, we propose a software transformation that divides a loop into two or three smaller loops so that the condition is evaluated only in the first loop, while the less frequent branch is executed in the second loop in a way that is much more efficient than in the original loop. To reduce the overhead of extra data transfer caused by the loop fission, we also present a hardware extension for a class of Coarse-Grained Reconfigurable Architectures (CGRAs). Our experiments using MiBench and computer vision benchmarks on a CGRA demonstrate that our techniques can improve the performance of loops over predicated execution by up to 65% (37.5%, on average), when the hardware extension is enabled. Without any hardware modification, our software-only version can improve performance by up to 64% (33%, on average), while simultaneously reducing the energy consumption of the entire CGRA including configuration and data memory by 22%, on average. Yeonghun Jeong, Seongseok Seo, Jongeun Lee |
ACM Trans. Archit. Code Optim. | 3 |
| 2013 | Software-based register file vulnerability reduction for embedded processorsabstractRegister File (RF) is extremely vulnerable to soft errors, and traditional redundancy based schemes to protect the RF are prohibitive not only because RF is often in the timing critical path of the processor, but also since it is one of the hottest blocks on the chip. Software approaches would be ideal in this case, but previous approaches based on instruction scheduling are only moderately effective due to local scope. In this article we present a compiler approach, based on interprocedural program analysis, to reduce the vulnerability of registers by temporarily writing live variables to protected memory. We formulate the problem as an integer linear programming problem and also present a very efficient heuristic algorithm. Further we present an iterative optimization method based on Kernighan-Lin's graph partitioning algorithm. Our experiments demonstrate that our proposed techniques can reduce the vulnerability of a RF by 33 ∼ 37% on average and up to 66%, with a small 2% increase in runtime. In addition, our overhead reduction optimization can effectively reduce the code size overhead, by more than 40% on average, to a mere 5 ∼ 6%, compared to highly optimized binaries. Jongeun Lee, Aviral Shrivastava |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2013 | Architecture customization of on-chip reconfigurable acceleratorsabstractIntegrating coarse-grained reconfigurable architectures (CGRAs) into a System-on-a-Chip (SoC) presents many benefits as well as important challenges. One of the challenges is how to customize the architecture for the target applications efficiently and effectively without performing explicit design space exploration. In this article we present a novel methodology for incremental interconnect customization of CGRAs that can suggest a new interconnection architecture which is able to maximize the performance for a given set of application kernels while minimizing the hardware cost. In our methodology, we translate the problem of interconnect customization into that of inexact graph matching, and we devised a heuristic for A* search algorithm to efficiently solve the inexact graph matching problem. Our experimental results demonstrate that our customization method can quickly find application-optimized interconnections that exhibit 80% higher performance on average compared to the base architecture which has mesh interconnections, with little energy and hardware increase in interconnections and muxes. Jonghee W. Yoon, Jongeun Lee, Jinyong Lee, Yunheung Paek, Doosan Cho |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2012 | Software-managed automatic data sharing for Coarse-Grained Reconfigurable coprocessorsabstractCoarse-Grained Reconfigurable Architecture (CGRA) in a hybrid system can significantly accelerate the execution of compute-intensive kernels of applications. However, the data communication overhead between the main processor (MP) and the CGRA may be huge and can negate the speed-up of the CGRA. In this paper we address the problem of reducing the data communication overhead in a hybrid system by offering a partially automatic data sharing technique using a special shared memory called Configurable Range Memory (CRM). Unlike the previous work the CRM architecture we use here is based on comparators, which gives much higher flexibility in terms of where an array can be placed within a CRM while it makes the runtime software management of a CRM much more challenging. We present an efficient runtime algorithm based on first-fit heuristic. Our experimental results demonstrate that our CRM-based system can reduce the amount of data transfer between a MP and a CGRA up to 89.5% compared to ScratchPad Memory (SPM)-based systems, while the software management overhead is only 1.20~1.34% on average (depending on CRM architecture parameters) of the kernel cycles in the MP-only execution. Overall our CRM-based system can achieve average kernel speedup of 3.47 times over the MP-only execution, which is about 20% improvement over the SPM-based system. Toan X. Mai, Jongeun Lee |
FPT | 2 |
| 2012 | Improving performance of nested loops on reconfigurable array processorsabstractPipelining algorithms are typically concerned with improving only the steady-state performance, or the kernel time. The pipeline setup time happens only once and therefore can be negligible compared to the kernel time. However, for Coarse-Grained Reconfigurable Architectures (CGRAs) used as a coprocessor to a main processor, pipeline setup can take much longer due to the communication delay between the two processors, and can become significant if it is repeated in an outer loop of a loop nest. In this paper we evaluate the overhead of such non-kernel execution times when mapping nested loops for CGRAs, and propose a novel architecture-compiler cooperative scheme to reduce the overhead, while also minimizing the number of extra configurations required. Our experimental results using loops from multimedia and scientific domains demonstrate that our proposed techniques can greatly increase the performance of nested loops by up to 2.87 times compared to the conventional approach of accelerating only the innermost loops. Moreover, the mappings generated by our techniques require only a modest number of configurations that can fit in recent reconfigurable architectures. Jongeun Lee, Toan X. Mai, Yunheung Paek |
ACM Trans. Archit. Code Optim. | 2 |
| 2012 | PICA: Processor Idle Cycle Aggregation for Energy-Efficient Embedded SystemsabstractProcessor Idle Cycle Aggregation (PICA) is a promising approach for low-power execution of processors, in which small memory stalls are aggregated to create large ones, enabling profitable switch of the processor into low-power mode. We extend the previous approach in three dimensions. First we develop static analysis for the PICA technique and present optimal parameters for five common types of loops based on steady-state analysis. Second, to remedy the weakness of software-only control in varying environment, we enhance PICA with minimal hardware extension that ensures correct execution for any loops and parameters, thus greatly facilitating exploration-based parameter tuning. Third, we demonstrate that our PICA technique can be applied to certain types of nested loops with variable bounds, thus enhancing the applicability of PICA. We validate our analytical model against simulation-based optimization and also show, through our experiments on embedded application benchmarks, that our technique can be applied to a wide range of loops with average 20% energy reductions, compared to executions without PICA. Jongeun Lee, Aviral Shrivastava |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2012 | Return Data Interleaving for Multi-Channel Embedded CMPs SystemsabstractUsing multi-channel memory subsystems is an efficient way of satisfying high volume memory requests from CMPs. At the same time, the imbalance between memory bandwidth and bus performance opens up new possibility of optimization before they are sent to bus. This paper presents a new memory controller design for embedded CMPs systems when the return data from the return buffer is sent back to bus. Our scheduling policy, called return data interleaving (RDI) interleaves the return data of each request in a round robin manner. Further, for each request, it sends the critical word first. To evaluate our technique, we model an Intel XScale-based CMPs using M5 simulator for CMPs simulation and DRAMsim for memory subsystem simulation and examine the performance of MiBench and SPEC2000 benchmarks. Simulation results show that for memory-bound benchmarks running on the CMPs systems with the number of cores from 6 to 16, RDI can improve the execution time by average 11% and up to 16.9%. Fei Hong, Aviral Shrivastava, Jongeun Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | I2CRF: Incremental interconnect customization for embedded reconfigurable fabricsabstractIntegrating coarse-grained reconfigurable architectures (CGRAs) into a System-on-a-Chip (SoC) presents many benefits as well as important challenges. One of the challenges is how to customize the architecture for the target applications efficiently and effectively without explicit design space exploration. In this paper we present a novel methodology for incremental interconnect customization of CGRAs that can suggest a new interconnection architecture that can maximize the performance for a given set of application kernels while minimizing the hardware cost. Applying the inexact graph matching analogy, we translate our problem into graph matching taking into account the cost of various graph edit operations, which we solve using the A* search algorithm with a heuristic tailored to our problem. Our experimental results demonstrate that our customization method can quickly find application-optimized interconnections that exhibit 70% higher performance on average compared to the base architecture, with relatively little hardware increase in interconnections and muxes. Jonghee W. Yoon, Jongeun Lee, Jaewan Jung, Yunheung Paek, Doosan Cho |
DATE | 2 |
| 2011 | Fast graph-based instruction selection for multi-output instructionsabstractAbstract A multi‐output instruction (MOI) is an instruction that produces multiple outputs to its destination locations. Such inherently parallel instructions are becoming more and more popular in embedded processors, due to the advances in application‐specific architectures. In order to provide high‐level programmability and thus guarantee widespread acceptance, sophisticated compiler support for these programmable cores is necessary. However, traditionaltree‐basedapproaches for instruction selection, although very fast, fail to exploit MOIs mainly because of the fundamental limitation of the tree representation. In fact, to generate optimal code with MOIs requires a more generalgraph‐basedformulation of the instruction selection problem, which is at least NP‐complete. In this paper we present a new methodology to automatically generate from simple instruction set descriptions, graph‐based code selectors that can effectively utilize all provided instructions including MOIs. Our experimental results using a set of benchmarks on a target processor with various MOIs of up to two outputs demonstrate that our generated code selectors can quickly and effectively exploit many MOIs at the application level, and therefore are highly desirable both for architecture exploration and as code generators after architecture is fixed. Copyright © 2010 John Wiley & Sons, Ltd. Jonghee M. Youn, Yunheung Paek, Jongeun Lee, Hanno Scharwächter, Rainer Leupers |
Softw. Pract. Exp. | 4 |
| 2011 | High Throughput Data Mapping for Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-grained reconfigurable arrays (CGRAs) are a very promising platform, providing both up to 10–100 MOps/mW of power efficiency and software programmability. However, this promise of CGRAs critically hinges on the effectiveness of application mapping onto CGRA platforms. While previous solutions have greatly improved the computation speed, they have largely ignored the impact of the local memory architecture on the achievable power and performance. This paper motivates the need for memory-aware application mapping for CGRAs, and proposes an effective solution for application mapping that considers the effects of various memory architecture parameters including the number of banks, local memory size, and the communication bandwidth between the local memory and the external main memory. Further we propose efficient methods to handle dependent data on a double-buffering local memory, which is necessary for recurrent loops. Our proposed solution achieves 59% reduction in the energy-delay product, which factors into about 47% and 22% reduction in the energy consumption and runtime, respectively, as compared to memory-unaware mapping for realistic local memory architectures. We also show that our scheme scales across a range of applications and memory parameters, and the runtime overhead of handling recurrent loops by our proposed methods can be less than 1%. Jongeun Lee, Aviral Shrivastava, Jonghee W. Yoon, Doosan Cho, Yunheung Paek |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2011 | Static Analysis of Register File VulnerabilityabstractWith continuous technology scaling, soft errors are becoming an increasingly important design concern even for earth-bound applications. While compiler approaches have the potential to mitigate the effect of soft errors with minimal runtime overheads, static vulnerability estimation-an essential part of compiler approaches-is lacking due to its inherent complexity. This paper presents a static analysis approach for register file (RF) vulnerability estimation. We decompose the vulnerability of a register into intrinsic and conditional basic-block vulnerabilities. This decomposition allows us to develop a fast, yet reasonably accurate RF vulnerability estimation mechanism. We validate and compare a linear equation based method and an iterative method. Also we demonstrate a practical application of RF vulnerability estimation to compiler optimizations. Our experimental results on benchmarks from MiBench suite indicate that not only our static RF vulnerability estimation is fast and accurate, but also compiler optimizations enabled by our static estimation can achieve very cost-effective protection of register files against soft errors. Jongeun Lee, Aviral Shrivastava |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2011 | Memory access optimization in compilation for coarse-grained reconfigurable architecturesabstractCoarse-grained reconfigurable architectures (CGRAs) promise high performance at high power efficiency. They fulfil this promise by keeping the hardware extremely simple, and moving the complexity to application mapping. One major challenge comes in the form of data mapping. For reasons of power-efficiency and complexity, CGRAs use multibank local memory, and a row of PEs share memory access. In order for each row of the PEs to access any memory bank, there is a hardware arbiter between the memory requests generated by the PEs and the banks of the local memory. However, a fundamental restriction remains in that a bank cannot be accessed by two different PEs at the same time. We propose to meet this challenge by mapping application operations onto PEs and data into memory banks in a way that avoids such conflicts. To further improve performance on multibank memories, we propose a compiler optimization for CGRA mapping to reduce the number of memory operations by exploiting data reuse. Our experimental results on kernels from multimedia benchmarks demonstrate that our local memory-aware compilation approach can generate mappings that are up to 53% better in performance (26% on average) compared to a memory-unaware scheduler. Jongeun Lee, Aviral Shrivastava, Yunheung Paek |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2010 | Memory-Aware Application Mapping on Coarse-Grained Reconfigurable Arrays
Jongeun Lee, Aviral Shrivastava, Jonghee W. Yoon, Yunheung Paek |
HiPEAC | 2 |
| 2010 | Operation and data mapping for CGRAs with multi-bank memoryabstractCoarse Grain Reconfigurable Architectures (CGRAs) promise high performance at high power efficiency. They fulfil this promise by keeping the hardware extremely simple, and moving the complexity to application mapping. One major challenge comes in the form of data mapping. For reasons of power-efficiency and complexity, CGRAs use multi-bank local memory, and a row of PEs share memory access. In order for each row of the PEs to access any memory bank, there is a hardware arbiter between the memory requests generated by the PEs and the banks of the local memory. However, a fundamental restriction remains that a bank cannot be accessed by two different PEs at the same time. We propose to meet this challenge by mapping application operations onto PEs and data into memory banks in a way that avoids such conflicts. Our experimental results on kernels from multimedia benchmarks demonstrate that our local memory-aware compilation approach can generate mappings that are up to 40% better in performance (17.3% on average) compared to a memory-unaware scheduler. Jongeun Lee, Aviral Shrivastava, Yunheung Paek |
LCTES | 2 |
| 2010 | Cache vulnerability equations for protecting data in embedded processor caches from soft errorsabstractContinuous technology scaling has brought us to a point, where transistors have become extremely susceptible to cosmic radiation strikes, or soft errors. Inside the processor, caches are most vulnerable to soft errors, and techniques at various levels of design abstraction, e.g., fabrication, gate design, circuit design, and microarchitecture-level, have been developed to protect data in caches. However, no work has been done to investigate the effect of code transformations on the vulnerability of data in caches. Data is vulnerable to soft errors in the cache only if it will be read by the processor, and not if it will be overwritten. Since code transformations can change the read-write pattern of program variables, they significantly effect the soft error vulnerability of program variables in the cache. We observe that often opportunity exists to significantly reduce the soft error vulnerability of cache data by trading-off a little performance. However, even if one wanted to exploit this trade-off, it is difficult, since there are no efficient techniques to estimate vulnerability of data in caches. To this end, this paper develops efficient static analysis method to estimate program vulnerability in caches, which enables the compiler to exploit the performance-vulnerability trade-offs in applications. Finally, as compared to simulation based estimation, static analysis techniques provide the insights into vulnerability calculations that provide some simple schemes to reduce program vulnerability. Aviral Shrivastava, Jongeun Lee, Reiley Jeyapaul |
LCTES | 2 |
| 2010 | A Compiler-Microarchitecture Hybrid Approach to Soft Error Reduction for Register FilesabstractFor embedded systems, where neither energy nor reliability can be easily sacrificed, this paper presents an energy efficient soft error protection scheme for register files (RFs). Unlike previous approaches, the proposed method explicitly optimizes for energy efficiency and can exploit the fundamental tradeoff between reliability and energy. While even simple compiler-managed RF protection scheme can be more energy efficient than hardware schemes, this paper formulates and solves further compiler optimization problems to significantly enhance the energy efficiency of RF protection schemes by an additional 30% on average, as demonstrated in our experiments on a number of embedded application benchmarks. Jongeun Lee, Aviral Shrivastava |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2009 | A software solution for dynamic stack management on scratch pad memoryabstractIn an effort to make processors more power efficient scratch pad memory (SPM) have been proposed instead of caches, which can consume majority of processor power. However, application mapping on SPMs remain a challenge. We propose a dynamic SPM management scheme for program stack data for processor power reduction. As opposed to previous efforts, our solution does not mandate any hardware changes, does not need profile information, and SPM size at compile-time, and seamlessly integrates support for recursive functions. Our technique manages stack frames on SPM using a scratch pad memory manager (SPMM), integrated into the application binary by the compiler. Our experiments on benchmarks from MiBench show average energy savings of 37% along with a performance improvement of 18%. Arun Kannan, Aviral Shrivastava, Amit Pabalkar, Jongeun Lee |
ASP-DAC | 4 |
| 2009 | Compiler-managed register file protection for energy-efficient soft error reductionabstractFor embedded systems where neither energy nor reliability can be easily sacrificed, we present an energy efficient soft error protection scheme for register files (RF). Unlike previous approaches, our method explicitly optimizes for energy efficiency and exploits the fundamental tradeoff between reliability and energy. While even simple compiler-managed RF protection scheme is more energy efficient than hardware schemes, this work formulates and solves further compiler optimization problems to significantly enhance the energy efficiency of RF protection schemes by an additional 24%. Jongeun Lee, Aviral Shrivastava |
ASP-DAC | 1 |
| 2009 | Static analysis to mitigate soft errors in register filesabstractWith continuous technology scaling, soft errors are becoming an increasingly important design concern even for earth-bound applications. While compiler approaches have the potential to mitigate the effect of soft errors with minimal runtime overheads, static vulnerability estimation-an essential part of compiler approaches-is lacking due to its inherent complexity. This paper presents a static analysis approach for Register File (RF) vulnerability estimation. We decompose the vulnerability of a register into intrinsic and conditional basic-block vulnerabilities. This decomposition allows us to develop a fast, yet reasonably accurate, linear equation-based RF vulnerability estimation mechanism. We demonstrate its practical application to compiler optimizations. Our experimental results on benchmarks from MiBench suite indicate that not only our static RF vulnerability estimation is fast and accurate, but also compiler optimizations enabled by our static estimation can achieve very cost-effective protection of register files against soft errors. Jongeun Lee, Aviral Shrivastava |
DATE | 1 |
| 2009 | FSAF: File system aware flash translation layer for NAND Flash MemoriesabstractNAND Flash Memories require Garbage Collection (GC) and Wear Leveling (WL) operations to be carried out by Flash Translation Layers (FTLs) that oversee flash management. Owing to expensive erasures and data copying, these two operations essentially determine application response times. Since file systems do not share any file deletion information with FTL, dead data is treated as valid by FTL, resulting in significant WL and GC overheads. In this work, we propose a novel method to dynamically interpret and treat dead data at the FTL level so as to reduce above overheads and improve application response times, without necessitating any changes to existing file systems. We demonstrate that our resource-efficient approach can improve application response times and memory write access times by 22% and reduce erasures by 21.6% on average. Sai Krishna Mylavarapu, Siddharth Choudhuri, Aviral Shrivastava, Jongeun Lee, Tony Givargis |
DATE | 4 |
| 2009 | A compiler optimization to reduce soft errors in register filesabstractRegister file (RF) is extremely vulnerable to soft errors, and traditional redundancy based schemes to protect the RF are prohibitive not only because RF is often in the timing critical path of the processor, but also since it is one of the hottest blocks on the chip, and therefore adding any extra circuitry to it is not desirable. Pure software approaches would be ideal in this case, but previous approaches that are based on program duplication have very significant runtime overheads, and others based on instruction scheduling are only moderately effective due to local scope. We show that the problem of protecting registers inherently requires inter-procedural analysis, and intra-procedural optimization are ineffective. This paper presents a pure compiler approach, based on inter-procedural code analysis to reduce the vulnerability of registers by temporarily writing live variables to protected memory. We formulate the problem as an integer linear programming problem and also present a very efficient heuristic algorithm. Our experiments demonstrate that our proposed technique can reduce the vulnerability of the RF by 33 similar to 37% on average and up to 66%, with a small 2% increase in runtime. In addition, our overhead reduction optimizations can effectively reduce the code size overhead, by more than 40% on average, to a mere 5 similar to 6%, as compared to highly optimized binaries Jongeun Lee, Aviral Shrivastava |
LCTES | 1 |
| 2009 | A Software-Only Solution to Use Scratch Pads for Stack DataabstractA dynamic scratch pad memory (SPM) management scheme for program stack data with the objective of processor power reduction is presented. Basic technique does not need the SPM size at compile time, does not mandate any hardware changes, does not need profile information, and seamlessly integrates support for recursive functions. Stack frames are managed using a software SPM manager, integrated into the application binary, and shows average energy savings of 32% along with a performance improvement of 13%, on benchmarks from MiBench. SPM management can be further optimized and made pointer safe, by knowing the SPM size. Aviral Shrivastava, Arun Kannan, Jongeun Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2008 | SDRM: Simultaneous Determination of Regions and Function-to-Region Mapping for Scratchpad Memories
Amit Pabalkar, Aviral Shrivastava, Arun Kannan, Jongeun Lee |
HiPC | 4 |
| 2007 | Instruction set synthesis with efficient instruction encoding for configurable processorsabstractApplication-specific instructions can significantly improve the performance, energy-efficiency, and code size of configurable processors. While generating new instructions from application-specific operation patterns has been a common way to improve the instruction set (IS) of a configurable processor, automating the design of ISs for given applications poses new challenges---how to create as well as utilize new instructions in a systematic manner, and how to choose the best set of application-specific instructions considering the various effects the new instructions may have on the data path and the compilation? To address these problems, we present a novel IS synthesis framework that optimizes the IS through an efficient instruction encoding for the given application as well as for the given data path architecture. We first build a library of new instructions created with various encoding alternatives taking into account the data path architecture constraints, and then select the best set of instructions while satisfying the instruction bitwidth constraint. We formulate the problem using integer linear programming and also present an effective heuristic algorithm. Experimental results using our technique generate ISs that show improvements of up to about 40% over the native IS for several application benchmarks running on typical embedded RISC processors. Jongeun Lee, Kiyoung Choi, Nikil Dutt |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2004 | Analysis on risk factors for cervical cancer using induction technique
Seung Hee Ho, Sun Ha Jee, Jongeun Lee, Jong Sup Park |
Expert Syst. Appl. | 3 |
| 2003 | Evaluating Memory Architectures for Media Applications on Coarse-Grained Recon.gurable Architectures
Jongeun Lee, Kiyoung Choi, Nikil Dutt |
ASAP | 1 |
| 2003 | Energy-efficient instruction set synthesis for application-specific processorsabstractSeveral techniques have been proposed to enhance the energy-efficiency of ASIPs (Application-Specific Instruction set Processors). While those techniques can reduce the energy consumption with a minimal change in the instruction set (IS), they fail to exploit the opportunity of designing the entire IS from the energy-efficiency perspective. In this paper, we present an energy-efficient IS synthesis technique that can comprehensively reduce the energy-delay product (EDP) of ASIPs through optimal instruction encoding, considering both the instruction bitwidth and the dynamic instruction count. Experimental results with a typical embedded RISC processor show that our technique can generate application-specific IS's that are up to 40% more energy-efficient over the native IS for several application benchmarks. Jongeun Lee, Kiyoung Choi, Nikil Dutt |
ISLPED | 1 |
| 2003 | An algorithm for mapping loops onto coarse-grained reconfigurable architecturesabstractWith the increasing demand for flexible yet highly efficient architecture platforms for media applications, there is a growing interest in the Coarse-grained Reconfigurable Architectures (CRAs). While many CRAs have demonstrated impressive performance improvement, the lack of compilation technology for such architectures causes a bottleneck in the current design process. In this paper, we present a novel mapping algorithm designed to support Reconfigurable ALU Array (RAA) architectures, that represent a significant class of CRAs. More specifically we present a core mapping algorithm that addresses the problem of placing and routing the operations of a loop body onto the ALU array, to be executed in a loop pipelined fashion. Experimental results using our mapping algorithm on a typical RAA show that our algorithm not only has very fast compilation time but can also generate quality mappings exhibiting high memory bandwidth utilization and low global interconnection requirements. Comparison with manual mapping also indicates that our algorithm can generate near-optimal mappings for several loops. Jongeun Lee, Kiyoung Choi, Nikil Dutt |
LCTES | 1 |
| 2002 | Efficient instruction encoding for automatic instruction set design of configurable ASIPsabstractApplication-specific instructions can significantly improve the performance, energy, and code size of configurable processors. A common approach used in the design of such instructions is to convert application-specific operation patterns into new complex instructions. However, processors with a fixed instruction bitwidth cannot accommodate all the potentially interesting operation patterns, due to the limited code space afforded by the fixed instruction bitwidth. We present a novel instruction set synthesis technique that employs an efficient instruction encoding method to achieve maximal performance improvement. We build a library of complex instructions with various encoding alternatives and select the best set of complex instructions while satisfying the instruction bitwidth constraint. We formulate the problem using integer linear programming and also present an effective heuristic algorithm. Experimental results using our technique generate instruction sets that show improvements of up to 38% over the native instruction set for several realistic benchmark applications running on a typical embedded RISC processor. Jongeun Lee, Kiyoung Choi, Nikil Dutt |
ICCAD | 1 |
| 2000 | Fast Hardware-Software Coverification by Optimistic Execution of Real ProcessorabstractTo achieve fast verification of the software part of an embedded system, we propose to run the target processor optimistically, which effectively reduces the synchronization overhead with other simulators. For the optimistic processor execution, we present a processor execution platform and state saving/restoration methods. We performed optimistic execution of ARM710A processor in the coverification of an IS-95 CDMA cellular phone system and obtained up to orders of magnitude higher performance compared with the case that the processor runs conservatively. Sungjoo Yoo, Jongeun Lee, Jinyong Jung, Kyoungseok Rha, Youngchul Cho, Kiyoung Choi |
DATE | 2 |