EDBT 2026 Demo / reviewers in the wild / expert
Jae-Joon Kim
dblp:79/6708
· DBLP profile ↗
74ranked-venue papers
1as first author
38since 2021 · last 2026
0000-0001-5175-8258ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 58 · 1 first-author · 28 since 2021Artificial intelligence and machine learning · 14 · 9 since 2021Software engineering, systems software and programming languages · 12 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 3D Integration of Hybrid IGZO/Si and IGZO eDRAMs for High-Density/High-Performance On-Chip MemoryabstractThe growing need for advanced memory architectures leveraging 3D integration has become increasingly critical in modern computing systems. In particular, memory architectures that match the performance of static random access memory (SRAM) while significantly increasing density are highly impactful. In this paper, we propose a 3D integration-based hybrid InGaZnO(IGZO)/Si embedded dynamic random access memory architecture (Hybrid-3D) and circuit design, which markedly increases on-chip memory density and enhances system performance. The superiority of Hybrid-3D is demonstrated through rigorous validation involving process integration verification, transistor-level modeling, and circuit-level memory design. Detailed evaluations of the vertically stacked memory operation confirm stable operations, enabling a 22× increase in on-chip memory density compared to SRAM. Integrating Hybrid-3D on-chip memory into neural processing unit (NPU) architectures results in substantial improvements in energy efficiency and processing speed. System-level evaluations across vision and natural language processing (NLP) tasks reveal a maximum energy efficiency improvement of 3.2× and a throughput increase of 2.6×. Munhyeon Kim, Sukhyun Choi, Yulhwa Kim, Jae-Joon Kim |
DATE | 4 |
| 2026 | MECA-CiM: A Shared-MicroExponent-aware Configurable Analog Compute-in-Memory Macro for Efficient Inference
Wonkyung Han, Juheun Lee, Sukhyun Choi, Wonjun Han, Jae-Joon Kim |
ISLPED | 7 |
| 2025 | L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language ModelsabstractDue to the high memory and computational costs associated with large language models (LLMs), model compression techniques such as quantization, which reduces inference costs, and parameter-efficient fine-tuning (PEFT) methods like Low-Rank Adaptation (LoRA), which reduce training costs, have gained significant popularity.This trend has spurred active research into quantization-aware PEFT techniques, aimed at maintaining model accuracy while minimizing memory overhead during both inference and training.Previous quantization-aware PEFT methods typically apply post-training quantization (PTQ) to pre-trained LLMs, followed by PEFT to recover accuracy loss.Meanwhile, this approach has limitations in recovering the accuracy loss.In this paper, we propose L4Q, a method that integrates Quantization-Aware Training (QAT) with LoRA.By employing a memory-optimized layer design, L4Q significantly reduces QAT's memory overhead, making its training cost comparable to LoRA, while preserving the advantage of QAT in producing fully quantized LLMs with high accuracy.Our experiments demonstrate that this combined approach to quantization and fine-tuning achieves superior accuracy compared to decoupled finetuning schemes, particularly in 4-bit and 3-bit quantization, positioning L4Q as an efficient QAT solution.Using the LLaMA and Mistral models with instructional datasets, we showcase L4Q's capabilities in language tasks and few-shot learning. Hyesung Jeon, Yulhwa Kim, Jae-Joon Kim |
ACL (1) | 3 |
| 2025 | Near-Memory LLM Inference Processor based on 3D DRAM-to-logic Hybrid BondingabstractLarge language model (LLM) inference poses dual challenges, demanding substantial memory bandwidth and computing resources. Recent advancements in near-memory accelerators leveraging 3D DRAM-to-logic hybrid-bonding (HB) interconnects have gained attention due to their highly parallel data transfer capabilities. We address limitations in previous HB-DRAM accelerators, such as those stemming from distributed controller designs, by introducing an architecture with a centralized controller and dual-IO scheme. This approach not only reduces the chip area overhead but also enables reconfigurable GEMV/GEMM operations, boosting the performance. Simulations for the OPT 66B model show that our proposed accelerator achieves 2.9X, 3.5X, and 2.5X higher performance compared to NPU, DRAM-PIM, and heterogeneous designs (DRAM-PIM + NPU), respectively. Sanghyeok Han, Byungkuk Yoon, Gyeonghwan Park, Choungki Song, Dongkyun Kim, Jae-Joon Kim |
DAC | 6 |
| 2025 | SplitSync: Bank Group-Level Split-Synchronization for High-Performance DRAM PIMabstractProcessing in Memory (PIM) architectures enhance memory bandwidth by utilizing bank-level parallelism, typically implemented with a SIMD structure where all banks operate simultaneously under a single command. However, this synchronous approach requires the activation of all banks before computation, leading to activation times that exceed computation times, limiting performance gain. Recently, asynchronous execution PIM has been proposed as an alternative, allowing banks to operate asynchronously and overlap activation with processing to hide the row activation overhead. While effective at reducing row activation overhead, the independent operation requires large shared accumulators for each bank group, increasing area overhead. To address the issues, we propose bank group (BG)-level split synchronization DRAM PIM, where each bank group operates asynchronously to hide row activation overhead while operating synchronously within the bank group to eliminate the need for shared accumulators. Evaluation results show that our proposed design achieves an average throughput improvement of $1.70 x$ and $1.06 x$ compared to conventional PIM and asynchronous execution PIM. Furthermore, the area overhead per processing unit (PU) increases by only $1.5 \%$ compared to conventional PIM and is significantly lower than that of asynchronous execution PIM. Byungkuk Yoon, Sanghyeok Han, Gyeonghwan Park, Jae-Joon Kim |
DAC | 4 |
| 2025 | Compute-in-Memory Array Design Using Stacked Hybrid IGZO/Si eDRAM cellsabstractTo effectively accelerate neural networks in compute-in-memory (CIM) based systems, higher memory cell density is essential to handle the increasing computational workload and number of parameters. While CMOS embedded dynamic random access memory (eDRAM) is being explored as an alternative, improving the short retention time$(t_{ret})$($t_{ret}$(> 100 s), but additional improvements are needed due to its substantial cell variability and slower operating speed compared to CMOS-based cells. This paper proposes a cell and array design for CIM using 3T-based stacked hybrid IGZO/Si eDRAM (Hybrid-3T) and performs a system-level deep neural network (DNN) evaluation. The Hybrid-3T cell, designed based on 7-nm FinFET technology, achieves$t_{ret}$that is 100 s longer compared to IGZO-based 3T eDRAM (IGZO-3T). The proposed Hybrid-3T offers a 3.4 × higher bit cell density compared to 8T SRAM bit cells and a 2 × higher density compared to CMOS-based 3T eDRAM (CMOS-3T), while demonstrating similar throughput and variability levels to CMOS eDRAM and SRAM-based systems. Furthermore, we evaluate DNN inference accuracy for vision and natural language processing (NLP) tasks using the proposed CIM design, examining the impact of improved cell variability and retention time on system-level characteristics. The retention time for ensuring CIM operation accuracy$(t_{ret,CIM})$is 108times longer in Hybrid-3T than CMOS-3T, and the$t_{ret,CIM}$considering variability$(t_{ret,CIM_{v}})$is more than 3 × longer than IGZO-3T eDRAM. As a result, the proposed Hybrid-3T eDRAM CIM leverages the advantages of both CMOS-3T and IGZO-3T CIM designs, enabling the development of high-performance, reliable systems. Munhyeon Kim, Yulhwa Kim, Jae-Joon Kim |
DATE | 3 |
| 2025 | Integer Unit-Based Outlier-Aware LLM Accelerator Preserving Numerical Accuracy of FP-FP GEMMabstractThe proliferation of large language models (LLMs) has significantly heightened the importance of quantization to alleviate the computational burden given the surge in the number of parameters. However, quantization often targets a subset of a LLM and relies on the floating-point (FP) arithmetic for matrix multiplication of specific subsets, leading to performance and energy overhead. Additionally, to compensate for the quality degradation incurred by quantization, retraining methods are frequently employed, demanding significant efforts and resources. This paper proposes OwL-P, an outlier-aware LLM inference accelerator which preserves the numerical accuracy of FP arithmetic while enhancing hardware efficiency with an integer (INT)-based arithmetic unit for general matrix multiplication (GEMM), through the use of a shared exponent and efficient management of outlier data. It also mitigates off-chip bandwidth requirements by employing a compressed number format. The proposed number format leverages outliers and shared exponents to facilitate the compression of both model weights and activations. We evaluate this work across 10 different transformer-based benchmarks, and the results demonstrate that the proposed integer-based LLM accelerator achieves an average 2.70x performance gain and 3.57 x energy savings while maintaining the numerical accuracy of the FP arithmetic. Jehun Lee, Jae-Joon Kim |
DATE | 2 |
| 2025 | An eDRAM Digital In-Memory Neural Network Accelerator for High-Throughput and Extended Data Retention TimeabstractComputing-in-Memory (CIM) optimizes multiply-and-accumulate (MAC) operations for energy-efficient acceleration of neural network models. While SRAM has been a popular choice for CIM designs due to its compatibility with logic processes, its large cell size restricts storage capacity for neural network parameters. Consequently, gain-cell eDRAM, featuring memory cells with only 2–4 transistors, has emerged as an alternative for CIM cells. While digital CIM (DCIM) structure has been actively adopted in SRAM-based CIMs for better accuracy and scalability than analog CIMs (ACIM), previous eDRAM-based CIMs still employed ACIM structure since the eDRAM CIM cells were not able to perform a complete digital logic operation. In this paper, we propose an eDRAM bit cell for more efficient DCIM operations using only 4 transistors. The proposed eDRAM DCIM structure also maintains consistent and accurate output values over time, improving retention times compared to previous eDRAM ACIM designs. We validate our approach by fabricating an eDRAM DCIM macro chip and conducting hardware validation experiments, measuring retention time and neural network accuracy. Experimental results show that the proposed eDRAM DCIM achieves 3× longer retention time than state-of-the-art eDRAM ACIM designs, along with higher throughput without accuracy loss. Inhwan Lee, Jehun Lee, Jaeyong Jang, Jae-Joon Kim |
DATE | 4 |
| 2025 | COMPASS: A Compiler Framework for Resource-Constrained Crossbar-Array Based In-Memory Deep Learning AcceleratorsabstractRecently, crossbar array based in-memory accelerators have been gaining interest due to their high throughput and energy efficiency. While software and compiler support for the in-memory accelerators has also been introduced, they are currently limited to the case where all weights are assumed to be on-chip. This limitation becomes apparent with the significantly increasing network sizes compared to the in-memory footprint. Weight replacement schemes are essential to address this issue. We propose COMPASS, a compiler framework for resource-constrained crossbar-based processing-in-memory (PIM) deep neural network (DNN) accelerators. COMPASS is specially targeted for networks that exceed the capacity of PIM crossbar arrays, necessitating access to external memories. We propose an algorithm to determine the optimal partitioning that divides the layers so that each partition can be accelerated on chip. Our scheme takes into account the data dependence between layers, core utilization, and the number of write instructions to minimize latency, memory accesses, and improve energy efficiency. Simulation results demonstrate that COMPASS can accommodate much more networks using a minimal memory footprint, while improving throughput by 1.78X and providing 1.28X savings in energy-delay product (EDP) over baseline partitioning methods. Jeongin Choe, Jae-Joon Kim |
DATE | 4 |
| 2025 | DOTS: DRAM-PIM Optimization for Tall and Skinny GEMM Operations in LLM InferenceabstractFor large language models (LLMs), increasing token lengths require smaller batch sizes due to increase in memory requirement for KV caching, leading to under-utilization of processing units and memory bandwidth bottleneck in NPUs. To address the challenge, we propose DOTS, a new DRAM-PIM architecture that can handle both GEMV and GEMM efficiently, even outperforming NPUs in GEMM operations when batch sizes are small. The proposed DRAM-PIM reduces power consumption and latency caused by frequent DRAM row activation switching in conventional DRAM PIMs with negligible hardware overhead. Simulation results show that our proposed design achieves throughput improvements of 1.83x, 1.92x, and 1.7x over GPU, NPU, and heterogeneous NPU/PIM systems, respectively, for models as large as or larger than OPT-175B. Gyeonghwan Park, Sanghyeok Han, Byungkuk Yoon, Jae-Joon Kim |
DATE | 4 |
| 2025 | LLM-on-the-Palm: Mobile LLM Inference with PIM-Enhanced NAND Flash MemoryabstractLarge Language Model (LLM) inference on edge devices is crucial for democratizing AI and addressing privacy and security concerns associated with cloud services. However, the large parameter sizes of LLMs pose significant challenges in deployment on resource-constrained devices, affecting user experience. Mobile DRAM typically cannot accommodate these parameters, necessitating the use of NAND flash memory, but it has considerably slower access speeds. To overcome the limitations, we propose LLM-on-the-Palm, an LLM inference system for smartphones that leverages Processing-in-Memory (PIM)-enhanced NAND flash memory to enable fast inference of billion-parameter LLMs with minimal hardware modifications. Our approach requires only a single multiply-accumulate (MAC) unit per plane in a NAND flash memory - 256 MAC units for 256 GB SSD - without compromising cell density. Simulation results show that our design achieves ∼100 ms per token generation on a 6.7B model under specifications similar to an iPhone-15. Sanghyeok Han, Jae-Joon Kim |
ICCAD | 3 |
| 2025 | Energy-Efficient Accelerator for Scalable Point Transformer Networks with Reduced Data AccessabstractTransformer-based architectures have recently emerged as the backbone for neural networks handling 3D point cloud applications, achieving state-of-the-art accuracy by avoiding the computational overhead of traditional k-nearest neighbor (k-NN) algorithms widely used in previous models. Despite their success, these transformer-based models often rely on submanifold convolution and attention layers, which introduce performance bottlenecks due to substantial weight parameters and high computational demands. In this paper, we propose two key hardware-friendly optimizations to address these challenges. First, we introduce a panel-level kernel mapping and weight-fetching strategy for weight-stationary computation, which minimizes data movement during submanifold convolution. Second, we propose a locality-based compression technique for attention queries and keys, allowing feature reuse from neighboring points to reduce the number of operations and memory requirements. Based on the optimizations, we develop a dedicated hardware tailored for 3D transformer-based 3D point clouds. Experimental results show that our approach significantly reduces memory accesses and enables efficient attention score computation with negligible accuracy loss. Hyunsung Yoon, Jehun Lee, Jae-Joon Kim |
ICCAD | 3 |
| 2025 | Partial-Sum Quantization Based on Pseudo-Quantization Noise for Variation-Tolerant Analog In-Memory ComputingabstractAnalog Computing-In-Memory (ACiM) accelerators with multi-level cells (MLCs) offer high density and area benefits for DNNs. To enhance efficiency, ADC resolution needs to be minimized, but this introduces significant quantization errors, lowering accuracy. Additionally, device noise and ADC integral nonlinearity (INL) noise further degrade accuracy. To address these challenges, we propose a training method that reduces ADC resolution while compensating for noise generated in ACiM arrays. By incorporating pseudo-quantization noise into Partial-Sum Training (PST), our approach not only stabilizes PST but also trains the model to become robust to ACiM-specific noise effects. Experimental results based on an industry ReRAM technology show that our PST scheme demonstrates robust noise tolerance across various ACiM configurations and maintains accuracy degradation within 1% even in the presence of cell conductance variability and ADC INL noise, while enabling low-resolution ADCs that reduces area and energy consumption by up to 16× and 31×, respectively. Nameun Kang, Eunhyeok Park, Sangsu Park, Jongil Kim, Jaeyun Yi, Jae-Joon Kim |
ISLPED | 6 |
| 2025 | CrossBit: Bitwise Computing in NAND Flash Memory with Inter-Bitline Data CommunicationabstractIn-flash processing (IFP), which involves performing data computation inside NAND flash memory, holds high potential for improving the performance and energy efficiency of data-intensive application by minimizing data movement.Recent research has introduced several IFP architectures enabling bulk bitwise operations inside NAND flash chips to demonstrate this potential.However, previous IFP designs were limited to performing bitwise operations on data sensed within the same bitline, thus lacking the capability to handle more complex functions requiring interactions between data from different bitlines.This paper presents CrossBit, a new IFP architecture that enables both intra-bitline and inter-bitline operations with minimal additional circuitry integrated into commodity NAND flash memory.With the capability for inter-bitline operations, CrossBit facilitates in-flash error correction code (IF-ECC) operations, thereby enabling reliable multi-level cell (MLC) IFP.Moreover, CrossBit efficiently processes fundamental database queries that were previously inefficient with existing work that only supports intra-bitline operations.Experimental results show that the implementation of IF-ECC in CrossBit results in a substantial reduction in bit error rate (BER) for MLC operations, leading to 1.8× increase in bit-density by using MLC compared to previous IFP designs which uses SLC only.When used for accelerating fundamental database queries, CrossBit achieves notable average speedup and energy efficiency improvement of 2.1× and 2.5× compared to the state-of-the-art (SOTA) IFP architecture.We further demonstrate the practicality by processing the full set of end-toend database queries from the widely used Star schema benchmark, where CrossBit achieves 1.7× speedup. Seunghwan Song, Sukhyun Choi, Jeongin Choe, Sanghyeok Han, Jisung Park 0001, Jinho Lee 0001, Jae-Joon Kim |
MICRO | 8 |
| 2025 | Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM ReasoningabstractRecent reasoning-focused language models achieve high accuracy by generating lengthy intermediate reasoning paths before producing final answers.
While this approach is effective in solving problems that require logical thinking, long reasoning paths significantly increase memory usage and reduce throughput of token generation, limiting the practical deployment of such models.
We propose Reasoning Path Compression (RPC), a training-free method that accelerates inference by leveraging the semantic sparsity of reasoning paths.
RPC periodically compresses the KV cache by retaining cache entries that receive high importance score, which are computed using a selector window composed of recently generated queries.
Experiments show that RPC improves generation throughput of QwQ-32B by up to 1.60$\times$ compared to the inference with full KV cache, with an accuracy drop of 1.2\% on the AIME 2024 benchmark.
Our findings demonstrate that semantic sparsity in reasoning traces can be effectively exploited for compression, offering a practical path toward efficient deployment of reasoning LLMs. Our code is available at https://github.com/jiwonsong-dev/ReasoningPathCompression. Jiwon Song, Dongwon Jo, Yulhwa Kim, Jae-Joon Kim |
NeurIPS | 4 |
| 2025 | NeuroTorch: DRAM-free nonvolatile memory-based hybrid training compute-in-memory system simulator
JongHyun Ko, Inseok Lee, Joon Hwang, Jiseong Im, Ryun-Han Koo, Jae-Joon Kim, Jong-Ho Lee 0002 |
Neurocomputing | 10 |
| 2024 | 4-Transistor Ternary Content Addressable Memory Cell Design using Stacked Hybrid IGZO/Si TransistorsabstractIn this paper, we propose a 4T-based paired orthogonally stacked transistors for random access memory (POST-RAM) cell structure and also suggest ternary content addressable memory (TCAM) applications. POST-RAM cells feature vertically stacked read and write transistors, maximizing area efficiency by utilizing only two transistors' space. POST-RAM employs InGaZnO (IGZO) channels for write transistors and single crystal silicon channels for read transistors, which results in both extremely long memory retention and fast reading performance. A comprehensive 3D-TCAD simulation is conducted to validate the procedural design of the proposed device structure. Furthermore, we introduced a self-clamped searching scheme (SC2S) designed to enhance the efficiency of TCAM operations. The results conclusively demonstrate that operating a TCAM based on the proposed POST-RAM architecture can lead to a 20% improvement in energy-delay product (EDP). Notably, the delay performance can be enhanced by up to 40% when compared to a 16T SRAM-based TCAM. Additionally, the proposed scheme enables a more than sixfold reduction in cell area, demonstrating an efficient use of space. Munhyeon Kim, Jae-Joon Kim |
DAC | 2 |
| 2024 | Fused Sampling and Grouping with Search Space Reduction for Efficient Point Cloud AccelerationabstractRecently, point-based deep neural networks (DNN) have demonstrated remarkable ability in analyzing point cloud data. However, challenges arise in sampling and grouping layers, particularly in terms of time and energy consumption due to the iterative access and computation of point cloud data for local feature extraction. In this paper, we introduce a Morton code-based data structure which stores point data with the shared upper bits together, enabling sequential access to the points within a specific voxel. We also propose a fused sampling and grouping approach with a reduced search space, which reuses the point data and the calculated distances for the farthest voxel and its neighbors. Additionally, a dedicated hardware architecture is introduced to maximize the efficiency of the proposed optimization technique. Experimental results show that our approach effectively reduces the number of distance calculations and data accesses with negligible accuracy loss, without requiring retraining of the network model. Hyunsung Yoon, Jae-Joon Kim |
DAC | 2 |
| 2024 | FIGNA: Integer Unit-Based Accelerator Design for FP-INT GEMM Preserving Numerical AccuracyabstractThe weight-only quantization has emerged as a promising technique for alleviating the computational burden of large language models (LLMs) by employing low-precision integer (INT) weights, while retaining full-precision floating point (FP) activations to ensure inference quality. Despite the memory footprint reduction achieved through decreased bit-precision of weight parameters, the actual computing performance is often not improved significantly due to FP-INT multiply-accumulation (MAC) operations being performed on the floating point unit (FPU) after de quantizing the INT weight values to FP values, owing to the lack of dedicated FP- INT arithmetic units. In this study, we investigate the impact of introducing a dedicated FP-INT unit on overall performance and find that such specialization does not yield substantial improvements. As an alternative approach, we propose FIGNA, an accelerator based on INT units designed specifically for FP- INT MAC operations. A key feature of FIGNA is its ability to achieve the same numerical accuracy as the FPU while relying solely on the integer-unit, a departure from prior methods that relied on integer-units with numerical approximations for FP arithmetic results, albeit claiming similar inference accuracy through dedicated network training. Through comprehensive experiments on FP- INT quantized networks for LLMs, including OPT and BLOOM, we demonstrate the superior performance of FIGNA compared to conventional FPUs in terms of performance per area ($TOPS/mm^{2}$) and energy efficiency (TOPS/W) across various input and weight precision combinations. For instance, in the FP16-INT4 case, FIGNA shows 6.34x higher$TOPS/ mm^{2}$and 2.19x higher TOPS/W compared to the baseline. Jaeyong Jang, Yulhwa Kim, Juheun Lee, Jae-Joon Kim |
HPCA | 4 |
| 2024 | SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer BlocksabstractLarge language models (LLMs) have proven to be highly effective across various natural language processing tasks. However, their large number of parameters poses significant challenges for practical deployment. Pruning, a technique aimed at reducing the size and complexity of LLMs, offers a potential solution by removing redundant components from the network. Despite the promise of pruning, existing methods often struggle to achieve substantial end-to-end LLM inference speedup. In this paper, we introduce SLEB, a novel approach designed to stream- line LLMs by eliminating redundant transformer blocks. We choose the transformer block as the fundamental unit for pruning, because LLMs exhibit block-level redundancy with high similarity between the outputs of neighboring blocks. This choice allows us to effectively enhance the processing speed of LLMs. Our experimental results demonstrate that SLEB outperforms previous LLM pruning methods in accelerating LLM inference while also maintaining superior perplexity and accuracy, making SLEB as a promising technique for enhancing the efficiency of LLMs. The code is available at: https://github.com/jiwonsong-dev/SLEB. Jiwon Song, Kyungseok Oh, Taesu Kim, Yulhwa Kim, Jae-Joon Kim |
ICML | 6 |
| 2024 | Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language ModelsabstractBinarization, which converts weight parameters to binary values, has emerged as an effective strategy to reduce the size of large language models (LLMs). However, typical binarization techniques significantly diminish linguistic effectiveness of LLMs.
To address this issue, we introduce a novel binarization technique called Mixture of Scales (BinaryMoS). Unlike conventional methods, BinaryMoS employs multiple scaling experts for binary weights, dynamically merging these experts for each token to adaptively generate scaling factors. This token-adaptive approach boosts the representational power of binarized LLMs by enabling contextual adjustments to the values of binary weights. Moreover, because this adaptive process only involves the scaling factors rather than the entire weight matrix, BinaryMoS maintains compression efficiency similar to traditional static binarization methods. Our experimental results reveal that BinaryMoS surpasses conventional binarization techniques in various natural language processing tasks and even outperforms 2-bit quantization methods, all while maintaining similar model size to static binarization techniques. Dongwon Jo, Taesu Kim, Yulhwa Kim, Jae-Joon Kim |
NeurIPS | 4 |
| 2024 | Mobileware: Distributed Architecture With Channel Stationary Dataflow for MobileNet AccelerationabstractThe depthwise separable convolution, a key feature of the MobileNet models, has a different input reuse pattern from the conventional standard convolution, and a smaller number of input/weight pairs are used for a dot product, thereby leading to extremely low MAC utilization. This paper proposes a Mobileware architecture for the high-performance acceleration of the MobileNet workloads. A new channel stationary dataflow architecture distributes the on-chip buffers, and the distributed SRAMs are placed near each PE. By doing so, PEs and SRAMs can communicate with high bandwidth. Our Mobileware architecture shows 1.4-29.5× higher throughput than conventional weight stationary-based hardware architecture, and the proposed design was verified on the Xilinx ZCU102 FPGA evaluation board. Sungju Ryu, Jaeyong Jang, Youngtaek Oh, Jae-Joon Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | In-Memory Neural Network Accelerator based on eDRAM Cell with Enhanced Retention TimeabstractLogic compatible eDRAM cell-based computing-in-memory (CIM) neural network accelerators have been actively studied as an energy-efficient neural network computing platform thanks to their small cell size and low static power compared to SRAM. However, previous eDRAM-based CIM accelerators suffer from significant accuracy degradation caused by process, voltage, temperature (PVT) variations and short retention time. To overcome the issues, we introduce a PVT-variation tolerant capacitive coupling-based eDRAM cell that has a much longer retention time than previous works. Simulation results show that the proposed eDRAM cell has up to 50× higher retention time compared to the state-of-the-art designs. Inhwan Lee, Eunhwan Kim, Nameun Kang, Hyunmyung Oh, Jae-Joon Kim |
DAC | 5 |
| 2023 | Efficient Sampling and Grouping Acceleration for Point Cloud Deep Learning via Single Coordinate ComparisonabstractWith the focus on three-dimensional (3D) applications, the importance of applying deep learning to point clouds have been growing recently. It is known that mapping operations including sampling and grouping play a critical role in extracting local features in point-based deep learning models. However, the mapping operations often become bottlenecks in terms of computing times due to the repetitive comparison of distances between input points. In this paper, we analyzed the characteristics of distance distribution during sampling and grouping operations, and discovered that substantial portion of the distance comparison does not need exact 3D Euclidean distance using all three coordinates. Based on the observations, we propose a technique called single coordinate comparison which selectively determines the comparison output with 1D-distance only. We also present a hardware architecture with a distance calculator capable of handling both 3D and 1D distance. The experimental results demonstrate the effectiveness of our approach in reducing both time and energy consumption, particularly as the number of points increases. Hyunsung Yoon, Jae-Joon Kim |
ICCAD | 2 |
| 2023 | INSTA-BNN: Binary Neural Network with INSTAnce-aware ThresholdabstractBinary Neural Networks (BNNs) have emerged as a promising solution for reducing the memory footprint and compute costs of deep neural networks, but they suffer from quality degradation due to the lack of freedom as activations and weights are constrained to the binary values. To compensate for the accuracy drop, we propose a novel BNN design called Binary Neural Network with INSTAnce-aware threshold (INSTA-BNN), which controls the quantization threshold dynamically in an input-dependent or instance-aware manner. According to our observation, higher-order statistics can be a representative metric to estimate the characteristics of the input distribution. INSTA-BNN is designed to adjust the threshold dynamically considering various information, including higher-order statistics, but it is also optimized judiciously to realize minimal overhead on a real device. Our extensive study shows that INSTA-BNN outperforms the baseline by 3.0% and 2.8% on the ImageNet classification task with comparable computing cost, achieving 68.5% and 72.2% top-1 accuracy on ResNet-18 and MobileNetV1 based models, respectively. Changhun Lee, Eunhyeok Park, Jae-Joon Kim |
ICCV | 4 |
| 2023 | Winning Both the Accuracy of Floating Point Activation and the Simplicity of Integer Arithmetic
Yulhwa Kim, Jaeyong Jang, Jehun Lee, Byeongwook Kim, Baeseong Park, Se Jung Kwon, Dongsoo Lee, Jae-Joon Kim |
ICLR | 10 |
| 2023 | Leveraging Early-Stage Robustness in Diffusion Models for Efficient and High-Quality Image SynthesisabstractWhile diffusion models have demonstrated exceptional image generation capabilities, the iterative noise estimation process required for these models is compute-intensive and their practical implementation is limited by slow sampling speeds. In this paper, we propose a novel approach to speed up the noise estimation network by leveraging the robustness of early-stage diffusion models. Our findings indicate that inaccurate computation during the early-stage of the reverse diffusion process has minimal impact on the quality of generated images, as this stage primarily outlines the image while later stages handle the finer details that require more sensitive information. To improve computational efficiency, we combine our findings with post-training quantization (PTQ) to introduce a method that utilizes low-bit activation for the early reverse diffusion process while maintaining high-bit activation for the later stages. Experimental results show that the proposed method can accelerate the early-stage computation without sacrificing the quality of the generated images. Yulhwa Kim, Dongwon Jo, Hyesung Jeon, Taesu Kim, Daehyun Ahn, Jae-Joon Kim |
NeurIPS | 7 |
| 2023 | Searching for Robust Binary Neural Networks via Bimodal Parameter PerturbationabstractBinary neural networks (BNNs) are advantageous in performance and memory footprint but suffer from low accuracy due to their limited expression capability. Recent works have tried to enhance the accuracy of BNNs via a gradient-based search algorithm and showed promising results. However, the mixture of architecture search and binarization induce the instability of the search process, resulting in convergence to the suboptimal point. To address this issue, we propose a BNN architecture search framework with bimodal parameter perturbation. The bimodal parameter perturbation can improve the stability of gradient-based architecture search by reducing the sharpness of the loss surface along both weight and architecture parameter axes. In addition, we refine the inverted bottleneck convolution block for having robustness with BNNs. The synergy of the refined space and the stabilized search process allows us to find out the accurate BNNs with high computation efficiency. Experimental results show that our framework finds the best architecture on CIFAR-100 and ImageNet datasets in the existing search space for BNNs. We also tested our framework on another search space based on the inverted bottleneck convolution block, and the selected BNN models using our approach achieved the highest accuracy on both datasets with a much smaller number of equivalent operations than previous works. Daehyun Ahn, Taesu Kim, Eunhyeok Park, Jae-Joon Kim |
WACV | 5 |
| 2023 | V-LSTM: An Efficient LSTM Accelerator Using Fixed Nonzero-Ratio Viterbi-Based PruningabstractLong short-term memory (LSTM) has been widely adopted in tasks with sequence data, such as speech recognition and language modeling. LSTM brought significant accuracy improvement by introducing additional parameters to recurrent neural network (RNN). However, increasing number of parameters and computations also led to inefficiency in computing LSTM on edge devices with limited on-chip memory size and DRAM bandwidth. In order to reduce the latency and energy of LSTM computations, there has been a pressing need for model compression schemes and suitable hardware accelerators. In this article, we first propose the Fixed Nonzero-ratio Viterbi-based Pruning, which can reduce the memory footprint of LSTM models by 96% with negligible accuracy loss. By applying additional constraints on the distribution of surviving weights in Viterbi-based Pruning, the proposed pruning scheme mitigates the load-imbalance problem and thereby increases the processing engine utilization rate. Then, we propose the V-LSTM, an efficient sparse LSTM accelerator based on the proposed pruning scheme. High compression ratio of the proposed pruning scheme allows the proposed accelerator to achieve 24.9% lower per-sample latency than that of state-of-the-art accelerators. The proposed accelerator is implemented on Xilinx VC-709 FPGA evaluation board running at 200 MHz for evaluation. Taesu Kim, Daehyun Ahn, Dongsoo Lee, Jae-Joon Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | Binaryware: A High-Performance Digital Hardware Accelerator for Binary Neural NetworksabstractBinary neural networks (BNNs) largely reduce the memory footprint and computational complexity, so they are gaining interests on various mobile applications. In the BNNs, the first layer often accounts for the largest part of the entire computing time because the layer usually uses multi-bit multiplications. However, traditional hardware designed for BNN computing focuses primarily on the rest layers, resulting in significant performance degradation. In this brief, we introduce Binaryware architecture which achieves the high-performance computation on both the first and rest layers. Experimental results show that our Binaryware improves the throughput per compute area by 1.5–$13.3\times $on various BNN workloads. Sungju Ryu, Youngtaek Oh, Jae-Joon Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2022 | TAIM: ternary activation in-memory computing hardware with 6T SRAM arrayabstractRecently, various in-memory computing accelerators for low precision neural networks have been proposed. While in-memory Binary Neural Network (BNN) accelerators achieved significant energy efficiency, BNNs show severe accuracy degradation compared to their full precision counterpart models. To mitigate the problem, we propose TAIM, an in-memory computing hardware that can support ternary activation with negligible hardware overhead. In TAIM, a 6T SRAM cell can compute the multiplication between ternary activation and binary weight. Since the 6T SRAM cell consumes no energy when the input activation is 0, the proposed TAIM hardware can achieve even higher energy efficiency compared to BNN case by exploiting input 0's. We fabricated the proposed TAIM hardware in 28nm CMOS process and evaluated the energy efficiency on various image classification benchmarks. The experimental results show that the proposed TAIM hardware can achieve ~ 3.61× higher energy efficiency on average compared to previous designs which support ternary activation. Nameun Kang, Hyunmyung Oh, Jae-Joon Kim |
DAC | 4 |
| 2022 | Workload-Balanced Graph Attention Network Accelerator with Top-K Aggregation CandidatesabstractGraph attention networks (GATs) are gaining attention for various transductive and inductive graph processing tasks due to their higher accuracy than conventional graph convolutional networks (GCNs). The power-law distribution of real-world graph-structured data, on the other hand, causes a severe workload imbalance problem for GAT accelerators. To reduce the degradation of PE utilization due to the workload imbalance, we present algorithm/hardware co-design results for a GAT accelerator that balances workload assigned to processing elements by allowing only K neighbor nodes to participate in aggregation phase. The proposed model selects the K neighbor nodes with high attention scores, which represent relevance between two nodes, to minimize accuracy drop. Experimental results show that our algorithm/hardware co-design of the GAT accelerator achieves higher processing speed and energy efficiency than the GAT accelerators using conventional workload balancing techniques. Furthermore, we demonstrate that the proposed GAT accelerators can be made faster than the GCN accelerators that typically process smaller number of computations. Naebeom Park, Daehyun Ahn, Jae-Joon Kim |
ICCAD | 3 |
| 2022 | Extreme Partial-Sum Quantization for Analog Computing-In-Memory Neural Network AcceleratorsabstractIn Analog Computing-in-Memory (CIM) neural network accelerators, analog-to-digital converters (ADCs) are required to convert the analog partial sums generated from a CIM array to digital values. The overhead from ADCs substantially degrades the energy efficiency of CIM accelerators so that previous works attempted to lower the ADC resolution considering the distribution of the partial sums. Despite the efforts, the required ADC resolution still remains relatively high. In this article, we propose the data-driven partial sum quantization scheme, which exhaustively searches for the optimal quantization range with little computational burden. We also report that analyzing the characteristics of the partial sum distributions at each layer gives an additional information to further reduce the ADC resolution compared to previous works that mostly used the characteristics of the partial sum distributions of the entire network. Based on the finer-level data-driven approach combined with retraining, we present a methodology for extreme partial-sum quantization. Experimental results show that the proposed method can reduce the ADC resolution to 2 to 3 bits for CIFAR-10 dataset, which is the smaller ADC bit resolution than any previous CIM-based NN accelerators. Yulhwa Kim, Jae-Joon Kim |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2021 | Improving Accuracy of Binary Neural Networks Using Unbalanced Activation DistributionabstractBinarization of neural network models is considered as one of the promising methods to deploy deep neural network models on resource-constrained environments such as mobile devices. However, Binary Neural Networks (BNNs) tend to suffer from severe accuracy degradation compared to the full-precision counterpart model. Several techniques were proposed to improve the accuracy of BNNs. One of the approaches is to balance the distribution of binary activations so that the amount of information in the binary activations becomes maximum. Based on extensive analysis, in stark contrast to previous work, we argue that unbalanced activation distribution can actually improve the accuracy of BNNs. We also show that adjusting the threshold values of binary activation functions results in the unbalanced distribution of the binary activation, which increases the accuracy of BNN models. Experimental results show that the accuracy of previous BNN models (e.g. XNOR-Net and Bi-Real-Net) can be improved by simply shifting the threshold values of binary activation functions without requiring any other modification. Changhun Lee, Jae-Joon Kim |
CVPR | 4 |
| 2021 | Mapping Binary ResNets on Computing-In-Memory Hardware with Low-bit ADCsabstractImplementing binary neural networks (BNNs) on computing-in-memory (CIM) hardware has several attractive features such as small memory requirement and minimal overhead in peripheral circuits such as analog-to-digital converters (ADCs). On the other hand, one of the downsides of using BNNs is that it degrades the classification accuracy. Recently, ResNet-style BNNs are gaining popularity with higher accuracy than conventional BNNs. The accuracy improvement comes from the high-resolution skip connection which binary ResNets use to compensate the information loss caused by binarization. However, the high-resolution skip connection forces the CIM hardware to use high-bit ADCs again so that area and energy overhead becomes larger. In this paper, we demonstrate that binary ResNets can be also mapped on CIM with low-bit ADCs via aggressive partial sum quantization and input-splitting combined with retraining. As a result, the key advantages of BNN CIM such as small area and energy consumption can be preserved with higher accuracy. Yulhwa Kim, Hyunmyung Oh, Jae-Joon Kim |
DATE | 5 |
| 2021 | SPRITE: Sparsity-Aware Neural Processing Unit with Constant Probability of Index-MatchingabstractSparse neural networks are widely used for memory savings. However, irregular indices of non-zero input activations and weights tend to degrade the overall system performance. This paper presents a scheme to maintain constant probability of index-matching for weight and input over a wide range of sparsity overcoming a critical limitation in previous works. A sparsity-aware neural processing unit based on the proposed scheme improves the system performance up to 6.1× compared to previous sparse convolutional neural network hardware accelerators. Sungju Ryu, Youngtaek Oh, Taesu Kim, Daehyun Ahn, Jae-Joon Kim |
DATE | 5 |
| 2021 | Mobileware: A High-Performance MobileNet Accelerator with Channel Stationary DataflowabstractMobileNet models have been gaining popularity with lighter computational load than previous convolutional neural network models. However, mapping depthwise convolutional layers in the MobileNets to the conventional hardware accelerators experiences significantly low utilization of multipliers, which leads to large performance degradation. The utilization becomes low because the depthwise convolutional layers have 1) different input/weight reuse patterns and 2) much smaller number of multiplications for each dot product computation compared to standard convolutional operations. To overcome such limitations, we propose an architecture called Mobileware which uses a channel stationary dataflow. Experimental results show that the Mobileware can achieve up to$1.4-29.5\times$higher overall system performance than the previous neural network accelerators. Sungju Ryu, Youngtaek Oh, Jae-Joon Kim |
ICCAD | 3 |
| 2021 | High-throughput Near-Memory Processing on CNNs with 3D HBM-like MemoryabstractThis article discusses the high-performance near-memory neural network (NN) accelerator architecture utilizing the logic die in three-dimensional (3D) High Bandwidth Memory– (HBM) like memory. As most of the previously reported 3D memory-based near-memory NN accelerator designs used the Hybrid Memory Cube (HMC) memory, we first focus on identifying the key differences between HBM and HMC in terms of near-memory NN accelerator design. One of the major differences between the two 3D memories is that HBM has the centralized through- silicon-via (TSV) channels while HMC has distributed TSV channels for separate vaults. Based on the observation, we introduce the Round-Robin Data Fetching and Groupwise Broadcast schemes to exploit the centralized TSV channels for improvement of the data feeding rate for the processing elements. Using synthesized designs in a 28-nm CMOS technology, performance and energy consumption of the proposed architectures with various dataflow models are evaluated. Experimental results show that the proposed schemes reduce the runtime by 16.4–39.3% on average and the energy consumption by 2.1–5.1% on average compared to conventional data fetching schemes. Naebeom Park, Sungju Ryu, Jaeha Kung 0001, Jae-Joon Kim |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2020 | Algorithm/Hardware Co-Design for In-Memory Neural Network Computing with Minimal Peripheral Circuit OverheadabstractWe propose an in-memory neural network accelerator architecture called MOSAIC which uses minimal form of peripheral circuits; 1-bit word line driver to replace DAC and 1-bit sense amplifier to replace ADC. To map multi-bit neural networks on MOSAIC architecture which has 1-bit precision peripheral circuits, we also propose a bit-splitting method to approximate the original network by separating each bit path of the multi-bit network so that each bit path can propagate independently throughout the network. Thanks to the minimal form of peripheral circuits, MOSAIC can achieve an order of magnitude higher energy and area efficiency than previous in-memory neural network accelerators. Yulhwa Kim, Sungju Ryu, Jae-Joon Kim |
DAC | 4 |
| 2020 | V-LSTM: An Efficient LSTM Accelerator Using Fixed Nonzero-Ratio Viterbi-Based PruningabstractLong Short-Term Memory (LSTM) has been widely adopted in tasks with sequence data, such as speech recognition and language modeling. LSTM brought significant accuracy improvement by introducing additional parameters to Recurrent Neural Network (RNN). However, increasing number of parameters and computations also led to inefficiency in computing LSTM on edge devices with limited on-chip memory size and DRAM bandwidth. In order to reduce the latency and energy of LSTM computations, there has been a pressing need for model compression schemes and suitable hardware accelerators. In this paper, we first propose the Fixed Nonzero-ratio Viterbi-based Pruning, which can reduce the memory footprint of LSTM models by 96% with negligible accuracy loss. By applying additional constraints on the distribution of surviving weights in Viterbi-based Pruning, the proposed pruning scheme mitigates the load-imbalance problem and thereby increases the processing engine utilization rate. Then, we propose the V-LSTM, an efficient sparse LSTM accelerator based on the proposed pruning scheme. High compression ratio of the proposed pruning scheme allows the proposed accelerator to achieve 24.9% lower per-sample latency than that of state-of-the-art accelerators. The proposed accelerator is implemented on Xilinx VC-709 FPGA evaluation board running at 200MHz for evaluation. Taesu Kim, Daehyun Ahn, Jae-Joon Kim |
FPGA | 3 |
| 2020 | Energy-efficient XNOR-free In-Memory BNN Accelerator with Input Distribution RegularizationabstractSRAM-based in-memory Binary Neural Network (BNN) accelerators are garnering interests as a platform for energy-efficient edge neural network computing thanks to their compactness in terms of hardware and neural network parameter size. However, previous works had to modify SRAM cells to support XNOR operations on memory array resulting in limited area and energy efficiencies. In this work, we present a conversion method which replaces the signed inputs (+1/-1) of BNN with the unsigned inputs (1/0) without computation error, and vice versa. The method enables BNN computing on conventional 6T SRAM arrays and improves area and energy efficiencies. We also demonstrate that further energy saving is possible by skewing the distribution of binary input data based on regularization during network training. Evaluation results show that the proposed techniques improve the inference energy efficiency by up to 9.4x for various benchmarks over previous works. Hyunmyung Oh, Jae-Joon Kim |
ICCAD | 3 |
| 2020 | BinaryDuo: Reducing Gradient Mismatch in Binary Activation Network by Coupling Binary Activations
Kyungsu Kim 0004, Jinseok Kim 0004, Jae-Joon Kim |
ICLR | 4 |
| 2020 | Time-step interleaved weight reuse for LSTM neural network computingabstractIn Long Short-Term Memory (LSTM) neural network models, a weight matrix tends to be repeatedly loaded from DRAM if the size of on-chip storage of the processor is not large enough to store the entire matrix. To alleviate heavy overhead of DRAM access for weight loading in LSTM computations, we propose a weight reuse scheme which utilizes the weight sharing characteristics in two adjacent time-step computations. Experimental results show that the proposed weight reuse scheme reduces the energy consumption by 28.4-57.3% and increases the overall throughput by 110.8% compared to the conventional schemes. Naebeom Park, Yulhwa Kim, Daehyun Ahn, Taesu Kim, Jae-Joon Kim |
ISLPED | 5 |
| 2020 | Unifying Activation- and Timing-based Learning Rules for Spiking Neural NetworksabstractFor the gradient computation across the time domain in Spiking Neural Networks (SNNs) training, two different approaches have been independently studied. The first is to compute the gradients with respect to the change in spike activation (activation-based methods), and the second is to compute the gradients with respect to the change in spike timing (timing-based methods). In this work, we present a comparative study of the two methods and propose a new supervised learning method that combines them. The proposed method utilizes each individual spike more effectively by shifting spike timings as in the timing-based methods as well as generating and removing spikes as in the activation-based methods. Experimental results showed that the proposed method achieves higher performance in terms of both accuracy and efficiency than the previous approaches. Jinseok Kim 0004, Kyungsu Kim 0004, Jae-Joon Kim |
NeurIPS | 3 |
| 2020 | Balancing Computation Loads and Optimizing Input Vector Loading in LSTM AcceleratorsabstractThe long short-term memory (LSTM) is a widely used neural network model for dealing with time-varying data. To reduce the memory requirement, pruning is often applied to the weight matrix of the LSTM, which makes the matrix sparse. In this paper, we present a new sparse matrix format, named rearranged compressed sparse column (RCSC), to maximize the inference speed of the LSTM hardware accelerator. The RCSC format speeds up the inference by: 1) evenly distributing the computation loads to processing elements (PEs) and 2) reducing the input vector load miss within the local buffer. We also propose a hardware architecture adopting hierarchical input buffer to further reduce the pipeline stalls which cannot be handled by the RCSC format alone. The simulation results for various datasets show that combined use of the RSCS format and the proposed hardware requires 2× smaller inference runtime on average compared to the previous work. Junki Park, Wooseok Yi, Daehyun Ahn, Jaeha Kung 0001, Jae-Joon Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | GeCo: Classification Restricted Boltzmann Machine Hardware for On-Chip Semisupervised Learning and Bayesian InferenceabstractThe probabilistic Bayesian inference of real-time input data is becoming more popular, and the importance of semisupervised learning is growing. We present a classification restricted Boltzmann machine (ClassRBM)-based hardware accelerator with on-chip semisupervised learning and Bayesian inference capability. ClassRBM is a specific type of Markov network that can perform classification tasks and reconstruct its input data. ClassRBM has several advantages in terms of hardware implementation compared to other backpropagation-based neural networks. However, its accuracy is relatively low compared to backpropagation-based learning. To improve the accuracy of ClassRBM, we propose the multi-neuron-per-class (multi-NPC) voting scheme. We also reveal that the contrastive divergence (CD) algorithm, which is commonly used to train RBM, shows poor performance in this multi-NPC ClassRBM. As an alternative, we propose an asymmetric contrastive divergence (ACD) training algorithm that improves the accuracy of multi-NPC ClassRBM. With the ACD learning algorithm, ClassRBM operates in the form of a combination of Markov Chain training and Bayesian inference. The experimental results on a field-programmable gate array (FPGA) board for a Modified National Institute of Standards and Technology data set confirm that the inference accuracy of the proposed ACD algorithm is 5.82% higher for a supervised learning case and 12.78% higher for a 1% labeled semisupervised learning case than the conventional CD algorithm. Also, the GeCo ver.2 hardware implemented on a Xilinx ZCU102 FPGA board was 349.04 times faster than the C simulation on CPU. Wooseok Yi, Junki Park, Jae-Joon Kim |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | In-memory batch-normalization for resistive memory based binary neural network hardwareabstractBinary Neural Network (BNN) has a great potential to be implemented on Resistive memory Crossbar Array (RCA)-based hardware accelerators because it requires only 1-bit precision for weights and activations. While general structures to implement convolution or fully-connected layers in RCA-based BNN hardware were actively studied in previous works, Batch-Normalization (BN) layer, which is another key layer of BNN, has not been discussed in depth yet. In this work, we propose in-memory batch-normalization schemes which integrate BN layers on RCA so that area/energy-efficiency of the BNN accelerators can be maximized. In addition, we also show that sense amp error due to device mismatch can be suppressed using the proposed in-memory BN design. Yulhwa Kim, Jae-Joon Kim |
ASP-DAC | 3 |
| 2019 | Peregrine: A Flexible Hardware Accelerator for LSTM with Limited Synaptic Connection PatternsabstractIn this paper, we present an integrated solution to design a high-performance LSTM accelerator. We propose a fast and flexible hardware architecture, named Peregrine, supported by a stack of innovations from algorithm to hardware design. Peregrine first minimizes the memory footprint by limiting the synaptic connection patterns within the LSTM network. Also, Peregrine provides parallel Huffman decoders with adaptive clocking to provide flexibility in dealing with a wide range of sparsity levels in the weight matrices. All these features are incorporated in a novel hardware architecture to maximize energy-efficiency. As a result, Peregrine improves performance by ~38% and energy-efficiency by ~33% in speech recognition compared to the state-of-the-art LSTM accelerator. Jaeha Kung 0001, Junki Park, Sehun Park, Jae-Joon Kim |
DAC | 4 |
| 2019 | BitBlade: Area and Energy-Efficient Precision-Scalable Neural Network Accelerator with Bitwise SummationabstractDeep Neural Networks (DNNs) have various performance requirements and power constraints depending on applications. To maximize the energy-efficiency of hardware accelerators for different applications, the accelerators need to support various bit-width configurations. When designing bit-reconfigurable accelerators, each PE must have variable shift-addition logic, which takes a large amount of area and power. This paper introduces an area and energy efficient precision-scalable neural network accelerator (BitBlade), which reduces the control overhead for variable shift-addition using bitwise summation method. The proposed BitBlade, when synthesized in a 28nm CMOS technology, showed reduction in area by 41% and in energy by 36-46% compared to the state-of-the-art precision-scalable architecture [14]. Sungju Ryu, Wooseok Yi, Jae-Joon Kim |
DAC | 4 |
| 2019 | Effect of Device Variation on Mapping Binary Neural Network to Memristor Crossbar ArrayabstractIn memristor crossbar array (MCA)-based neural network hardware, it is generally assumed that entire word-lines (WLs) are simultaneously enabled for parallel matrix-vector multiplication (MxV) operation. However, the error probability of MxV in a memristor crossbar array (MCA) increases as the resistance ratio (R-ratio) of a memristor decreases and the resistance variation and the number of simultaneously activated WLs increase. In this paper, we analyze the effect of R-ratio and variation of memristor devices on read sense margin and inference accuracy of MCA-based Binary Neural Network (BNN) hardware. We first show that only a limited number of WLs should be enabled to ensure correct MxV output when the R-ratio is small. On the other hand, we also show that, if the resistance variation becomes higher than a certain level, simultaneous activation of large number of WLs produces the higher accuracy even when R-ratio is small. Based on the analysis, we propose the Accuracy Estimation (AE) factor to find the optimal number of word lines that are simultaneously activated. Wooseok Yi, Yulhwa Kim, Jae-Joon Kim |
DATE | 3 |
| 2019 | Double Viterbi: Weight Encoding for High Compression Ratio and Fast On-Chip Reconstruction for Deep Neural Network
Daehyun Ahn, Dongsoo Lee, Taesu Kim, Jae-Joon Kim |
ICLR (Poster) | 4 |
| 2019 | Feedforward-Cutset-Free Pipelined Multiply-Accumulate Unit for the Machine Learning AcceleratorabstractMultiply-accumulate (MAC) computations account for a large part of machine learning accelerator operations. The pipelined structure is usually adopted to improve the performance by reducing the length of critical paths. An increase in the number of flip-flops due to pipelining, however, generally results in significant area and power increase. A large number of flip-flops are often required to meet the feedforward-cutset rule. Based on the observation that this rule can be relaxed in machine learning applications, we propose a pipelining method that eliminates some of the flip-flops selectively. The simulation results show that the proposed MAC unit achieved a 20% energy saving and a 20% area reduction compared with the conventional pipelined MAC. Sungju Ryu, Naebeom Park, Jae-Joon Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Maximizing system performance by balancing computation loads in LSTM acceleratorsabstractThe LSTM is a popular neural network model for modeling or analyzing the time-varying data. The main operation of LSTM is a matrix-vector multiplication and it becomes sparse (spMxV) due to the widely-accepted weight pruning in deep learning. This paper presents a new sparse matrix format, named CBSR, to maximize the inference speed of the LSTM accelerator. In the CBSR format, speed-up is achieved by balancing out the computation loads over PEs. Along with the new format, we present a simple network transformation to completely remove the hardware overhead incurred when using the CBSR format. Also, the detailed analysis on the impact of network size or the number of PEs is performed, which lacks in the prior work. The simulation results show 16~38% improvement in the system performance compared to the well-known CSC/CSR format. The power analysis is also performed in 65nm CMOS technology to show 9~22% energy savings. Junki Park, Jaeha Kung 0001, Wooseok Yi, Jae-Joon Kim |
DATE | 4 |
| 2018 | Viterbi-based Pruning for Sparse Matrix with Fixed and High Index Compression Ratio
Dongsoo Lee, Daehyun Ahn, Taesu Kim, Pierce Chuang, Jae-Joon Kim |
ICLR (Poster) | 5 |
| 2018 | Input-Splitting of Large Neural Networks for Power-Efficient Accelerator with Resistive Crossbar Memory ArrayabstractResistive Crossbar memory Arrays (RCA) have been gaining interest as a promising platform to implement Convolutional Neural Networks (CNN). One of the major challenges in RCA-based design is that the number of rows in an RCA is often smaller than the number of input neurons in a layer. Previous works used high-resolution Analog-to-Digital Converters (ADCs) to compute the partial weighted sum in each array and merged partial sums from multiple arrays outside the RCAs. However, such approach suffers from significant power consumption due to the need for high-resolution ADCs. In this paper, we propose a methodology to more efficiently construct a large CNN with multiple RCAs. By splitting the input feature map and retraining the CNN with proper initialization, we demonstrate that any CNN model can be represented with multiple arrays without using intermediate partial sums. The experimental results show that the ADC power of the proposed design is 32x smaller and the total chip power of the proposed design is 3x smaller than those of the baseline design. Yulhwa Kim, Daehyun Ahn, Jae-Joon Kim |
ISLPED | 4 |
| 2018 | Compact Convolution Mapping on Neuromorphic Hardware using Axonal DelayabstractMapping Convolutional Neural Network (CNN) to a neuromorphic hardware has been inefficient in synapse memory usage because both kernel/input reuse are not exploited well. We propose a method to enable kernel reuse by utilizing axonal delay, which is a biological parameter for a spiking neuron. Using IBM TrueNorth as a test platform, we demonstrate that the number of cores, neurons, synapses, and synaptic operations per time step can be reduced by up to 20.9x, 27.9x, 88.4x, and 1586x, respectively, compared to the conventional scheme, which raises the possibility of implementing large-scale CNN on neuromorphic hardware. Jinseok Kim 0004, Yulhwa Kim, Jae-Joon Kim |
ISLPED | 4 |
| 2018 | Deep Neural Network Optimized to Resistive Memory with Nonlinear Current-Voltage CharacteristicsabstractArtificial Neural Network computation relies on intensive vector-matrix multiplications. Recently, the emerging nonvolatile memory (NVM) crossbar array showed a feasibility of implementing such operations with high energy efficiency. Thus, there have been many works on efficiently utilizing emerging NVM crossbar arrays as analog vector-matrix multipliers. However, nonlinear I-V characteristics of NVM restrain critical design parameters, such as the read voltage and weight range, resulting in substantial accuracy loss. In this article, instead of optimizing hardware parameters to a given neural network, we propose a methodology of reconstructing the neural network itself to be optimized to resistive memory crossbar arrays. To verify the validity of the proposed method, we simulated various neural networks with MNIST and CIFAR-10 dataset using two different Resistive Random Access Memory models. Simulation results show that our proposed neural network produces inference accuracies significantly higher than conventional neural network when the network is mapped to synapse devices with nonlinear I-V characteristics. Taesu Kim, Jinseok Kim 0004, Jae-Joon Kim |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2017 | Low design overhead timing error correction scheme for elastic clock methodologyabstractThe elastic clock scheme is a robust design methodology to ensure timing closure under PVT variation using locally generated clocks and handshaking protocol. However, it still has a chance of timing errors due to delay mismatch between the data-path and delay replica. In this paper, we propose a low design overhead timing error correction scheme tailored to elastic clock. In the proposed scheme, a timing error can be corrected within a cycle using clock stretching. The proposed scheme shows 40.3× and 4.6× reduction in timing margin with 9.1% and 9.0% area overhead over the synchronous baseline and elastic clock design, respectively. Sungju Ryu, Jongeun Koo, Jae-Joon Kim |
ISLPED | 3 |
| 2017 | GeCo: classification restricted Boltzmann machine hardware for on-chip learningabstractWe present a Classification Restricted Boltzmann Machine (Class-RBM) hardware for embedded machines with on-chip learning capability. The RBM is a kind of the generative model, and has been used as one of the most popular feature extractors and image pre-processors. The ClassRBM is a variant of the RBM that is adapted to classification tasks. We propose the multi-Neuron-Per-Class (multi-NPC) voting scheme for improving accuracy of ClassRBM. We also show that the Contrastive Divergence (CD), which is one of the most popular algorithms to train RBM, has limitations in multi-NPC ClassRBM learning and propose a modified CD algorithm to overcome the limitation. Experimental results on FPGA flatform for MNIST datasets confirm that classification accuracy of the proposed algorithm is ∼ 2.12% higher than the conventional CD. Wooseok Yi, Junki Park, Jae-Joon Kim |
RSP | 3 |
| 2016 | One-Cycle Correction of Timing Errors in Pipelines With Standard Clocked ElementsabstractOne of the most aggressive uses of dynamic voltage scaling is timing speculation, which in turn requires fast correction of timing errors. The fastest existing error correction technique imposes a one-cycle time penalty only, but it is restricted to two-phase transparent latch-based pipelines. We perform one-cycle error correction by gating only the main latch in each stage of the pipeline that precedes a failed stage. This new method is applicable to widely used clocking elements, such as flip-flops and pulsed latches. Because it prevents inputs arriving at a stage, which is stalled, it can also be used in pipelines with multiple fan-in, fan-out, and looping. Simulations show an energy saving of 8%-12% with a target throughput of 0.9 instructions per cycle, and 15%-18% when the target is 0.8. Insup Shin, Jae-Joon Kim, Yu-Shiang Lin, Youngsoo Shin |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Power minimization of pipeline architecture through 1-cycle error correction and voltage scalingabstractWe present a new 1-cycle timing error correction method, which enables aggressive voltage scaling in a pipelined architecture. The proposed method differs from the state-of-the-art in that the pipeline stage where the timing error occurs can continue to receive input data without halting to avoid data collision. The feature allows the pipeline to avoid recurring clock gating when timing errors happen at multiple stages or timing errors continue to occur at a certain stage. Compared to a state-of-the-art method, the proposed method shows 2-6% energy reduction for a 5-stage pipeline and 7-11% reduction for a 10-stage pipeline. In addition, the proposed logic to propagate clock gating signal is much simpler than that of the previous method [1] by eliminating reverse propagation path of clock gating signal. Insup Shin, Jae-Joon Kim, Youngsoo Shin |
ASP-DAC | 2 |
| 2014 | Coarse-grained Bubble Razor to exploit the potential of two-phase transparent latch designsabstractTiming margin to cover process variation is one of the most critical factors that limit the amount of supply voltage reduction thereby power consumption. To remove too conservative timing margin, Bubble Razor was introduced to dynamically detect and correct errors in two-phase transparent latch designs [13]. However, it does not fully exploit the potential of two-phase transparent latch design, e.g. time borrowing. Thus, especially at low supply voltage where the effect of process variation becomes significant, the existing Bubble Razor can suffer from significant overhead in performance and power consumption due to too frequent occurrence of bubble generations. We present a design methodology for coarse-grained Bubble Razor which exploits the time-borrowing characteristic of two-phase transparent latch design. By selectively inserting error checkpoints, i.e., shadow latches and error management logic, in the circuit, time borrowing can be applied between error checkpoints thereby avoiding bubbles which could occur in the existing Bubble Razor design with a checkpoint at every latch on the critical path. We present a methodology to choose the grain size (the number of stages between error checkpoints) based on 3-sigma delay distribution. We also verify the benefits of coarse-grained Bubble Razor with a real microprocessor, Core-A design [15] using 20nm Predictive Technology Model (PTM) [16]. The proposed methodology offers 62% improvement in performance (MIPS) and 49% less energy consumption (per instruction) at 0.6V operation (zero frequency margin) over the original Bubble Razor scheme. In addition, it gives 25% area reduction in core design. Hayoung Kim, Dongyoung Kim, Jae-Joon Kim, Sungjoo Yoo, Sunggu Lee |
DATE | 3 |
| 2013 | A pipeline architecture with 1-cycle timing error correction for low voltage operationsabstractWe present a new timing error correction scheme which allows each pipeline stage to halt for one cycle only. The small timing penalty for the error correction operation in the proposed scheme makes it possible to eliminate the extra timing guardband that was needed to accommodate timing uncertainty due to process variations. As a result, lower supply voltage can be used with the proposed scheme for low power operations. Compared to the previous 1-cycle error correction scheme which uses two-phase transparent latch based pipeline [1], the proposed scheme can be applied to the pipeline based on more popular clocking elements such as flip-flop or pulsed latch. Insup Shin, Jae-Joon Kim, Yu-Shiang Lin, Youngsoo Shin |
ISLPED | 2 |
| 2013 | Slew-Rate Monitoring Circuit for On-Chip Process Variation DetectionabstractThe need for efficient and accurate detection schemes to assess the impact of process variations on the parametric yield of integrated circuits has increased in the nanometer design era. In this paper, the difference of rise and fall slew is presented as another process-variation metric along with the delay in determining the relative mismatch between the drive strengths of nMOS and pMOS devices. The importance of considering both of these metrics is illustrated, and a new slew-rate monitoring circuit is presented for measuring the difference of rise and fall slew of a signal on the critical path of a circuit. Sensitivity analysis with multiple pulses as input has also been investigated. Bias generator circuits that track nMOS and pMOS threshold voltages have been incorporated, which makes the design less susceptible to process variation. Design considerations, simulation results, and characteristics of the slew-rate monitor circuitry in a 65-nm IBM CMOS process are presented, and a sensitivity of 50 MHz/50 ps for single pulse input is achieved. The measurement sensitivity of a fabricated slew-rate monitor in a 65-nm IBM CMOS technology is 0.11 V/μs, with 1089 pF as the output load of the slew-rate monitor. Amlan Ghosh, Rahul M. Rao, Jae-Joon Kim, Ching-Te Chuang, Richard B. Brown |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Column-selection-enabled 8T SRAM array with ~1R/1W multi-port operation for DVFS-enabled processors
Sang Phill Park, Dongsoo Lee, Jae-Joon Kim, W. Paul Griffin, Kaushik Roy 0001 |
ISLPED | 4 |
| 2011 | Robust Level Converter for Sub-Threshold/Super-Threshold Operation: 100 mV to 2.5 VabstractFor ultra low power application, digital sub-threshold logic design has been explored. Extremely low power supply (VDD) of sub-threshold logic results in significant power reduction. However, it is difficult to convert signals from core logic to input/output (I/O) circuits since core VDD is vastly different from high I/O supply voltage. In this work, we propose a level converter based on dynamic logic style for sub-threshold I/O part, having a large dynamic range of conversion. For the level converter, high voltage clock signal needs to be delivered through separate clock path from core logic, leading to clock synchronization problem between high voltage and low voltage clocks. To overcome this issue, we employed a Clock Synchronizer. A test chip is fabricated in 130-nm CMOS technology in order to verify the proposed technique. Hardware measurement results show that the level converter successfully converts 0.3 V 8 MHz pulse to 2.5 V signal. Ik Joon Chang, Jae-Joon Kim, Keejong Kim, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | SRAM Write-Ability Improvement With Transient Negative Bit-Line VoltageabstractIncreasing variations in device parameters significantly degrades the write-ability of SRAM cells in deep sub-100 nm CMOS technology. In this paper, a transient negative bit-line voltage technique is presented to improve write-ability of SRAM cell. Capacitive coupling is used to generate a transient negative voltage at the low-going bit-line during Write operation without using any on-chip or off-chip negative voltage source. Statistical simulations in a 45-nm PD/SOI technology show a 103X reduction in the Write-failure probability with the proposed method. Saibal Mukhopadhyay, Rahul M. Rao, Jae-Joon Kim, Ching-Te Chuang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | Self-Repairing SRAM Using On-Chip Detection and CompensationabstractIn nanometer scale static-RAM (SRAM) arrays, systematic inter-die and random within-die variations in process parameters can cause significant parametric failures, severely degrading parametric yield. In this paper, we investigate the interaction between the inter-die and intra-dieV tvariations on SRAM read and write failures. To improve the robustness of the SRAM cell, we propose a closed-loop compensation scheme using on-chip monitors that directly sense the global read stability and writability of the cell. Simulations based on 45-nm partially depleted silicon-on-insulator technology demonstrate the viability and the effectiveness of the scheme in SRAM yield enhancement. Niladri Narayan Mojumder, Saibal Mukhopadhyay, Jae-Joon Kim, Ching-Te Chuang, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2008 | Capacitive coupling based transient negative bit-line voltage (Tran-NBL) scheme for improving write-ability of SRAM design in nanometer technologiesabstractIncreasing process variation can significantly degrade the write-ability of an SRAM. In this paper, we propose negative bit- line voltage technique to improve cell write-ability without using any on-chip or off-chip negative voltage source. Capacitive coupling is used to generate a transient negative voltage at the low bit-line during write operation. Simulations in 45 nm PD/SOI technology show a 103times reduction in the write-failure probability with the proposed technique. Saibal Mukhopadhyay, Rahul M. Rao, Jae-Joon Kim, Ching-Te Chuang |
ISCAS | 3 |
| 2008 | Design and Analysis of a Self-Repairing SRAM with On-Chip Monitor and Compensation CircuitryabstractIn an SRAM array, the systematic inter-die and the random within-die variations in process parameters cause significant number of parametric failures, to degrade process yield in the nanometer technology regime. In this paper, we investigate the interaction between the inter-die and intra-die Vt variations on SRAM read and write failures. To improve robustness of SRAM cell, we propose a closed-loop compensation scheme using on-chip monitors that directly sense the global read stability and writability of the cell directly. Computer simulations based on 45nm PD/SOI technology demonstrate the viability and effectiveness of the scheme in SRAM yield enhancement. Niladri Narayan Mojumder, Saibal Mukhopadhyay, Jae-Joon Kim, Ching-Te Chuang, Kaushik Roy 0001 |
VTS | 3 |
| 2006 | Robust level converter design for sub-threshold logicabstractThe large supply voltage difference between sub-threshlold core logic and I/O makes it extremely challenging to convert signals from core circuit to I/O circuit. In this paper, we propose two novel circuits, Clock Synchronizer and Reduced Swing Inverter to design dynamic and static level converters for sub-threshold logic. Circuit simulations shows that our level converters work at frequency > 500Khz between 20®C and 40®C with a supply voltage of 0.25V. Ik Joon Chang, Jae-Joon Kim, Kaushik Roy 0001 |
ISLPED | 2 |
| 2006 | A Leakage-Tolerant Low-Swing Circuit Style in Partially Depleted Silicon-on-Insulator CMOS TechnologiesabstractThe parasitic bipolar leakage and the large subthreshold leakage due to high floating-body voltage reduce the noise margin and increase the delay of the circuits in the partially depleted silicon-on-insulator (PD/SOI). Differential cascode voltage switch logic (DCVSL) has circuit topologies susceptible to the leakage currents. In this paper, we propose a new circuit style to effectively handle the leakage problems in PD/SOI DCVSL. The proposed low-swing DCVSL (LS-DCVSL) uses the small internal swing to prevent the body of evaluation transistors from being charged to high voltage and, hence, suppress the leakages in DCVSL. Simulation results show that the proposed LS-DCVSL five-input XOR circuit is 33% faster than DCVSL five-input XOR circuit. In addition, the proposed circuit does not experience noise margin reduction due to pass-gate leakage. Jae-Joon Kim, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2005 | A forward body-biased low-leakage SRAM cache: device, circuit and architecture considerationsabstractThis paper presents a forward body-biasing (FBB) technique for active and standby leakage power reduction in cache memories. Unlike previous low-leakage SRAM approaches, we include device level optimization into the design. We utilize super high Vt (threshold voltage) devices to suppress the cache leakage power, while dynamically FBB only the selected SRAM cells for fast operation. In order to build a super high Vt device, the two-dimensional (2-D) halo doping profile was optimized considering various nanoscale leakage mechanisms. The transition latency and energy overhead associated with FBB was minimized by waking up the SRAM cells ahead of the access and exploiting the general cache access pattern. The combined device-circuit-architecture level techniques offer 64% total leakage reduction and 7.3% improvement in bit line delay compared to a previous state-of-the-art low-leakage SRAM technique. Static noise margin of the proposed SRAM cell is comparable to conventional SRAM cells. Chris H. Kim, Jae-Joon Kim, Saibal Mukhopadhyay, Kaushik Roy 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | A forward body-biased low-leakage SRAM cache: device and architecture considerationsabstractThis paper presents a forward body-biasing (FBB) scheme for active leakage power reduction in cache memories. We utilize super high VT (threshold voltage) devices to suppress the leakage power in unselected portions of a cache while fast operation is achieve by dynamically forward body-biasing the selected SRAM cells. In order to generate a super high VT device, the 2-D halo doping profile was optimized considering different nanometer regime leakage mechanisms. The transition latency and energy overhead associated with FBB could be minimized by (i) waking up the SRAM cells ahead of the access and (ii) exploiting the cache access pattern. The combined device-circuit-architecture level techniques offer 64% total leakage reduction and 7.3% improvement in bitline delay compared to a previous state-of-the-art low-leakage SRAM technique. Chris H. Kim, Jae-Joon Kim, Saibal Mukhopadhyay, Kaushik Roy 0001 |
ISLPED | 2 |