VLDB 2026 Research / reviewers in the wild / expert
Lee-Sup Kim
dblp:75/34
· DBLP profile ↗
112ranked-venue papers
1as first author
26since 2021 · last 2025
0000-0001-9585-4591ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 98 · 1 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14Applied, interdisciplinary, general and emerging computing · 5Software engineering, systems software and programming languages · 4 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FLAG: An FPGA-Based System for Low-Latency GNN Inference Service Using Vector QuantizationabstractEnabling real-time GNN inference services requires low end-to-end latency to meet service level agreements. However, intensive preparation steps and the neighborhood explosion problem pose significant challenges to efficient GNN inference serving. In this paper, we propose FLAG, an FPGA-based GNN inference serving system using vector quantization. To reduce preparation overhead, we introduce offline preprocessing to precompute and compress hidden embeddings for serving. A dedicated FPGA accelerator leverages the precomputed data to enable lightweight aggregation. As a result, FLAG achieves average speedups of $154 \times 176 \times$, and $333 \times$ on three GNN models compared to the baseline system. Yunki Han, Taehwan Kim 0011, Seohye Ha, Lee-Sup Kim |
DAC | 5 |
| 2025 | SAFF: Scalable Acceleration of GNN-based Machine Learning Force Fields using Tensor-Aware Hardware for Molecular SimulationabstractMachine Learning Force Fields (MLFFs) have emerged as a key technique in molecular dynamics (MD) simulations, enabling accurate prediction of molecular energies and forces with a level of precision comparable to that of Density Functional Theory (DFT), while significantly reducing computational cost. However, applying MLFFs to real-world scenarios, such as semiconductor process simulations involving massive atomic scales and long simulation times, exposes severe performance limitations, resulting in current MLFF models being inadequate for practical use and emphasizing the necessity of significant computational acceleration. Graph Neural Networks (GNNs) are frequently utilized to construct MLFFs, with tensor product operations constituting a substantial portion of the computational workload. These operations involve high-dimensional tensors whose sizes are determined by the number of channels and the order of rotation. This results in frequent memory accesses and low data reuse. This leads to memory bottlenecks that limit the efficient use of GPU computational resources. In this work, we propose a tensor-based function fusion technique to improve data continuity between functions, thereby increasing on-chip data reuse and reducing off-chip memory access. Furthermore, we identified inefficient utilization of hardware resources and addressed this by restructuring the computation flow and removing redundant operations to enhance computational performance. To optimize performance, we have also designed a dedicated tensor product hardware architecture that is optimized for GNN-based MLFFs. The experimental results demonstrate that the proposed system achieves an average speedup of 6.26× and a 424× improvement in energy efficiency compared to GPU-based execution, while maintaining the original model accuracy. Seohye Ha, Yunki Han, Taehwan Kim 0011, Gunhee Park, Lee-Sup Kim |
ICCAD | 7 |
| 2025 | EOD: Enabling Low Latency GNN Inference via Near-Memory Concatenate AggregationabstractAs online services based on graph databases increasingly integrate with machine learning, serving low-latency Graph Neural Network (GNN) inference for individual requests has become a critical challenge.Real-time GNN inference services operate in an inductive setup, which can handle newly added, previously unseen nodes and their edges.In this setup, the system must prepare the computational graph and input node features for target nodes, followed by GNN inference using the given input data.However, the workflow of a GNN serving system presents two key challenges that hinder low-latency inference.The first challenge arises from the extensive preparation step, which involves heavy memory access in host memory and significant data I/O to devices, constituting the largest portion of end-to-end inference latency.The second challenge is the well-known neighborhood explosion problem in GNN research.As the receptive field for target nodes increases exponentially with the number of layers, this issue exacerbates overall latency.To address these challenges, we propose a Near-Memory Processing (NMP) based low-latency GNN inference serving system named EOD.To ensure low-latency real-time GNN service, we co-design the algorithm and hardware to tackle the aforementioned issues.First, to mitigate the neighborhood explosion problem, we propose a precomputation method for the training node set, reducing memory access and computational complexity from exponential to linear growth.Additionally, we introduce a concatenated ZVC compression method to minimize the overhead of storing precomputed hidden features.Finally, to alleviate heavy host-side memory access and data I/O, we design an NMP architecture that enables efficient aggregation on concatenated ZVC-compressed data.As a result, EOD achieves a geometric mean of 981.1× and 912.0× aggregation speedup over the baseline and the existing architecture for GNN aggregation.Additionally, EOD achieves a geometric mean of 17.9×, and up to 74.3× end-to-end latency speedup over the GPU baseline. Taehwan Kim 0011, Yunki Han, Seohye Ha, Lee-Sup Kim |
ISCA | 5 |
| 2025 | AToM: Adaptive Token Merging for Efficient Acceleration of Vision TransformerabstractRecently, Vision Transformers (ViTs) have set a new standard in computer vision (CV), showing unparalleled image processing performance. However, their substantial computational requirements hinder practical deployment, especially on resource-limited devices common in CV applications. Token merging has emerged as a solution, condensing tokens with similar features to cut computational and memory demands. Yet, existing applications on ViTs often miss the mark in token compression, with rigid merging strategies and a lack of in-depth analysis of ViT merging characteristics. To overcome these issues, this paper introduces Adaptive Token Merging (AToM), a comprehensive algorithm-architecture co-design for accelerating ViTs. The AToM algorithm employs an image-adaptive, fine-grained merging strategy, significantly boosting computational efficiency. We also optimize the merging and unmerging processes to minimize overhead, employing techniques like First-Come-First-Merge mapping and Linear Distance Calculation. On the hardware side, the AToM architecture is tailor-made to exploit the AToM algorithm's benefits, with specialized engines for efficient merge and unmerge operations. Our pipeline architecture ensures end-to-end ViT processing, minimizing latency and memory overhead from the AToM algorithm. Across various hardware platforms including CPU, EdgeGPU, and GPU, AToM achieves average end-to-end speedups of 10.9$\boldsymbol{\times}$, 7.7$\boldsymbol{\times}$, and 5.4$\boldsymbol{\times}$, alongside energy savings of 24.9$\boldsymbol{\times}$, 1.8$\boldsymbol{\times}$, and 16.7$\boldsymbol{\times}$. Moreover, AToM offers 1.2$\boldsymbol{\times}$1.9$\boldsymbol{\times}$higher effective throughput compared to existing transformer accelerators. Jaekang Shin, Myeonggu Kang, Yunki Han, Lee-Sup Kim |
IEEE Trans. Computers | 5 |
| 2024 | Token-Picker: Accelerating Attention in Text Generation with Minimized Memory Transfer via Probability EstimationabstractThe attention mechanism in text generation is memory-bounded due to its sequential characteristics. Therefore, off-chip memory accesses should be minimized for faster execution. Although previous methods addressed this by pruning unimportant tokens, they fall short in selectively removing tokens with near-zero attention probabilities in each instance. Our method estimates the probability before the softmax function, effectively removing low probability tokens and achieving an 12.1x pruning ratio without fine-tuning. Additionally, we present a hardware design supporting seamless on-demand off-chip access. Our approach shows 2.6x reduced memory accesses, leading to an average 2.3x speedup and a 2.4x energy efficiency. Myeonggu Kang, Yunki Han, Yanggon Kim 0001, Jaekang Shin, Lee-Sup Kim |
DAC | 6 |
| 2024 | CoCoA: Algorithm-Hardware Co-Design for Large-Scale GNN Training using Compressed GraphabstractScaling Graph Neural Network (GNN) training on large-scale graph data poses a critical challenge for implementing GNN applications on real-world giant graphs. The size of real-world graphs often exceeds the memory capacity of accelerator devices, necessitating the use of multiple devices or host memory for training. While expanding memory space alleviates the out-of-memory problem, this approach introduces another bottleneck through heavy communication via low-bandwidth interconnection. Therefore, achieving efficient, scalable GNN training on large-scale graphs requires addressing both capacity and communication issues. Yunki Han, Jaekang Shin, Gunhee Park, Lee-Sup Kim |
ICCAD | 4 |
| 2024 | ToEx: Accelerating Generation Stage of Transformer-Based Language Models via Token-Adaptive Early ExitabstractTransformer-based language models have recently gained popularity in numerous natural language processing (NLP) applications due to their superior performance compared to traditional algorithms. These models involve two execution stages: summarization and generation. The generation stage accounts for a significant portion of the total execution time due to its auto-regressive property, which necessitates considerable and repetitive off-chip accesses. Consequently, our objective is to minimize off-chip accesses during the generation stage to expedite transformer execution. To achieve the goal, we propose a token-adaptive early exit (ToEx) that generates output tokens using fewer decoders, thereby reducing off-chip accesses for loading weight parameters. Although our approach has the potential to minimize data communication, it brings two challenges: 1) inaccurate self-attention computation, and 2) significant overhead for exit decision. To overcome these challenges, we introduce a methodology that facilitates accurate self-attention by lazily performing computations for previously exited tokens. Moreover, we mitigate the overhead of exit decision by incorporating a lightweight output embedding layer. We also present a hardware design to efficiently support the proposed work. Evaluation results demonstrate that our work can reduce the number of decoders by 2.6× on average. Accordingly, it achieves 3.2× speedup on average compared to transformer execution without our work. Myeonggu Kang, Hyein Shin, Jaekang Shin, Lee-Sup Kim |
IEEE Trans. Computers | 5 |
| 2024 | ADC-Free ReRAM-Based In-Situ Accelerator for Energy-Efficient Binary Neural NetworksabstractWith the ever-increasing parameter size of deep learning models, conventional ASIC-based accelerators in mobile environments suffer from low energy budget due to limited memory capacity and frequent data movements. Binary neural networks (BNNs) deployed in ReRAM-based in-situ accelerators provide a promising solution, and various related architectures have been proposed recently. However, their performances are largely compromised by the tremendous cost of domain conversion via analog-to-digital converters (ADCs), essential for mixed-signal processing in ReRAM. This paper identifies two root causes of the need for such ADCs and proposes effective solutions to address them. First, we minimize redundant operations in BNNs and reduce the number of ReRAM arrays with ADCs approximately by half. We also propose a partial-sum range adjustment technique based on a layer remapping to deal with the remaining ADCs. Proper handling of the partial-sum distribution allows ReRAM-based in-situ processing without domain conversion, completely bypassing the need for ADCs. Experimental results show that the proposed architecture achieves a 3.44x speedup and 91.5% energy savings, making it an attractive solution for on-device AI at the edge. Hyeonuk Kim, Youngbeom Jung, Lee-Sup Kim |
IEEE Trans. Computers | 3 |
| 2023 | OptimStore: In-Storage Optimization of Large Scale DNNs with On-Die ProcessingabstractTraining deep neural network (DNN) models is a resource-intensive, iterative process. For this reason, nowadays, complex optimizers like Adam are widely adopted as it increases the speed and efficiency of training. These optimizers, however, employ additional variables and raise the memory demand 2× to 3× of model parameters, worsening the memory capacity bottleneck. Moreover, as the size of DNN models is projected to grow even further, it is not practical to assume that the future models will fit in accelerator memory. This has triggered various efforts to offload models to flash-based storage. However, when the model, especially the optimizer, is offloaded to flash, the limited I/O bandwidth severely slows down the overall training process. To this end, we present OptimStore, a solid-state drive (SSD) system with on-die processing (ODP) architectures for gradient descent-based machine learning models. OptimStore accelerates the training process of such large-scale models by processing model optimization in the storage device, specifically inside the flash dies. ODP capability of OptimStore eliminates the heavy data movement over external interconnect and internal flash channels. Overall, OptimStore achieves, on average, a 2.8× speedup and a 3.6× improved energy efficiency in the weight update stage over baseline SSD offloading. Junkyum Kim, Myeonggu Kang, Yunki Han, Yanggon Kim 0001, Lee-Sup Kim |
HPCA | 5 |
| 2023 | MGen: A Framework for Energy-Efficient In-ReRAM Acceleration of Multi-Task BERTabstractRecently, multiple transformer models, such as BERT, have been utilized together to support multiple natural language processing (NLP) tasks in a system, also known as multi-task BERT. Multi-task BERT with very high weight parameters increases the area requirement of a processing in resistive memory (ReRAM) architecture, and several works have attempted to address this model size issue. Despite the reduced parameters, the number of multi-task BERT computations remains the same, leading to massive energy consumption in ReRAM-based deep neural network (DNN) accelerators. Therefore, we suggest a framework for better energy efficiency during the ReRAM acceleration of multi-task BERT. First, we analyze the inherent redundancies of multi-task BERT and the computational properties of the ReRAM-based DNN accelerator, after which we propose what is termed the model generator, which produces optimal BERT models supporting multiple tasks. The model generator reduces multi-task BERT computations while maintaining the algorithmic performance. Furthermore, we present task scheduler, which adjusts the execution order of multiple tasks, to run the produced models efficiently. As a result, the proposed framework achieves maximally 4.4× higher energy efficiency over the baseline, and it can also be combined with the previous multi-task BERT works to achieve both a smaller area and higher energy efficiency. Myeonggu Kang, Hyein Shin, Junkyum Kim, Lee-Sup Kim |
IEEE Trans. Computers | 4 |
| 2023 | Fault-Free: A Framework for Analysis and Mitigation of Stuck-at-Fault on Realistic ReRAM-Based DNN AcceleratorsabstractResistive RAM(ReRAM) is gaining attention as a suitable memory platform for accelerating deep neural networks(DNNs) in an energy-efficient way. However, energy-efficient ReRAM-based DNN accelerators suffer from serious Stuck-At-Fault(SAF) issues that significantly degrade the inference accuracy. SAF is a device-level non-ideality, and the problems of SAF worsen in the realistic ReRAM with low cell resolution. To address the problem in the realistic ReRAM, we present a framework for mitigating SAF on ReRAM-based accelerators(Fault-free). We first analyze the impact of SAF on low-resolution cells. Based on the analysis, we present offline compilation, which drastically reduces the impact of SAF on inference accuracy. At the first stage, we extract indices of distorted weights due to SAF. For extracted weights, fault-aware weight decomposition and closest value mapping are applied to minimize the error of weights. In the online phase, the target DNN model is executed on the ReRAM-based accelerator along with lightweight compensation units. The online compensation is selectively performed for a small portion of weights to reduce the hardware overhead. With the proposed framework, the ReRAM-based accelerator successfully ensures the inference accuracy of various DNN models with an average area of 5% and an energy overhead of 0.8% from an ideal ReRAM-based accelerator. Hyein Shin, Myeonggu Kang, Lee-Sup Kim |
IEEE Trans. Computers | 3 |
| 2023 | Accelerating On-Device DNN Training Workloads via Runtime Convergence MonitorabstractWith the growing demand for processing deep learning applications on edge devices, on-device DNN training has become a major workload to execute a variety of vision tasks suited for users. Therefore, architectures employed with algorithm co-design to accelerate the training process have been steadily studied. However, previous solutions are mostly supported by extended versions of the inference studies, such as sparsity, data flow, quantization, etc. Moreover, most works examine their schemes on from-the-scratch training that cannot tolerate inaccurate computing. Accordingly, there are still factors that hinder the overall speed of the DNN training process that has not been addressed in practical workloads. In this work, we propose a runtime convergence monitor to achieve massive computational savings in the practical on-device training workloads (i.e., transfer-learning-based task adaptation). By monitoring the network output data, we determine the training intensity of incoming tasks and adaptively detect the convergence in iteration intervals for training diverse datasets. Furthermore, we enable the computation skip of converged images determined by the monitored prediction probability to enhance the training speed within an iteration. As a result, we perform an accurate but fast convergence in model training for the task adaptation with minimal overhead. Unlike the previous approximation methods, our monitoring system enables runtime optimization and can be easily applicable to any type of accelerator attaining significant speedup. Evaluation results on various datasets show geomean of$2.2\times $speedup when applied in any systolic architectures and further enhancement of$3.6\times $when applied in accelerators dedicated for on-device training. Seungkyu Choi, Jaekang Shin, Lee-Sup Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Energy-Efficient CNN Personalized Training by Adaptive Data ReformationabstractTo adopt deep neural networks in resource-constrained edge devices, various energy- and memory-efficient embedded accelerators have been proposed. However, most off-the-shelf networks are well trained with vast amounts of data, but unexplored users’ data or accelerator’s constraints can lead to unexpected accuracy loss. Therefore, a network adaptation suitable for each user and device is essential to make a high confidence prediction in given environment. We propose simple but efficient data reformation methods that can effectively reduce the communication cost with off-chip memory during the adaptation. Our proposal utilizes the data’s zero-centered distribution and spatial correlation to concentrate the sporadically spread bit-level zeros to the units of value. Consequently, we reduced the communication volume by up to 55.6% per task with an area overhead of 0.79% during the personalization training. Youngbeom Jung, Hyeonuk Kim, Seungkyu Choi, Jaekang Shin, Lee-Sup Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Algorithm/architecture co-design for energy-efficient acceleration of multi-task DNNabstractReal-world AI applications, such as augmented reality or autonomous driving, require processing multiple CV tasks simultaneously. However, the enormous data size and the memory footprint have been a crucial hurdle for deep neural networks to be applied in resource-constrained devices. To solve the problem, we propose an algorithm/architecture co-design. The proposed algorithmic scheme, named SqueeD, reduces per-task weight and activation size by 21.9x and 2.1x, respectively, by sharing those data between tasks. Moreover, we design architecture and dataflow to minimize DRAM access by fully utilizing benefits from SqueeD. As a result, the proposed architecture reduces the DRAM access increment and energy consumption increment per task by 2.2x and 1.3x, respectively. Jaekang Shin, Seungkyu Choi, Jongwoo Ra, Lee-Sup Kim |
DAC | 4 |
| 2022 | Re2fresh: A Framework for Mitigating Read Disturbance in ReRAM-Based DNN AcceleratorsabstractA severe read disturbance problem degrades the inference accuracy of a resistive RAM (ReRAM) based deep neural network (DNN) accelerator. Refresh, which reprograms the ReRAM cells, is the most obvious solution for the problem, but programming ReRAM consumes huge energy. To address the issue, we first analyze the resistance drift pattern of each conductance state and the actual read stress applied to the ReRAM array by considering the characteristics of ReRAM-based DNN accelerators. Based on the analysis, we cluster ReRAM cells into a few groups for each layer of DNN and generate a proper refresh cycle for each group in the offline phase. The individual refresh cycles reduce energy consumption by reducing the number of unnecessary refresh operations. In the online phase, the refresh controller selectively launches refresh operations according to the generated refresh cycles. ReRAM cells are selectively refreshed by minimally modifying the conventional structure of the ReRAM-based DNN accelerator. The proposed work successfully resolves the read disturbance problem by reducing 97% of the energy consumption for the refresh operation while preserving inference accuracy. Hyein Shin, Myeonggu Kang, Lee-Sup Kim |
ICCAD | 3 |
| 2022 | A Deep Neural Network Training Architecture With Inference-Aware Heterogeneous Data-TypeabstractAs deep learning applications often encounter accuracy degradation due to the distorted inputs from a variety of environmental conditions, training with personal data has become essential for the edge devices. Hence, ‘training on edge’ by supporting a trainable deep learning accelerator has been actively studied. Nevertheless, previous research does not consider the fundamental datapath for training and the importance of retaining the high performance for inference tasks. In this work, we propose NeuroFlix, a deep neural network training accelerator supporting heterogeneous data-type of floating- and fixed-point for input operands. From two perspectives: 1) separate precision decision for each input data, 2) maintenance of high performance on inference, we configure the data with low-bit fixed-point of activation/weight and floating-point based error gradient securing up to half-precision. A novel MAC architecture is designed to compute low- and high-precision modes for the different input combinations. By substituting a high-cost floating-point based addition to brick-level separate accumulations, we realize both area-efficient architecture and high throughput for low-precision computation. Consequently, NeuroFlix outperforms the previous architectures of state-of-the-art configurations proving its high efficiency in both training and inference. By also comparing with the off-the-shelf bfloat16-based accelerator, it achieves 1.2 ×/2.0 × of speedup/energy-efficiency at training and further enhancement of 3.6 ×/4.5 × at inference. Seungkyu Choi, Jaekang Shin, Lee-Sup Kim |
IEEE Trans. Computers | 3 |
| 2022 | EGCN: An Efficient GCN Accelerator for Minimizing Off-Chip Memory AccessabstractAs Graph Convolutional Networks (GCNs) have emerged as a promising solution for graph representation learning, designing specialized GCN accelerators has become an important challenge. An analysis of GCN workloads shows that the main bottleneck of GCN processing is not computation but the memory latency of intensive off-chip data transfer. Therefore, minimizing off-chip data transfer is the primary challenge for designing an efficient GCN accelerator. To address this challenge, optimization is initialized by considering GCNs as tiled matrix multiplication. In this paper, we optimize off-chip memory access from both the in- and out-of-tile perspectives. From the out-of-tile perspective, we find optimal tile configurations of given datasets and on-chip buffer capacity, then observe the dataflow across phases and layers. Inter-layer phase fusion dataflow with optimal tile configuration reduces data transfer of intermediate outputs. From the in-tile perspective, due to the sparsity of tiles, tiles have redundant data which does not participate in computation. Redundant data load is eliminated with hardware support. Finally, we introduce an efficient GCN inference accelerator, EGCN, specialized for minimizing off-chip memory access. EGCN achieves 41.9% off-chip DRAM access reduction, 1.49× speedup, and 1.95× energy efficiency improvement on average over the state-of-the-art accelerators. Yunki Han, Kangkyu Park, Youngbeom Jung, Lee-Sup Kim |
IEEE Trans. Computers | 4 |
| 2022 | S-FLASH: A NAND Flash-Based Deep Neural Network Accelerator Exploiting Bit-Level SparsityabstractThe processing in-memory (PIM) approach that combines memory and processor appears to solve the memory wall problem. NAND flash memory, which is widely adopted in edge devices, is one of the promising platforms for PIM with its high-density property and the intrinsic ability for analog vector-matrix multiplication. Despite its potential, the domain conversion process, which converts an analog current to a digital value, accounts for most energy consumption on the NAND flash-based accelerator. It restricts the NAND flash memory usage for PIM compared to the other platforms. In this paper, we propose a NAND flash-based DNN accelerator to achieve both large memory density and energy efficiency among various platforms. As the NAND flash memory already shows higher memory density than other memory platforms, we aim to enhance energy efficiency by reducing the domain conversion process burden. Firstly, we optimize the bit width of partial multiplication by considering the analog-to-digital converter (ADC) resource. For further optimization, we propose a methodology to exploit many zero partial multiplication results for enhancing both energy efficiency and throughput. The proposed work successfully exploits the bit-level sparsity of DNN, which results in achieving up to 8.6/8.2 larger energy efficiency/throughput over the provisioned baseline. Myeonggu Kang, Hyeonuk Kim, Hyein Shin, Jaehyeong Sim, Kyeonghan Kim, Lee-Sup Kim |
IEEE Trans. Computers | 6 |
| 2022 | Rare Computing: Removing Redundant Multiplications From Sparse and Repetitive Data in Deep Neural NetworksabstractRecent research shows that 4-bit data precision is sufficient for Deep Neural Network (DNN) inference without accuracy degradation. Due to the low bit-width, a large amount of data is repeated. In this article, we propose a hardware architecture, named Rare Computing Architecture (RCA), that skips redundant computations due to repetitive data in the networks. By exploiting redundancy, RCA is not significantly affected by data-sparsity and maintains great improvements in performance and energy efficiency, while the improvements of existing DNN accelerators are vulnerable to variations in sparsity. In the RCA, repeated data in a window for censoring repetition are detected by a Redundancy Censoring Unit (RCU) and processed at a time, achieving high effective throughput. Additionally, we present a dataflow that exploits abundant data-reusability in DNNs, which enables the high-throughput computations to be ceaselessly performed without an increase of bandwidth for data-read. The proposed architecture is evaluated in two ways of exploiting weight- and activation-repetition. In the evaluation, RCA is compared to a value-agnostic computation and UCNN that is the state-of-the-art accelerator exploiting weight-repetition. Additionally, RCA is compared to Bit-pragmatic that exploits bit-level sparsity. Both evaluations demonstrate that the RCA shows steadily high improvements in performance and energy-efficiency. Kangkyu Park, Seungkyu Choi, Yeongjae Choi, Lee-Sup Kim |
IEEE Trans. Computers | 4 |
| 2022 | A Framework for Accelerating Transformer-Based Language Model on ReRAM-Based ArchitectureabstractTransformer-based language models have become thede-factostandard model for various natural language processing (NLP) applications given the superior algorithmic performances. Processing a transformer-based language model on a conventional accelerator induces the memory wall problem, and the ReRAM-based accelerator is a promising solution to this problem. However, due to the characteristics of the self-attention mechanism and the ReRAM-based accelerator, the pipeline hazard arises when processing the transformer-based language model on the ReRAM-based accelerator. This hazard issue greatly increases the overall execution time. In this article, we propose a framework to resolve the hazard issue. First, we propose the concept of window self-attention to reduce the attention computation scope by analyzing the properties of the self-attention mechanism. After that, we present a window-size search algorithm, which finds an optimal window size set according to the target application/algorithmic performance. We also suggest a hardware design that exploits the advantages of the proposed algorithm optimization on the general ReRAM-based accelerator. The proposed work successfully alleviates the hazard issue while maintaining the algorithmic performance, leading to a$5.8\times $speedup over the provisioned baseline. It also delivers up to$39.2\times /643.2\times $speedup/higher energy efficiency over GPU, respectively. Myeonggu Kang, Hyein Shin, Lee-Sup Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Fault-free: A Fault-resilient Deep Neural Network Accelerator based on Realistic ReRAM DevicesabstractEnergy-efficient Resistive RAM (ReRAM) based deep neural network (DNN) accelerator suffers from severe Stuck-At-Fault (SAF) problem that drastically degrades the inference accuracy. The SAF problem gets even worse in realistic ReRAM devices with low cell resolution. To address the issue, we propose a fault-resilient DNN accelerator based on realistic ReRAM devices. We first analyze the SAF problem in a realistic ReRAM device and propose a 3-stage offline fault-resilient compilation and lightweight online compensation. The proposed work enables the reliable execution of DNN with only 5% area and 0.8% energy overhead from the ideal ReRAM-based DNN accelerator. Hyein Shin, Myeonggu Kang, Lee-Sup Kim |
DAC | 3 |
| 2021 | Optimizing ADC Utilization through Value-Aware Bypass in ReRAM-based DNN AcceleratorabstractReRAM-based Processing-In-Memory (PIM) has been widely studied as a promising approach for Deep Neural Networks (DNN) accelerator with its energy-efficient analog operations. However, the domain conversion process for the analog operation requires frequent accesses to power-hungry Analog-to-Digital Converter (ADC), hindering the overall energy efficiency. Although previous research has been suggested to address this problem, the ADC cost has not been sufficiently reduced because of its unsuitable approach for ReRAM. In this paper, we propose mixed-signal-based value-aware bypass techniques to optimize the ADC utilization of the ReRAM-based PIM. By utilizing the property of bit-line (BL) level value distribution, the proposed work bypasses the redundant ADC operations depending on the magnitude of value. Evaluation results show that our techniques successfully reduce ADC access and improve overall energy efficiency by 2.48 × -3.07 × compared to ISAAC. HanCheon Yun, Hyein Shin, Myeonggu Kang, Lee-Sup Kim |
DAC | 4 |
| 2021 | A Convergence Monitoring Method for DNN Training of On-Device Task AdaptationabstractDNN training has become a major workload in on-device situations to execute various vision tasks with high performance. Accordingly, training architectures accompanying approximate computing have been steadily studied for efficient acceleration. However, most of the works examine their scheme on from-the-scratch training where inaccurate computing is not tolerable. Moreover, previous solutions are mostly provided as an extended version of the inference works, e.g., sparsity/pruning, quantization, dataflow, etc. Therefore, unresolved issues in practical workloads that hinder the total speed of the DNN training process remain still. In this work, with targeting the transfer learning-based task adaptation of the practical on-device training workload, we propose a convergence monitoring method to resolve the redundancy in massive training iterations. By utilizing the network's output value, we detect the training intensity of incoming tasks and monitor the prediction convergence with the given intensity to provide early-exits in the scheduled training iteration. As a result, an accurate approximation over various tasks is performed with minimal overhead. Unlike the sparsity-driven approximation, our method enables runtime optimization and can be easily applicable to off-the-shelf accelerators achieving significant speedup. Evaluation results on various datasets show a geomean of$2.2\times$speedup over baseline and$1.8\times$speedup over the latest convergence-related training method. Seungkyu Choi, Jaekang Shin, Lee-Sup Kim |
ICCAD | 3 |
| 2021 | A Framework for Area-efficient Multi-task BERT Execution on ReRAM-based AcceleratorsabstractWith the superior algorithmic performances, BERT has become the de-facto standard model for various NLP tasks. Accordingly, multiple BERT models have been adopted on a single system, which is also called multi-task BERT. Although the ReRAM-based accelerator shows the sufficient potential to execute a single BERT model by adopting in-memory computation, processing multi-task BERT on the ReRAM-based accelerator extremely increases the overall area due to multiple fine-tuned models. In this paper, we propose a framework for area-efficient multi-task BERT execution on the ReRAM-based accelerator. Firstly, we decompose the fine-tuned model of each task by utilizing the base-model. After that, we propose a two-stage weight compressor, which shrinks the decomposed models by analyzing the properties of the ReRAM-based accelerator. We also present a profiler to generate hyper-parameters for the proposed compressor. By sharing the base-model and compressing the decomposed models, the proposed framework successfully reduces the total area of the ReRAM-based accelerator without an additional training procedure. It achieves a 0.26 x area than baseline while maintaining the algorithmic performances. Myeonggu Kang, Hyein Shin, Jaekang Shin, Lee-Sup Kim |
ICCAD | 4 |
| 2021 | Deferred Dropout: An Algorithm-Hardware Co-Design DNN Training Method Provisioning Consistent High Activation SparsityabstractThis paper proposes a deep neural network training method that provisions consistent high activation sparsity and the ability to adjust the sparsity. To improve training performance, prior work reduces the memory footprint for training by exploiting input activation sparsity which is observed due to the ReLU function. However, the previous approach relies solely on the inherent sparsity caused by the function, and thus the footprint reduction is not guaranteed. In particular, models for natural language processing tasks like BERT do not use the function, so the models have almost zero activation sparsity and the previous approach loses its efficiency. In this paper, a new training method, Deferred Dropout, and its hardware architecture are proposed. With the proposed method, input activations are dropped out after the conventional forward-pass computation. In contrast to the conventional dropout where activations are zeroed before forward-pass computation, the dropping timing is deferred until the completion of the computation. Then, the sparsified activations are compressed and stashed in memory. This approach is based on our observation that networks preserve training quality even if only a few high magnitude activations are used in the backward pass. The hardware architecture enables designers to exploit the tradeoff between training quality and activation sparsity. Evaluation results demonstrate that the proposed method achieves 1.21-3.60 × memory footprint reduction and 1.06-1.43 x speedup on the TPUv3 architecture, compared to the prior work. Kangkyu Park, Yunki Han, Lee-Sup Kim |
ICCAD | 3 |
| 2021 | Amnesiac DRAM: A Proactive Defense Mechanism Against Cold Boot AttacksabstractDRAMs in modern computers or hand-held devices store private or often security-sensitive data. Unfortunately, one known attack vector, called a cold boot attack, remains threatening and easy-to-exploit, especially when attackers have physical access to the device. It exploits the fundamental property of current DRAMs: remanence effects that retain the stored contents for a certain period of time even after powering off. To magnify the remanence effect, cold boot attacks typically freeze the victim DRAM, thereby providing a chance to detach, move, and reattach it to an attacker's computer. Once power is on, attackers can steal all the security-critical information from the victim's DRAM, such as a master decryption key for an encrypted disk storage. Two types of defenses were proposed in the past: 1) CPU-bound cryptography, where keys are stored in CPU registers and caches instead of in DRAMs, and 2) full or partial memory encryption, where sensitive data are stored encrypted. However, both methods impose non-negligible performance or energy overheads to the running systems, and worse, significantly increase the hardware and software manufacturing costs. We found that these proposed solutions attempted to address the cold boot attacks passively: either by avoiding or by indirectly addressing the root cause of the problem, the remanence effect. In this article, we propose and evaluate a proactive defense mechanism, Amnesiac DRAM, that comprehensively prevents the cold boot attacks. The key idea is to discard the contents in the DRAM when attackers attempt to retrieve (i.e., power on) them from the stolen DRAM. When Amnesiac DRAM senses a physical separation, it locks itself and deletes all the remaining contents, making it amnesiac. The Amnesiac DRAM causes neither performance nor energy overhead in ordinary operations (e.g., load and store) and can be easily implemented with negligible area overhead in commodity DRAM architectures. Hoseok Seol, Minhye Kim, Taesoo Kim, Yongdae Kim, Lee-Sup Kim |
IEEE Trans. Computers | 5 |
| 2020 | A Pragmatic Approach to On-device Incremental Learning System with Selective Weight UpdatesabstractIncremental learning is drawing attention to widen capabilities of device-AI. Previous works have researched to reduce numerous computations and memory accesses required for the training process of IL, but they could not show a noticeable improvement in the weight gradient computation (WGC) phase. Therefore, we propose a selective weight update technique that searches for critical weights to be updated by applying the IL algorithm that training per-task binary masks. Also, we introduce a novel dataflow for the implementation of selective WGC on typical NPUs with minimum overheads. On average, our system shows a 2.9× speed up and 2.5× energy efficiency in WGC without degrading training quality. Jaekang Shin, Seungkyu Choi, Yeongjae Choi, Lee-Sup Kim |
DAC | 4 |
| 2020 | A Thermal-aware Optimization Framework for ReRAM-based Deep Neural Network AccelerationabstractResistive RAM (ReRAM) is widely regarded as a promising platform for deep neural network (DNN) acceleration. However, the ReRAM device suffers from severe thermal problems that degrade the lifetime and inference accuracy of the ReRAM-based DNN accelerator. To address the issues, we propose a thermal-aware optimization framework for accelerating DNN on ReRAM (TOPAR). TOPAR includes 3-stage offline thermal optimization and online thermal-aware error compensation. Offline thermal optimization consists of thermal-aware weight decomposition, thermal-aware column reordering, and fine-grained weight adjustment to reduce the temperature of the ReRAM-based DNN accelerator. For online thermal-aware error compensation, we compensate conductance change according to the temperature variation. With TOPAR, the endurance degradation due to temperature rise improves up to 2.39×, and inference accuracy is preserved without harming the performance of the ReRAM-based DNN accelerator. Hyein Shin, Myeonggu Kang, Lee-Sup Kim |
ICCAD | 3 |
| 2020 | An Energy-Efficient Deep Convolutional Neural Network Inference Processor With Enhanced Output Stationary Dataflow in 65-nm CMOSabstractWe propose a deep convolutional neural network (CNN) inference processor based on a novel enhanced output stationary (EOS) dataflow. Based on the observation that some activations are commonly used in two successive convolutions, the EOS dataflow employs dedicated register files (RFs) for storing such reused activation data to eliminate redundant memory accesses for highly energy-consuming SRAM banks. In addition, processing elements (PEs) are split into multiple small groups such that each group covers a tile of input activation map to increase the usability of activation RFs (ARFs). The processor has two different voltage/frequency domains. The computation domain with 512 PEs operates at near-threshold voltage (NTV) (0.4 V) and 60-MHz frequency to increase energy efficiency, while the rest of the processors including 848-KB SRAMs run at 0.7 V and 120-MHz frequency to increase both on-chip and off-chip memory bandwidths. The measurement results show that our processor is capable of running AlexNet at 831 GOPS/W, VGG-16 at 1151 GOPS/W, ResNet-18 at 1004 GOPS/W, and MobileNet at 948 GOPS/W energy efficiency. Jaehyeong Sim, Somin Lee, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | An Optimized Design Technique of Low-bit Neural Network Training for Personalization on IoT DevicesabstractPersonalization by incremental learning has become essential for IoT devices to enhance the performance of the deep learning models trained with global datasets. To avoid massive transmission traffic in the network, exploiting on-device learning is necessary. We propose a software/hardware co-design technique that builds an energy-efficient low-bit trainable system: (1) software optimizations by local low-bit quantization and computation freezing to minimize the on-chip storage requirement and computational complexity, (2) hardware design of a bit-flexible multiply-and-accumulate (MAC) array sharing the same resources in inference and training. Our scheme saves 99.2% on on-chip buffer storage and achieves 12.8x higher peak energy efficiency compared to previous trainable accelerators. Seungkyu Choi, Jaekang Shin, Yeongjae Choi, Lee-Sup Kim |
DAC | 4 |
| 2019 | NAND-Net: Minimizing Computational Complexity of In-Memory Processing for Binary Neural NetworksabstractPopular deep learning technologies suffer from memory bottlenecks, which significantly degrade the energy-efficiency, especially in mobile environments. In-memory processing for binary neural networks (BNNs) has emerged as a promising solution to mitigate such bottlenecks, and various relevant works have been presented accordingly. However, their performances are severely limited by the overheads induced by the modification of the conventional memory architectures. To alleviate the performance degradation, we propose NAND-Net, an efficient architecture to minimize the computational complexity of in-memory processing for BNNs. Based on the observation that BNNs contain many redundancies, we decomposed each convolution into sub-convolutions and eliminated the unnecessary operations. In the remaining convolution, each binary multiplication (bitwise XNOR) is replaced by a bitwise NAND operation, which can be implemented without any bit cell modifications. This NAND operation further brings an opportunity to simplify the subsequent binary accumulations (popcounts). We reduced the operation cost of those popcounts by exploiting the data patterns of the NAND outputs. Compared to the prior state-of-the-art designs, NAND-Net achieves 1.04-2.4x speedup and 34-59% energy saving, thus making it a suitable solution to implement efficient in-memory processing for BNNs. Hyeonuk Kim, Jaehyeong Sim, Yeongjae Choi, Lee-Sup Kim |
HPCA | 4 |
| 2019 | eSRCNN: A Framework for Optimizing Super-Resolution Tasks on Diverse Embedded CNN AcceleratorsabstractCNN-based Super-Resolution (SR), the most representative of low-level vision task, is a promising solution to improve users' QoS on IoT devices that suffer from limited network bandwidth and storage capacity by effectively enhancing image/video resolution. Although prior accelerators to embed CNN show tremendous performance and energy efficiency, they are not suitable for SR tasks regarding off-chip memory accesses. In this work, we present eSRCNN, a framework that enables performing energy-efficient SR tasks on diverse embedded CNN accelerators by decreasing off-chip memory accesses. To reduce off-chip memory accesses, our framework consists of three steps: a network reformation using a cross-layer weight scaling, a precision minimization with priority-based quantization, and an activation map compression exploiting a data locality. As a result, the energy consumption of off-chip memory accesses is reduced up to 71.89% with less than 3.52% area overhead. Youngbeom Jung, Yeongjae Choi, Jaehyeong Sim, Lee-Sup Kim |
ICCAD | 4 |
| 2019 | An Energy-efficient Processing-in-memory Architecture for Long Short Term Memory in Spin Orbit Torque MRAMabstractMany recent studies have focused on Processing-in-memory (PIM) architectures for neural networks to resolve the memory bottleneck problem. Especially, an increased interest in Spin Orbit Torque (SOT)-MRAMs has emerged due to its low latency, high energy efficiency, and non-volatility. However, the previous work added extra computing circuits to support complicated computations, which results in large energy overheads. In this work, we propose a new PIM architecture with relatively small peripheral circuit, which produces the highest energy efficiency for processing a Long Short Term Memory (LSTM) among the PIM architectures. We improve the efficiency with a new computing method for logical operations, which exploits characteristics of SOT-MRAMs. We reduce the number of word lines (WLs) activated concurrently to one from two in the previous works. As a result, the energy for driving WLs is saved, and the sensing current for computation is reduced. Moreover, we propose efficient methods for additions, multiplications and non-linear activation functions in memory to process an LSTM. Accordingly, we achieve 1.26x energy efficiency with the proposed computing method for logical operations compared to the previous study based on SOT-MRAMs and up to 5.54x energy efficiency over the previous PIM architectures based on other memories. Kyeonghan Kim, Hyein Shin, Jaehyeong Sim, Myeonggu Kang, Lee-Sup Kim |
ICCAD | 5 |
| 2019 | A PVT-robust Customized 4T Embedded DRAM Cell Array for Accelerating Binary Neural NetworksabstractDeep neural networks (DNNs) are widely used for real-world applications. However, large amount of kernel and intermediate data incur a memory wall problem in resource-limited edge devices. The recent advances of a binary deep neural network (BNN) and a computing in-memory (CIM) have effectively alleviated this bottleneck especially when they are combined together. However, previous CIM-based accelerators for BNN are highly vulnerable to process/supply voltage/temperature (PVT) variation, resulting in severe accuracy degradation which makes them impractical to be employed in real-world edge devices. To address this vulnerability, we propose a PVT-robust accelerator architecture for BNN with a computable 4T embedded DRAM (eDRAM) cell array. First, we implement the XNOR operation of BNN in a time-multiplexed manner by utilizing the fundamental read operation of the conventional eDRAM cell. Next, a PVT-robust bit-count based on charge sharing is proposed with a computable 4T eDRAM cell array. In result, the proposed architecture achieves 6.9× less variation in PVT-variant environments which guarantees a stable accuracy and 2.03-49.4× improvement of energy efficiency over previous CIM-based accelerators. Hyein Shin, Jaehyeong Sim, Daewoong Lee, Lee-Sup Kim |
ICCAD | 4 |
| 2019 | Compressing Sparse Ternary Weight Convolutional Neural Networks for Efficient Hardware AccelerationabstractReducing the bit-width of weights is an attractive solution for decreasing the large size of CNN models embedded in IoT devices. In the most extreme case, the bit per weight can be reduced to 1-bit in binary weight CNNs. However, this network cannot be compressed further due to the lack of sparsity because the weight distribution of the well-trained model is not biased to either +1 or -1. On the other hand, sparse ternary weight CNNs can be compressed to less than 1-bit per weight, maintaining higher accuracy than binary weight CNNs. Therefore, we propose the following for a weight compression methodology in sparse ternary weight CNNs to minimize the model size: (1) an encoding scheme exploiting high sparsity, (2) two elaborate compression techniques based on encoding direction exploration and layer-wise optimization. To verify the efficiency of hardware acceleration, we design an accelerator that fully exploits our compression scheme. Moreover, a layer rearrangement technique is presented to address a load imbalance problem that occurs during hardware acceleration. As a result, we reduce the effective bit per weight to 0.67-0.80 bit and achieve 4.52-7.70x and 1.52-2.21x improvement of performance and energy efficiency respectively, with higher accuracy compared to previous binary weight CNN work. Hyeonwook Wi, Hyeonuk Kim, Seungkyu Choi, Lee-Sup Kim |
ISLPED | 4 |
| 2019 | DC-PCM: Mitigating PCM Write Disturbance with Low Performance Overhead by Using Detection CellsabstractAs DRAM scaling becomes ever more difficult, Phase Change Memory (PCM) is attracting attention as a new memory or storage class memory. Unfortunately, PCM cell data can be changed by frequently writing `0' to adjacent cells. This phenomenon is called Write Disturbance (WD). To mitigate WD errors with low performance overhead, we propose a Detection Cell PCM (DC-PCM). In the DC-PCM, additional cells called Detection Cells (DC) are allocated to a memory-line to pre-detect WD errors. For pre-detection, we propose schemes that give DCs higher WD-vulnerability than normal cells. However, additional time is needed to verify DCs. To hide the time needed to perform the verifications during a WRITE, DC-PCM enables the local word-lines of DCs to operate independently (Decoupled Word-line), and verifies different directions in parallel (Parallel DC-Verification). After verification, the DC-PCM increases the WD-vulnerability of the DCs, or restores the memory-line data (DC-Correction). In our simulation, DC-PCMs showed performance comparable to a WD-free PCM for all workloads. Jungwhan Choi, Jaemin Jang, Lee-Sup Kim |
IEEE Trans. Computers | 3 |
| 2019 | Sparse-Insertion Write Cache to Mitigate Write Disturbance Errors in Phase Change MemoryabstractAs the number of datasets processed in computing systems has increased in recent years, there is growing demand for high capacity main memory subsystems. However, further increases in the capacity of conventional DRAM-based main memory systems have stalled due to scaling limitations. Recent studies have shown that PCM, which can provide greater capacity than DRAM, is emerging as a candidate for high capacity memory. However, PCM suffers from problems related to the thermal mechanisms employed for storing data. The Write Disturbance (WD) phenomenon occurs when the thermal mechanisms of the PCM severely damage the data reliability of proximate cells. WD in PCM has become more significant below 20 nm. In this paper, we propose Sparse-Insertion Write Cache (SIWC), a practical, low-cost approach to mitigate WD errors. In PCM, repeated writes gradually degrade the validity of data in neighboring cells. SIWC uses a private write cache for the PCM write data to prevent repeated writes to the same address. The sparse-insertion technique can reduce cache eviction and minimize increases in the total write count. Our experimental results show that SIWC effectively reduces repeated writes and reduces the number of WD-vulnerable addresses across a wide range of applications. Jaemin Jang, Wongyu Shin, Jungwhan Choi, Yongju Kim, Lee-Sup Kim |
IEEE Trans. Computers | 5 |
| 2019 | A 0.9-V 12-Gb/s Two-FIR Tap Direct DFE With Feedback-Signal Common-Mode ControlabstractIn this brief, a 0.9-V 12-Gb/s quarter-rate two finite impulse response tap direct decision feedback equalizer (DFE) is presented. For a high-speed incorporated operation of both DFE summing and slicing, a common-mode-controlled charge-based latch (CMCCBL) is proposed. In a CMCCBL-based DFE, the common mode of the first tap feedback signal is changed to adjust the first tap feedback weight. Therefore, CMCCBL does not need the conventional first tap weighting transistors which increase the DFE feedback delay. The DFE core compensates for the FR4 printed circuit board channel loss of −22 dB at 6 GHz and consumes 4.16 mW achieving 0.347 pJ/bit for PRBS7 input. The active area of DFE is 0.0036 mm2in a 65-nm CMOS process. Daewoong Lee, Dongil Lee, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | NID: processing binary convolutional neural network in commodity DRAMabstractRecent large-scale CNNs suffer from a severe memory wall problem as their number of weights range from tens to hundreds of millions. Processing in-memory (PIM) and binary CNN have been proposed to alleviate the number of memory accesses and footprints, respectively. By combining the two separate concepts, we propose a novel processing in-DRAM framework for binary CNN, called NID, where dominant convolution operations are processed using in-DRAM bulk bitwise operations. We first identify the problem that the bitcount operations with only bulk bitwise AND/OR/NOT incur significant overhead in terms of delay when the size of kernels gets larger. Then, we not only optimize the performance by efficiently allocating inputs and kernels to DRAM banks for both convolutional and fully-connected layers through design space explorations, but also mitigate the overhead of bitcount operations by splitting kernels into multiple parts. Partial sum accumulations and tasks of the other layers such as max-pooling and normalization layers are processed in the peripheral area of DRAM with negligible overheads. In results, our NID framework achieves 19×-36× performance and 9×-14× EDP improvements for convolutional layers, and 9×-17× performance and 1.4×-4.5× EDP improvements for fully-connected layers over previous PIM technique in four large-scale CNN models. Jaehyeong Sim, Hoseok Seol, Lee-Sup Kim |
ICCAD | 3 |
| 2018 | TrainWare: A Memory Optimized Weight Update Architecture for On-Device Convolutional Neural Network TrainingabstractTraining convolutional neural network on device has become essential where it allows applications to consider user's individual environment. Meanwhile, the weight update operation from the training process is the primary factor of high energy consumption due to its substantial memory accesses. We propose a dedicated weight update architecture with two key features: (1) a specialized local buffer for the DRAM access deduction (2) a novel dataflow and its suitable processing element array structure for weight gradient computation to optimize the energy consumed by internal memories. Our scheme achieves 14.3%-30.2% total energy reduction by drastically eliminating the memory accesses. Seungkyu Choi, Jaehyeong Sim, Myeonggu Kang, Lee-Sup Kim |
ISLPED | 4 |
| 2018 | Elaborate Refresh: A Fine Granularity Retention Management for Deep Submicron DRAMsabstractAs the DRAM cell size continues to shrink, the proportion of leaky cells is increasing. As a result, the prior approaches, called retention aware refresh, which skip unnecessary refresh operations for non-leaky cells, are unable to skip as many refresh operations as before. The large granularity of the DRAM refresh mechanism makes this problem more serious. Specifically, even when there are only a small number of leaky cells in a particular retention group, that group is classified as a leaky group. Because of that, many non-leaky cells that also belong to that group are refreshed at an unnecessarily frequent rate. Since the granularity of the retention group is larger, this inefficiency becomes huge. To solve this problem, we propose a novel retention aware refresh approach called Elaborate Refresh, to reduce the granularity of the retention group further. The key idea of the Elaborate Refresh is to store leaky row addresses per each chip, and refresh different leaky row in each chip simultaneously. By doing so, Elaborate Refresh reduces the overhead of the leaky group refresh 16 times. In addition, Elaborate Refresh stores retention information in the DRAM chip, thus saving the refresh energy, even in the self-refresh mode when the memory controller cannot control the DRAM. Hoseok Seol, Wongyu Shin, Jaemin Jang, Jungwhan Choi, Hakseung Lee, Lee-Sup Kim |
IEEE Trans. Computers | 6 |
| 2017 | A Kernel Decomposition Architecture for Binary-weight Convolutional Neural NetworksabstractThe binary-weight CNN is one of the most efficient solutions for mobile CNNs. However, a large number of operations are required to process each image. To reduce such a huge operation count, we propose an energy-efficient kernel decomposition architecture, based on the observation that a large number of operations are redundant. In this scheme, all kernels are decomposed into sub-kernels to expose the common parts. By skipping the redundant computations, the operation count for each image was consequently reduced by 47.7%. Furthermore, a low cost bit-width quantization technique was implemented by exploiting the relative scales of the feature data. Experimental results showed that the proposed architecture achieves a 22% energy reduction. Hyeonuk Kim, Jaehyeong Sim, Yeongjae Choi, Lee-Sup Kim |
DAC | 4 |
| 2017 | SENIN: An energy-efficient sparse neuromorphic system with on-chip learningabstractApplying highly accurate neural networks to mobile devices encounters energy problems in battery-limited mobile environments. To resolve these problems, neuromorphic hardware solutions that enable event-driven operation have been proposed. In this work, we present a novel sparse neuromorphic system that implements an E-I Net algorithm to further improve energy efficiency. We introduce a neuron clock-gating technique that significantly reduces energy consumption by predicting future neuron spike activity without any loss of accuracy. We also propose synaptic pruning to save additional energy with minimal impact on classification accuracy. For fast adaptation to a changing environment, a learning algorithm is implemented in the proposed system. Compared to prior studies, our experimental results illustrate that the proposed system achieves 5.3×-11.4× energy efficiency improvement with comparable accuracy. Myung-Hoon Choi, Seungkyu Choi, Jaehyeong Sim, Lee-Sup Kim |
ISLPED | 4 |
| 2017 | Hardware-Centric Vision Processing for Mobile IoT Environment Exploiting Approximate Graph Cut in Resistor GridabstractThe Internet of things (IoT) has become a general trend in the electronic world. In this IoT era, various types of sensors gathering data are placed everywhere, and visual data is one of the most valuable information. While extracting important features from visual images, complex vision algorithms are often processed by high speed computing units like CPU or GPU. However, most IoT devices are mobile, so hardware resources are severely limited, which disturbs usage of the high performance processors in IoT appliances. Therefore, efficient vision computing becomes a crucial concern in this resource poor environment. Graph cut is a popular and widespread algorithm for image segmentation, image denoising, and stereo matching. However, it suffers from inefficient memory access patterns and massive digital computations. Due to the small hardware capacity not enough to overcome these obstacles, it is hardly feasible to run the algorithm in the IoT devices. In this paper, we propose an approximate version of the graph cut with the aid of the dedicated hardware design. By exploiting the analogy between the max flow and electric currents, we design the theoretical model of an approximate max flow solution and suggest the resistor grid as its processing circuit. Because this work utilizes electric potentials for the min cut computation, it is possible to find the cut at a low energy and a high speed. A prototype circuit is designed and simulated for evaluation. Using public GrabCut benchmark, algorithm evaluations are performed. As a result, we can verify a fast and an energy-efficient image processing with an acceptable accuracy. For supervised binary image segmentation, speed efficiency is 4.57 times higher, and energy-efficiency is 9.11 times better when compared to other image/video segmentation works. Yeongjae Choi, Jun-Seok Park, Lee-Sup Kim |
WACV | 3 |
| 2017 | Refresh-Aware Write Recovery Memory ControllerabstractCurrent computer systems require large memory capacities to manage the tremendous volume of datasets. A DRAM cell consists of a transistor and a capacitor, and their size has a direct impact on DRAM density. While technology scaling can provide higher density, this benefit comes at the expense of low drivability, due to the increase in series resistance of the smaller transistor, which slows the process of restoring the charge in cells. DRAM operations require recovery processes due to the destructive nature of DRAM cells. Among such operations, the write recovery process has the most difficulty in meeting the timing constraints. In this paper, we explore an intrinsic mechanism in the DRAM write operation, and find a relation between restoration and retention times. Based on our observation, we propose a practical mechanism, Relaxed Refresh with Compensated Write Recovery (RRCW), which efficiently mitigates refresh overheads by providing longer restoration periods. Furthermore, to minimize the penalty of the longer restoration, we also introduce another mechanism, Refresh-Aware Write Recovery (RAWR), which appropriately curtails longer recovery time according to the waiting time until being refreshed. Lastly, we introduce a scheduling policy to efficiently utilize RAWR. Evaluations show that the benefits of our mechanisms increase as memory intensity increases. Jaemin Jang, Wongyu Shin, Jungwhan Choi, Jinwoong Suh, Yongkee Kwon, Yongju Kim, Lee-Sup Kim |
IEEE Trans. Computers | 7 |
| 2017 | Bank-Group Level ParallelismabstractDDR4 SDRAM introduced a new hierarchy in DRAM organization: bank-group (BG). The main purpose of BG is to increase I/O bandwidth without growing DRAM-internal bus-width. We, however, found that other benefits can be derived from the new hierarchy. To achieve the benefits, we propose a new DRAM architecture using the BG-hierarchy, leading to a creation of BG-Level Parallelism (BGLP). By exploiting BGLP, the overall parallelism grows in DRAM operations. We also argue that BGLP is a feasible solution in the cost-sensitive DRAM industry because the additional cost is negligible and only cost-insensitive area needs to be modified. Wongyu Shin, Jaemin Jang, Jungwhan Choi, Jinwoong Suh, Lee-Sup Kim |
IEEE Trans. Computers | 5 |
| 2017 | Rank-Level Parallelism in DRAMabstractDRAM systems are hierarchically organized: Channel-Rank-Bank. A channel is connected to multiple ranks, and each rank has multiple banks. This hierarchical structure facilitates creating parallelisms in DRAM. The current DRAM architecture supports bank-level parallelism; as many rows as banks can be moved simultaneously at bank-level. However, rank-level parallelism is not supported. For this reason, only one column can be accessed at a time, although each rank has its own data bus that can carry a column. Namely, current DRAM operations do not exploit the structural opportunity created by multiple ranks. We, therefore, propose a novel DRAM architecture supporting rank-level parallelism. Thereby, as many columns as ranks can be moved concurrently at rank-level. In this paper, we illustrate the rank-level parallelism and its benefit in DRAM operations. Wongyu Shin, Jaemin Jang, Jungwhan Choi, Jinwoong Suh, Yongkee Kwon, Youngsuk Moon, Lee-Sup Kim |
IEEE Trans. Computers | 7 |
| 2017 | A 5-Gb/s Digital Clock and Data Recovery Circuit With Reduced DCO Supply Noise Sensitivity Utilizing Coupling NetworkabstractA digital clock and data recovery (CDR) is presented, which employs a low supply sensitivity scheme for a digitally controlled oscillator (DCO). A coupling network comprising capacitors, resistors, and coupling buffers enhances the supply variation immunity of the DCO and mitigates the jitter performance degradation. A supply variation-dependent bias generator produces the corresponding bias voltage to alleviate the supply variation with minimal area and power penalty. The proposed scheme improves 29.3 ps of peak-to-peak jitter and 11.5 dB of spur level, at 6 and 5 MHz 50 mVppsinusoidal supply noise tone, respectively. Fabricated in a 65-nm CMOS process, the proposed CDR operates at 5-Gb/s data rate with BER$<10^{-12}$for PRBS 31 and consumes 15.4 mW. The CDR occupies an active die area of 0.075 mm2. Taeho Lee 0001, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | In-DRAM Data InitializationabstractInitializing memory with zero data is essential for safe memory management. However, initializing a large memory area slows down the system significantly. The most likely cause for initialization to slow down the system is the limited DRAM initialization method. At present, the only way to initialize DRAM area is to execute multiple WRITE commands. However, the WRITE command slows the initialization because of its small granularity and data bus occupancy. In this brief, we propose an efficient in-DRAM initialization method inspired by the internal structure and operation of DRAM. The proposed method, called row reset, uses a DRAM row buffer to zero out a single DRAM row at a time. Row Reset allows for parallel initialization on multiple DRAM banks without using off-chip data transfer, thus reducing initialization time by up to 63 times. Row reset is a practical approach, because it can be implemented with existing circuitry in DRAM without additional area overhead. Hoseok Seol, Wongyu Shin, Jaemin Jang, Jungwhan Choi, Jinwoong Suh, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2016 | Energy Efficient Data Encoding in DRAM Channels Exploiting Data Value SimilarityabstractAs DRAM data bandwidth increases, tremendous energy is dissipated in the DRAM data bus. To reduce the energy consumed in the data bus, DRAM interfaces with symmetric termination, such as Pseudo Open Drain (POD) and Low Voltage Swing Terminated Logic (LVSTL), have been adopted in modern DRAMs. In interfaces using asymmetric termination, the amount of termination energy is proportional to the hamming weight of the data words. In this work, we propose Bitwise Difference Encoding (BD-Encoding), which decreases the hamming weight of data words, leading to a reduction in energy consumption in the modern DRAM data bus. Since smaller hamming weight of the data words also reduces switching activity, switching energy and power noise are also both reduced. BD-Encoding exploits the similarity in data words in the DRAM data bus. We observed that similar data words (i.e. data words whose hamming distance is small) are highly likely to be sent over at similar times. Based on this observation, BD-coder stores the data recently sent over in both the memory controller and DRAMs. Then, BD-coder transfers the bitwise difference between the current data and the most similar data. In an evaluation using SPEC 2006, BD-Encoding using 64 recent data reduced termination energy by 58.3% and switching energy by 45.3%. In addition, 55% of the LdI/dt noise was decreased with BD-Encoding. Hoseok Seol, Wongyu Shin, Jaemin Jang, Jungwhan Choi, Jinwoong Suh, Lee-Sup Kim |
ISCA | 6 |
| 2016 | Crosstalk avoidance code for direct pass-through architectureabstractAs modern high performance computing systems have various ICs associated together, one of the most appreciated architecture is direct path passing through the intermediate chips. Inter-chip and intra-chip communication are linked together in this architecture. Therefore to remove crosstalk induced noise and improve signal integrity, a novel CAC and corresponding CODEC which can be applied for both on-chip and off-chip channels is proposed. The proposed code does not require additional encoder and decoder at chip borders. Minhye Kim, Soochang Chae, Young-Ju Kim 0001, Seung-Jun Bae, Lee-Sup Kim |
ISCAS | 5 |
| 2016 | Q-DRAM: Quick-Access DRAM with Decoupled Restoring from Row-ActivationabstractThe relatively high latency of DRAM is mostly caused by the long row-activation time which in fact consists of sensing and restoring time. Memory controllers cannot distinguish between them since they are performed consecutively by a single row-activation command. If these two steps are separated, the restoring can be delayed until DRAM access is uncongested. Hence, we propose Quick-Access DRAM (Q-DRAM) which discriminates between sensing and restoring. Our approach is to allow destructive access (i.e., only sensing is performed without restoring by a row-activation command) using per-bank multiple row-buffers. We call the destructive access and per-bank multiple row-buffers quick-access and quick-buffers (q-buffers) respectively. In addition, we propose Quick-access Trigger (Q-TRIGGER) and RESTORER to utilize Q-DRAM. Q-TRIGGER makes a decision whether quick-access is required or not, and RESTORER decides when to restore the data at the destructed cell. Specifically, RESTORER detects the proper timing to hide restoring time by predicting data bus occupation and by exploiting bank-level locality. Evaluations show that Q-DRAM significantly improved performance for both single- and multi-core systems. Wongyu Shin, Jungwhan Choi, Jaemin Jang, Jinwoong Suh, Yongkee Kwon, Youngsuk Moon, Hongsik Kim, Lee-Sup Kim |
IEEE Trans. Computers | 8 |
| 2016 | DRAM-Latency Optimization Inspired by Relationship between Row-Access Time and Refresh TimingabstractIt is widely known that relatively long DRAM latency forms a bottleneck in computing systems. However, DRAM vendors are strongly reluctant to decrease DRAM latency due to the additional manufacturing cost. Therefore, we set our goal to reduce DRAM latency without any modification in the existing DRAM structure. To accomplish our goal, we focus on an intrinsic phenomenon in DRAM: electric charge variation in DRAM cell capacitors. Then, we draw two key insights: i) DRAM row-access latency of a row is a function of the elapsed time from when the row was last refreshed, and ii) DRAM row-access latency of a row is also a function of the remaining time until the row is next refreshed. Based on these two insights, we propose two mechanisms to reduce DRAM latency: NUAT-1 and NUAT-2. NUAT-1 exploits the first key insight and NUAT-2 exploits the second key insight. For evaluation, circuit- and system-level simulations are performed, which show the performance improvement for various environments. Wongyu Shin, Jungwhan Choi, Jaemin Jang, Jinwoong Suh, Youngsuk Moon, Yongkee Kwon, Lee-Sup Kim |
IEEE Trans. Computers | 7 |
| 2016 | A Vision Processor With a Unified Interest-Point Detection and Matching Hardware for Accelerating a Stereo-Matching AlgorithmabstractIn this paper, a unified interest-point detection and matching hardware with an optimized memory architecture is proposed for a real-time stereo-matching system. In order to support a stereo-matching algorithm, the unified datapath in the hardware performs not only interest-point detection and matching algorithms such as features from the accelerated segment test (FAST) and binary robust independent elementary feature (BRIEF) in real time but also a census transform, which is widely used in stereo matching. To achieve maximum performance, we propose two special memory architectures: 1) reconfigurable image memory (RIM) and 2) point cloud index memory system (PCIM). RIM is a unified memory architecture that loads pixel values from a raw image patch. Since FAST, BRIEF, and the census transform have different and complex memory access patterns, the miss rate of the memory access may increase. To optimize the memory operation, RIM can change its memory configuration according to the algorithm. PCIM is a dedicated memory system that utilizes the geometric information of the cameras in order to reduce the off-chip memory bandwidth. Based on the geometric information, PCIM removes most of the redundant candidates. Since PCIM minimizes the off-chip memory bandwidth using a dedicated cache, the performance degradation is negligible compared with the exact-nearest-neighbor method. Area-based stereo matching is accelerated based on the general-purpose computing on graphics processing units (GPGPU) architecture because the search range is adaptively reduced according to the disparity of the matched correspondences. The overall hardware consists of 1.20-M logic gates and consumes a maximum of 185 mW. Interest point detection and matching accelerator achieves 106 frames/s in 1080p full high-definition video (HD) resolution at a 200-MHz operating frequency with 3500 descriptors per image. Jun-Seok Park, Hyo-Eun Kim, Hong-Yun Kim, Lee-Sup Kim |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | A 21-Gbit/s 1.63-pJ/bit Adaptive CTLE and One-Tap DFE With Single Loop Spectrum Balancing MethodabstractThis brief presents an adaptive continuous-time linear equalizer (CTLE) and one-tap decision feedback equalizer (DFE) using the spectrum balancing (SB) method. The SB method is extended for not only CTLE but also DFE with the aid of gain characteristics of one-tap DFE. Thus, adaptation loops for each equalizer type are merged to a single loop. As a result, the complexity and power consumption of the adaptation circuits are reduced significantly. The test chip consumes 34.2 mW from 1.2 V supply with 65-nm CMOS process. Young-Ju Kim 0001, Taeho Lee 0001, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | A 5-Gb/s 2.67-mW/Gb/s Digital Clock and Data Recovery With Hybrid Dithering Using a Time-Dithered Delta-Sigma ModulatorabstractA digital clock and data recovery (CDR) employing a time-dithered delta–sigma modulator (TDDSM) is presented. By enabling hybrid dithering of a sampling period as well as an output bit of the TDDSM, the proposed CDR enhances the resolution of digitally controlled oscillator, removes a low-pass filter in the integral path, and reduces jitter generation. Fabricated in a 65-nm CMOS process, the proposed CDR operates at 5-Gb/s data rate with$\textrm {BER}<10^{-12}$for PRBS 31. The CDR consumes 13.32 mW at 5 Gb/s and achieves 2.14 and 29.7 ps of a long-term rms and peak-to-peak jitter, respectively. Taeho Lee 0001, Jaehyeong Sim, Jun-Seok Park, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2015 | Multiple clone row DRAM: a low latency and area optimized DRAMabstractSeveral previous works have changed DRAM bank structure to reduce memory access latency and have shown performance improvement. However, changes in the area-optimized DRAM bank can incur large area-overhead. To solve this problem, we propose Multiple Clone Row DRAM (MCR-DRAM), which uses existing DRAM bank structure without any modification. Jungwhan Choi, Wongyu Shin, Jaemin Jang, Jinwoong Suh, Yongkee Kwon, Youngsuk Moon, Lee-Sup Kim |
ISCA | 7 |
| 2015 | An integrated time register and arithmetic circuit with combined operation for time-domain signal processingabstractThis paper describes a novel integrated time register and arithmetic circuit (TRAC) for time-domain signal processing (TDSP). TRAC has four basic functions: time register, time adder, time subtractor, and time amplifier. One TRAC is able to accept multiple inputs and perform the combination of basic functions. Thus, complex arithmetic operations can be performed synchronously with a TRAC. In a synchronous time-to-digital converter (TDC), a high-gain residue amplifier with the suppressed offset-time component at the output can be implemented by exploiting the advantages of TRAC. The functions of TRAC are with post-layout data simulated in 110nm CMOS technology. Daewoong Lee, Dongil Lee, Taeho Lee 0001, Lee-Sup Kim |
ISCAS | 5 |
| 2015 | A 9.6-Gb/s 1.22-mW/Gb/s Data-Jitter Mixing Forwarded-Clock Receiver in 65-nm CMOSabstractIn this paper, a data-jitter mixing (DJM) forwarded-clock receiver is proposed that achieves high jitter correlation between data and a clock for high speed and small power consumption. The first-stage injection-locked oscillator (ILO) filters out high-frequency clock jitter that loses the correlation due to a latency mismatch between data and the clock. Then, a data-jitter mixer in the second stage of the proposed receiver further increases the jitter correlation reduced by nonoptimal jitter filtering in ILO. Moreover, the DJM reduces power supply noise induced jitter from a clock distribution network, while the conventional jitter filter cannot track the high-frequency jitter because of filtering it out. A prototype receiver implemented in 1-V 65-nm CMOS process achieves 9.6 Gb/s with 1.22-mW/Gb/s in spite of a 1.92-ns latency mismatch between data and a clock. Sang-Hye Chung, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | A Forwarded Clock Receiver Based on Injection-Locked Oscillator With AC-Coupled Clock Multiplication Unit in 0.13~µm CMOSabstractThis brief presents a forwarded clock receiver based on an injection-locked oscillator with a simple clock multiplication unit (CMU) to reduce the clock jitter and power consumption of the CMU. In addition, an optimal clock multiplication factor is considered to optimally multiply the clock frequency without serious degradation of the jitter correlation between data and clock. The proposed CMU employs ac coupling and a superposition technique to generate first-harmonic injection pulses. The measured power efficiency of the proposed receiver is 1.69 mW/Gb/s at a 7.4-Gb/s data rate in a 1.2 V 0.13-μm CMOS process. Young-Ju Kim 0001, Sang-Hye Chung, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | An 11.5 Gb/s 1/4th Baud-Rate CTLE and Two-Tap DFE With Boosted High Frequency Gain in 110-nm CMOSabstractAs the data rate has been increased over 10 Gb/s with copper interconnect, the intersymbol interference (ISI) caused from the channel loss should be compensated. While a decision feedback equalizer (DFE), which is widely used in the receiver can compensate the ISI, its ability to enhance the signal-to-noise ratio (SNR) is limited especially for high frequency data patterns (alternating data patterns). Even a DFE with a large number of taps, which has powerful compensation capacity for ISI, can improve only limited amount of SNR. To improve SNR, this brief presents a 1/4th baud-rate continuous-time linear equalizer (CTLE) and a two-tap DFE. The 1/4th baud-rate CTLE recovers data to be adequate for the DFE, and remaining ISI is removed by the DFE. To boost the high frequency gain of the DFE, exclusive or merged adders are introduced, so the proposed equalizer is more tolerant against channel noise. It compensates 21.7-dB channel loss and operates with 11.5-Gb/s data rate as it consumes 25.35 mW from 1.3 V supply in a 110-nm CMOS technology. Young-Ju Kim 0001, Taeho Lee 0001, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2014 | NUAT: A non-uniform access time memory controllerabstractWith rapid development of micro-processors, off-chip memory access becomes a system bottleneck. DRAM, a main memory in most computers, has concentrated only on capacity and bandwidth for decades to achieve high performance computing. However, DRAM access latency should also be considered to keep the development trend in multi-core era. Therefore, we propose NUAT which is a new memory controller focusing on reducing memory access latency without any modification of the existing DRAM structure. We only exploit DRAM's intrinsic phenomenon: electric charge variation in DRAM cell capacitors. Given the cost-sensitive DRAM market, it is a big advantage in terms of actual implementation. NUAT gives a score to every memory access request and the request with the highest score obtains a priority. For scoring, we introduce two new concepts: Partitioned Bank Rotation (PBR) and PBR Page Mode (PPM). First, PBR is a mechanism that draws information of access speed from refresh timing and position; the request which has faster access speed gains higher score. Second, PPM selects a better page mode between open- and close-page modes based on the information from PBR. Evaluations show that NUAT decreases memory access latency significantly for various environments. Wongyu Shin, Jeongmin Yang, Jungwhan Choi, Lee-Sup Kim |
HPCA | 4 |
| 2014 | Timing error masking by exploiting operand value locality in SIMD architectureabstractA significant amount of energy is consumed by a voltage guardband to ensure error-free operations under the worsening PVT variations in modern processors. Circuit-level timing speculation has become a popular approach that increases energy efficiency by removing such guardband and tolerating occasional timing errors. However, SIMD processors suffer from a large throughput and energy efficiency loss induced by a conventional error correction mechanism which requires several extra cycles for each timing error. In this paper, we present an error masking scheme to eliminate the chances of performing the error correction. The error masking is done by allowing potential erroneous addition instructions to reuse the partial result of previous operations. We show that reuse can be applied to a large number of addition instructions by exploiting the observations that SIMD applications exhibit high levels of temporal operand value locality and operand value locality across SIMD lanes. Our implementation of the proposed masking scheme is augmented with the conventional pipeline logics. Simulation results verify that our scheme achieves up to 5.1% improvement in energy efficiency and 30% improvement in EDP (Energy-Delay-Product) over the baseline design. Jaehyeong Sim, Jun-Seok Park, Seungwook Paek, Lee-Sup Kim |
ICCD | 4 |
| 2014 | An area-efficient on-chip temperature sensor with nonlinearity compensation using injection-locked oscillator (ILO)abstractThis paper describes CMOS time-domain temperature sensors. A principle of this type of sensors is CMOS inverter's time-delay variation with temperature. The variation, however, has nonlinearity which is a fundamental error source. Therefore, we propose a new temperature sensor that improves linearity using an injection-locked oscillator (ILO). Since the ILO has the opposite curvature of an inverter delay line in temperature domain, nonlinear error induced by the CMOS inverters can be eliminated. Integral nonlinearity (INL) error is reduced from 3.6 LSB to 0.56 LSB (84% reduction), resulting in under 1-bit error by nonlinearity. The proposed design has -0.19°C~0.2°C inaccuracy which is only quantization error over 0°C~100°C. The power consumption is 2.88mW including an 8-bit temperature-to-digital code conversion block at the sampling rate of 15.6M samples/s. Wongyu Shin, Seungwook Paek, Lee-Sup Kim |
ISCAS | 3 |
| 2013 | A 7 mW 2.5 GHz spread spectrum clock generator using switch-controlled injection-locked oscillatorabstractAs electric components are integrated in the smaller area, the electromagnetic interference (EMI) becomes a serious problem. To reduce an EMI problem and increase noise immunity of electric circuit, the clock spreading is one of the most popular methods these days. This work proposes a switch-controlled injection-locked oscillator (SCILO)-based spread spectrum clock generator (SSCG) with the features including low power consumption, simple operation mechanism, and small core area. In addition, sub-harmonic injection locking is used to provide high frequency output signal. The test-chip fabricated in 0.13 μm CMOS technology provides 2.5 GHz output clock with 1 % spreading ratio. Peak reduction is 13.5 dB at 1 % spreading ratio. The proposed SSCG consumes 7 mW from 1.2 V supply and only occupies 0.034 mm2. Jeongmin Yang, Young-Ju Kim 0001, Lee-Sup Kim |
ISCAS | 3 |
| 2013 | PowerField: A Probabilistic Approach for Temperature-to-Power Conversion Based on Markov Random Field TheoryabstractTemperature-to-power technique is useful for post-silicon power model validation. However, the previous works were applicable only to the steady-state analysis. In this paper, we propose a new temperature-to-power technique, named PowerField, supporting both transient and steady-state analysis based on a probabilistic approach. Unlike the previous works, PowerField uses two consecutive thermal images to find the most feasible power distribution that causes the change between the two input images. To obtain the power map with the highest probability, we adopted maximum a posteriori Markov random field (MAP-MRF). For MAP-MRF framework, we modeled the spatial thermal system as a set of thermal nodes and derived an approximated transient heat transfer equation that requires only the local information of each thermal node. Experimental results with a thermal simulator show that PowerField outperforms the previous method in transient analysis reducing the error by half on average. We also show that our framework works well for steady-state analysis by using two identical steady-state thermal maps as inputs. Lastly, an application to determining the binary power patterns of an FPGA device is presented achieving 90.7% average accuracy. Seungwook Paek, Wongyu Shin, Jaehyeong Sim, Lee-Sup Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2013 | A 182 mW 94.3 f/s in Full HD Pattern-Matching Based Image Recognition Accelerator for an Embedded Vision System in 0.13-µm CMOS TechnologyabstractA pattern-matching based image recognition accelerator (PRA) is presented for embedded vision applications. It is a hardware accelerator that performs interest point detection and matching for image-based recognition applications in real time in both mobile devices and vehicles. The proposed system is implemented as a small IP, and it has eight times higher throughput than state-of-the-art object recognition processors, which are implemented based on a heterogeneous many-core system. PRA has three key features: joint algorithm-architecture optimizations for exploiting bit-level parallelism; a low-power unified hardware platform for interest point detection and matching; and scalable hardware architecture. PRA achieves 9.5× performance improvement with only 30% of logic gates including static random-access memory (SRAM) compared to the state-of-the-art object recognition processors. It consists of 78.3 k logic gates and 128 kB SRAM, which are integrated in a test chip implemented for PRA verification. It achieves 94.3 frames per second (fps) in 1080 p full HD resolution at 200-MHz operating frequency while consuming 182 mW. Each complete operation for interest point detection and matching requires 2.09 cycles and 8 cycles on average, respectively, based on a unified bit-level matching accelerator, which is implemented only with 680 logic gates. Jun-Seok Park, Hyo-Eun Kim, Lee-Sup Kim |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | A Unified Graphics and Vision Processor With a 0.89 µW/fps Pose Estimation Engine for Augmented RealityabstractA unified vision and graphics processor with three layers is shown to provide a fast pipeline for augmented reality. In the image-level layer, a 153.6 GOPS massively parallel processing unit with eight SIMD processors, each containing 128 processing elements, performs highly data-parallel operations. In the sub-image layer, a rasterizer and a pixel arranger respectively generate and reduce data-level parallelism. In the descriptor-level layer, a pose estimation engine executes sequential programs. Our processor can provide images for augmented reality at 100 fps, for a power consumption of 413 mW. This is 39% faster than a comparable smartphone implementation. Our chip is fabricated in a 0.18 μm CMOS process and contains 0.95 M gates. Jae-Sung Yoon, Jeong-Hyun Kim, Hyo-Eun Kim, Won-Young Lee, Seok-Hoon Kim, Kyusik Chung, Jun-Seok Park, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2012 | PowerField: a transient temperature-to-power technique based on Markov random field theoryabstractTransient temperature-to-power conversion is as important as steady-state analysis since power distributions tend to change dynamically. In this work, we propose PowerField framework to find the most probable power distribution from consecutive thermal images. Since the transient analysis is vulnerable to spatio-temporal thermal noise, we adopted a maximum-a-posteriori Markov random field framework to enhance the noise immunity. The most probable power map is obtained by minimizing the energy function which is calculated using an approximated transient thermal equation. Experimental results with a thermal simulator shows that PowerField outperforms the previous method in transient analysis reducing the error by half on average. We also applied our method to a real silicon achieving 90.7% accuracy. Seungwook Paek, Seok-Hwan Moon, Wongyu Shin, Jaehyeong Sim, Lee-Sup Kim |
DAC | 5 |
| 2012 | A 20 Gbps 1-tap decision feedback equalizer with unfixed tap coefficientabstractThis paper describes a 1-tap DFE with unfixed tap coefficient. According to the data patterns, the inter-symbol interference (ISI) is changed. The data patterns will be predicted using the DC level detector and the tap coefficient of the DFE will be adjusted to achieve high gain against high attenuation. The proposed 1-tap DFE with unfixed tap coefficient solves a high channel loss problem while it consumes low power. Designed in 65-nm CMOS technology, a 20 Gb/s receiver is composed of 1-tap half rate DFE with unfixed tap coefficient which compensates 20 dB attenuation, and it consumes 22.65 mW for 1-V supply. Lee-Sup Kim |
ISCAS | 2 |
| 2012 | A Reconfigurable Heterogeneous Multimedia Processor for IC-Stacking on Si-InterposerabstractThis paper presents a heterogeneous multimedia processor for embedded media applications such as image processing, vision, 3-D graphics and augmented reality (AR), assuming integrated circuit (IC)-stacking on Si-interposer. This processor embeds reconfigurable output drivers for external memory interface to increase memory bandwidth even in a mobile environment. The implemented output driver reconfigures its driving strength according to channel loss between the implemented processor and the memory, so it enables highspeed data communication while achieving 8× higher memory bandwidth compared to previous embedded media processors. The implemented processor includes three main programmable intellectual properties, mode-configurable vector processing units (MCVPUs), a unified filtering unit (UFU), and a unified shader. MCVPUs have 32 integer (16 bit) cores in order to support dual-mode operations between image-level processing and graphics processing. This mode-configuration enables a frame-level pipelining in AR application, so the proposed processor achieves 1.7× higher frame rate compared to the sequential AR processing. UFU supports 16 types of filtering operations only with a single instruction. Most image-level processing consists of various types of filtering operations, so UFU can improve media processing performance and energy-efficiency. UFU also supports texture filtering which is performance bottleneck of common graphics pipeline. A memory-access-efficient (off-chip memory) texturing algorithm named as an adaptive block selection is proposed to enhance texturing performance in 3-D graphics pipeline. UFU has two-level on-chip memory hierarchies, a 512B level-0 (L0) data buffer, and an 8kB level-1 (L1) static random-access memory (SRAM) cache. The small-sized L0 data buffer limits direct references to the large-sized L1 SRAM cache to reduce energy consumed in on-chip memories. Unified shader consists of four homogeneous scalar processing elements (SPEs) for geometry operations in 3-D graphics. Each SPE has single-precision floating-point data-paths, since precision of geometry operations in 3-D graphics is important in today's handheld devices (high resolution). The proposed media processor is fabricated in 0.13 μm CMOS technology with 4 mm × 4 mm chip size, and dissipates 275 mW for full AR operation. Hyo-Eun Kim, Jae-Sung Yoon, Kyu-Dong Hwang, Young-Jun Kim 0001, Jun-Seok Park, Lee-Sup Kim |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2012 | Homogeneous Stream Processors With Embedded Special Function Units for High-Utilization Programmable ShadersabstractWe embed special function units (SFUs) in homogeneous stream processors (SPs) within a graphics processing unit (GPU), to improve its performance in running modern programmable shaders, which make poor use of a single-instruction multiple-data (SIMD) architecture. We also compact instructions, so as to reduce the size of the instruction memory, and reduce area requirements by using a partial SFU in SPs, and a lookup table which is shared between multiple SFUs. The result is an increase of 88% in utilization and a reduction in the normalized area-delay product of 27%, compared to a baseline SIMD architecture. We verified our architecture on an field-programmable gate-array evaluation platform with an ARM9 host processor and a full 3-D graphics pipeline. Young-Jun Kim 0001, Hyo-Eun Kim, Seok-Hoon Kim, Jun-Seok Park, Seungwook Paek, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2012 | A Mobile 3-D Display Processor With A Bandwidth-Saving SubdividerabstractA mobile 3-D display processor with a subdivider is presented for higher visual quality on handhelds. By combining a subdivision technique with a 3-D display, the processor can support viewers see realistic smooth surfaces in the air. However, both the subdivision and the 3-D display processes require a high number of memory operations to mobile memory architecture. Therefore, we make efforts to save the bandwidth between the processor and off-chip memory. In the subdivider, we propose a recomputing based depth-first scheme that has much smaller working set than prior works. The proposed scheme achieves about 100:1 bandwidth reduction over the prior subdivision methods. Also the designed 3-D display engine reduces the bandwidth to 27% by reordering the operation sequence of the 3-D display process. This bandwidth saving translates into reductions of off-chip access energy and time. Consequently the overall bandwidth of both the subdivision and the 3-D display processes is affordable to a commercial mobile bus. In addition to saving bandwidth, our work provides enough visual quality and performance. Overall the 3-D display engine achieves 325 fps for 480×320 display resolution. Seok-Hoon Kim, Sung-Eui Yoon, Sang-Hye Chung, Young-Jun Kim 0001, Hong-Yun Kim, Kyusik Chung, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2012 | An Adaptive Equalizer With the Capacitance Multiplication for DisplayPort Main Link in 0.18-µm CMOSabstractAn adaptive equalizer with the capacitance multiplication for DisplayPort main link has been proposed. The proposed equalizing filter is based on Miller's theorem and composed of metal-insulator-metal capacitors and a sub-amplifier. The active source degeneration capacitor achieves low cost and area saving with the capacitance multiplication. The equalizer satisfies the specification of DisplayPort version 1.1a. The measured eye widths of 2.7 Gb/s data are 0.6 and 0.5 UI for 5 and 8 m cables, respec- tively. The core area is 286 × 380 μm2and power consumption is 22.3 mW at 2.7 Gb/s at 1.8 V. Won-Young Lee, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | A 5.4 Gb/s clock and data recovery circuit using the seamless loop transition scheme without phase noise degradationabstractThis paper presents a 5.4 Gb/s clock and data recovery circuit using the seamless loop transition scheme which has no phase noise degradation. The controllable loop filter enables the CDR circuit to change the operation mode without the output phase noise degradation and the stability problem. The modified half-rate linear phase detector reduces the phase error between the data and clock. A tested chip is manufactured using 0.13 μm CMOS technology. The RMS jitter of the proposed CDR circuit is 5.98 ps-rms, which is 2.61 ps lower than the CDR circuit with the conventional scheme. The measured power dissipation is 138 mW with off-chip drivers at 5.4 Gb/s data rate. Won-Young Lee, Lee-Sup Kim |
ISCAS | 2 |
| 2011 | Area-efficient dynamic thermal management unit using MDLL with shared DLL scheme for many-core processorsabstractAn area-efficient dynamic thermal management (DTM) unit using multiplying delay-locked loop (MDLL) with shared DLL scheme is proposed for per-core DTM in many-core processors. The proposed DTM unit consists of a MDLL and a shared-DLL-based temperature sensor. The shared DLL takes part in both temperature sensing and frequency scaling while reducing the size of whole DTM unit. The area is reduced by 54.4% compared to the design in which a phase-locked loop (PLL) is used without any optimization scheme. Seungwook Paek, Jiehwan Oh, Sang-Hye Chung, Lee-Sup Kim |
ISCAS | 4 |
| 2011 | A Memory-Efficient Unified Early Z-TestabstractThe Unified Early Z-Test (U-EZT) is proposed to examine the visibility of pixels during tile-based rasterization in a mobile 3D graphics processor. U-EZT combines the advantages of the Z-max and Z-min EZT algorithms: the Z-max algorithm is improved by the independently updatable z-max tiles and the use of mask bits; and the Z-min algorithm is improved by reusing the mask bits from the z-max test to update the z-min tiles after tile rasterizing. As a result, storage requirements are reduced to 3 bits per pixel, and simulations suggest that U-EZT requires 20 percent to 57 percent less memory bandwidth than previous EZT algorithms. Hong-Yun Kim, Chang-Hyo Yu, Lee-Sup Kim |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2011 | A Dual-Shader 3-D Graphics Processor With Fast 4-D Vector Inner Product Units and Power-Aware Texture CacheabstractThis paper presents a fully programmable 3-D graphics processor using unified shaders for mobile environment. In the system level, we adopted dual-core, dual-issue VLIW, and multithreading methods to utilize instruction, data, and task level parallelism in the graphics applications. In the shader core level, a novel IEEE-754 compliant 4-D vector inner product arithmetic unit and a configurable texture cache are proposed. Using these methods, the proposed processor achieves 143 Mvertices/s and 2.3 Gtexels/s consuming the power of 367 mW. The evaluation shows significant performance and power-delay product benefits. For real graphics applications, test results indicate 2.07 times improvement in performance and 34% reduction in power-delay product compared to previous mobile 3-D graphics processors. The proposed 3-D graphics processor is implemented in 4.5× 4.52mm using 0.18 μm CMOS technology. Jae-Sung Yoon, Chang-Hyo Yu, Donghyun Kim 0014, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2010 | A high resolution metastability-independent two-step gated ring oscillator TDC with enhanced noise shapingabstractThis paper presents a high resolution two-step gated-ring oscillator (TSGRO) time-to-digital converter (TDC) in an all digital phase-locked loop (ADPLL). TSGRO-TDC consists of a coarse step and a fine step gated-ring oscillator (GRO) TDC to achieve a high resolution. An edge aligner is used in the fine step GRO-TDC to enhance a first-order noise shaping property. A meta-stability free selection logic (MSFSL) is proposed in this paper which achieves low power and small area. In 0.13μm CMOS process, TSGRO-TDC has a 3ps raw resolution and first-order noise shaping. Sang-Hye Chung, Kyu-Dong Hwang, Won-Young Lee, Lee-Sup Kim |
ISCAS | 4 |
| 2010 | An area efficient asynchronous gated ring oscillator TDC with minimum GRO stagesabstractAn 8-bit, 3-stage asynchronous gated ring oscillator (GRO) time-to-digital converter (TDC) is presented. It employs asynchronous techniques to achieve minimum GRO stages. This lead to about 40% to 70% gate count reduction compared to synchronous GRO-TDC. Count-missing, glitch, and unnecessary addition are eliminated. The uncorrupted noise shaping characteristic is obtained. The chip is implemented in a 0.18 μm CMOS technology. It occupies small area (140μm×310μm) and consumes low power (4mW to 13mW). Kyu-Dong Hwang, Lee-Sup Kim |
ISCAS | 2 |
| 2009 | Bank-partition and Multi-fetch Scheme for Floating-point Special Function units in Multi-core SystemsabstractA table loader unit with bank-partition and multi-fetch feature is proposed for multiple special function units in multi-core systems. By sharing the look-up tables among special function units, the proposed scheme reduces look-up table size by 54% and read power consumption by 75% for 8-core/4-bank configuration. However, the performance loss is less than 25% for various benchmark applications. As the number of cores increases for the future multi-core systems, the effects of area and power saving by the proposed schemes increase, but the performance loss is reduced. Young-Jun Kim 0001, Kyusik Chung, Lee-Sup Kim, Seong Mo Park |
ISCAS | 3 |
| 2009 | A Spread Spectrum Clock Generator with Spread Ratio Error Reduction Scheme for DisplayPort Main LinkabstractIn this paper, a spread spectrum clock generator (SSCG) with a process variation compensator for DisplayPort main link is presented. The process variation compensator not only reduces the error of spread ratio but also guarantees the reliability of the operation of an SSCG against process variation. The proposed SSCG has been implemented in 0.18-mum CMOS process and supports 10-phase 270 MHz and 162 MHz output clock. The experimental results show that the average rms jitter of 270 MHz output clock is 4.7 ps without spread spectrum clocking. 8.75 dBm of the peak reduction and 5000 ppm of spread ratio with the process variation compensator are achieved. Won-Young Lee, Lee-Sup Kim |
ISCAS | 2 |
| 2009 | Shader-based tessellation to save memory bandwidth in a mobile multimedia processor
Kyusik Chung, Chang-Hyo Yu, Donghyun Kim 0014, Lee-Sup Kim |
Comput. Graph. | 4 |
| 2009 | A Floating-Point Unit for 4D Vector Inner Product with Reduced LatencyabstractThis paper presents the algorithm and implementation of a new high-performance functional unit for floating-point four-dimensional vector inner product (4D dot product; DP4), which is most frequently performed in 3D graphics application. The proposed IEEE-compliant DP4 unit computes Z = AB + CD + EF + GH in one path and keeps the intermediate rounding by IEEE-754 rounding to nearest even. The intermediate rounding is merged with shift alignment, and intermediate carry-propagated addition and normalization are omitted to reduce latency in the proposed architecture. The proposed DP4 unit is implemented with 0.18-mum CMOS technology and has 12.8-ns critical path delay, which is reduced by 45.5 percent compared to a previous DP4 implementation using discrete multipliers and adders. The proposed DP4 unit also reduces the cycle time of 3D graphics applications by 12.4 percent on the average compared to the usual 3D graphics FPU based on four-way multiply-add-fused units. Donghyun Kim 0014, Lee-Sup Kim |
IEEE Trans. Computers | 2 |
| 2009 | A 186-Mvertices/s 161-mW Floating-Point Vertex Processor With Optimized Datapath and Vertex CachesabstractIn this paper, a power efficient vertex processor for mobile graphics applications is presented. A four-threaded and four-issue expanded VLIW datapath with a quad-float vertex texture fetcher is proposed by exploiting graphics specific characteristics after evaluation of several candidate architectures. Instruction-level power control methods such as operand sharing and writeback re-allocation along with operand isolations and gated clocks result in 40.4% and 82% reduction in energy dissipation and energy delay product compared to the most widely used single threaded SIMD. The proposed processor with the optimized datapath and vertex caches implemented in a 0.18- mum 1P4M CMOS process achieves 186-Mvertices/s geometry performance which is the best result among the processors that are IEEE-754 compliant. Chang-Hyo Yu, Kyusik Chung, Donghyun Kim 0014, Seok-Hoon Kim, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2008 | High speed serial interface for mobile LCD driver ICabstractA high speed serial interface is proposed for a mobile VGA-resolution TFT-LCD driver IC. This high speed serial interface is intended to replace the legacy RGB interface for the LCD driver IC's. The transmission channel consists of one clock and two data differential lines, and thereby we can reduce the number of signal lines going through the hinge of a mobile phone from 28 down to 6. The simple synchronization encoding scheme facilitates the pixel synchronization. All channels conform to sub low-voltage differential signaling (SubLVDS) convention. The total data transfer rate can be up to 800 Mbps for two data channels. Hyun-Kyu Jeon, Hye-Ran Kim, Jung-Min Choi, Ju-Pyo Hong, Hyung-Seog Oh, Dae-Keun Han, Lee-Sup Kim |
ISCAS | 8 |
| 2008 | Clipping-ratio-independent 3D graphics clipping engine by dual-thread algorithmabstractIn this paper, we propose a clipping engine (CE) in 3D graphics. A conventional CE shows lower performance as clipping-ratio increases, because the process time of a clipped triangle is much longer than that of a non-clipped triangle. We focus on conserving performance, while clipping-ratio changes. Proposed architecture separates datapath into clip datapath and perspective division & viewport mapping (PDVM) datapath, which makes it possible for CE to handle two triangles at a time. Its performance gain is up to four times of clipping-ratio. It is implemented with 79 k logic gates and 3 kB SRAM in 0.18 um CMOS technology. Jeong-Hyun Kim, Kyusik Chung, Young-Jun Kim 0001, Seok-Hoon Kim, Lee-Sup Kim |
ISCAS | 5 |
| 2008 | Area-efficient pixel rasterization and texture coordinate interpolation
Donghyun Kim 0014, Lee-Sup Kim |
Comput. Graph. | 2 |
| 2007 | Triangle-Level Depth Filter Method for Bandwidth Reduction in 3D Graphics HardwareabstractIn this paper, we proposed an early depth test method which extends a pixel-level depth filter to a triangle-level depth filter. This method decides whether the triangles are removed as well as pixels by a mask plane called depth filter (DF). The proposed method evaluates the DF in the triangle-level by using a hierarchical map of the conventional DF; thus the method removes the invisible triangles before the rasterization. Therefore, the rasterization performance can be increased and the bandwidth is further saved than the conventional DF. As the simulation results using complex scenes, the memory bandwidth is reduced by up to 31% compared with the conventional pixel-level DF method Jae-Sung Yoon, Chang-Hyo Yu, Donghyun Kim 0014, Lee-Sup Kim |
ISCAS | 4 |
| 2006 | An efficient texture cache for programmable vertex shadersabstractVertex texturing is state-of-the-art functionality of the 3D geometry processor. However, it aggravates the bandwidth problem between 3D graphics hardware and external memory besides the per-pixel texturing. Since vertex texturing does not guarantee high locality all over the texture data unlike the per-pixel texturing, we propose a caching scheme that adaptively adjusts its operation by estimating the amount of the locality of access pattern. The proposed cache improves 27.0% of texture loading performance for general test scenes with only 9.6% of hardware overhead. Seunghyun Cho, Chang-Hyo Yu, Lee-Sup Kim |
ISCAS | 3 |
| 2006 | Vertex cache of programmable geometry processor for mobile multimedia applicationabstractVertex cache of programmable geometry processor is proposed and implemented. The proposed vertex cache is organized into pre-TnL vertex cache and post-TnL vertex cache. The pre-TnL vertex cache reduces 32.8% geometry bandwidth by reusing fetched vertices with 32 entries. The post-TnL vertex cache improves the performance of a geometry processor by reusing recently processed vertices with 16 entries. Its performance gain is 1.69 on the average. 1.9k logic gates and 12kB single port SRAM are required for their implementation in 0.18 /spl mu/m CMOS technology. Kyusik Chung, Chang-Hyo Yu, Lee-Sup Kim |
ISCAS | 3 |
| 2006 | Charge-pump reducing current mismatch in DLLs and PLLsabstractConventional CMOS charge-pump circuits have some current mismatch problems. The current mismatch induces a phase offset which deteriorates the performance of PLLs or DLLs. This paper investigates causes of current mismatch of conventional CMOS charge-pump and proposes new charge-pump circuits to reduce the current mismatch. The DLL with proposed charge-pump is simulated in a 0.18 /spl mu/m CMOS process. Kyung-Soo Ha, Lee-Sup Kim |
ISCAS | 2 |
| 2006 | A 0.18µm CMOS 10Gb/s 1: 4 DEMUX using replica-bias circuits for optical receiverabstractThis paper presents a 0.18mum CMOS 10Gb/s 1:4 demultiplexer using window blocking for stable level swing and replica bias circuit for specific swing level maintenance. A modified current mode logic (CML) is proposed for the high-speed operation of demultiplexer. It prevents holding incorrect data during data transition. Replica bias circuit consists of feedback circuit using simple comparator. The 1:4 demultiplexer is a binary tree type. It consumes 12.24mW in 10Gb/s data rates with 1.8V supply voltage Ju-Pyo Hong, Kyung-Soo Ha, Lee-Sup Kim |
ISCAS | 3 |
| 2006 | A low power SoC bus with low-leakage and low-swing techniqueabstractA novel low power SoC bus with low-leakage and low swing technique is proposed. The repeater used in the bus lines effectively reduces leakage power through stacking effect, not losing its logic values even in sleep mode. The proposed SoC bus reduces not only the leakage power in the normal active/sleep mode but also both dynamic and leakage power in the low swing operation mode. The proposed scheme reduces the total power by 44.5% compared to the conventional SoC bus architecture. Kwang-Il Oh, Seunghyun Cho, Lee-Sup Kim |
ISCAS | 3 |
| 2006 | A cost-effective VLSI architecture for anisotropic texture filtering in limited memory bandwidthabstractTexture mapping is one of the techniques that express realism in three-dimensional (3-D) graphics. To produce high-quality images, various anisotropic filtering methods have been proposed for texture mapping. These methods require more texels than isotropic (trilinear) filtering method. In spite of increases to texture memory bandwidth, however, texture memory bandwidth is still a bottleneck of texture-filtering hardware. Consequently, an exact filtering method is required for good-quality images in a limited texture memory bandwidth. In this paper, we propose anisotropic texture filtering based on edge functions. Our method proposes an exact footprint-shape approximation with edge functions for generating weights. For real-time filtering, the weight plays a key role in effective filtering of the restricted texels loaded from memory. The normalized value of the edge function gives the distance relative to the contribution of texels to a final intensity. Calculating a Gaussian filter using this normalized value, generates a good weight. The quality of rendered images is superior to other anisotropic filtering methods with the same restricted number of texels. For images of the same quality, our method requires less than half the texels of other methods. Consequently, the improvement in performance is more than twice that of other methods. With low hardware overheads, our method can be implemented at a reasonable cost. In practice, the algorithm is demonstrated through VLSI implementation. The hardware, which is described by verilog and synthesized with a 0.35-/spl mu/m 3.3-V standard cell library, is operated at 100 MHz and it generates 100 M texture-filtered RGB pixel-color values per second. Hyunchul Shin, Jin-Aeon Lee, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2006 | A low-power ROM using single charge-sharing capacitor and hierarchical bit lineabstractThis paper describes a low-power read-only memory (ROM) using a single charge-sharing capacitor (SCSC) and hierarchical bit line (HBL). The SCSC-ROM reduces the power consumption in bit lines. It lowers the swing voltage of bit lines to a minimal voltage by using a charge-sharing technique with a single capacitor. It implements the capacitor with dummy bit lines to improve noise immunity and to make it easier to design. Furthermore, the HBL saves power by reducing the capacitance and leakage current in bit lines. The SCSC-ROM also reduces the power consumption in control unit and predecoder by using the hierarchical word line decoder. The simulation result shows that the SCSC-ROM with 4 K/spl times/32 bits consumes only 37% power of a conventional ROM. An SCSC-ROM chip is fabricated in a 0.25-/spl mu/m CMOS process. It consumes 8.2 mW at 240 MHz with 2.5 V. Byung-Do Yang, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2005 | Geometry engine architecture with early backface culling hardware
Chang-Young Han, Yeon-Ho Im, Lee-Sup Kim |
Comput. Graph. | 3 |
| 2005 | A Method to Generate Soft Shadows Using a Layered Depth Image and WarpingabstractWe present an image-based method for propagating area light illumination through a Layered Depth Image (LDI) to generate soft shadows from opaque and nonrefractive transparent objects. In our approach, using the depth peeling technique, we render an LDI from a reference light sample on a planar light source. Light illumination of all pixels in an LDI is then determined for all the other sample points via warping, an image-based rendering technique, which approximates ray tracing in our method. We use an image-warping equation and McMillan's warp ordering algorithm to find the intersections between rays and polygons and to find the order of intersections. Experiments for opaque and nonrefractive transparent objects are presented. Results indicate our approach generates soft shadows fast and effectively. Advantages and disadvantages of the proposed method are also discussed. Yeon-Ho Im, Chang-Young Han, Lee-Sup Kim |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2004 | Adaptive Selection of an Index in a Texture CacheabstractFor a specified application, there is an opportunity to improve cache performance by smart choosing of index bits of a cache. A texture cache for texture mapping of 3D computer graphics is an example. Texel (texture pixel)-access characteristics for texture mapping are dependent on the rasterization order of a polygon. In this paper, we introduce A-index, a method to reduce memory bandwidth required for fetching texture image data by reducing cache miss through adaptive selection of an index according to span direction of a rasterizer. By designing a texture mapping hardware including a texture cache, it is verified that the number of cache misses and the number of total cycles are reduced by 21.6% and 8.8% for rendering of textured scenes. In addition, we show that the number of cache misses in a texture cache can be estimated more accurately by calculating intra-span replacement when the proposed A-index is used. Chun-Ho Kim, Lee-Sup Kim |
ICCD | 2 |
| 2003 | A clock delayed sleep mode domino logic for wide dynamic OR gateabstractA high performance and low power clock delayed sleep mode (CDSM) domino logic is proposed for wide fan-in domino logic. The CDSM-domino logic not only improves the robustness but also reduces the active and stand-by power. The proposed scheme reduces delay by 21%, dynamic power by 16%, and leakage power by 91% respectively compared to the typical wide fan-in domino logic in 0.18? CMOS technology. In addition, the sleep mode entrance power is reduced to 10-5 of the HS-domino logic [3]. Kwang-Il Oh, Lee-Sup Kim |
ISLPED | 2 |
| 2003 | Winscale: an image-scaling algorithm using an area pixel modelabstractWe propose a new scaling algorithm, winscale, which performs the scale up/down transform using an area pixel model rather than a point pixel model. The proposed algorithm has low complexity: the algorithm uses a maximum of four pixels of an original image to calculate one pixel of a scaled image. Nevertheless, the algorithm has good characteristics such as fine-edge and changeable smoothness. We implemented a hardware design of winscale using an FPGA and displayed some test scenes in an liquid crystal display panel using a digital visual interface. The hardware cost and the image quality were compared with those of the conventional image scaling algorithms. It is proved that winscale has good scale property with low complexity. Winscale can be used in various digital display devices that need image scaling, especially in applications that require good image quality with low hardware cost. Chun-Ho Kim, Si-Mun Seong, Jin-Aeon Lee, Lee-Sup Kim |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2003 | A low-power charge-recycling ROM architectureabstractThis paper describes a newly proposed low-power charge-recycling read-only memory (CR-ROM) architecture. The CR-ROM reduces the power consumption in bit lines, word lines, and precharge lines by recycling the previously used charge. In the proposed CR-ROM, bit-line swing voltage is lowered by the charge recycling between bit lines. When N bit lines recycle their charges, the swing voltage and the power of the bit lines become 1/N and 1/N/sup 2/ compared to the conventional ROMs, respectively. As the number of N increases, the power saving in bit lines becomes salient. Also, power consumption in word lines and precharge lines can be reduced theoretically to half by the proposed charge-recycling techniques. The simulation results show that the CR-ROM consumes 60%/spl sim/85% of the conventional low-power ROMs with 1 K /spl times/ 32 b. A CR-ROM with 32 Kb was implemented in a 0.35-/spl mu/m CMOS process. The power dissipation is 6.60 mW at 100 MHz with 3.3 V and the maximum operating clock frequency is 150 MHz. Byung-Do Yang, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2001 | An advanced contrast enhancement using partially overlapped sub-block histogram equalizationabstractAn advanced histogram-equalization algorithm for contrast enhancement is presented. Histogram equalization is the most popular algorithm for contrast enhancement due to its effectiveness and simplicity. It can be classified into two branches according to the transformation function used: global or local. Global histogram equalization is simple and fast, but its contrast-enhancement power is relatively low. Local histogram equalization, on the other hand, can enhance overall contrast more effectively, but the complexity of computation required is very high due to its fully overlapped sub-blocks. In this paper, a low-pass filter-type mask is used to get a nonoverlapped sub-block histogram-equalization function to produce the high contrast associated with local histogram equalization but with the simplicity of global histogram equalization. This mask also eliminates the blocking effect of nonoverlapped sub-block histogram-equalization. The low-pass filter-type mask is realized by partially overlapped sub-block histogram-equalization (POSHE). With the proposed method, since the sub-blocks are much less overlapped, the computation overhead is reduced by a factor of about 100 compared to that of local histogram equalization while still achieving high contrast. The proposed algorithm can be used for commercial purposes where high efficiency is required, such as camcorders, closed-circuit cameras, etc. Joung-Youn Kim, Lee-Sup Kim, Seung-Ho Hwang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2001 | A hardware cost minimized fast Phong shaderabstractOne of the most successful algorithms that bring realism to the world of three-dimensional (3-D) image generation is Phong shading. With the continuous improvement in VLSI technology and the demand for higher realism, this algorithm is amenable to the commercially available hardware implementation for real-time rendering in 3-D graphics. Taylor series approximation is appropriate for the hardware implementation of fast Phong shading. However, in this method, the exponentiation of the cosine term requires a very large ROM table. This paper describes the minimization of this overhead in terms of hardware size by proposing an adaptive-compressed nonuniform quantization method. With this method, the ROM table is reduced to 1/64th of the size required for a uniform quantization method while the picture quality is maintained. Due to the reduced ROM table size, the size of the total hardware required for fast Phong shading is minimized to 1/56th of the original size. Hyunchul Shin, Jin-Aeon Lee, Lee-Sup Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2000 | A Memory Architecture with 4-Address Configurations for Video Signal ProcessingabstractA memory architecture with four-address configurations is proposed for video signal processing. The implemented 8-words /spl times/ 64-bits 8-port SRAM has 256-bit simultaneous data accessibility by horizontal and vertical address configurations and has 25.6 Gbits/s of high bandwidth. Sunho Chang, Jong-Sun Kim, Lee-Sup Kim |
DATE | 3 |
| 2000 | An advanced contrast enhancement using partially overlapped sub-block histogram equalizationabstractIn this paper, an advanced histogram equalization algorithm for contrast enhancement is presented. Histogram equalization is the most popular algorithm for contrast enhancement due to its effectiveness and simplicity. Global histogram equalization is simple and fast, but its contrast enhancement power is relatively low. Local histogram equalization, on the other hand, can enhance overall contrast more effectively, but the computational complexity is very high due to its fully overlapped sub-blocks. For high contrast and simple calculation, a low pass filter type mask is proposed. The low pass filter type mask is realized by partially overlapped sub-block histogram equalization (POSHE). With the proposed method, the computation overhead is reduced by a factor of about one hundred compared to that of local histogram equalization while still achieving high contrast. Joung-Youn Kim, Lee-Sup Kim, Seung-Ho Hwang |
ISCAS | 2 |
| 2000 | SPARP: a single pass antialiased rasterization processor
Jin-Aeon Lee, Lee-Sup Kim |
Comput. Graph. | 2 |
| 2000 | A programmable 3.2-GOPS merged DRAM logic for video signal processingabstractThis paper proposes a programmable high-performance architecture of datapath in the merged DRAM logic (MDL) for video signal processing. A model of a datapath in the programmable MDL is generated, and two basic parameters, total required clock cycles (TRCC) and DRAM access rate (DAR), are defined by analysis of the model. Design guidelines are suggested for the optimized video signal processor based on the modeling and analysis of the MDL. The inverse discrete cosine transform (IDCT) and motion compensation (MC) of the video signal processing are analyzed in the MDL architecture. Two measures, TRCC and DAR, are determined such that the data bandwidth between DRAM and logic is not a bottleneck in the MDL architecture. The efficient datapath is designed based on these design guidelines. The datapath has processing units (ALU, MAC, and barrel shifter) with splittabilities of data and multi-port SRAM. The maximum performance of the proposed datapath with 200-MHz clock frequency is 3.2 GOPS for 8-bit video signals, which can deal with a decoding high-level (1920/spl times/1080) in MPEG. The proposed MDL architecture has 2.1-4.8 times higher performance compared with conventional dedicated hardware chips. It can also be used for other multimedia signal processing due to its programmability. Sunho Chang, Bum-Sik Kim, Lee-Sup Kim |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2000 | A real-time wavelet vector quantization algorithm and its VLSI architectureabstractA real-time wavelet image compression algorithm using vector quantization and its VLSI architecture are proposed. The proposed zerotree wavelet vector quantization (WVQ) algorithm focuses on the problem of how to reduce the computation time to encode wavelet images with high coding efficiency. A conventional wavelet image-compression algorithm exploits the tree structure of wavelet coefficients coupled with scalar quantization. However, they can not provide the real-time computation because they use iterative methods to decide zerotrees. In contrast, the zerotree WVQ algorithm predicts in real-time zero-vector trees of insignificant wavelet vectors by a noniterative decision rule and then encodes significant wavelet vectors by the classified VQ. These cause the zerotree WVQ algorithm to provide the best compromise between the coding performance and the computation time. The noniterative decision rule was extracted by the simulation results. Moreover, the zerotree WVQ exploits the multistage VQ to encode the lowest frequency subband, which is generally known to be robust to wireless channel errors. The proposed WVQ VLSI architecture has only one VQ module to execute in real-time the proposed zerotree WVQ algorithm by utilizing the vacant cycles for zero-vector trees which are not transmitted. And the VQ module has only L+1 processing elements (PEs) for the real-time minimum distance calculation, where the codebook size is L. L PEs are for Euclidean distance calculation and a PE is for parallel distance comparison. Seung-Kwon Paek, Lee-Sup Kim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1998 | A minimized hardware architecture of fast Phong shader using Taylor series approximation in 3D graphicsabstractOne of the most successful algorithms that bring realism to the world of 3D-image generation is Phong shading. But, Gouraud shading has been used instead of Phong shading because of per pixel computation and hardware costs. However, with continuous improvement of VLSI technologies and request for higher realism, real-time Phong shading will be the next technology-push in 3D graphics. Taylor series approximation is an algorithm of fast Phong shading algorithms that is appropriate for hardware implementation. But, the hardware implementation of this algorithm requires a large ROM table that induces an overhead in terms of hardware size. We reduced this overhead by minimizing the ROM table size while keeping visual quality through visual comparison. We minimized a large ROM table size of a uniform quantization method to 1/64 using an adaptive-compressed non-uniform quantization method. By minimizing the ROM table size, we could minimize the total hardware size to 1/56. Hyunchul Shin, Jin-Aeon Lee, Lee-Sup Kim |
ICCD | 3 |
| 1997 | A new 4-2 adder and booth selector for low power MAC unitabstractArticle A new 4-2 adder and booth selector for low power MAC unit Share on Authors: Bum-Sik Kim Department of Electrical Engineering, Korea Advanced Institute of Science and Technology, 373-1 Kusong-dong, Yusong-gu, Taejon 305701 Korea Department of Electrical Engineering, Korea Advanced Institute of Science and Technology, 373-1 Kusong-dong, Yusong-gu, Taejon 305701 KoreaView Profile , Dae-Hyum Chung View Profile , Lee-Sup Kim View Profile Authors Info & Claims ISLPED '97: Proceedings of the 1997 international symposium on Low power electronics and designAugust 1997 Pages 100–103https://doi.org/10.1145/263272.263295Published:01 August 1997 2citation379DownloadsMetricsTotal Citations2Total Downloads379Last 12 Months0Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Bum-Sik Kim, Dae-Hyum Chung, Lee-Sup Kim |
ISLPED | 3 |
| 1989 | Modeling of the distributed gate RC effect in MOSFET'sabstractScattering parameters of wide MOSFET devices have been measured at the wafer level in the frequency range up to 1 GHz. These scattering parameters are converted to Y-parameters for device characterization and compared with SPICE simulations of L-, Pi -, and T-ladder distributed RC circuits taking into account the effect of distributed gates. Good experimental agreement shows that this effect can be modeled appropriately by multisection RC ladders, especially T-ladder circuits. Differences among ladder circuits in modeling this effect are discussed qualitatively.> Lee-Sup Kim, Robert W. Dutton |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |