Sungjoo Yoo

dblp:82/6218 · DBLP profile ↗
← Back
115ranked-venue papers
6as first author
16since 2021 · last 2025
0000-0002-5853-0675ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 90 · 6 first-author · 3 since 2021Software engineering, systems software and programming languages · 31 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 10 since 2021Artificial intelligence and machine learning · 14 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2025 Dense-SfM: Structure from Motion with Dense Consistent Matching
abstract
We present Dense-SfM, a novel Structure from Motion (SfM) framework designed for dense and accurate 3D reconstruction from multi-view images. Sparse keypoint matching, which traditional SfM methods often rely on, limits both accuracy and point density, especially in textureless areas. Dense-SfM addresses this limitation by integrating dense matching with a Gaussian Splatting (GS) based track extension which gives more consistent, longer feature tracks. To further improve reconstruction accuracy, Dense-SfM is equipped with a multi-view kernelized matching module leveraging transformer and Gaussian Process architectures, for robust track refinement across multi-views. Evaluations on the ETH3D and Texture-Poor SfM datasets show that Dense-SfM offers significant improvements in accuracy and density over state-of-the-art methods. Project page: https://icetea-cv.github.io/densesfm/.
Sungjoo Yoo
CVPR2
2025 NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
abstract
Large Language Model (LLM) inference is typically memory-intensive, especially when processing large batch sizes and long sequences, due to the large size of key-value (KV) cache. Vector Quantization (VQ) is recently adopted to alleviate this issue, but we find that the existing approach is susceptible to distribution shift due to its reliance on calibration datasets. To address this limitation, we introduce $\textbf{NSNQuant}$, a calibration-free Vector Quantization (VQ) technique designed for low-bit compression of the KV cache. By applying a three-step transformation—$\textbf{1)}$ a token-wise normalization ($\textbf{N}$ormalize), $\textbf{2)}$ a channel-wise centering ($\textbf{S}$hift), and $\textbf{3)}$ a second token-wise normalization ($\textbf{N}$ormalize)—with Hadamard transform, NSNQuant effectively aligns the token distribution with the standard normal distribution. This alignment enables robust, calibration-free vector quantization using a single reusable codebook. Extensive experiments show that NSNQuant consistently outperforms prior methods in both 1-bit and 2-bit settings, offering strong generalization and up to 3$\times$ throughput gains over full-precision baselines.
Donghyun Son, Euntae Choi, Sungjoo Yoo
NeurIPS3
2024 MetaMix: Meta-State Precision Searcher for Mixed-Precision Activation Quantization
abstract
Mixed-precision quantization of efficient networks often suffer from activation instability encountered in the exploration of bit selections. To address this problem, we propose a novel method called MetaMix which consists of bit selection and weight training phases. The bit selection phase iterates two steps, (1) the mixed-precision-aware weight update, and (2) the bit-search training with the fixed mixed-precision-aware weights, both of which combined reduce activation instability in mixed-precision quantization and contribute to fast and high-quality bit selection. The weight training phase exploits the weights and step sizes trained in the bit selection phase and fine-tunes them thereby offering fast training. Our experiments with efficient and hard-to-quantize networks, i.e., MobileNet v2 and v3, and ResNet-18 on ImageNet show that our proposed method pushes the boundary of mixed-precision quantization, in terms of accuracy vs. operations, by outperforming both mixed- and single-precision SOTA methods.
Han-Byul Kim, Joo Hyung Lee, Sungjoo Yoo
AAAI3
2024 MFOS: Model-Free & One-Shot Object Pose Estimation
abstract
Existing learning-based methods for object pose estimation in RGB images are mostly model-specific or category based. They lack the capability to generalize to new object categories at test time, hence severely hindering their practicability and scalability. Notably, recent attempts have been made to solve this issue, but they still require accurate 3D data of the object surface at both train and test time. In this paper, we introduce a novel approach that can estimate in a single forward pass the pose of objects never seen during training, given minimum input. In contrast to existing state-of-the-art approaches, which rely on task-specific modules, our proposed model is entirely based on a transformer architecture, which can benefit from recently proposed 3D-geometry general pretraining. We conduct extensive experiments and report state-of-the-art one-shot performance on the challenging LINEMOD benchmark. Finally, extensive ablations allow us to determine good practices with this relatively new type of architecture in the field.
Yohann Cabon, Romain Brégier, Sungjoo Yoo, Jérôme Revaud
AAAI4
2024 Geometry Transfer for Stylizing Radiance Fields
abstract
Shape and geometric patterns are essential in defining stylistic identity. However, current 3D style transfer methods predominantly focus on transferring colors and textures, often overlooking geometric aspects. In this paper, we introduce Geometry Transfer, a novel method that leverages geometric deformation for 3D style transfer. This technique employs depth maps to extract a style guide, subsequently applied to stylize the geometry of radiance fields. Moreover, we propose new techniques that utilize geometric cues from the 3D scene, thereby enhancing aesthetic expressiveness and more accurately reflecting intended styles. Our extensive experiments show that Geometry Transfer enables a broader and more expressive range of stylizations, thereby significantly expanding the scope of 3D style transfer.
Hyunyoung Jung 0001, Seonghyeon Nam, Nikolaos Sarafianos, Sungjoo Yoo, Alexander Sorkine-Hornung
CVPR4
2024 TCP: A Tensor Contraction Processor for AI Workloads Industrial Product
abstract
We introduce a novel tensor contraction processor (TCP) architecture that offers a paradigm shift from traditional architectures that rely on fixed-size matrix multiplications. TCP aims at exploiting the rich parallelism and data locality inherent in tensor contractions, thereby enhancing both efficiency and performance of AI workloads.TCP is composed of coarse-grained processing elements (PEs) to simplify software development. In order to efficiently process operations with diverse tensor shapes, the PEs are designed to be flexible enough to be utilized as a large-scale single unit or a set of small independent compute units.We aim at maximizing data reuse on both levels of inter and intra compute units. To do that, we propose a circuit switch-based fetch network to flexibly connect compute units to enable inter-compute unit data reuse. We also exploit input broadcast to multiple contraction engines and input buffer based reuse to further exploit reuse behavior in tensor contraction. Our compiler explores the design space of tensor contractions considering tensor shapes and the order of their associated loop operations as well as the underlying accelerator architecture.A TCP chip was designed and fabricated in 5nm technology as the second-generation product of Furiosa AI, offering 256/512/1024 TOPS (BF16/FP8 or INT8/INT4) with 256 MB SRAM and 1.5 TB/s 48 GB HBM3 under 150 W TDP. Commercialization will start in August 2024.We performed an extensive case study of running the LLaMA-2 7B model and evaluated its performance and power efficiency on various configurations of sequence length and batch size. For this model, TCP is 2.7 × and 4.1 × better than H100 and L40s, respectively, in terms of performance per watt.
Hanjoon Kim, Byeongwook Bae, Hyunmin Jeong, Sang Min Lee 0014, Jeseung Yeon, Changjae Park, Boncheol Gu, Changman Lee, Jaeick Bae, SungGyeong Bae, Yojung Cha, Wooyoung Choe, Jonguk Choi, Juho Ha, Hyuck Han, Namoh Hwang, Seokha Hwang, Kiseok Jang, Haechan Je, Hojin Jeon, Jaewoo Jeon, Hyunjun Jeong, Yeonsu Jung, Dongok Kang, Hyewon Kim, Muhwan Kim, Sewon Kim, Suhyung Kim, Yong Kim, Youngsik Kim, Younki Ku, Jeong Ki Lee, Juyun Lee, Seokho Lee, Minwoo Noh, Hyuntaek Oh, Gyunghee Park, Jimin Seo, Jungyoung Seong, June Paik, Nuno P. Lopes, Sungjoo Yoo
ISCA49
2024 NeRF-PIM: PIM Hardware-Software Co-Design of Neural Rendering Networks
abstract
Neural radiance field (NeRF) has emerged as a state-of-the-art technique, offering unprecedented realism in rendering. Despite its advancements, the adoption of NeRF is constrained by high computational cost, leading to slow rendering speed. Voxel-based optimization of NeRF addresses this by reducing the computational cost, but it introduces substantial memory overheads. To address this problem, we propose NeRF-PIM, a hardware-software co-design approach. In order to address the problem of the memory accesses to the large model (of the voxel grid) with poor locality and low compute density, we propose exploiting processing-in-memory (PIM) together with PIM-aware software optimizations in terms of the data layout, redundancy removal, and computation reuse. Our PIM hardware aims to accelerate the trilinear interpolation and dot product operations. Specifically, to address the low utilization of internal bandwidth due to the random accesses to the voxels, we propose a data layout that judiciously exploits the characteristics of the interpolation operation on the voxel grid, which helps remove bank conflicts in voxel accesses and also improves the efficiency of PIM command issue by exploiting the all-bank mode in the existing PIM device. As PIM-aware software optimizations, we also propose occupancy-grid-aware pruning and one-voxel two-sampling (1V2S) methods, which contribute to compute the efficiency improvement (by avoiding the redundant computation on the empty space) and memory traffic reduction (by reusing the per-voxel dot product results). We conduct experiments using an actual baseline HBM-PIM device. Our NeRF-PIM demonstrates a speedup of 7.4 and$5.0\times $compared to the baseline on the two datasets, Synthetic-NeRF and Tanks and Temples, respectively.
Jaeyoung Heo, Sungjoo Yoo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 AnyFlow: Arbitrary Scale Optical Flow with Implicit Neural Representation
abstract
To apply optical flow in practice, it is often necessary to resize the input to smaller dimensions in order to reduce computational costs. However, downsizing inputs makes the estimation more challenging because objects and motion ranges become smaller. Even though recent approaches have demonstrated high-quality flow estimation, they tend to fail to accurately model small objects and precise boundaries when the input resolution is lowered, restricting their applicability to high-resolution inputs. In this paper, we introduce AnyFlow, a robust network that estimates accurate flow from images of various resolutions. By representing optical flow as a continuous coordinate-based representation, AnyFlow generates outputs at arbitrary scales from low-resolution inputs, demonstrating superior performance over prior works in capturing tiny objects with detail preservation on a wide range of scenes. We establish a new state-of-the-art performance of cross-dataset generalization on the KITTI dataset, while achieving comparable accuracy on the online benchmarks to other SOTA methods.
Hyunyoung Jung 0001, Zhuo Hui, Feng Liu 0015, Sungjoo Yoo, Denis Demandolx
CVPR6
2023 NIPQ: Noise proxy-based Integrated Pseudo-Quantization
abstract
Straight-through estimator (STE), which enables the gradient flow over the non-differentiable function via approximation, has been favored in studies related to quantization-aware training (QAT). However, STE incurs unstable convergence during QAT, resulting in notable quality degradation in low precision. Recently, pseudo-quantization training has been proposed as an alternative approach to updating the learnable parameters using the pseudo-quantization noise instead of STE. In this study, we propose a novel noise proxy-based integrated pseudo-quantization (NIPQ) that enables unified support of pseudo-quantization for both activation and weight by integrating the idea of truncation on the pseudo-quantization framework. NIPQ updates all of the quantization parameters (e.g., bit-width and truncation boundary) as well as the network parameters via gradient descent without STE instability. According to our extensive experiments, NIPQ outperforms existing quantization algorithms in various vision and language applications by a large margin.
Juncheol Shin, Junhyuk So, Sein Park, Seungyeop Kang, Sungjoo Yoo, Eunhyeok Park
CVPR5
2023 Masked Token Similarity Transfer for Compressing Transformer-Based ASR Models
abstract
Recent self-supervised automatic speech recognition (ASR) models based on transformers are showing best performance, but their footprint is too large to be trained on low-resource environments or deployed to edge devices. Knowledge distillation (KD) can be employed to reduce the model size. However, setting embedding dimension of teacher and student network to different values makes it difficult to transfer token embeddings for better performance. To mitigate this issue, we present a novel KD method in which student mimics the prediction vector of teacher under our proposed masked token similarity transfer (MTST) loss where the temporal relation between a token and the other unmasked ones is encoded into a dimension-agnostic token similarity vector. Under our transfer learning setting with a fine-tuned teacher, our proposed methods reduce the model size of student to 28.3% of teacher’s while word error rate on test-clean subset in LibriSpeech corpus is 4.93%, which surpasses prior works. Our source code will be made available.
Euntae Choi, Youshin Lim, Byeong-Yeol Kim, Hyung Yong Kim, Hanbin Lee, Yunkyu Lim, Seung Woo Yu, Sungjoo Yoo
ICASSP8
2022 BASQ: Branch-wise Activation-clipping Search Quantization for Sub-4-bit Neural Networks
Han-Byul Kim, Eunhyeok Park, Sungjoo Yoo
ECCV (12)3
2022 TernaryNeRF: Quantizing Voxel Grid-based NeRF Models
abstract
Photo-realistic neural rendering, represented by neural radiance field (NeRF), is considered to be a key technology for AR/VR applications and has been actively studied in recent years. In order to enable widespread adoptions of AR/VR, it is critical to enable low-cost and high-quality rendering on mobile and server systems. In our work, we investigate the feasibility of low-precision representation on the two state-of-the-art NeRF models, InstantNeRF and TensoRF. Our proposed quantization is based on our observation on the characteristics of trained NeRF models. In order to reduce the model size while limiting the loss of rendering quality due to model compression, we propose quantizing the portion of model which dominates the total model size while being robust to aggressive quantization. In our experiments, we demonstrate our proposed ternary quantization can reduce by$7 \times \sim 15\times$the model sizes of state-of-the-art NeRF models at a negligible loss of rendering quality, which, we consider, will contribute to the AR/VR adoptions on mobile and server systems.
Seungyeop Kang, Sungjoo Yoo
RSP2
2022 Introduction to the Special Section on Energy-Efficient AI Chips
abstract
No abstract available.
Vikas Chandra, Yiran Chen 0001, Sungjoo Yoo
ACM Trans. Design Autom. Electr. Syst.3
2021 Exploiting Spatial Dimensions of Latent in GAN for Real-Time Image Editing
abstract
Generative adversarial networks (GANs) synthesize realistic images from random latent vectors. Although manipulating the latent vectors controls the synthesized outputs, editing real images with GANs suffers from i) time-consuming optimization for projecting real images to the latent vectors, ii) or inaccurate embedding through an encoder. We propose StyleMapGAN: the intermediate latent space has spatial dimensions, and a spatially variant modulation replaces AdaIN. It makes the embedding through an encoder more accurate than existing optimization-based methods while maintaining the properties of GANs. Experimental results demonstrate that our method significantly outperforms state-of-the-art models in various image manipulation tasks such as local editing and image interpolation. Last but not least, conventional editing methods on GANs are still valid on our StyleMapGAN. Source code is available at https://github.com/naver-ai/StyleMapGAN.
Hyunsu Kim, Yunjey Choi, Sungjoo Yoo, Youngjung Uh
CVPR4
2021 Fine-grained Semantics-aware Representation Enhancement for Self-supervised Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation has been widely studied, owing to its practical importance and recent promising improvements. However, most works suffer from limited supervision of photometric consistency, especially in weak texture regions and at object boundaries. To overcome this weakness, we propose novel ideas to improve self-supervised monocular depth estimation by leveraging cross-domain information, especially scene semantics. We focus on incorporating implicit semantic knowledge into geometric representation enhancement and suggest two ideas: a metric learning approach that exploits the semantics-guided local geometry to optimize intermediate depth representations and a novel feature fusion module that judiciously utilizes cross-modality between two heterogeneous feature representations. We comprehensively evaluate our methods on the KITTI dataset and demonstrate that our method outperforms state-of-the-art methods. The source code is available at https://github.com/hyBlue/FSRE-Depth.
Hyunyoung Jung 0001, Eunhyeok Park, Sungjoo Yoo
ICCV3
2021 FPGA Prototyping of Systolic Array-based Accelerator for Low-Precision Inference of Deep Neural Networks
abstract
In this study, we aim to design an energy-efficient computation system for deep neural networks on edge devices. To maximize energy efficiency, we design a novel hardware accelerator that supports low-precision computation and sparsity-aware structured zero-skipping on top of the well-known systolic-array structure. In addition, we introduce a full-stack software platform, including a model optimizer, instruction compiler, and host interface, to translate the pre-trained PyTorch model to the proposed accelerator and orchestrate it automatically. We validate the entire system by prototyping the accelerator on the Xilinx Alveo U250 FPGA board and demonstrating the inference of the 4-bit ResNet-50 model through the software stack. According to our experiment, our platform shows 317 GOPS inference speed and 51.96 GOPS/W energy efficiency for ResNet-50 on Xilinx Alveo U250 FPGA at 108 MHz, which is comparable to the advanced commercial acceleration system in terms of energy efficiency.
Soobeom Kim, Seunghwan Cho, Eunhyeok Park, Sungjoo Yoo
RSP4
2020 PROFIT: A Novel Training Method for sub-4-bit MobileNet Models
Eunhyeok Park, Sungjoo Yoo
ECCV (6)2
2020 MEANTIME: Mixture of Attention Mechanisms with Multi-temporal Embeddings for Sequential Recommendation
abstract
Recently, self-attention based models have achieved state-of-the-art performance in sequential recommendation task. Following the custom from language processing, most of these models rely on a simple positional embedding to exploit the sequential nature of the user’s history. However, there are some limitations regarding the current approaches. First, sequential recommendation is different from language processing in that timestamp information is available. Previous models have not made good use of it to extract additional contextual information. Second, using a simple embedding scheme can lead to information bottleneck since the same embedding has to represent all possible contextual biases. Third, since previous models use the same positional embedding in each attention head, they can wastefully learn overlapping patterns. To address these limitations, we propose MEANTIME (MixturE of AtteNTIon mechanisms with Multi-temporal Embeddings) which employs multiple types of temporal embeddings designed to capture various patterns from the user’s behavior sequence, and an attention structure that fully leverages such diversity. Experiments on real-world data show that our proposed method outperforms current state-of-the-art sequential recommendation methods, and we provide an extensive ablation study to analyze how the model gains from the diverse positional information.
Sung Min Cho, Eunhyeok Park, Sungjoo Yoo
RecSys3
2020 $Q$ -Value Prediction for Reinforcement Learning Assisted Garbage Collection to Reduce Long Tail Latency in SSD
abstract
Garbage collection (GC), an essential operation for flash storage systems, causes long tail latency which is one of the key problems in real-time and quality-critical systems. In this article, we take advantage of reinforcement learning (RL) to reduce long tail latency. Especially, we propose two novel techniques which are: 1) Q-table cache (QTC) and 2) Q-value prediction. The QTC allows us to utilize appropriate and frequently recurring key states at a small memory cost. We propose a neural network called Q-value prediction network (QP Net) that predicts the initial Q-value of a new state in the QTC. The integrated solution of QTC and QP Net enables us to benefit from both short-term (by QTC) and long-term (by QP Net) history of system behavior to reduce the long tail latency. The experimental results demonstrate that the proposed scheme offers significant (by 25%-37%) reductions in the long tail latency of storageintensive workloads compared with the state-of-the-art solution that adopts an RL-assisted GC scheduler.
Won-Kyung Kang, Sungjoo Yoo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 Machine Learning at Facebook: Understanding Inference at the Edge
abstract
At Facebook, machine learning provides a wide range of capabilities that drive many aspects of user experience including ranking posts, content understanding, object detection and tracking for augmented and virtual reality, speech and text translations. While machine learning models are currently trained on customized data-center infrastructure, Facebook is working to bring machine learning inference to the edge. By doing so, user experience is improved with reduced latency (inference time) and becomes less dependent on network connectivity. Furthermore, this also enables many more applications of deep learning with important features only made available at the edge. This paper takes a data-driven approach to present the opportunities and design challenges faced by Facebook in order to enable machine learning inference locally on smart phones and other edge platforms.
Carole-Jean Wu, David Brooks 0001, Douglas Chen, Sy Choudhury, Marat Dukhan, Kim M. Hazelwood, Eldad Isaac, Yangqing Jia, Bill Jia, Tommer Leyvand, Yang Lu 0013, Lin Qiao, Brandon Reagen, Joe Spisak, Fei Sun 0002, Andrew Tulloch, Peter Vajda, Xiaodong Wang 0020, Yanghan Wang, Bram Wasti, Yiming Wu 0013, Ran Xian, Sungjoo Yoo, Peizhao Zhang
HPCA25
2019 Tag2Pix: Line Art Colorization Using Text Tag With SECat and Changing Loss
abstract
Line art colorization is expensive and challenging to automate. A GAN approach is proposed, called Tag2Pix, of line art colorization which takes as input a grayscale line art and color tag information and produces a quality colored image. First, we present the Tag2Pix line art colorization dataset. A generator network is proposed which consists of convolutional layers to transform the input line art, a pre-trained semantic extraction network, and an encoder for input color information. The discriminator is based on an auxiliary classifier GAN to classify the tag information as well as genuineness. In addition, we propose a novel network structure called SECat, which makes the generator properly colorize even small features such as eyes, and also suggest a novel two-step training method where the generator and discriminator first learn the notion of object and shape and then, based on the learned notion, learn colorization, such as where and how to place which color. We present both quantitative and qualitative evaluations which prove the effectiveness of the proposed method.
Hyunsu Kim, Ho Young Jhoo, Eunhyeok Park, Sungjoo Yoo
ICCV4
2019 Machine Learning-Based Automatic Generation of eFuse Configuration in NAND Flash Chip
abstract
Post fabrication process is becoming more and more important as memory technology becomes complex, in the bid to satisfy target performance and yield across diverse business domains, such as servers, PCs, automotive, mobiles, and embedded devices, etc. Electronic fuse adjustment (eFuse optimization and trimming) is a traditional method used in the post fabrication processing of memory chips. Engineers adjust eFuse to compensate for wafer inter-chip variations or guarantee the operating characteristics, such as reliability, latency, power consumption, and I/O bandwidth. These require highly skilled expert engineers and yet take significant time. This paper proposes a novel machine learning-based method of automatic eFuse configuration to meet the target NAND flash operating characteristics. The proposed techniques can maximally reduce the expert engineer's workload. The techniques consist of two steps: initial eFuse generation and eFuse optimization. In the first step, we apply the variational autoencoder (VAE) method to generate an initial eFuse configuration that will probably satisfy the target characteristics. In the second step, we apply the genetic algorithm (GA), which attempts to improve the initial eFuse configuration and finally achieve the target operating characteristics. We evaluate the proposed techniques with Samsung 64-Stacked vertical NAND (VNAND) in mass production. The automatic eFuse configuration takes only two days to complete the implementation.
Jisuk Kim, Jin-Yub Lee, Sungjoo Yoo
ITC3
2018 Dynamic management of key states for reinforcement learning-assisted garbage collection to reduce long tail latency in SSD
abstract
Garbage collection (GC) is one of main causes of the long-tail latency problem in storage systems. Long-tail latency due to GC is more than 100 times greater than the average latency at the 99th percentile. Therefore, due to such a long tail latency, real-time systems and quality-critical systems cannot meet the system requirements. In this study, we propose a novel key state management technique of reinforcement learning-assisted garbage collection. The purpose of this study is to dynamically manage key states from a significant number of state candidates. Dynamic management enables us to utilize suitable and frequently recurring key states at a small area cost since the full states do not have to be managed. The experimental results show that the proposed technique reduces by 22--25% the long-tail latency compared to a state-of-the-art scheme with real-world workloads.
Won-Kyung Kang, Sungjoo Yoo
DAC2
2018 Joint optimization of speed, accuracy, and energy for embedded image recognition systems
abstract
This paper presents the image recognition system that won the first prize in the LPIRC (Low Power Image Recognition Challenge) in 2017. The goal of the challenge is to maximize the ratio between the accuracy and energy consumption within a time limit of 10 minutes for the processing of 20,000 images. Among three conflicting goals of accuracy, speed, and energy consumption, we considered the trade-off between accuracy and speed first to select Nvidia Jetson TX2 as the hardware platform and Tiny YOLO as the image recognition algorithm. Next, we applied a series of software optimization techniques to improve throughput, such as pipelining, multithreading, Tucker decomposition, and 16-bit quantization. Lastly, we explored the CPU and GPU frequencies to minimize the total energy consumption. As a result, we could achieve an accuracy of 0.24 mAP with energy consumption of 2.08Wh, which corresponds to the score of 0.11931, 2.7 times higher than the winner of LPIRC 2016.
Duseok Kang, Jintaek Kang, Sungjoo Yoo, Soonhoi Ha
DATE4
2018 Value-Aware Quantization for Training and Inference of Neural Networks
Eunhyeok Park, Sungjoo Yoo, Peter Vajda
ECCV (4)2
2018 Energy-Efficient Neural Network Accelerator Based on Outlier-Aware Low-Precision Computation
abstract
Owing to the presence of large values, which we call outliers, conventional methods of quantization fail to achieve significantly low precision, e.g., four bits, for very deep neural networks, such as ResNet-101. In this study, we propose a hardware accelerator, called the outlier-aware accelerator (OLAccel). It performs dense and low-precision computations for a majority of data (weights and activations) while efficiently handling a small number of sparse and high-precision outliers (e.g., amounting to 3% of total data). The OLAccel is based on 4-bit multiply-accumulate (MAC) units and handles outlier weights and activations in a different manner. For outlier weights, it equips SIMD lanes of MAC units with an additional MAC unit, which helps avoid cycle overhead for the majority of outlier occurrences, i.e., a single occurrence in the SIMD lanes. The OLAccel performs computations using outlier activation on dedicated, high-precision MAC units. In order to avoid coherence problem due to updates from low- and high-precision computation units, both units update partial sums in a pipelined manner. Our experiments show that the OLAccel can reduce by 43.5% (27.0%), 56.7% (36.3%), and 62.2% (49.5%) energy consumption for AlexNet, VGG-16, and ResNet-18, respectively, compared with a 16-bit (8-bit) state-of-the-art zero-aware accelerator. The energy gain mostly comes from the memory components, the DRAM, and on-chip memory due to reduced precision.
Eunhyeok Park, Dongyoung Kim, Sungjoo Yoo
ISCA3
2018 FPGA Prototyping of Low-Precision Zero-Skipping Accelerator for Neural Networks
abstract
Owing to the ever-increasing usage of neural networks in various applications from mobile devices to data centers, hardware accelerators for neural networks have been widely studied in recent years. Recently proposed accelerators mainly utilize sparsity and/or reduced precision approach to accelerate neural networks. Since neural networks go deeper to be applied to more complex applications, future hardware accelerators for neural networks may require to utilize both sparsity and very low-precision methods, which however, have not been analyzed quantitatively. In this paper, we first introduce an end-to-end FPGA prototyping flow and apply it to a neural network accelerator which supports both fine-grained zero-skipping and very low-precision. We report our analyses of resource usage and performance by varying bit-width and zero data ratio. We summarize the paper with our lessons learned from our prototyping study on future zero-aware very-low-precision accelerators.
Dongyoung Kim, Soobeom Kim, Sungjoo Yoo
RSP3
2018 McDRAM: Low Latency and Energy-Efficient Matrix Computations in DRAM
abstract
We propose a novel memory architecture for in-memory computation called McDRAM, where DRAM dies are equipped with a large number of multiply accumulate (MAC) units to perform matrix computation for neural networks. By exploiting high internal memory bandwidth and reducing offchip memory accesses, McDRAM realizes both low latency and energy efficient computation. In our experiments, we obtained the chip layout based on the state-of-the-art memory, LPDDR4 where McDRAM is equipped with 2048 MACs in a single chip package with a small area overhead (4.7%). Compared with the state-ofthe-art accelerator, TPU and the power-efficient GPU, Nvidia P4, McDRAM offers 9.5× and 14.4× speedup, respectively, in the case that the large-scale MLPs and RNNs adopt the batch size of 1. McDRAM also gives 2.1× and 3.7× better computational efficiency in TOPS/W than TPU and P4, respectively, for the large batches.
Hyunsung Shin, Dongyoung Kim, Eunhyeok Park, Yongsik Park, Sungjoo Yoo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2018 Nonvolatile Write Buffer-Based Journaling Bypass for Storage Write Reduction in Mobile Devices
abstract
In mobile systems, such as smartphones, most of storage writes are incurred by the SQLite database (DB) system. These writes consist of two parts: writes to original data (e.g., SQLite DB file) and journaling-induced writes. In this paper, we first report our observations on the characteristics of these writes: (1) journaling writes dominate storage write, (2) overwrites are frequent in mobile storage, and most importantly, and (3) those overwrites change only a small portion of data. Based on these observations, we propose to employ small capacitor-backed nonvolatile write buffer and use it to durably store small writes without journaling. This allows us to completely avoid journal writes for those small writes, which contributes to significant reduction in journaling-induced writes while still providing the same level of data consistency and durability. In our mechanism, it is critical to minimize the capacity of nonvolatile write buffer so that the backing capacitor does not violate stringent resource/size requirements of mobile systems. Therefore, we propose three optimizations that make the best use of small write buffer: (1) write buffer that manages only the difference between old and new data, (2) a dynamic method to determine which data to store in the write buffer, and (3) an incremental flush policy, which controls the number of write buffer entries to be flushed. The experimental results show 92.4% and 85.1% storage write reductions on average when running single and multiple mobile applications, respectively, with only 8 KB write buffer and a tiny capacitor.
Mungyu Son, Junwhan Ahn, Sungjoo Yoo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 Weighted-Entropy-Based Quantization for Deep Neural Networks
abstract
Quantization is considered as one of the most effective methods to optimize the inference cost of neural network models for their deployment to mobile and embedded systems, which have tight resource constraints. In such approaches, it is critical to provide low-cost quantization under a tight accuracy loss constraint (e.g., 1%). In this paper, we propose a novel method for quantizing weights and activations based on the concept of weighted entropy. Unlike recent work on binary-weight neural networks, our approach is multi-bit quantization, in which weights and activations can be quantized by any number of bits depending on the target accuracy. This facilitates much more flexible exploitation of accuracy-performance trade-off provided by different levels of quantization. Moreover, our scheme provides an automated quantization flow based on conventional training algorithms, which greatly reduces the design-time effort to quantize the network. According to our extensive evaluations based on practical neural network models for image classification (AlexNet, GoogLeNet and ResNet-50/101), object detection (R-FCN with 50-layer ResNet), and language modeling (an LSTM network), our method achieves significant reductions in both the model size and the amount of computation with minimal accuracy loss. Also, compared to existing quantization schemes, ours provides higher accuracy with a similar resource constraint and requires much lower design effort.
Eunhyeok Park, Junwhan Ahn, Sungjoo Yoo
CVPR3
2017 Making DRAM Stronger Against Row Hammering
abstract
Modern DRAM suffers from a new problem called row hammering. The problem is expected to become more severe in future DRAMs mostly due to increased inter-row coupling at advanced technology. In order to address this problem, we present a probabilistically managed table (called PRoHIT) implemented on the DRAM chip. The table keeps track of victim row candidates in a probabilistic way and, in case of auto-refresh, the topmost entry is additionally refreshed thereby mitigating the row hammering problem. Our experiments with PARSEC benchmark and synthetic traces show that PRoHIT outperforms the state-of-the-art method, PARA, by 35.7% (PARSEC) in terms of the reduction ratio of row-hammer cases. Our proposed method also shows constantly superior performance to PARA for synthetic traces.
Mungyu Son, Hyunsun Park, Junwhan Ahn, Sungjoo Yoo
DAC4
2017 A novel zero weight/activation-aware hardware architecture of convolutional neural network
abstract
It is imperative to accelerate convolutional neural networks (CNNs) due to their ever-widening application areas from server, mobile to IoT devices. Based on the fact that CNNs can be characterized by a significant amount of zero values in both kernel weights and activations, we propose a novel hardware accelerator for CNNs exploiting zero weights and activations. We also report a zero-induced load imbalance problem, which exists in zero-aware parallel CNN hardware architectures, and present a zero-aware kernel allocation as a solution. According to our experiments with a cycle-accurate simulation model, RTL, and layout design of the proposed architecture running two real deep CNNs, pruned AlexNet [1] and VGG-16 [2], our architecture offers 4x/1.8x (AlexNet) and 5.2x/2.1x (VGG-16) speedup compared with state-of-the-art zero-agnostic/zero-activation-aware architectures.
Dongyoung Kim, Junwhan Ahn, Sungjoo Yoo
DATE3
2017 ExtraV: Boosting Graph Processing Near Storage with a Coherent Accelerator
abstract
In this paper, we propose ExtraV, a framework for near-storage graph processing. It is based on the novel concept of graph virtualization , which efficiently utilizes a cache-coherent hardware accelerator at the storage side to achieve performance and flexibility at the same time. ExtraV consists of four main components: 1) host processor, 2) main memory, 3) AFU (Accelerator Function Unit) and 4) storage. The AFU, a hardware accelerator, sits between the host processor and storage. Using a coherent interface that allows main memory accesses, it performs graph traversal functions that are common to various algorithms while the program running on the host processor (called the host program) manages the overall execution along with more application-specific tasks. Graph virtualization is a high-level programming model of graph processing that allows designers to focus on algorithm-specific functions. Realized by the accelerator, graph virtualization gives the host programs an illusion that the graph data reside on the main memory in a layout that fits with the memory access behavior of host programs even though the graph data are actually stored in a multi-level, compressed form in storage. We prototyped ExtraV on a Power8 machine with a CAPI-enabled FPGA. Our experiments on a real system prototype offer significant speedup compared to state-of-the-art software only implementations.
Jinho Lee 0001, Heesu Kim, Sungjoo Yoo, Kiyoung Choi, H. Peter Hofstee, Gi-Joon Nam, Mark Nutter, Damir A. Jamsek
Proc. VLDB Endow.3
2017 Reinforcement Learning-Assisted Garbage Collection to Mitigate Long-Tail Latency in SSD
abstract
NAND flash memory is widely used in various systems, ranging from real-time embedded systems to enterprise server systems. Because the flash memory has erase-before-write characteristics, we need flash-memory management methods, i.e., address translation and garbage collection. In particular, garbage collection (GC) incurs long-tail latency, e.g., 100 times higher latency than the average latency at the 99 th percentile. Thus, real-time and quality-critical systems fail to meet the given requirements such as deadline and QoS constraints. In this study, we propose a novel method of GC based on reinforcement learning. The objective is to reduce the long-tail latency by exploiting the idle time in the storage system. To improve the efficiency of the reinforcement learning-assisted GC scheme, we present new optimization methods that exploit fine-grained GC to further reduce the long-tail latency. The experimental results with real workloads show that our technique significantly reduces the long-tail latency by 29--36% at the 99.99 th percentile compared to state-of-the-art schemes.
Won-Kyung Kang, Dongkun Shin, Sungjoo Yoo
ACM Trans. Embed. Comput. Syst.3
2016 AIM: Energy-Efficient Aggregation Inside the Memory Hierarchy
abstract
In this article, we propose Aggregation-in-Memory (AIM), a new processing-in-memory system designed for energy efficiency and near-term adoption. In order to efficiently perform aggregation, we implement simple aggregation operations in main memory and develop a locality-adaptive host architecture for in-memory aggregation, called cache-conscious aggregation. Through this, AIM executes aggregation at the most energy-efficient location among all levels of the memory hierarchy. Moreover, AIM minimally changes existing sequential programming models and provides fully automated compiler toolchain, thereby allowing unmodified legacy software to use AIM. Evaluations show that AIM greatly improves the energy efficiency of main memory and the system performance.
Junwhan Ahn, Sungjoo Yoo, Kiyoung Choi
ACM Trans. Archit. Code Optim.2
2016 Prediction Hybrid Cache: An Energy-Efficient STT-RAM Cache Architecture
abstract
Spin-transfer torque RAM (STT-RAM) has emerged as an energy-efficient and high-density alternative to SRAM for large on-chip caches. However, its high write energy has been considered as a serious drawback. Hybrid caches mitigate this problem by incorporating a small SRAM cache for write-intensive data along with an STT-RAM cache. In such architectures, choosing cache blocks to be placed into the SRAM cache is the key to their energy efficiency. This paper proposes a new hybrid cache architecture called prediction hybrid cache. The key idea is to predict write intensity of cache blocks at the time of cache misses and determine block placement based on the prediction. We design a write intensity predictor that realize the idea by exploiting a correlation between write intensity of blocks and memory access instructions that incur cache misses of those blocks. It includes a mechanism to dynamically adapt the predictor to application characteristics. We also design a hybrid cache architecture in which write-intensive blocks identified by the predictor are placed into the SRAM region. Evaluations show that our scheme reduces energy consumption of hybrid caches by 28 percent (31 percent) on average compared to the existing hybrid cache architecture in a single-core (quad-core) system.
Junwhan Ahn, Sungjoo Yoo, Kiyoung Choi
IEEE Trans. Computers2
2016 Array Organization and Data Management Exploration in Racetrack Memory
abstract
As the descendant ofspin-transfer random access memory(STT-RAM), racetrack memory technology saves data in magnetic domains along nanoscopic wires. Such a unique structure can achieve unprecedentedly high storage density meanwhile inheriting the promising features of STT-RAM, such as fast access speed, non-volatility, zero standby power, hardness to soft errors, and compatibility with CMOS technology. Moreover, the recent success in planar racetrack nanowire promised its fabrication feasibility and continuous scalability. In this paper, we investigate the design and optimization of racetrack memory as last-level cache by embracing design considerations across multiple abstraction layers, including the cell design, the array structure, the architecture organization, and the data management. The cross-layer optimization makes racetrack memory based last-level cache achieve 6.4$\times$reduction in area, 25 percent enhancement in system performance, and 62 percent saving in energy consumption, compared to STT-RAM cache design. Its benefit over SRAM technology is even more significant.
Zhenyu Sun 0001, Xiuyuan Bi, Sungjoo Yoo, Hai Li 0001
IEEE Trans. Computers4
2016 Memory Access Scheduling for a Smart TV
abstract
A smart TV system-on-chip (SoC) has very heavy computation and memory demands that must be met with low-cost components. As a result, there is potentially an extremely high utilization of the channel between the SoC and its memory chip. This paper presents the design of a new memory access scheduler customized for the type of memory traffic typically encountered with smart TVs. This includes special accumulated hard real-time graphics requirements, user response-sensitive soft real-time requirements, and the need to provide high memory throughput and priority-handling capabilities even under extremely heavy memory traffic conditions. The simulation results show that the proposed memory access scheduler is able to achieve up to 98% of the ideal upper bound memory throughput when faced with extremely heavy memory traffic-this is a significant improvement over previous schedulers. Novel future prediction and light-handed priority handling methods are used to achieve these results while satisfying the unique real-time requirements of smart TVs.
Cheul-Hee Hahm, Sunggu Lee, Sungjoo Yoo
IEEE Trans. Circuits Syst. Video Technol.4
2016 Improving Write Performance by Controlling Target Resistance Distributions in MLC PRAM
abstract
Multi-level cell (MLC) phase change RAM (PRAM) is expected to offer lower cost main memory than DRAM. However, poor write performance is one of the most critical problems for practical applications of MLC PRAM. In this article, we present two schemes to improve write performance by controlling the target resistance distribution of MLC PRAM cells. First, we propose multiple RESET/SET operations that relax the target resistance bands of intermediate logic levels with additional RESET/SET operations, which reduces the program time of intermediate logic levels, thereby improving write performance. Second, we propose a two-step write scheme consisting of lightweight write and idle-time completion write that exploits the fact that hot dirty data tend to be overwritten in a short time period and the MLC PRAM often has long idle times. Experimental results show that the multiple RESET/SET and two-step write schemes result in an average IPC improvement of 15.7% and 10.4%, respectively, on a hybrid DRAM/PRAM main memory subsystem. Furthermore, their integrated solution results in an average IPC improvement of 23.2% (up to 46.4%).
Youngsik Kim, Sungjoo Yoo, Sunggu Lee
ACM Trans. Design Autom. Electr. Syst.2
2016 Differential Write-Conscious Software Design on Phase-Change Memory: An SQLite Case Study
abstract
Phase-change memory (PCM) has several benefits including low cost, non-volatility, byte-addressability, etc., and limitations such as write endurance. There have been several hardware approaches to exploit the benefits while minimizing the negative impact of limitations. Software approaches could give further improvements, when used together with hardware approaches, by taking advantage of write behavior present in the program, e.g., write behavior on dynamically allocated data, which is hardly captured by hardware approaches. This work proposes a software design methodology to reduce costly PCM writes. First, on top of existing hardware approach such as Flip-N-Write, we advocate exploiting the capability of PCM bit-level differential write in the software by judiciously reusing previously allocated memory resource. In order to avoid wear-out incurred by the reuse, we present software-based wear-leveling methods that distribute writes across PCM cells. In order to further reduce PCM writes, we propose identifying data, the loss of which does not affect the functionality of the underlying software, and then diverting write traffic for those data items to volatile memory. To evaluate the effectiveness of these methods, as a case study, we applied the proposed methods to the design of journaling in SQLite, which is an important database application commonly used in smartphones. For the experiments, we used an in-house PCM-based prototype board. Our experiments with four representative mobile applications show that the proposed design methods, which is applied on top of the hardware approach, Flip-N-Write, result in 75.2% further reduction in total bit updates in PCM, on average, without aggravating wear-out compared with the baseline of PCM-based journaling, which is based only on the hardware approach. Also, the proposed design methods result in 49.4% reduction in energy consumption and 52.3% reduction in runtime compared to a typical FIFO management of free resources.
Sungkwang Lee, Taemin Lee, Hyunsun Park, Junwhan Ahn, Sungjoo Yoo, Youjip Won, Sunggu Lee
ACM Trans. Design Autom. Electr. Syst.5
2016 Low-Power Hybrid Memory Cubes With Link Power Management and Two-Level Prefetching
abstract
The hybrid memory cube (HMC) is a 3-D-stacked DRAM architecture designed for substantially improved memory bandwidth. In particular, its I/O interface achieves up to 320 GB/s of external bandwidth through high-speed serial links. However, it comes at the cost of large static power of off-chip links, which dominates total power consumption of HMCs. In this paper, we propose an adaptive mechanism to partially disable off-chip links of HMCs to reduce the energy consumption of the off-chip links. In order to determine the number of the links to be disabled upon application loads, we develop a simple hardware module called link delay monitor to simulate all different link configurations at the same time and find the largest number of the links to be disabled while satisfying the given performance constraint. We also present two-level prefetching with in-HMC prefetch buffers to further improve the efficiency of our link power management scheme in the presence of prefetching. Evaluations show that our scheme reduces the energy consumption of HMCs by 52% on average with a negligible performance degradation.
Junwhan Ahn, Sungjoo Yoo, Kiyoung Choi
IEEE Trans. Very Large Scale Integr. Syst.2
2015 Memory fast-forward: a low cost special function unit to enhance energy efficiency in GPU for big data processing
Eunhyeok Park, Junwhan Ahn, Sungpack Hong, Sungjoo Yoo, Sunggu Lee
DATE4
2015 A small non-volatile write buffer to reduce storage writes in smartphones
Mungyu Son, Sungkwang Lee, Kyungho Kim, Sungjoo Yoo, Sunggu Lee
DATE4
2015 A scalable processing-in-memory accelerator for parallel graph processing
abstract
The explosion of digital data and the ever-growing need for fast data analysis have made in-memory big-data processing in computer systems increasingly important. In particular, large-scale graph processing is gaining attention due to its broad applicability from social science to machine learning. However, scalable hardware design that can efficiently process large graphs in main memory is still an open problem. Ideally, cost-effective and scalable graph processing systems can be realized by building a system whose performance increases proportionally with the sizes of graphs that can be stored in the system, which is extremely challenging in conventional systems due to severe memory bandwidth limitations.
Junwhan Ahn, Sungpack Hong, Sungjoo Yoo, Onur Mutlu, Kiyoung Choi
ISCA3
2015 PIM-enabled instructions: a low-overhead, locality-aware processing-in-memory architecture
abstract
Processing-in-memory (PIM) is rapidly rising as a viable solution for the memory wall crisis, rebounding from its unsuccessful attempts in 1990s due to practicality concerns, which are alleviated with recent advances in 3D stacking technologies. However, it is still challenging to integrate the PIM architectures with existing systems in a seamless manner due to two common characteristics: unconventional programming models for in-memory computation units and lack of ability to utilize large on-chip caches.
Junwhan Ahn, Sungjoo Yoo, Onur Mutlu, Kiyoung Choi
ISCA2
2015 Locality-aware vertex scheduling for GPU-based graph computation
abstract
Graph computation is becoming more and more popular in machine learning, big data analytics, etc. For such workloads, GPU is considered as an efficient execution platform since graph computation is characterized by massively parallel computation and high demand of memory bandwidth. In our investigation, existing GPU programming methods for graph computation do not fully exploit high memory bandwidth as well as high computing power in GPU. We propose a novel optimization called locality-aware vertex scheduling, which aims at minimizing memory requests by adjusting the order of vertex computations to improve temporal locality of vertex data stored in on-chip caches. Experiments with nine real-world graphs and three graph algorithms on the recent GPU platform show that the proposed method offers a significant speedup (average 46%) over the state-of-the-art graph algorithm implementation on GPUs.
Hyunsun Park, Junwhan Ahn, Eunhyeok Park, Sungjoo Yoo
VLSI-SoC4
2015 Filtering dirty data in DRAM to reduce PRAM writes
abstract
Phase-change RAM (PRAM) is a promising candidate of emerging memory technologies which provides large capacity and low leakage power to compensate for the limitations of DRAM in the hybrid DRAM/PRAM memory subsystem. However, for practical applications of PRAM in the hybrid main memory, we need to reduce write traffics to PRAM in order to overcome the write-related limitations in PRAM such as write endurance. In our work, we propose a concept called in-DRAM write buffer for the hybrid main memory to reduce PRAM write traffics. The cache line-level dirty data are filtered out from the evicted DRAM row and stored in the write buffer which occupies a portion of DRAM by sacrificing the capacity of DRAM cache. In order to reduce PRAM writes, the write buffer tries to maximize write coalescing by avoiding the PRAM write-back of soon-to-be-accessed dirty data. In order to adapt to the dynamically changing program behavior in PRAM writes, we also propose a method to adjust the write buffer size dynamically during runtime. Experimental results show that the proposed dynamic method offers up to 91.92% reduction in PRAM writes and gives results (average 14.81% and 9.47% reduction in PRAM writes and program runtime, respectively) comparable to the best of static write buffer size cases.
Hyunsun Park, Chanha Kim, Sungjoo Yoo, Chanik Park
VLSI-SoC3
2015 Time slot assignment for convergecast in wireless sensor networks
Sunggu Lee, Sungjoo Yoo
J. Parallel Distributed Comput.3
2015 Dynamic Wear Leveling for Phase-Change Memories With Endurance Variations
abstract
Phase change memory (PCM) has a write endurance problem. This problem is exacerbated due to endurance variations (EVs) when using advanced process technology (e.g., sub-20 nm), where PCM is expected to provide scaling benefits over dynamic random access memory (RAM). Wear leveling can solve this problem by dynamically changing the mapping from memory addresses to PCM physical addresses such that all PCM cells are evenly written, thereby extending the effective lifetime of such devices. PCM permits fine-grained writes, i.e., even bit level updates are allowed. To allow fine-grained wear leveling, this capability must be exploited. However, previous wear leveling approaches do not fully exploit fine-grained writes since fine-grained writes cause them to suffer from high data copy (called swap) overhead for address remapping, and/or high area and runtime overhead for the management of write frequency and address mapping information. This paper proposes a dynamic wear leveling method for PCMs that addresses all of these issues. The method: 1) uses bloom filters to enable low-cost write counters for fine-grained writes and 2) exploits the EV of PCM cells to avoid mapping hot data onto weak cells. To improve the effectiveness of the bloom filters, dynamic bloom filter management (write counts, hash functions, and write counter thresholds) and hot-cold address lists are used. The proposed method was evaluated using simulations and a hardware implementation. Using a small amount of PCM capacity overhead (0.3%), the proposed method extended the lifetime of a PCM device by 2.8-4.6 times over the existing methods when there were significant EVs.
Joosung Yun, Sunggu Lee, Sungjoo Yoo
IEEE Trans. Very Large Scale Integr. Syst.3
2014 Dynamic Power Management of Off-Chip Links for Hybrid Memory Cubes
abstract
The Hybrid Memory Cube (HMC) is a 3D-stacked DRAM architecture designed for substantially improved memory bandwidth. In particular, its I/O interface achieves up to 320 GB/s of external bandwidth through high-speed serial links. However, it comes at a cost of large static power of off-chip links, which dominates total power consumption of HMCs. Therefore, we propose an adaptive mechanism to partially disable off-chip links of HMCs with a minimal performance impact. We also present two-level prefetching with in-HMC prefetch buffers to further improve its efficiency in the presence of prefetching. Evaluations show that our scheme reduces energy consumption of HMCs by 51% on average.
Junwhan Ahn, Sungjoo Yoo, Kiyoung Choi
DAC2
2014 Coarse-grained Bubble Razor to exploit the potential of two-phase transparent latch designs
abstract
Timing margin to cover process variation is one of the most critical factors that limit the amount of supply voltage reduction thereby power consumption. To remove too conservative timing margin, Bubble Razor was introduced to dynamically detect and correct errors in two-phase transparent latch designs [13]. However, it does not fully exploit the potential of two-phase transparent latch design, e.g. time borrowing. Thus, especially at low supply voltage where the effect of process variation becomes significant, the existing Bubble Razor can suffer from significant overhead in performance and power consumption due to too frequent occurrence of bubble generations. We present a design methodology for coarse-grained Bubble Razor which exploits the time-borrowing characteristic of two-phase transparent latch design. By selectively inserting error checkpoints, i.e., shadow latches and error management logic, in the circuit, time borrowing can be applied between error checkpoints thereby avoiding bubbles which could occur in the existing Bubble Razor design with a checkpoint at every latch on the critical path. We present a methodology to choose the grain size (the number of stages between error checkpoints) based on 3-sigma delay distribution. We also verify the benefits of coarse-grained Bubble Razor with a real microprocessor, Core-A design [15] using 20nm Predictive Technology Model (PTM) [16]. The proposed methodology offers 62% improvement in performance (MIPS) and 49% less energy consumption (per instruction) at 0.6V operation (zero frequency margin) over the original Bubble Razor scheme. In addition, it gives 25% area reduction in core design.
Hayoung Kim, Dongyoung Kim, Jae-Joon Kim, Sungjoo Yoo, Sunggu Lee
DATE4
2014 Accelerating graph computation with racetrack memory and pointer-assisted graph representation
abstract
The poor performance of NAND Flash memory, such as long access latency and large granularity access, is the major bottleneck of graph processing. This paper proposes an intelligent storage for graph processing which is based on fast and low cost racetrack memory and a pointer-assisted graph representation. Our experiments show that the proposed intelligent storage based on racetrack memory reduces total processing time of three representative graph computations by 40.2%~86.9% compared to the graph processing, GraphChi, which exploits sequential accesses based on normal NAND Flash memory-based SSD. Faster execution also reduces energy consumption by 39.6%~90.0%. The in-storage processing capability gives additional 10.5%~16.4% performance improvements and 12.0%~14.4% reduction of energy consumption.
Eunhyuk Park, Sungjoo Yoo, Sunggu Lee, Hai Li 0001
DATE2
2014 DASCA: Dead Write Prediction Assisted STT-RAM Cache Architecture
abstract
Spin-Transfer Torque RAM (STT-RAM) has been considered as a promising candidate for on-chip last-level caches, replacing SRAM for better energy efficiency, smaller die footprint, and scalability. However, it also introduces several new challenges into last-level cache design that need to be overcome for feasible deployment of STT-RAM caches. Among other things, mitigating the impact of slow and energy-hungry write operations is of the utmost importance. In this paper, we propose a new mechanism to reduce write activities of STT-RAM last-level caches. The key observation is that a significant amount of data written to last-level caches is not actually re-referenced again during the lifetime of the corresponding cache blocks. Such write operations, which we call dead writes, can bypass the cache without incurring extra misses by definition. Based on this, we propose Dead Write Prediction Assisted STT-RAM Cache Architecture (DASCA), which predicts and bypasses dead writes for write energy reduction. For this purpose, we first propose a novel classification of dead writes, which is composed of dead-on-arrival fills, dead-value fills, and closing writes, as a theoretical model for redundant write elimination. On top of that, we present a dead write predictor based on a state-of-the-art dead block predictor. Evaluations show that our architecture achieves an energy reduction of 68% (62%) in last-level caches and an additional energy reduction of 10% (16%) in main memory and even improves system performance by 6% (14%) on average compared to the STT-RAM baseline in a single-core (quad-core) system.
Junwhan Ahn, Sungjoo Yoo, Kiyoung Choi
HPCA2
2014 FPGA-based prototyping systems for emerging memory technologies
abstract
As DRAM faces scaling limit, several new memory technologies are considered as candidates for replacing or complementing DRAM main memory. Compared to DRAM, the new memories have two major differences, non-volatility and write overhead in terms of endurance, latency and power. We built two different FPGA-based evaluation boards to evaluate hardware and software designs for new-memory based main memory; one with a DRAM subsystem having parameterizable latency and non-volatility emulation, and the other with the real chips of new memory namely phase-change RAM (PRAM). We experimented primitive functions and SQLite-based benchmarks on Linux, verifying the workings of new functionalities, e.g., nonvolatility and evaluating the impacts of new memory on software performance. In our experiments, we also demonstrated the impact of new memory-aware software/hardware designs on program performance on a DRAM/PRAM hybrid memory.
Taemin Lee, Dongki Kim, Hyunsun Park, Sungjoo Yoo, Sunggu Lee
RSP4
2014 An Adaptive Idle-Time Exploiting Method for Low Latency NAND Flash-Based Storage Devices
abstract
The market share of NAND flash-based storage devices (NFSDs) has rapidly grown in recent years since many characteristics, such as non-volatility, low latency, and high reliability, meet the requirements for various types of storage devices. However, the unique characteristic of NAND flash memories (NFMs), erase-before-write, causes problems for NFSDs from a performance perspective. Specifically, performance degradation is incurred by extra operations that serve to hide the bad characteristics of NFMs. In order to resolve this problem, many attractive methods have been proposed. Various algorithms for flash translation layers (FTLs) are representative methods that provide space redundancy to NFSDs for better performance. However, the amount of space redundancy is limited by the capacity of NFMs and thus, space redundancy is still insufficient for improving the performance of NFSDs. Consequently, a new type of redundancy, termed temporal redundancy, has recently been introduced for NFSDs. More precisely, the idleness of NFSDs is exploited so as to precede extra operations for NFSDs while minimizing the overhead of extra operations. In this paper, we propose an adaptive time-out method based on the Hidden-Markov Model (HMM) to efficiently utilize idle periods. In addition, we also suggest a simple scheduling scheme for extra operations that can be customized for general FTLs. The experimental results demonstrate that the proposed method yields performance improvements in terms of average write latency and peak latency, 74% and 76% better than the existing method, respectively, and approaching within average 9% and 5% of the optimal case, respectively.
Sang-Hoon Park, Donggun Kim 0005, Kwanhu Bang, Hyuk-Jun Lee, Sungjoo Yoo, Eui-Young Chung
IEEE Trans. Computers5
2014 A Memory-Efficient Architecture of Full HD Around View Monitor Systems
abstract
The around view monitor (AVM) is one of the representative features of smart vision systems adopted in various application areas, e.g., advanced driver assistance systems. The design of AVM systems with full high-definition (HD)-level resolution presents significant technical challenges. In particular, a high-memory performance is required to process full HD images obtained from multiple cameras. Specifically, a full HD AVM system requires a high-memory bandwidth (six times higher than D1 image-based systems) and is characterized by a significant amount of single writes, which degrades the effective performance of modern DRAM. To address these problems, two methods are proposed in this paper. The first method reduces the required memory bandwidth by storing in the memory only the input pixel data that will be used in the final processing step, and the second improves the single-write performance of a DRAM subsystem by DRAM-aware data mapping. The effectiveness of the proposed methods is proved by designing an AVM system incorporating field-programmable gate arrays and DDR2-200 SDRAMs. The proposed methods reduce the memory bandwidth requirement by 51%, allowing a full HD AVM system to run at over 24 frames per second.
Byeongchan Jeon, Gyuro Park, Sungjoo Yoo, Hong Jeong
IEEE Trans. Intell. Transp. Syst.4
2013 Selectively protecting error-correcting code for area-efficient and reliable STT-RAM caches
abstract
Recent researches on STT-RAM revealed that device scaling makes its write operations unreliable. To mitigate the impact of this problem, this paper proposes a low-cost, ECC-based solution for STT-RAM caches. In particular, it proposes to share storage for ECC among different blocks within a set and to use them only for unsuccessful write operations. Experimental results show that our scheme reduces 74% to 98% of area overhead incurred by the conventional per-block ECC while maintaining system performance and reliability.
Junwhan Ahn, Sungjoo Yoo, Kiyoung Choi
ASP-DAC2
2013 Write intensity prediction for energy-efficient non-volatile caches
abstract
This paper presents a novel concept called write intensity prediction for energy-efficient non-volatile caches as well as the architecture that implements the concept. The key idea is to correlate write intensity of cache blocks with addresses of memory access instructions that incur cache misses of those blocks. The predictor keeps track of instructions that tend to load write-intensive blocks and utilizes that information to predict write intensity of blocks. Based on this concept, we propose a block placement strategy driven by write intensity prediction for SRAM/STT-RAM hybrid caches. Experimental results show that the proposed approach reduces write energy consumption by 55% on average compared to the existing hybrid cache architecture.
Junwhan Ahn, Sungjoo Yoo, Kiyoung Choi
ISLPED2
2013 A network congestion-aware memory subsystem for manycore
abstract
The network-on-chip (NoC) plays a crucial role in memory performance due to the fact that it can handle the majority of traffics from/to the DRAM memory controllers. However, there has been little work on the interplay between the NoC and memory controllers. In this article, we address a problem called network congestion-induced memory blocking and propose a novel memory controller, which performs memory access scheduling and network entry control in a network congestion-aware manner. In case of network congestion, in order to avoid performance degradation due to the blocking caused by data bound for congested regions in the NoC, the proposed memory controller favors requests and data associated with uncongested regions. In addition, in order to avoid the fairness problem of such a policy, we also propose a gradual method, which enables a trade-off between performance (in memory utilization) and fairness (in memory access latency). Experimental results show that the proposed method can offer up to 1.76 ∼ 2.99 times improvement in memory utilization in the latency-tolerant designs.
Dongki Kim, Sungjoo Yoo, Sunggu Lee
ACM Trans. Embed. Comput. Syst.2
2013 MAEPER: Matching Access and Error Patterns With Error-Free Resource for Low Vcc L1 Cache
abstract
Large SRAMs are the practical bottleneck to achieve a low supply voltage, because they suffer from process variation-induced bit errors at a low supply voltage. In this paper, we present an error-resilient cache architecture that resolves the drawback of previous approaches, i.e., the performance degradation at a low supply voltage which is caused by cache misses in accesses to faulty resources. We utilize cache access locality and error-free resources in a cost-effective manner. First, we classify cache lines into fully and partially accessed groups and apply appropriate methods to each group. For the partially accessed group, we propose a method of matching memory access behavior and error locations with intra-cache line word-level remapping. In order to reduce the area overhead used to store the cache access information history, we present an access pattern-learning line-fill buffer (LFB). For the fully accessed group, we propose the utilization of error-free assist functions in the cache, i.e., a LFB and victim cache with no process variation-induced error at the target minimum supply voltage. We also present an error-aware prefetch method that allows us to utilize the error-free victim cache to achieve a further reduction in cache misses due to faulty resources. Experimental results show that the proposed method gives an average 32.6% reduction in cycles per instruction at an error rate of 0.2% with a small area overhead of 8.2%.
Young-Geun Choi, Sungjoo Yoo, Sunggu Lee, Jung Ho Ahn, Kangmin Lee
IEEE Trans. Very Large Scale Integr. Syst.2
2012 Hybrid DRAM/PRAM-based main memory for single-chip CPU/GPU
abstract
Single-chip CPU/GPU architecture is being adopted in high-end (embedded) systems, e.g., smartphones and tablet PCs. Main memory subsystem is expected to consist of hybrid DRAM and phase-change RAM (PRAM) due to the difficulties in DRAM scaling. In this work, we address the performance optimization of the hybrid DRAM/PRAM main memory for single chip CPU/GPU. Based on the tight requirements of low latency from CPU and the relative tolerance to long latency from GPU, DRAM is first allocated to CPU while PRAM with longer write latency is allocated to GPU. Then, in order to improve the write performance of GPU traffic, we propose (1) an in-DRAM write buffer to accommodate GPU write traffics, (2) dynamic hot data management to improve the efficiency of write buffer, (3) runtime-adaptive adjustment of write buffer size to meet the given CPU performance bound, and (4) CPU-aware DRAM access scheduling to give low latency to CPU traffics. The experiments show that the proposed method gives 1.02~44.2 times performance improvement in GPU performance with modest (negligible) CPU performance overhead (when compute-intensive CPU programs run).
Dongki Kim, Sungkwang Lee, Jaewoong Chung, Daehyun Kim 0001, Dong Hyuk Woo, Sungjoo Yoo, Sunggu Lee
DAC6
2012 Write performance improvement by hiding R drift latency in phase-change RAM
abstract
Phase-change RAM (PRAM) is considered to be one of the most promising candidates to complement or replace DRAM in the near future. However, it is imperative to overcome the limitations of PRAM, especially, long write latency for its widespread applications. R drift latency occupies a significant portion in PRAM write latency thereby adversely affecting system performance. In this paper, we propose a novel method called write status holding register (WSHR) to reduce the write latency due to R drift latency. The WSHR allows for non-blocking accesses to PRAM during R drift latency thereby improving system performance. Our experiments with SPEC benchmarks show that the proposed WSHR gives 53.6%~0% performance improvements in the hybrid DRAM/PRAM main memory (256MB DRAM and 14nm PRAM).
Youngsik Kim, Sungjoo Yoo, Sunggu Lee
DAC2
2012 A case study on the application of real phase-change RAM to main memory subsystem
abstract
Phase-change RAM (PCM) has the advantages of better scaling and non-volatility compared with the DRAM which is expected to face its scaling limit in the near future. There have been many studies on applying the PCM to main memory in order to complement or replace the DRAM. One common limitation of these studies is that they are based on synthetic PCM models. In our study, we investigate the feasibility and issues of applying a real PCM to main memory. In this paper, we report our case study of characterizing the PCM and evaluating its usefulness in the main memory. Our results show that the PCM/DRAM hybrid main memory with a modest DRAM size can give comparable performance to that of the DRAM only main memory. However, the hybrid memory with small DRAMs or large footprint programs can suffer from performance degradation due to the long latency of both PCM writes and write preemption penalty, which requires architectural innovations for exploiting the full potential of PCM write performance.
Suknam Kwon, Dongki Kim, Youngsik Kim, Sungjoo Yoo, Sunggu Lee
DATE4
2012 Bloom filter-based dynamic wear leveling for phase-change RAM
abstract
Phase Change RAM (PCM) is a promising candidate of emerging memory technology to complement or replace existing DRAM and NAND Flash memory. A key drawback of PCMs is limited write endurance. To address this problem, several static wear-leveling methods that change logical to physical address mapping periodically have been proposed. Although these methods have low space overhead, they suffer from unnecessary data migrations thereby failing to exploit the full lifetime potential of PCMs. This paper proposes a new dynamic wear-leveling method that reduces unnecessary data migrations by adopting a hot/cold swapping-based dynamic method. Compared with the conventional hot/cold swapping-based dynamic method, the proposed method requires only a small amount of space overhead by applying Bloom filters to the identification of hot and cold data. We simulate our method using SPEC2000 benchmark traces and compare with previous methods. Simulation results show that the proposed method reduces unnecessary data migrations by 58~92% and extends the memory lifetime by 2.18~2.30 times over previous methods with a negligible area overhead of 0.3%.
Joosung Yun, Sunggu Lee, Sungjoo Yoo
DATE3
2012 Optimal wake-up scheduling of data gathering trees for wireless sensor networks
Ungjin Jang, Sunggu Lee, Sungjoo Yoo
J. Parallel Distributed Comput.3
2012 Active Memory Processor for Network-on-Chip-Based Architecture
abstract
Memory-intensive operations and their memory access latency are often the performance bottleneck in parallel applications. In this paper, we investigate the concept of active memory operation which is an active data processing operation performed on the memory side. Utilizing the active memory operation, we can replace multiple transactions of memory accesses over the on-chip network and related computations on the processor side with a smaller number of high-level transactions and computations on the memory side. To realize the concept, we have designed a special-purpose processor called active memory processor which is tightly coupled with the memory and executes the active memory operations. In our case studies, we have applied the concept to five real-world applications (parallelized JPEG, FFT, text indexing for data mining, histogram, and eikonal equation solver) running on a 36--tile architecture with 64 cores and four memory tiles and found that the proposed approach can improve performance by 20.5 \sim 259.3 percent.
Jun-hee Yoo, Sungjoo Yoo, Kiyoung Choi
IEEE Trans. Computers2
2012 A Multistep Tag Comparison Method for a Low-Power L2 Cache
abstract
Tag comparison in a highly associative cache consumes a significant portion of the cache energy. Existing methods for tag comparison reduction are based on predicting either cache hits or cache misses. In this paper, we present novel ideas for both cache hit and miss predictions. We present a partial tag-enhanced Bloom filter to improve the accuracy of the cache miss prediction method and hot/cold checks that control data liveness to reduce the tag comparisons of the cache hit prediction method. We also combine both methods so that their order of application can be dynamically adjusted to adapt to changing cache access behavior, which further reduces tag comparisons. To overcome the common limitation of multistep tag comparison methods, we propose a method that reduces tag comparisons while meeting the given performance bound. Experimental results showed that the proposed method reduces the energy consumption of tag comparison by an average of 88.40%, which translates to an average reduction of 35.34% (40.19% with low-power data access) in the total energy consumption of the L2 cache and a further reduction of 8.86% (10.07% with low-power data access) when compared with existing methods.
Hyunsun Park, Sungjoo Yoo, Sunggu Lee
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2012 Optimizing Video Application Design for Phase-Change RAM-Based Main Memory
abstract
Video applications including video codecs place a large traffic demand on main memory. Emerging memory technology, such as phase-change RAM (PRAM) tends to suffer from the write endurance problem, in which the maximum number of writes is limited. Thus, it is required to improve video application designs to adapt to the new requirements of emerging memory technology, i.e., to minimize the number of writes in terms of bit updates. In this paper, we present a way to optimize video application design for PRAM-based main memory. We propose two methods to resolve the write endurance problem: inter-block differential data encoding and inter-frame multiple experts. Experimental results show an average of 18.4% reduction in bit updates when compared to the best existing data encoding methods for PRAM.
Suknam Kwon, Sungjoo Yoo, Sunggu Lee, Jinpyo Park
IEEE Trans. Very Large Scale Integr. Syst.2
2011 Matching cache access behavior and bit error pattern for high performance low Vcc L1 cache
abstract
Cache is a roadblock towards low supply voltage (Vcc). It is mainly because low Vcc incurs process variation-induced bit errors in large SRAM in cache. Existing approaches for low Vcc cache suffer from low performance due to reduced effective capacity, long latency to correct errors, and increased misses due to accesses to faulty words. In our work, we propose a word-level sub-block disable-based method which increases the utilization of available cache capacity. Our key idea is to minimize accesses to faulty words. To do that, we propose utilizing access behavior history in allocating cache resource with faulty words. In addition, we propose remapping cache words inside of cache line in order to better match both access and error patterns. Experimental results show that the proposed method gives average 21.8% (up to 34.0%) performance improvement with a small area overhead in L1 and L2 caches.
Young-Geun Choi, Sungjoo Yoo, Sunggu Lee, Jung Ho Ahn
DAC2
2011 FlexiBuffer: reducing leakage power in on-chip network routers
abstract
The increasing number of integrated components on a single chip has increased the importance of on-chip networks. A significant part of on-chip network routers is the buffer, as it occupies a large area and consumes a significant amount of power. In this work, we propose FlexiBuffer, a microarchitecture in which we minimize buffer leakage power by using fine-grained power gating and adjusting the size of the active buffers adaptively. We propose two microarchitecture techniques to support fine-grained power gating -- early credit in credit-based flow control and new buffer organizations to overcome the limitation of circular buffers. Our results show that, with minimal loss in performance, we can reduce the leakage power of on-chip network router buffers by up to 61% and overall router power consumption by up to 39%.
Gwangsun Kim, John Kim 0001, Sungjoo Yoo
DAC3
2011 Power management of hybrid DRAM/PRAM-based main memory
abstract
Hybrid main memory consisting of DRAM and non-volatile memory is attractive since the non-volatile memory can give the advantage of low standby power while DRAM provides high performance and better active power. In this work, we address the power management of such a hybrid main memory consisting of DRAM and phase-change RAM (PRAM). In order to reduce DRAM refresh energy which occupies a significant portion of total memory energy, we present a runtime-adaptive method of DRAM decay. In addition, we present two methods, DRAM bypass and dirty data keeping, for further reduction in refresh energy and memory access latency, respectively. The experiments show that by reducing DRAM refreshes, we can obtain 23.5%~94.7% reduction in the energy consumption with negligible performance overhead compared with the conventional DRAM-only main memory.
Hyunsun Park, Sungjoo Yoo, Sunggu Lee
DAC2
2011 A quantitative analysis of performance benefits of 3D die stacking on mobile and embedded SoC
abstract
3D stacked DRAM improves peak memory performance. However, its effective performance is often limited by the constraints of row-to-row activation delay (tRRD), four active bank window (tFAW), etc. In this paper, we present a quantitative analysis of the performance impact of such constraints. In order to resolve the problem, we propose balancing the budget of DRAM row activation across DRAM channels. In the proposed method, an inter-memory controller coordinator receives the current demand of row activation from memory controllers and re-distributes the budget to the memory controllers in order to improve DRAM performance. Experimental results show that sharing the budget of row activation between memory channels can give average 4.72% improvement in the utilization of 3D stacked DRAM.
Dongki Kim, Sungjoo Yoo, Sunggu Lee, Jung Ho Ahn, Hyunuk Jung
DATE2
2011 A novel tag access scheme for low power L2 cache
abstract
Tag comparisons occupy a significant portion of cache power consumption in the highly associative cache such as L2 cache. In our work, we propose a novel tag access scheme which applies a partial tag-enhanced Bloom filter to reduce tag comparisons by detecting per-way cache misses. The proposed scheme also classifies cache data into hot and cold data and the tags of hot data are compared earlier than those of cold data exploiting the fact that most of cache hits go to hot data. In addition, the power consumption of each tag comparison can be further reduced by dividing the tag comparison into two micro-steps where a partial tag comparison is performed first and, only if the partial tag comparison gives a partial hit, then the remaining tag bits are compared. We applied the proposed scheme to an L2 cache with 10 programs from SPEC2000 and SPEC2006. Experimental results show average 23.69% and 8.58% reduction in cache energy consumption compared with the conventional serial tag-data access and the other existing methods, respectively.
Hyunsun Park, Sungjoo Yoo, Sunggu Lee
DATE2
2011 Runtime Power Management of 3-D Multi-Core Architectures Under Peak Power and Temperature Constraints
abstract
3-D integration is a new technology that overcomes the limitations of 2-D integrated circuits, e.g., power and delay induced from long interconnect wires, by stacking multiple dies to increase logic integration density. However, chip-level power and peak temperature are the major performance limiters in 3-D multi-core architectures. In this paper, we propose a runtime power management method for both peak power and temperature-constrained 3-D multi-core systems in order to maximize the instruction throughput. The proposed method exploits dynamic temperature slack (defined as peak temperature constraint minus current temperature) and workload characteristics (e.g., instructions per cycle and memory-boundness) as well as thermal characteristics of 3-D stacking architectures. Compared with existing thermal-aware power management solutions for 3-D multi-core systems, our method yields up to 34.2% (average 18.5%) performance improvement in terms of instructions per second without significant additional energy consumption.
Kyungsu Kang, Jungsoo Kim, Sungjoo Yoo, Chong-Min Kyung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2011 Program Phase-Aware Dynamic Voltage Scaling Under Variable Computational Workload and Memory Stall Environment
abstract
Most complex software programs are characterized by program phase behavior and runtime distribution. Dynamism of the two characteristics often makes the design-time workload prediction difficult and inefficient. Especially, memory stall time whose variation is significant in memory-bound applications has been mostly neglected or handled in a too simplistic manner in previous works. In this paper, we present a novel online dynamic voltage and frequency scaling (DVFS) method which takes into account both program phase behavior and runtime distribution of memory stall time, as well as computational workload. The online DVFS problem is addressed in two ways: intraphase workload prediction and program phase detection. The intraphase workload prediction is to predict the workload based on the runtime distribution of computational workload and memory stall time in the current program phase. The program phase detection is to identify to which program phase the current instant belongs and then to obtain the predicted workload corresponding to the detected program phase, which is used to set voltage and frequency during the program phase. The proposed method considers leakage power consumption as well as dynamic power consumption by a temperature-aware combined Vdd/Vbbscaling. Compared to a conventional method, experimental results show that the proposed method provides up to 34.6% and 17.3% energy reduction for two multimedia applications, MPEG4 and H.264 decoder, respectively.
Jungsoo Kim, Sungjoo Yoo, Chong-Min Kyung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2010 An analytical dynamic scaling of supply voltage and body bias exploiting memory stall time variation
abstract
Success of workload prediction, which is critical in achieving low energy consumption via dynamic voltage and frequency scaling (DVFS), depends on the accuracy of modeling the major sources of workload variation. Among them, memory stall time, whose variation is significant especially in case of memory-bound applications, has been mostly neglected or handled in too simplistic assumptions in previous works. In this paper, we present an analytical DVFS method which takes into account variations in both computation and memory stall cycles. The proposed method reduces leakage power consumption as well as switching power consumption through combined Vdd/Vbbscaling. Experimental results on MPEG4 and H.264 decoder have shown that, compared to previous methods and, our method achieves up to additional 30.0% and 15.8% energy reductions, respectively.
Jungsoo Kim, Younghoon Lee, Sungjoo Yoo, Chong-Min Kyung
ASP-DAC3
2010 Event statistics and criticality-aware bitrate allocation to minimize energy consumption of memory-constrained wireless surveillance system
abstract
An event criticality-aware wireless surveillance system tries to minimize the energy consumption by adjusting the image distortion (or quality) requirement according to event criticality. In this paper, we present a novel video encoding method which scales bitrate so as to minimize the energy consumption of the wireless surveillance system while satisfying the given memory constraint. Given event statistics, memory size, and distortion requirement, the presented method gives an energy-optimal bitrate based on an analytic formulation of energy-rate-distortion (E-R-D) relationship for the target system consisting of H.264 video encoder, event detector, transceiver and memory. Experimental results show that the proposed method offers up to 59.6% (29.1% on average) energy savings compared to an existing bitrate allocation method which does not consider event statistics.
Jungsoo Kim, Jaemoon Kim, Giwon Kim, Sangkwon Na, Sungjoo Yoo, Chong-Min Kyung
ICME5
2010 A Network Congestion-Aware Memory Controller
abstract
Network-on-chip and memory controller become correlated with each other in case of high network congestion since the network port of memory controller can be blocked due to the (back-propagated) network congestion. We call such a problem network congestion-induced memory blocking. In order to resolve the problem, we present a novel idea of network congestion-aware memory controller. Based on the global information of network congestion, the memory controller performs (1) congestion-aware memory access scheduling and (2) congestion-aware network entry control of read data. The experimental results obtained from a 5×5 tile architecture show that the proposed memory controller presents up to 18.9% improvement in memory utilization.
Dongki Kim, Sungjoo Yoo, Sunggu Lee
NOCS2
2010 Temperature-Aware Integrated DVFS and Power Gating for Executing Tasks With Runtime Distribution
abstract
At high-operating temperature, chip cooling is crucial due to the exponential temperature dependence of leakage current. However, traditional cooling methods, e.g., power/clock gating applied when a temperature threshold is reached, often cause excessive performance degradation. In this paper, we propose a method for delivering lower energy consumption by integrating the cooling and running in a temperature-aware manner without incurring performance penalty. In order to further reduce the energy consumption, we exploited the runtime distribution of each sub-segment of a task called “bin” in an analytical manner such that time budget for cooling in each bin is allocated in proportion to the probability of the occurrence of the bin. We apply the proposed method to two realistic software programs, H.264 decoder and ray tracing and a benchmark program, equake. The experimental results show that the proposed method yields additional 19.4%-27.2% reduction in energy consumption compared with existing methods.
Kyungsu Kang, Jungsoo Kim, Sungjoo Yoo, Chong-Min Kyung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2010 Dual Motion Estimation for Frame Rate Up-Conversion
abstract
In this letter, we present a new motion estimation algorithm for frame rate up-conversion. The proposed dual motion estimation algorithm enhances the estimation accuracy of motion vectors by using the unidirectional and bidirectional matching ratios of blocks in the previous and current frames. In addition, the proposed motion estimation approach uses motion vector validity to evaluate the accuracy of motion vectors thereby avoiding false motion vectors. In experiments using benchmark image sequences, the proposed motion estimation algorithm improved the average peak signal-to-noise ratio of interpolated frames by up to 2.272 dB, when compared to conventional motion estimation algorithms. For the comparison of the perceptual image quality using the structural similarity, the average value of the proposed dual motion estimation was by up to 0.062 higher than those of the conventional algorithms.
Suk-Ju Kang, Sungjoo Yoo, Young Hwan Kim
IEEE Trans. Circuits Syst. Video Technol.2
2009 Multiprocessor System-on-Chip designs with active memory processors for higher memory efficiency
abstract
Memory access latency and memory-related operations are often the performance bottleneck in parallel applications. In this paper, we present a concept of active memory operations which is an on-chip network transaction that operates based on the microcode provided by the software designer. Utilizing the active memory operation, we can replace multiple transactions of memory accesses over the on-chip network and related local processing element computation with a smaller number of high-level transactions and near-memory computation. We implemented a processor called active memory processor which is located near the memory and executes the active memory operations. In our case studies, we applied the concept to three real-world applications (parallelized JPEG, FFT, and text indexing for data mining) running on a 36-tile architecture with 32 cores and 4 memories and found that the programmable transaction approach can improve performance by 34.3% to 618% at the cost of additional design effort.
Jun-hee Yoo, Sungjoo Yoo, Kiyoung Choi
DAC2
2009 Program phase and runtime distribution-aware online DVFS for combined Vdd/Vbb scaling
abstract
Complex software programs are mostly characterized by phase behavior and runtime distributions. Due to the dynamism of the two characteristics, it is not efficient to make workload predictions during design-time. In our work, we present a novel online DVFS method that exploits both phase behavior and runtime distribution during runtime in combined Vdd/Vbb scaling. The presented method performs a bi-modal analysis of runtime distribution, and then a runtime distribution-aware workload prediction based on the analysis. In order to minimize the runtime overhead of the sophisticated workload prediction method, it performs table lookups to the pre-characterized data during runtime without compromising the quality of energy reduction. It also offers a new concept of program phase suitable for DVFS. Experiments show the effectiveness of the presented method in the case of H.264 decoder with two sets of long-term scenarios consisting of total 4655 frames. It offers 6.6% ∼ 33.5% reduction in energy consumption compared with existing offline and online solutions.
Jungsoo Kim, Sungjoo Yoo, Chong-Min Kyung
DATE2
2009 In-network reorder buffer to improve overall NoC performance while resolving the in-order requirement problem
abstract
Data-intensive functions on chip, e.g., codec, 3D graphics, pixel processing, etc. need to make best use of the increased bandwidth of multiple memories enabled by 3D die stacking via accessing multiple memories in parallel. Parallel memory accesses with originally in-order requirements necessitate reorder buffers to avoid deadlock. Reorder buffers are expensive in terms of area and power consumption. In addition, conventional reorder buffers suffer from a problem of low resource utilization. In our work, we present a novel idea, called in-network reorder buffer, to increase the utilization of reorder buffer resource. In our method, we move the reorder buffer resource and related functions from network entry/exit points to network routers. Thus, the in-network reorder buffers can be better utilized in two ways. First, they can be utilized by other packets without in-order requirements while there are no in-order packets. Second, even in-order packets can benefit from in-network reorder buffers by enjoying more shares of reorder buffers than before. Such an increase in reorder buffer utilization enables NoC performance improvement while supporting the original in-order requirements. Experimental results with an industrial strength DTV SoC example show that the presented idea improves the total execution cycle by 16.9%.
Woo-Cheol Kwon, Sungjoo Yoo, Junhyung Um, Seh-Woong Jeong
DATE2
2009 Topology Synthesis of Cascaded Crossbar Switches
abstract
Performance requirements of on-chip network increase as system-on-chips (SoCs) are becoming more and more complex. For high-performance applications, crossbar switch-based networks are replacing the traditional shared buses as the backbone networks in SoCs. In this paper, we tackle the topology design of on-chip networks with crossbar switches in a cascaded fashion. We also resolve the unacceptable complexity of our previous method based on mixed integer linear programming by a heuristic method. Experimental results show that the proposed method overcomes the frequency limitation of the single crossbar-based design, particularly when the wire delay effect is considered. The proposed heuristic method also achieves more area reduction (up to 69.5%) over the existing methods, and finds as good solutions as the exact method while the synthesis time is saved by orders of magnitude.
Minje Jun, Sungjoo Yoo, Eui-Young Chung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2009 An Analytical Dynamic Scaling of Supply Voltage and Body Bias Based on Parallelism-Aware Workload and Runtime Distribution
abstract
Dynamic voltage and frequency scaling (DVFS) for a parallel software program is crucial for lowering the ever-increasing power consumption of multiprocessor systems-on-chips (SoCs). In this paper, we propose an analytical DVFS method that judiciously exploits slack by considering the varying parallelism over each path in a task graph. The proposed method overcomes the conventional pessimistic assumption on the remaining workload, i.e., worst-case execution cycle. It yields minimum average energy consumption by utilizing the runtime distribution of a software program while satisfying the deadline constraints. The proposed method tackles leakage power consumption as well as dynamic power consumption by combinedVdd/Vbbscaling. Compared to conventional method , experimental results show that the proposed method provides up to 49.20% energy reduction for a set of synthetic task graphs and yields 23.93% and 27.15% energy reductions for two multimedia applications, namely, the H.264 encoder and decoder, respectively.
Jungsoo Kim, Seungyong Oh, Sungjoo Yoo, Chong-Min Kyung
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2009 Topology/Floorplan/Pipeline Co-Design of Cascaded Crossbar Bus
abstract
On-chip bus design has a significant impact on the die area, power consumption, performance and design cycle of complex system-on-chips (SoCs). Especially, for high frequency systems having on-chip buses pipelined extensively to cope with long wire delay, a naive bus design may yield a significant area/power cost mostly due to bus pipeline cost. The topology, floorplan, and pipeline are the most important design factors that affect the cost and frequency of the on-chip bus. Since they are strongly correlated with each other, it is imperative to codesign all of the three. In this paper, we present an automated codesign method for cascaded crossbar bus design. We present CADBUS (CAscadeD crossbar BUS design tool), an automated tool for AXI-based cascaded crossbar bus architecture design. The primary objective of this study is to design a cascaded crossbar bus, including the topology/floorplan/bus pipelines, having minimum area/power cost while satisfying the given constraints of communication bandwidth/latency or frequency. Experimental results of the three industrial strength SoCs show that, compared to the existing approach, the proposed method gives as much as 11.6%-34.2% (9.9%-33.5%) savings in bus area (power consumption).
Jun-hee Yoo, Sungjoo Yoo, Kiyoung Choi
IEEE Trans. Very Large Scale Integr. Syst.2
2008 An industrial perspective of power-aware reliable SoC design
abstract
Reliable SoC design is becoming one of important real design problems since the fast pace of semiconductor scaling and the introduction of new device structures and materials incur more reliability problems than can be solved in the given time frame (2 years/technology node). Reliable design mostly requires resource overhead (additional power consumption, silicon area, and execution time) to recover from errors. Minimizing the overhead in the reliable SoC design will give SoC industries a competitive edge. Especially, in the case of mobile SoC, mastering the overhead of power consumption is absolutely imperative. In this paper, we investigate reliable SoC design in terms of reducing the overhead of power consumption. First, we review the current practice of reliable SoC design and assess its impact on power consumption. Then, we present our perspective on new design methodology towards power-aware reliable SoC design.
Soo-Kwan Eo, Sungjoo Yoo, Kyu-Myung Choi
ASP-DAC2
2008 Mixed integer linear programming-based optimal topology synthesis of cascaded crossbar switches
abstract
We present a topology synthesis method for high performance System-on-Chip (SoC) design. Our method provides an optimal topology of on-chip communication network for the given bandwidth, latency, frequency and/or area constraints. The optimal topology consists of multiple crossbar switches and some of them can be connected in a cascaded fashion for higher clock frequency and/or area efficiency. Compared to previous works, the major contribution of our work is the exactness of the solution from two aspects. First, the solving method of our work is exact by employing the mixed integer linear programming (MILP) method. Second, we generalize the crossbar switch representation in MILP in order that the optimal topology can include any arbitrary sizes of crossbar switches together. The experimental results show that the topologies optimized for the clock frequency (area) give up to 37.3% (12.7%) improvements compared to the conventional single large crossbar switch networks for two industrial strength SoC designs.
Minje Jun, Sungjoo Yoo, Eui-Young Chung
ASP-DAC2
2008 A practical approach of memory access parallelization to exploit multiple off-chip DDR memories
abstract
3D stacked memory enables more off-chip DDR memories. Redesigning existing IPs to exploit the increased memory parallelism will be prohibitively costly. In our work, we propose a practical approach to exploit the increased bandwidth and reduced latency of multiple off-chip DDR memories while reusing existing IPs without modification. The proposed approach is based on two new concepts: transaction id renaming and distributed soft arbitration. We present two on-chip network components, request parallelizer and read data serializer, to realize the concepts. Experiments with synthetic test cases and an industrial strength DTV SoC design show that the proposed approach gives significant improvements in total execution cycle (21.6%) and average memory access latency (31.6%) in the DTV case with a small area overhead (30.1% in the on-chip network, and less than 1.4% in the entire chip).
Woo-Cheol Kwon, Sungjoo Yoo, Sung-Min Hong, Byeong Min, Kyu-Myung Choi, Soo-Kwan Eo
DAC2
2008 Dynamic Voltage Scaling of Supply and Body Bias Exploiting Software Runtime Distribution
abstract
This paper presents a method of dynamic voltage scaling (DVS) that tackles both switching and leakage power with combined Vdd/Vbsscaling and gives minimum average energy consumption exploiting the runtime distribution of software execution. We present a mathematical formulation of the DVS problem and an efficient numerical solution. Experimental results show that the presented method shows up to 44% further reduction in energy consumption compared with existing methods. Especially, when the leakage power consumption is significant, i.e. when temperature is high, the presented method is proven to be the most effective.
Sungpack Hong, Sungjoo Yoo, Byeong Bin, Kyu-Myung Choi, Soo-Kwan Eo, Taehwan Kim 0007
DATE2
2008 An Open-Loop Flow Control Scheme Based on the Accurate Global Information of On-Chip Communication
abstract
3D stacked memory is being adopted as a promising solution to offer high bandwidth and low latency in memory access. Compared with the on-chip network design with conventional off chip memory, it gives a new problem of minimizing communication conflicts since multiple concurrent high bandwidth data transfers will flow through the on-chip network. In order to tackle this problem, we propose applying an open-loop flow control scheme based on the accurate global information (destination and status) of on-chip communication. The proposed open-loop flow control scheme exploits the information and selectively buffers and arbitrates data transfers to remove conflicts at destinations in a preventive manner. As an implementation of the presented scheme, we present on-chip buffers called Buf3D's that share the global information with each other to perform the selective buffering and arbitration of data transfers. Experiments with synthetic test cases and an industrial strength DTV design show that the proposed method improves aggregate memory bandwidth significantly (average 19.0 %~25.8 % in the synthetic cases and up to 18.4 % in the DTV case) with a small area overhead (15.2 % in the DTV case) of on-chip network.
Woo-Cheol Kwon, Sung-Min Hong, Sungjoo Yoo, Byeong Min, Kyu-Myung Choi, Soo-Kwan Eo
DATE3
2008 Entry control in network-on-chip for memory power reduction
abstract
As high-end mobile embedded systems become data-intensive, the off-chip memory is becoming a major contributor to the total energy consumption. Especially, high-end mobile chips accommodate dedicated hardware blocks, e.g., codec and 3D graphics IP's, required for both performance and power consumption reasons. Those IP's usually do not have a large shared memory on chip. Thus, they communicate with each other via the off-chip DDR memory increasing off-chip memory accesses, which increases memory energy consumption during read/write operations. In this paper, we present a method of reducing memory energy consumption during read/write operations. It aims at minimizing the number of row opens and closes, which are the major source of energy consumption during read/write operations. The basic idea is to apply network entry control to prioritize consecutive open row memory accesses. The experimental results show up to 35% reduction in memory energy consumption with an industrial strength multimedia mobile SoC.
Sungjoo Yoo, Kiyoung Choi
ISLPED2
2007 Communication Architecture Synthesis of Cascaded Bus Matrix
abstract
For high frequency on-chip communication architecture design, we propose cascaded bus matrix-based solutions. Due to the huge design space in cascaded bus matrix design, it is crucial to perform an efficient design space exploration. In our work, we present a simulated annealing-based design space exploration method. For an efficient representation of bus topology, we propose an encoding method called traffic group encoding and apply it to AMBA3 AXI-based bus system design. In addition, we propose a method of two-step simulated annealing to improve the quality of results. Experimental results show that the proposed methods allow designing complex communication architectures (ones with up to 31 masters and 71 slaves) with high frequency constraints to which existing methods could not give solutions.
Jun-hee Yoo, Sungjoo Yoo, Kiyoung Choi
ASP-DAC3
2006 PowerViP: Soc power estimation framework at transaction level
abstract
In this work, we propose a SoC power estimation framework built on our system-level simulation environment. Our framework provides designers with the system-level power profile in a cycle-accurate manner. We target the framework to run fast and accurately, which is enabled by adopting different modeling techniques depending on the power characteristics of various IP blocks. The framework can be applied to any target SoC design
Ikhwan Lee, Sungjoo Yoo, Eui-Young Chung, Kyu-Myung Choi, Jeong-Taek Kong, Soo-Kwan Eo
ASP-DAC4
2006 Runtime distribution-aware dynamic voltage scaling
abstract
We propose a new intra-task dynamic voltage scaling (DVS) method to capture an important fact of ‘software runtime distribution ’ and integrate it into DVS effectively. Specifically, the proposed method finds performance levels, for a given software runtime distribution, i.e. statistical profiling of execution cycles (neither the execution cycle of worst-case execution path nor the worst-case execution cycles of basic blocks), which leads to a minimal energy consumption while satisfying the given deadline constraints. Experimental results report that the proposed method gives 19.2%~33.3 % further energy reduction compared with the best-known methods for two industrial multimedia software programs, H.264 decoder and MPEG4 decoder. 1.
Sungpack Hong, Sungjoo Yoo, HoonSang Jin, Kyu-Myung Choi, Jeong-Taek Kong, Soo-Kwan Eo
ICCAD2
2005 Scheduler implementation in MP SoC design
abstract
In the design of a heterogeneous multiprocessor system on chip, we face a new design problem; scheduler implementation. In this paper, we present an approach to implementing a static scheduler, which controls all the task executions and communication transactions of a system according to a pre-determined schedule. For the scheduler implementation, we consider both intra-processor and inter-processor synchronization. We also consider scheduler overhead, which is often neglected. In particular, we address the issue of centralized implementation versus distributed implementation. We investigate the pros and cons of the two different scheduler implementations. Through experiments with synthetic examples and a real world multimedia application, we show the effectiveness of our approach.
Youngchul Cho, Sungjoo Yoo, Kiyoung Choi, Nacer-Eddine Zergainoh, Ahmed Amine Jerraya
ASP-DAC2
2004 Fast and accurate timed execution of high level embedded software using HW/SW interface simulation model
Aimen Bouchhima, Sungjoo Yoo, Ahmed Amine Jerraya
ASP-DAC2
2004 Debugging HW/SW interface for MPSoC: video encoder system design case study
abstract
This paper reports a case study of multiprocessor SoC (MPSoC) design of a complex video encoder, namely OpenDivX. OpenDivX is a popular version of MPEG4. It requires massive computation resources and deals with complex data structures to represent video streams. In this study, the initial specification is given in sequential C code that had to be parallelized to be executed on four different processors. High level programming model, namely Message Passing Interface (MPI) was used to enable inter-task communication among parallelized C code. A four processor hardware prototyping platform was used to debug the parallelized software before final SoC hardware is ready. The targeting of abstract parallel code using MPI to the multiprocessor architecture required the design of an additional hardware-dependent software layer to refine the abstract programming model. The design was made by a team work of three types of designer: application software, hardware-dependent software and hardware platform designers. The collaboration was necessary to master the whole flow from the specification to the platform.The study showed that HW/SW interface debug was the most time-consuming step. This is identified as a potential killer for application-specific MPSoC design. To further investigate the ways to accelerate the HW/SW interface debug, we analyzed bugs found in the case study and the available debug environments. Finally, we address a debug strategy that exploits efficiently existing debug environments to reduce the time for HW/SW interface debug.
Mohamed-Wassim Youssef, Sungjoo Yoo, Arif Sasongko, Yanick Paviot, Ahmed Amine Jerraya
DAC2
2004 Multi-Processor SoC Design Methodology Using a Concept of Two-Layer Hardware-Dependent Software
abstract
In conventional multiprocessor SoC (MPSoC) design methods, we find two problems: lack of SW code portability and lack of early SW validation. The problems cause a long design cycle. To resolve them, we present a concept of two-layer hardware-dependent software (HdS). The presented HdS consists of hardware abstraction layer to abstract the sub-system architecture and SoC abstraction layer to abstract the global MPSoC architecture. During the exploration of global and sub-system architectures, the application programming interfaces of presented two-layer HdS allow to keep the SW independent from architectural change. The simulation models of two-layer HdS enable to validate the entire system including the SW and HW design early in the design steps. We show the effectiveness of the presented methodology in the MPSoC architecture exploration of an OpenDiVX encoder system design.
Sungjoo Yoo, Mohamed-Wassim Youssef, Aimen Bouchhima, Ahmed Amine Jerraya, Mario Diaz-Nava
DATE1
2003 Scheduling and Timing Analysis of HW/SW On-Chip Communication in MP SoC Design
abstract
On-chip communication design includes designing software (SW) parts (operating system, device drivers, interrupt service routines, etc.) as well as hardware (HW) parts (on-chip communication network, communication interfaces of processor/IP/memory, on-chip memory, etc.). For an efficient exploration of its design space, we need fast scheduling and timing analysis. In this work, we tackle two problems (one for SW and the other for HW) in on-chip communication design. One is to incorporate the dynamic behavior of SW (interrupt processing and context switching) into on-chip communication scheduling. The other is to reduce on-chip data storage required for on-chip communication, by sharing physical communication buffers with different communication transactions. To solve the problems, we present both ILP (integer linear programming) formulation and heuristic algorithm, which enable the designer to perform efficient onchip communication scheduling and obtain accurate timing information. Experimental results show the effectiveness of our work.
Youngchul Cho, Ganghee Lee, Sungjoo Yoo, Kiyoung Choi, Nacer-Eddine Zergainoh
DATE3
2003 Building Fast and Accurate SW Simulation Models Based on Hardware Abstraction Layer and Simulation Environment Abstraction Layer
Sungjoo Yoo, Iuliana Bacivarov, Aimen Bouchhima, Yanick Paviot, Ahmed Amine Jerraya
DATE1
2003 Introduction to Hardware Abstraction Layers for SoC
Sungjoo Yoo, Ahmed Amine Jerraya
DATE1
2002 Component-based design approach for multicore SoCs
abstract
This paper presents a high-level component-based methodology and design environment for application-specific multicore SoC architectures. Component-based design provides primitives to build complex architectures from basic components. This bottom-up approach allows design-architects to explore efficient custom solutions with best performances. This paper presents a high-level component-based methodology and design environment for application-specific multicore SoC architectures. The system specifications are represented as a virtual architecture described in a SystemC-like model and annotated with a set of configuration parameters. Our component-based design environment provides automatic wrapper-generation tools able to synthesize hardware interfaces, device drivers, and operating systems that implement a high-level interconnect API. This approach, experimented over a VDSL system, shows a drastic design time reduction without any significant efficiency loss in the final circuit.
Wander O. Cesário, Amer Baghdadi, Lovic Gauthier, Damien Lyonnard, Gabriela Nicolescu, Yanick Paviot, Sungjoo Yoo, Ahmed Amine Jerraya, Mario Diaz-Nava
DAC7
2002 Automatic Generation of Fast Timed Simulation Models for Operating Systems in SoC Design
abstract
To enable fast and accurate evaluation of HW/SW implementation choices of on-chip communication, we present a method to automatically generate timed OS simulation models. The method generates the OS simulation models with the simulation environment as a virtual processor Since the generated OS simulation models use final OS code, the presented method can mitigate the OS code equivalence problem. The generated model also simulates different types of processor exceptions. This approach provides two orders of magnitude higher simulation speedup compared to the simulation using instruction set simulators for SW simulation.
Sungjoo Yoo, Gabriela Nicolescu, Lovic Gauthier, Ahmed Amine Jerraya
DATE1
2002 An intra-task dynamic voltage scaling method for SoC design with hierarchical FSM and synchronous dataflow model
abstract
This paper presents a method of intra-task dynamic voltage scaling (DVS) for SoC design with hierarchical FSM and synchronous dataflow model (in short, HFSM-SDF model). To have an optimal intra-task DVS, exact execution paths need to be determined in compile time or runtime. In general programs, since determining exact execution paths in compile time or runtime is not possible, existing methods assume worst/average-case execution paths and take static voltage scaling approaches. In our work, we exploit a property of HFSM-SDF model to calculate exact execution paths in runtime. With the information of exact execution paths, our DVS method can calculate exact remaining workload. The exact workload enables to calculate optimal voltage level which gives optimal energy consumption while satisfying the given timing constraint. Experiments show the effectiveness of the presented method in low-power design of an MPEG4 decoder system.
Sunghyun Lee 0002, Kiyoung Choi, Sungjoo Yoo
ISLPED3
2001 Scalable and flexible cosimulation of SoC designs with heterogeneous multi-processor target architectures
abstract
In this paper, we present a cosimulation environment that provides modularity, scalability, and flexibility in cosimulation of SoC designs with heterogeneous multi-processor target architectures. Our cosimulation environment is based on an object-oriented simulation environment, SystemC. Exploiting the object orientation in SystemC representation, we achieve modularity and scalability of cosimulation by developing modular cosimulation interfaces. The object orientation also enables mixed-level cosimulation to be easily implemented thereby the designer can have flexibility in trade off between simulation performance and accuracy. Experiments with an IS-95 CDMA cellular phone system design show the effectiveness of the cosimulation environment.
Patrice Gerin, Sungjoo Yoo, Gabriela Nicolescu, Ahmed Amine Jerraya
ASP-DAC2
2001 Automatic Generation of Application-Specific Architectures for Heterogeneous Multiprocessor System-on-Chip
abstract
We present a design flow for the generation of application-specific multiprocessor architectures. In the flow, architectural parameters are first extracted from a high-level system specification. Parameters are used to instantiate architectural components, such as processors, memory modules and communication networks. The flow includes the automatic generation of communication coprocessor that adapts the processor to the communication network in an application-specific way. Experiments with two system examples show the effectiveness of the presented design flow.
Damien Lyonnard, Sungjoo Yoo, Amer Baghdadi, Ahmed Amine Jerraya
DAC2
2001 Automatic generation and targeting of application specific operating systems and embedded systems software
abstract
We propose a method of automatic generation of application specific operating systems (OS's) and automatic targeting of application software. OS generation starts from a very small bur yet flexible OS kernel. OS services, which are specific to the application and deduced from dependencies between services, are added to the kernel to construct the whole OS. Communication and synchronization functions in the application code are adapted to the generated OS. As a preliminary experiment, we applied the proposed method to a system example called token ring system.
Lovic Gauthier, Sungjoo Yoo, Ahmed Amine Jerraya
DATE2
2001 Performance improvement of multi-processor systems cosimulation based on SW analysis
abstract
We propose a method for performance improvement of multi-processor systems cosimulation by reducing synchronization overhead between multiple simulators. To reduce the amount of simulator synchronization, we predict synchronization time points based on a static analysis of application software running on each processor. In experiments with real embedded systems, we obtained orders of magnitude higher performance in cosimulation runtimes.
Jinyong Jung, Sungjoo Yoo, Kiyoung Choi
DATE2
2001 Mixed-level cosimulation for fine gradual refinement of communication in SoC design
abstract
In this paper we propose a method of mixed-level cosimulation that enables gradual refinement of SoC communication from protocol-neutral communication to protocol-fixed communication. For fine granularity in refinement, the method enables the designer to perform channel refinement and module refinement. Thus, the designer can perform more extensive design space exploration in communication refinement. We show the effectiveness of the proposed method in a case study of communication refinement is an IS-95 CDMA cellular phone system design.
Gabriela Nicolescu, Sungjoo Yoo, Ahmed Amine Jerraya
DATE2
2001 Automatic generation and targeting of application-specificoperating systems and embedded systems software
abstract
Software (SW) parts become crucial in embedded systems. Operating systems (OSs) are often used to handle SW concurrency and communication. We propose a method of automatic generation of application-specific OSs and automatic targeting of application SW. OS generation starts from a very small, but yet flexible OS kernel. OS services, which are specific to the application and deduced from dependencies created by the system specification, are added to the kernel to construct the whole OS. Communication and synchronization functions in the application code are adapted to the generated OS. As experiments, we applied the proposed method to two system examples: a token-ring system and a very high data-rate digital subscriber line framer.
Lovic Gauthier, Sungjoo Yoo, Ahmed Amine Jerraya
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2000 Hardware-software cosynthesis for run-time incrementally reconfigurable FPGAs
abstract
This paper presents a method for hardware-softw are cosynthesis with run-time incrementally recon gurable FPGAs.T o reduce the run-time o v erhead of recon guring FPGAs, we present a concept called early partial recon guration (EPR) which minimizes the ov erhead by performing recon guration for an operation (or a task in our terms) mapped to an FPGA as early as possible so that the operation is ready to start when its execution is requested.F or further reduction of the ov erhead,w ein tegrate the incremental reconguration (IR) of FPGAs with the EPR concept.We present an ILP formulation and an eÆcient heuristic algorithm based on the EPR and IR concepts.Experiments on embedded system examples and syn thetic examples show the eÆciency of the proposed method
Byungil Jeong, Sungjoo Yoo, Sunghyun Lee 0002, Kiyoung Choi
ASP-DAC2
2000 Fast Hardware-Software Coverification by Optimistic Execution of Real Processor
abstract
To achieve fast verification of the software part of an embedded system, we propose to run the target processor optimistically, which effectively reduces the synchronization overhead with other simulators. For the optimistic processor execution, we present a processor execution platform and state saving/restoration methods. We performed optimistic execution of ARM710A processor in the coverification of an IS-95 CDMA cellular phone system and obtained up to orders of magnitude higher performance compared with the case that the processor runs conservatively.
Sungjoo Yoo, Jongeun Lee, Jinyong Jung, Kyoungseok Rha, Youngchul Cho, Kiyoung Choi
DATE1
2000 Performance improvement of geographically distributed cosimulation by hierarchically grouped messages
abstract
To improve the performance of geographically distributed cosimulation, we propose a concept called hierarchically grouped message. The concept improves cosimulation performance, preserving the cosimulation accuracy, by hierarchically grouping messages transferred between simulators in a short period of simulated time into a single physical message, thereby reducing the number of physical messages. Applying the concept to hybrid and optimistic cosimulation, we can reduce the number of rollbacks as well as the communication overhead accompanying the message transfer. Experimental results show the efficiency of the proposed method for practical examples in an internationally distributed cosimulation environment.
Sungjoo Yoo, Kiyoung Choi, Dong Sam Ha
IEEE Trans. Very Large Scale Integr. Syst.1
1999 Exploiting Early Partial Reconfiguration of Run-Time Reconfigurable FPGAs in Embedded Systems Design
abstract
No abstract available.
Byungil Jeong, Sungjoo Yoo, Kiyoung Choi
FPGA2