EDBT 2026 Demo / reviewers in the wild / expert
Kiyoung Choi
dblp:53/2459
· DBLP profile ↗
133ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0001-6138-6697ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 122 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 16Applied, interdisciplinary, general and emerging computing · 11Artificial intelligence and machine learning · 8 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSecurity and privacy · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Low Impedance Rendering Toward Safe Human-Robot InteractionabstractThis study proposes a novel approach for low-impedance rendering in robots to ensure safe human-robot interaction. Low-impedance control enhances safety by enabling robots to respond flexibly to physical contact with humans. However, real-world disturbances such as friction often necessitate high impedance, creating a trade-off between safety and task precision. To address this challenge, a modified disturbance observer (DOB)-based control framework is introduced, designed to prevent external contact forces from being treated as disturbances. Experimental results demonstrate significant improvements in impedance rendering performance under various conditions, ensuring precise and safe robot operations in dynamic shared environments. Wonbum Yun, Kiyoung Choi, Junyoung Kim 0003, Sehoon Oh, Hyun-Joon Chung |
HRI | 2 |
| 2025 | Robust Orientation Control of Robot Manipulator Using Orientation Disturbance ObserverabstractThis paper presents a robust control algorithm for precise orientation control of robot manipulators using a disturbance observer (DOB) specifically designed for orientation dynamics. Our approach addresses the challenges of 3D orientation control by incorporating various orientation representations, such as Euler angles, quaternions, and exponential coordinates, and analyzing their impact on DOB performance. Through theoretical analysis and experimental validation, we demonstrate the effectiveness of our method in achieving high-precision orientation control under uncertainties and disturbances. This work offers a comprehensive framework for robust orientation control, advancing the application of DOB in complex robotic tasks. Kiyoung Choi, Wonbum Yun, Sehoon Oh |
ICRA | 1 |
| 2024 | Identification of Flexible Joint Robot Inertia Matrix Using Frequency Response AnalysisabstractThis paper presents a novel, nonlinearity robust identification method for deriving the inertia matrix of multi-DOF Flexible Joint Robots (FJR), utilizing resonance and anti-resonance frequencies in the Frequency Response Functions (FRF). Our proposed method overcomes the limitations of conventional approaches, which are susceptible to mechanical nonlinearities, leading to inaccurate models. By leveraging frequency domain techniques, our approach effectively mitigates the influence of nonlinear characteristics, providing a more accurate and reliable means of robot control. Furturmore, the paper highlights the benefits of frequency domain system identification, including nonlinear robustness and the ability to decompose the flexible joint into motor and load components. Finally, a novel sequential excitation algorithm is proposed to obtain the inertia matrix of a multi-DOF robot manipulator without relying on complex theories or optimizations. The effectiveness of the proposed algorithm is verified through simulation and experiment. Kiyoung Choi, Wonbum Yun, Deokjin Lee, Sehoon Oh |
IROS | 1 |
| 2022 | Human-Robot Interaction Force based Power Assistive Algorithm of Upper Limb Exoskeleton Robots Driven by a Series Elastic ActuatorabstractUpper limb exoskeleton robots have been widely used to assist humans in industry and rehabilitation. Human-robot interaction control consists of two modes. A compliant mode and an assistive mode. The compliant mode is manipulating the robot compliantly according to the movement of the human. At this time, the human must not receive impedance from the robot. Therefore, in this paper, a novel external force control which makes the robot follow human movement without causing resistance is proposed. The assistive mode is to generate assistive power according to human intention. In this paper, the algorithm for recognizing the human intention for primitive movement and generating assistive forces adaptive to real-time work environments is proposed. The proposed two modes are also integrated into a novel human-robot interaction control framework. The performance of the proposed methods is verified through experiments. Deokjin Lee, Kiyoung Choi, Wonbum Yun, Sehoon Oh |
IECON | 2 |
| 2022 | ComPreEND: Computation Pruning through Predictive Early Negative Detection for ReLU in a Deep Neural Network AcceleratorabstractA vast amount of activation values of DNNs are zeros due to ReLU (Rectified Linear Unit), which is one of the most common activation functions used in modern neural networks. Since ReLU outputs zero for all negative inputs, the inputs to ReLU do not need to be determined exactly as long as they are negative. However, many accelerators usually do not consider such aspects of DNNs, losing a huge amount of opportunities for speedups and energy savings. To exploit such opportunities, we propose early negative detection (END), a computation pruning technique that detects the negative results at an early stage. The key to the early negative detection is the adoption of inverted two's complement representation for filter parameters. This ensures that as soon as the intermediate results become negative, the final results are guaranteed to be negative. Upon detection, the remaining computation can be skipped and the following ReLU output can be simply set to zero. We also propose a DNN accelerator architecture (ComPreEND) that takes advantage of such skipping. ComPreEND with END significantly improves both the energy efficiency and the performance according to the evaluation. Compared to the baseline, we obtain 20.5 and 29.3 percent speedup with accurate mode and predictive mode, and energy savings by 28.4 and 41.4 percent, respectively. Namhyung Kim, Hanmin Park, Sungbum Kang, Jinho Lee 0001, Kiyoung Choi |
IEEE Trans. Computers | 6 |
| 2021 | GradPIM: A Practical Processing-in-DRAM Architecture for Gradient DescentabstractIn this paper, we present GradPIM, a processingin-memory architecture which accelerates parameter updates of deep neural networks training. As one of processing-in-memory techniques that could be realized in the near future, we propose an incremental, simple architectural design that does not invade the existing memory protocol. Extending DDR4 SDRAM to utilize bank-group parallelism makes our operation designs in processing-in-memory (PIM) module efficient in terms of hardware cost and performance. Our experimental results show that the proposed architecture can improve the performance of DNN training and greatly reduce memory bandwidth requirement while posing only a minimal amount of overhead to the protocol and DRAM area. Heesu Kim, Hanmin Park, Kwanheum Cho, Eojin Lee, Soojung Ryu, Kiyoung Choi, Jinho Lee 0001 |
HPCA | 8 |
| 2019 | Network Recasting: A Universal Method for Network Architecture Transformation
Joonsang Yu, Sungbum Kang, Kiyoung Choi |
AAAI | 3 |
| 2019 | Cell division: weight bit-width reduction technique for convolutional neural network hardware acceleratorsabstractThe datapath bit-width of hardware accelerators for convolutional neural network (CNN) inference is generally chosen to be wide enough, so that they can be used to process upcoming unknown CNNs. Here we introduce the cell division technique, which is a variant of function-preserving transformations. With this technique, it is guaranteed that CNNs that have weights quantized to fixed-point format of arbitrary bit-widths, can be transformed to CNNs with less bit-widths of weights without any accuracy drop (or any accuracy change). As a result, CNN hardware accelerators are released from the weight bit-width constraint, which has been preventing them from having narrower datapaths. In addition, CNNs that have wider weight bit-widths than those assumed by a CNN hardware accelerator can be executed on the accelerator. Experimental results on LeNet-300-100, LeNet-5, AlexNet, and VGG-16 show that weights can be reduced down to 2--5 bits with 2.5X--5.2X decrease in weight storage requirement and of course without any accuracy drop. Hanmin Park, Kiyoung Choi |
ASP-DAC | 2 |
| 2019 | Acceleration of DNN Backward Propagation by Selective Computation of GradientsabstractThe training process of a deep neural network commonly consists of three phases: forward propagation, backward propagation, and weight update. In this paper, we propose a hardware architecture to accelerate the backward propagation. Our approach applies to neural networks that use rectified linear unit. Considering that the backward propagation results in a zero activation gradient when the corresponding activation is zero, we can safely skip the gradient calculation. Based on this observation, we design an efficient hardware accelerator for training deep neural networks by selectively computing gradients. We show the effectiveness of our approach through experiments with various network models. Hanmin Park, Namhyung Kim, Joonsang Yu, Sujeong Jo, Kiyoung Choi |
DAC | 6 |
| 2019 | Aging Gracefully with ApproximationabstractThis paper presents a design methodology to turn aging-induced chip slowdown into approximation without adding reliability guardband or increasing supply voltage. It guarantees always-best quality while the system is under aging. It is based on run-time monitoring of critical path delay. If the delay increases due to aging, the proposed approach curtails the critical path at the cost of precision reduction. We evaluate our approach at the component level as well as microarchitecture level. The evaluation results show that the approach reduces the dynamic and static power consumptions by 19.8% and 10.2%, respectively, with minimal area overhead and quality degradation. Heesu Kim, Hussam Amrouch, Jörg Henkel, Andreas Gerstlauer, Kiyoung Choi |
ISCAS | 6 |
| 2019 | VCAM: Variation Compensation through Activation Matching for Analog Binarized Neural NetworksabstractWe propose an energy-efficient analog implementation of binarized neural network with a novel technique called VCAM, variation compensation through activation matching. The architecture consists of 1T1R ReRAM arrays and differential amplifiers for implementing synapses and neurons, respectively. To restore classification test accuracy degraded by process variation, we adjust the biases of the neurons to match their average output activations with those of ideal neurons. Experimental results show that the proposed approach recovers the accuracy to 98.55% on MNIST and 89.63% on CIFAR-10 even in the presence of 50% threshold voltage and 15% resistance variations at 3-sigma point. This result corresponds to the accuracy degradation of only 0.05% and 1.35%, respectively, compared to the ideal case. Chaeun Lee, Yumin Kim, Cheol Seong Hwang, Kiyoung Choi |
ISLPED | 6 |
| 2018 | Architectures and algorithms for user customization of CNNsabstractIn this paper we present a convolutional neural network architecture that supports user customization through incremental transfer learning. The architecture consists of a large basic inference engine and a small augmenting engine. After training the basic inference engine and augmenting engine on a large general dataset, the basic inference engine is fixed. For user customization, only the augmenting engine is re-trained on-device using a small user specific dataset provided by the user. To accelerate the training of the augmenting engine we map this to a coarsegrained reconfigurable array processor. The complete network architecture is evaluated using the Caffe framework, and a C-code equivalent network is implemented and tested on a CGRA processor. Experiments with NIST'19 and our user-specific datasets show an increase in accuracy of the system from 76.3% to 93.2% after user customization. Mapping this code to a CGRA gives us a speed up of 45x and a 49-and 3-fold reduced energy consumption over an ARMv7 processor and a 3-way VLIW processor, respectively, showing the potential of CGRAs as DNN processors. Barend Harris, Mansureh S. Moghaddam, Duseok Kang, Inpyo Bae, Euiseok Kim, Hyemi Min, Hansu Cho, Sukjin Kim, Bernhard Egger 0002, Soonhoi Ha, Kiyoung Choi |
ASP-DAC | 11 |
| 2018 | Training Neural Networks with Low Precision Dynamic Fixed-PointabstractDynamic fixed-point (DFP) is one of the most successful attempts to reduce bit-widths in training neural networks. It has been reported that DFP can reduce the bit-widths of the training operations to 16 bits, except for parameter update operations; parameter updates for general networks need higher precision and are usually done with 32-bit floating-point operations. In this paper, we propose two methods of using 16-bit DFP for all the training operations including parameter updates; weight clipping and gradual batch size increase. Lastly, we combine the two methods to further explore their potentials. We successfully apply 16-bit DFP operations on the parameter updates of LeNet-5 and VGG-16 networks using CIFAR10 and CIFAR100 datasets. Sujeong Jo, Hanmin Park, Kiyoung Choi |
ICCD | 4 |
| 2018 | ComPEND: Computation Pruning through Early Negative Detection for ReLU in a Deep Neural Network AcceleratorabstractWhile negative inputs for ReLU are useless, it consumes a lot of computing power to calculate them for deep neural networks. We propose a computation pruning technique that detects at an early stage that the result of a sum of products will be negative by adopting an inverted two's complement expression for weights and a bit-serial sum of products. Therefore, it can skip a large amount of computations for negative results and simply set the ReLU outputs to zero. Moreover, we devise a DNN accelerator architecture that can efficiently apply the proposed technique. The evaluation shows that the accelerator using the computation pruning through early negative detection technique significantly improves the energy efficiency and the performance. Sungbum Kang, Kiyoung Choi |
ICS | 3 |
| 2018 | Deep neural networks with weighted spikes
Heesu Kim, Subin Huh, Jinho Lee 0001, Kiyoung Choi |
Neurocomputing | 5 |
| 2018 | Benzene: An Energy-Efficient Distributed Hybrid Cache Architecture for Manycore SystemsabstractThis article proposes Benzene, an energy-efficient distributed SRAM/STT-RAM hybrid cache for manycore systems running multiple applications. It is based on the observation that a naïve application of hybrid cache techniques to distributed caches in a manycore architecture suffers from limited energy reduction due to uneven utilization of scarce SRAM. We propose two-level optimization techniques: intra-bank and inter-bank. Intra-bank optimization leverages highly associative cache design, achieving more uniform distribution of writes within a bank. Inter-bank optimization evenly balances the amount of write-intensive data across the banks. Our evaluation results show that Benzene significantly reduces energy consumption of distributed hybrid caches. Namhyung Kim, Junwhan Ahn, Kiyoung Choi, Daniel Sánchez 0003, Donghoon Yoo, Soojung Ryu |
ACM Trans. Archit. Code Optim. | 3 |
| 2018 | An Efficient and Accurate Stochastic Number Generator Using Even-Distribution CodingabstractStochastic computing (SC) is a promising approach for low-power and low-cost applications with the added benefit of high error tolerance. However, the high overhead of generating stochastic bitstreams can offset the advantages of SC especially when a large number of bitstreams are needed. In this paper, we propose a new stochastic number generator (SNG) that significantly reduces area and energy while improving accuracy. Experimental results show that the proposed SNG can reduce energy by more than 72% compared with the state-of-the-art designs. Aidyn Zhakatayev, Kyounghoon Kim, Kiyoung Choi, Jongeun Lee |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Delay Monitoring System With Multiple Generic Monitors for Wide Voltage Range OperationabstractAs the semiconductor process technology continuously scales down, circuit delay variations due to manufacturing and environmental variations become more and more serious. These delay variations are hardly predictable and thus require an additional design margin, which impedes the chance to reduce the area and power consumption of a chip. One of the best solutions to alleviate this problem is to measure circuit delays at run time and control the supply voltage accordingly through a closed-loop dynamic voltage and frequency scaling (DVFS) scheme. The key issue of this scheme is the delay mismatch between the monitoring circuit and the target block. A large delay mismatch might lose the advantage of the closed-loop DVFS. It becomes much worse as a circuit block operates in wider voltage range, from near-threshold voltage to super-overdrive voltage. This paper proposes novel delay monitoring systems with multiple generic monitors for wide voltage range operation, which provide a better delay correlation between the monitoring circuit and the target block compared to conventional monitoring approaches. The proposed approaches reduce the maximum error by up to 91% for a popular processor core in a 14-nm FinFET process technology, thereby bring a decrease of design margin, lower-power, and/or lower-cost design. Kiyoung Choi, Wook Kim, Kyung Tae Do, Jung-Hwan Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Scalable stochastic-computing accelerator for convolutional neural networksabstractStochastic Computing (SC) is an alternative design paradigm particularly useful for applications where cost is critical. SC has been applied to neural networks, as neural networks are known for their high computational complexity. However previous work in this area has critical limitations such as the fully-parallel architecture assumption, which prevent them from being applicable to recent ones such as convolutional neural networks, or ConvNets. This paper presents the first SC architecture for ConvNets, shows its feasibility, with detailed analyses of implementation overheads. Our SC-ConvNet is a hybrid between SC and conventional binary design, which is a marked difference from earlier SC-based neural networks. Though this might seem like a compromise, it is a novel feature driven by the need to support modern ConvNets at scale, which commonly have many, large layers. Our proposed architecture also features hybrid layer composition, which helps achieve very high recognition accuracy. Our detailed evaluation results involving functional simulation and RTL synthesis suggest that SC-ConvNets are indeed competitive with conventional binary designs, even without considering inherent error resilience of SC. Hyeon Uk Sim, Dong Nguyen 0001, Jongeun Lee, Kiyoung Choi |
ASP-DAC | 4 |
| 2017 | Incremental training of CNNs for user customization: work-in-progressabstractThis paper presents a convolutional neural network architecture that supports transfer learning for user customization. The architecture consists of a large basic inference engine and a small augmenting engine. Initially, both engines are trained using a large dataset. Only the augmenting engine is tuned to the user-specific dataset. To preserve the accuracy for the original dataset, the novel concept of quality factor is proposed. The final network is evaluated with the Caffe framework, and our own implementation on a coarse-grained reconfigurable array (CGRA) processor. Experiments with MNIST, NIST'19, and our user-specific datasets show the effectiveness of the proposed approach and the potential of CGRAs as DNN processors. Mansureh S. Moghaddam, Barend Harris, Duseok Kang, Inpyo Bae, Euiseok Kim, Hyemi Min, Hansu Cho, Sukjin Kim, Bernhard Egger 0002, Soonhoi Ha, Kiyoung Choi |
CASES | 11 |
| 2017 | A space- and energy-efficient code Compression/Decompression technique for coarse-grained reconfigurable architectures
Bernhard Egger 0002, Duseok Kang, Mansureh S. Moghaddam, Youngchul Cho, Yeonbok Lee, Sukjin Kim, Soonhoi Ha, Kiyoung Choi |
CGO | 9 |
| 2017 | Design space exploration of FPGA accelerators for convolutional neural networksabstractThe increasing use of machine learning algorithms, such as Convolutional Neural Networks (CNNs), makes the hardware accelerator approach very compelling. However the question of how to best design an accelerator for a given CNN has not been answered yet, even on a very fundamental level. This paper addresses that challenge, by providing a novel framework that can universally and accurately evaluate and explore various architectural choices for CNN accelerators on FPGAs. Our exploration framework is more extensive than that of any previous work in terms of the design space, and takes into account various FPGA resources to maximize performance including DSP resources, on-chip memory, and off-chip memory bandwidth. Our experimental results using some of the largest CNN models including one that has 16 convolutional layers demonstrate the efficacy of our framework, as well as the need for such a high-level architecture exploration approach to find the best architecture for a CNN model. Atul Rahman, Sangyun Oh, Jongeun Lee, Kiyoung Choi |
DATE | 4 |
| 2017 | FPGA implementation of convolutional neural network based on stochastic computingabstractThere has been a body of research to use stochastic computing (SC) for the implementation of neural networks, in the hope that it will reduce the area cost and energy consumption. However, no working neural network system based on stochastic computing has been demonstrated to support the viability of SC-based deep neural networks in terms of both recognition accuracy and cost/energy efficiency. In this demonstration we present an SC-based deep nenural network system that is highly accurate and efficient. Our system takes an input image and processes it with a convolutional neural network implemented on an FPGA using stochastic computing to recognize the input image, with nearly the same accuracy as conventional binary implementations. Daewoo Kim, Mansureh S. Moghaddam, Hossein Moradian, Hyeon Uk Sim, Jongeun Lee, Kiyoung Choi |
FPT | 6 |
| 2017 | Accurate and Efficient Stochastic Computing Hardware for Convolutional Neural NetworksabstractThis paper presents an efficient unipolar stochastic computing hardware for convolutional neural networks (CNNs). It includes stochastic ReLU and optimized max function, which are key components in a CNN. To avoid the range limitation problem of stochastic numbers and increase the signal-to-noise ratio, we perform weight normalization and upscaling. In addition, to reduce the overhead of binary-to-stochastic conversion, we propose a scheme for sharing stochastic number generators among the neurons in a CNN. Experimental results show that our approach outperforms the previous ones based on stochastic computing in terms of accuracy, area, and energy consumption. Joonsang Yu, Kyounghoon Kim, Jongeun Lee, Kiyoung Choi |
ICCD | 4 |
| 2017 | Synthesis of multi-variate stochastic computing circuitsabstractStochastic computing (SC) is a promising technique to enhance computing efficiency in terms of area, power, and error tolerance with slight compromise of the accuracy. This paper presents a novel approach to automatic synthesis of an SC circuit from a given arithmetic expression with multiple variables. It first extracts building blocks called iSC kernels from the given expressions and then synthesizes an SC circuit by using the iSC kernels. Experimental results demonstrate the efficiency of the proposed technique in terms of synthesis time and the quality of the generated circuit. Kyounghoon Kim, Kiyoung Choi |
VLSI-SoC | 2 |
| 2017 | ExtraV: Boosting Graph Processing Near Storage with a Coherent AcceleratorabstractIn this paper, we propose ExtraV, a framework for near-storage graph processing. It is based on the novel concept of graph virtualization , which efficiently utilizes a cache-coherent hardware accelerator at the storage side to achieve performance and flexibility at the same time. ExtraV consists of four main components: 1) host processor, 2) main memory, 3) AFU (Accelerator Function Unit) and 4) storage. The AFU, a hardware accelerator, sits between the host processor and storage. Using a coherent interface that allows main memory accesses, it performs graph traversal functions that are common to various algorithms while the program running on the host processor (called the host program) manages the overall execution along with more application-specific tasks. Graph virtualization is a high-level programming model of graph processing that allows designers to focus on algorithm-specific functions. Realized by the accelerator, graph virtualization gives the host programs an illusion that the graph data reside on the main memory in a layout that fits with the memory access behavior of host programs even though the graph data are actually stored in a multi-level, compressed form in storage. We prototyped ExtraV on a Power8 machine with a CAPI-enabled FPGA. Our experiments on a real system prototype offer significant speedup compared to state-of-the-art software only implementations. Jinho Lee 0001, Heesu Kim, Sungjoo Yoo, Kiyoung Choi, H. Peter Hofstee, Gi-Joon Nam, Mark Nutter, Damir A. Jamsek |
Proc. VLDB Endow. | 4 |
| 2017 | Dirty-Block Tracking in a Direct-Mapped DRAM Cache with Self-Balancing DispatchabstractRecently, processors have begun integrating 3D stacked DRAMs with the cores on the same package, and there have been several approaches to effectively utilizing the on-package DRAMs as caches. This article presents an approach that combines the previous approaches in a synergistic way by devising a module called the dirty-block tracker to maintain the dirtiness of each block in a dirty region. The approach avoids unnecessary tag checking for a write operation if the corresponding block in the cache is not dirty. Our simulation results show that the proposed technique achieves a 10.3% performance improvement on average over the state-of-the-art DRAM cache technique. Sangheon Lee 0006, Soojung Ryu, Kiyoung Choi |
ACM Trans. Archit. Code Optim. | 4 |
| 2017 | Excavating the Hidden Parallelism Inside DRAM Architectures With Buffered ComparesabstractWe propose an approach called buffered compares, a less-invasive processing-in-memory solution that can be used with existing processor memory interfaces such as DDR3/4 with minimal changes. The approach is based on the observation that multibank architecture, a key feature of modern main memory DRAM devices, can be used to provide huge internal bandwidth without any major modification. We place a small buffer and a simple ALU per bank, define a set of new DRAM commands to fill the buffer and feed data to the ALU, and return the result for a set of commands (not for each command) to the host memory controller. By exploiting the under-utilized internal bandwidth using `compare-n-op' operations, which are frequently used in various applications, we not only reduce the amount of energy-inefficient processor-memory communication, but also accelerate the computation of big data processing applications by utilizing parallelism of the buffered compare units in DRAM banks. We present two versions of buffered compare architecture-full-scale architecture and reduced architecture-in trade of performance and energy. The experimental results show that our solution significantly improves the performance and efficiency of the system on the tested workloads. Jinho Lee 0001, Jongwook Chung, Jung Ho Ahn, Kiyoung Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | An energy-efficient random number generator for stochastic circuitsabstractStochastic circuits provide very high efficiency in terms of gate area and power consumption compared with conventional binary logic. However, they require random bit streams generated by stochastic number generators (SNGs), which account for a significant portion of area and energy offsetting their merits. In this paper, we propose a new SNG that significantly reduces area and energy while improving accuracy in progressive precision. Experimental results show that the proposed SNG reduces energy by more than 72% compared to the state-of-the-art designs. Kyounghoon Kim, Jongeun Lee, Kiyoung Choi |
ASP-DAC | 3 |
| 2016 | Dynamic energy-accuracy trade-off using stochastic computing in deep neural networksabstractThis paper presents an efficient DNN design with stochastic computing. Observing that directly adopting stochastic computing to DNN has some challenges including random error fluctuation, range limitation, and overhead in accumulation, we address these problems by removing near-zero weights, applying weight-scaling, and integrating the activation function with the accumulator. The approach allows an easy implementation of early decision termination with a fixed hardware design by exploiting the progressive precision characteristics of stochastic computing, which was not easy with existing approaches. Experimental results show that our approach outperforms the conventional binary logic in terms of gate area, latency, and power consumption. Kyounghoon Kim, Jungki Kim, Joonsang Yu, Jungwoo Seo, Jongeun Lee, Kiyoung Choi |
DAC | 6 |
| 2016 | Adaptive delay monitoring for wide voltage-range operation
Kiyoung Choi, Wook Kim, Kyung Tae Do, Jung Yun Choi |
DATE | 3 |
| 2016 | Buffered compares: Excavating the hidden parallelism inside DRAM architectures with lightweight logic
Jinho Lee 0001, Jung Ho Ahn, Kiyoung Choi |
DATE | 3 |
| 2016 | Efficient FPGA acceleration of Convolutional Neural Networks using logical-3D compute array
Atul Rahman, Jongeun Lee, Kiyoung Choi |
DATE | 3 |
| 2016 | Dynamic clock synchronization scheme between voltage domains in multi-core architectureabstractUsing independent voltage (and frequency) domains for cores and caches allows us to achieve high energy efficiency since it enables operating the cores and caches at their own optimal voltages. However, it incurs a clock synchronization problem between the core and cache voltage domains. One of the conventional solutions is to add asynchronous FIFOs on the domain crossing boundary, but it degrades performance due to the increased latency. This paper presents a dynamic clock synchronization scheme between two different voltage domains. It uses a fast clock phase detector to monitor the clock phase difference between two voltage domains, a variable delay element consisting of multiple-stage thyristor-like delay circuits to support a wide delay range, and a small hardware module that uses the phase detector output to control the variable delay element. In this way, the scheme automatically adjusts the delay of a clock until the phase is aligned with that of the other clock. Also presented is an algorithm that speeds up the synchronization process. Experimental results show 1.4× speedup and 13% energy saving on average compared to the conventional approach. Kiyoung Choi, Sangheon Lee 0006, Soojung Ryu |
VLSI-SoC | 2 |
| 2016 | Exploration of trade-offs in the design of volatile STT-RAM cache
Namhyung Kim, Kiyoung Choi |
J. Syst. Archit. | 2 |
| 2016 | A design framework for hierarchical ensemble of multiple feature extractors and multiple classifiers
Kyounghoon Kim, Helin Lin, Kiyoung Choi |
Pattern Recognit. | 4 |
| 2016 | AIM: Energy-Efficient Aggregation Inside the Memory HierarchyabstractIn this article, we propose Aggregation-in-Memory (AIM), a new processing-in-memory system designed for energy efficiency and near-term adoption. In order to efficiently perform aggregation, we implement simple aggregation operations in main memory and develop a locality-adaptive host architecture for in-memory aggregation, called cache-conscious aggregation. Through this, AIM executes aggregation at the most energy-efficient location among all levels of the memory hierarchy. Moreover, AIM minimally changes existing sequential programming models and provides fully automated compiler toolchain, thereby allowing unmodified legacy software to use AIM. Evaluations show that AIM greatly improves the energy efficiency of main memory and the system performance. Junwhan Ahn, Sungjoo Yoo, Kiyoung Choi |
ACM Trans. Archit. Code Optim. | 3 |
| 2016 | Prediction Hybrid Cache: An Energy-Efficient STT-RAM Cache ArchitectureabstractSpin-transfer torque RAM (STT-RAM) has emerged as an energy-efficient and high-density alternative to SRAM for large on-chip caches. However, its high write energy has been considered as a serious drawback. Hybrid caches mitigate this problem by incorporating a small SRAM cache for write-intensive data along with an STT-RAM cache. In such architectures, choosing cache blocks to be placed into the SRAM cache is the key to their energy efficiency. This paper proposes a new hybrid cache architecture called prediction hybrid cache. The key idea is to predict write intensity of cache blocks at the time of cache misses and determine block placement based on the prediction. We design a write intensity predictor that realize the idea by exploiting a correlation between write intensity of blocks and memory access instructions that incur cache misses of those blocks. It includes a mechanism to dynamically adapt the predictor to application characteristics. We also design a hybrid cache architecture in which write-intensive blocks identified by the predictor are placed into the SRAM region. Evaluations show that our scheme reduces energy consumption of hybrid caches by 28 percent (31 percent) on average compared to the existing hybrid cache architecture in a single-core (quad-core) system. Junwhan Ahn, Sungjoo Yoo, Kiyoung Choi |
IEEE Trans. Computers | 3 |
| 2016 | Low-Power Hybrid Memory Cubes With Link Power Management and Two-Level PrefetchingabstractThe hybrid memory cube (HMC) is a 3-D-stacked DRAM architecture designed for substantially improved memory bandwidth. In particular, its I/O interface achieves up to 320 GB/s of external bandwidth through high-speed serial links. However, it comes at the cost of large static power of off-chip links, which dominates total power consumption of HMCs. In this paper, we propose an adaptive mechanism to partially disable off-chip links of HMCs to reduce the energy consumption of the off-chip links. In order to determine the number of the links to be disabled upon application loads, we develop a simple hardware module called link delay monitor to simulate all different link configurations at the same time and find the largest number of the links to be disabled while satisfying the given performance constraint. We also present two-level prefetching with in-HMC prefetch buffers to further improve the efficiency of our link power management scheme in the presence of prefetching. Evaluations show that our scheme reduces the energy consumption of HMCs by 52% on average with a negligible performance degradation. Junwhan Ahn, Sungjoo Yoo, Kiyoung Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | THOR: Orchestrated thermal management of cores and networks in 3D many-core architecturesabstractMost previous researches on thermal management of many-core architectures focus on the control of either core resources or network resources only, even though both have significant thermal impacts. This paper proposes a holistic thermal management that applies dynamic voltage/frequency scaling to cores and routers together to maximize system performance under temperature constraint. The proposed method first determines a power budget given in aggregate weighted power for every pillar of vertically adjacent tiles. Then it performs voltage/frequency assignment under the budget while exploiting the characteristics of the applications. Experiments show that our approach outperforms existing methods. Jinho Lee 0001, Junwhan Ahn, Kiyoung Choi, Kyungsu Kang |
ASP-DAC | 3 |
| 2015 | A scalable processing-in-memory accelerator for parallel graph processingabstractThe explosion of digital data and the ever-growing need for fast data analysis have made in-memory big-data processing in computer systems increasingly important. In particular, large-scale graph processing is gaining attention due to its broad applicability from social science to machine learning. However, scalable hardware design that can efficiently process large graphs in main memory is still an open problem. Ideally, cost-effective and scalable graph processing systems can be realized by building a system whose performance increases proportionally with the sizes of graphs that can be stored in the system, which is extremely challenging in conventional systems due to severe memory bandwidth limitations. Junwhan Ahn, Sungpack Hong, Sungjoo Yoo, Onur Mutlu, Kiyoung Choi |
ISCA | 5 |
| 2015 | PIM-enabled instructions: a low-overhead, locality-aware processing-in-memory architectureabstractProcessing-in-memory (PIM) is rapidly rising as a viable solution for the memory wall crisis, rebounding from its unsuccessful attempts in 1990s due to practicality concerns, which are alleviated with recent advances in 3D stacking technologies. However, it is still challenging to integrate the PIM architectures with existing systems in a seamless manner due to two common characteristics: unconventional programming models for in-memory computation units and lack of ability to utilize large on-chip caches. Junwhan Ahn, Sungjoo Yoo, Onur Mutlu, Kiyoung Choi |
ISCA | 4 |
| 2015 | Message from the general chairsabstractOn behalf of the Organizing Committee, we welcome you to the 23rd IFIP/IEEE International Conference on Very Large Scale Integration (VLSI-SoC) in Daejeon, the City of Science and Technology of Korea. Naehyuck Chang, Kiyoung Choi |
VLSI-SoC | 2 |
| 2015 | Energy-efficient exclusive last-level hybrid caches consisting of SRAM and STT-RAMabstractThis paper presents an energy-efficient exclusive last-level cache design based on STT-RAM, which is an emerging memory technology that has higher density and lower static power compared to SRAM. Exclusive caches are known to provide higher effective cache capacity than inclusive caches by removing duplicated copies of cache blocks across hierarchies. However, in exclusive cache hierarchies, every block evicted from the lower-level cache is written back to the last-level cache regardless of its dirtiness thereby incurring extra write overhead. This makes it challenging to use STT-RAM for exclusive last-level caches due to its high write energy and long write latency. To mitigate this problem, we design an SRAM/STT-RAM hybrid cache architecture based on reuse distance prediction. In our architecture, for the cache blocks evicted from the lower-level cache, blocks that are likely to be accessed again soon are inserted into the SRAM region and blocks that are unlikely to be reused are forced to bypass the last-level cache. Evaluation results show that the proposed architecture significantly reduces energy consumption of the last-level cache while slightly improving the system performance. Namhyung Kim, Junwhan Ahn, Woong Seo, Kiyoung Choi |
VLSI-SoC | 4 |
| 2015 | Dynamic error tracking and supply voltage adjustment for low powerabstractAmongst many techniques to reduce power consumption of chips, lowering the supply voltage is known to be the most effective one. However, lowering the supply voltage of chips too much down to near the threshold voltage of transistors causes the logic delay to vary exponentially with intrinsic and extrinsic variations such as process variations, temperature variations, and aging, and thus forces the designer to set increased timing margin or use more advanced techniques such as adaptive voltage scaling, where the supply voltage is adjusted by tracking timing errors due to the variations. This paper proposes a technique for adaptive supply voltage adjustment that minimizes power consumption in the near-threshold voltage region while satisfying a given constraint on error rate. It can be used in signal processing applications where intermittent errors are tolerated. The technique employs a current sensing completion detector for tracking errors and increases/decreases the voltage if the error rate is too high/low. We show that, for the average case, our approach tracks errors better than other approaches. We also show that it achieves 42% and 54% power savings for an error rate of 0.1% and 5%, respectively, for the TT corner at 25°C. Pierre Nicolas-Nicolaz, Kiyoung Choi |
VLSI-SoC | 2 |
| 2015 | REDELF: An Energy-Efficient Deadlock-Free Routing for 3D NoCs with Partial Vertical Connectionsabstract3D integrated circuits (3D ICs) using through-silicon vias (TSVs) allow to envision the stacking of dies with different functions and technologies, using as an interconnect backbone a 3D network-on-chip (NoC). However, partial vertical connection in 3D NoCs seems unavoidable because of the large overhead of TSV itself (e.g., large footprint, low fabrication yield, additional fabrication processes) as well as the heterogeneity in dimension. This article proposes an energy-efficient deadlock-free routing algorithm for 3D mesh topologies where vertical connections partially exist. By introducing some rules for selecting elevators (i.e., vertical links between dies), the routing algorithm can eliminate the dedicated virtual channel requirement. In this article, the rules themselves as well as the proof of deadlock freedom are given. By eliminating the virtual channels for deadlock avoidance, the proposed routing algorithm reduces the energy consumption by 38.9% compared to a conventional routing algorithm. When the virtual channel is used for reducing the head-of-line blocking, the proposed routing algorithm increases performance by up to 23.1% and 6.9% on average. Jinho Lee 0001, Kyungsu Kang, Kiyoung Choi |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2014 | Leveraging parallelism in the presence of control flow on CGRAsabstractCoarse-Grained Reconfigurable Architectures (CGRAs) are suitable for accelerating data-intensive applications in embedded systems due to high performance and power efficiency. However, as application programs become complex having more control flows in them, it becomes harder to accelerate such programs on CGRAs. Previous researches on this issue have focused on correct execution of control flows rather than their acceleration. This paper reveals how control flows degrade the performance of programs and proposes a software approaches to accelerating control flows by exploiting parallelism residing in each conditionals as well as among conditionals. Experiments show that our proposed techniques improve performance by 2.51 times on average. Jihyun Ryoo, Kyuseung Han, Kiyoung Choi |
ASP-DAC | 3 |
| 2014 | Dynamic Power Management of Off-Chip Links for Hybrid Memory CubesabstractThe Hybrid Memory Cube (HMC) is a 3D-stacked DRAM architecture designed for substantially improved memory bandwidth. In particular, its I/O interface achieves up to 320 GB/s of external bandwidth through high-speed serial links. However, it comes at a cost of large static power of off-chip links, which dominates total power consumption of HMCs. Therefore, we propose an adaptive mechanism to partially disable off-chip links of HMCs with a minimal performance impact. We also present two-level prefetching with in-HMC prefetch buffers to further improve its efficiency in the presence of prefetching. Evaluations show that our scheme reduces energy consumption of HMCs by 51% on average. Junwhan Ahn, Sungjoo Yoo, Kiyoung Choi |
DAC | 3 |
| 2014 | DASCA: Dead Write Prediction Assisted STT-RAM Cache ArchitectureabstractSpin-Transfer Torque RAM (STT-RAM) has been considered as a promising candidate for on-chip last-level caches, replacing SRAM for better energy efficiency, smaller die footprint, and scalability. However, it also introduces several new challenges into last-level cache design that need to be overcome for feasible deployment of STT-RAM caches. Among other things, mitigating the impact of slow and energy-hungry write operations is of the utmost importance. In this paper, we propose a new mechanism to reduce write activities of STT-RAM last-level caches. The key observation is that a significant amount of data written to last-level caches is not actually re-referenced again during the lifetime of the corresponding cache blocks. Such write operations, which we call dead writes, can bypass the cache without incurring extra misses by definition. Based on this, we propose Dead Write Prediction Assisted STT-RAM Cache Architecture (DASCA), which predicts and bypasses dead writes for write energy reduction. For this purpose, we first propose a novel classification of dead writes, which is composed of dead-on-arrival fills, dead-value fills, and closing writes, as a theoretical model for redundant write elimination. On top of that, we present a dead write predictor based on a state-of-the-art dead block predictor. Evaluations show that our architecture achieves an energy reduction of 68% (62%) in last-level caches and an additional energy reduction of 10% (16%) in main memory and even improves system performance by 6% (14%) on average compared to the STT-RAM baseline in a single-core (quad-core) system. Junwhan Ahn, Sungjoo Yoo, Kiyoung Choi |
HPCA | 3 |
| 2014 | Concept-aware ensemble system for pedestrian detectionabstractFor pedestrian detection in ADAS, using multiple classifiers generally performs better than using a single classifier in terms of accuracy since the classifiers can be made to complement one another. On the other hand, such a pedestrian detector needs to be tuned dynamically to the variation of real-world environment such as different poses of pedestrians and variable background. Thus the system is requested to incrementally accept new information while retaining the old one. This paper presents an environment-adaptive ensemble system that performs incremental learning for pedestrian detection. It combines a pedestrian detector comprised of multiple classifiers with a front-end concept recognizer that selectively turns on and off the member classifiers adaptively according to the recognized concept of the input image. It adopts an incremental learning algorithm to add a new classifier, which is trained with a newly added batch of dataset, to the existing ensemble. With the intervention of the front-end concept recognizer, the system can retain good accuracy for old environments while not losing the focus on current environment. Helin Lin, Kyounghoon Kim, Kiyoung Choi |
Intelligent Vehicles Symposium | 3 |
| 2014 | Energy-efficient partitioning of hybrid caches in multi-core architectureabstractThis paper proposes a technique for reducing energy consumed by hybrid caches that have both SRAM and STT-RAM (Spin-Transfer Torque RAM) in multi-core architecture. It is based on dynamic partitioning of the SRAM cache as well as the STT-RAM cache. It assigns cache blocks to a specific region of a cache based on an existing technique called read-write aware region-based hybrid cache architecture. Thus, when a store operation from a core causes a write miss, the block is assigned to the SRAM cache. When a load operation from a core causes a read miss and thus causes a block fill, the block is assigned to the STT-RAM cache. However, if the core is already using maximum cache ways allocated to it in the SRAM, then the block fill is done into the SRAM. The partitioning is updated periodically. Simulation results show that the proposed technique improves the performance of the multi-core architecture and significantly reduces energy consumption in the hybrid caches compared to the state-of-the-art migration-based hybrid cache management. Kiyoung Choi |
VLSI-SoC | 2 |
| 2014 | Design of a coarse-grained reconfigurable architecture with floating-point support and comparative study
Manhwee Jo, Kyuseung Han, Kiyoung Choi |
Integr. | 4 |
| 2014 | Software-Level Approaches for Tolerating Transient Faults in a Coarse-GrainedReconfigurable ArchitectureabstractCoarse-grained reconfigurable architectures have drawn increasing attention due to their merits in performance and flexibility. Typically, they have many processing elements in the form of an array, which is suitable for implementing spatial redundancy used for fault-tolerant systems design. This paper presents a purely software-level approach to implementing transient-fault-tolerance on an existing processing element array without any modification to the architecture. It includes automated design flow to construct a fault-tolerant system and mathematical modeling for analyzing system reliability. Experiments with real-world applications show the effectiveness of the proposed approaches in terms of yield enhancement and system reliability. Kyuseung Han, Ganghee Lee, Kiyoung Choi |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2014 | Critical-path-aware high-level synthesis with distributed controller for fast timing closureabstractCentralized controllers commonly used in high-level synthesis often require long wires and cause high load capacitance, and that is why critical paths typically occur on paths from controllers to data registers instead of paths from data registers to data registers. However, conventional high-level synthesis has focused on delays within a datapath, making it difficult to solve the timing closure problem during physical synthesis. This article presents hardware architecture with a distributed controller, which makes the timing closure problem much easier. A novel critical-path-aware high-level synthesis flow is also presented for synthesizing such hardware through datapath partitioning, register binding, and controller optimization. We explore the design space related to the number of partitions, which is an important design parameter for target architecture. According to our experiments, the proposed approach reduces the critical path delay excluding FUs by 29.3% and that including FUs by 10.0%, with 2.2% area overhead on average compared to centralized controller architecture. Seokhyun Lee, Kiyoung Choi |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2014 | Configurable range memory for effective data reuse on programmable acceleratorsabstractWhile programmable accelerators such as application-specific processors and reconfigurable architectures can dramatically speed up compute-intensive kernels of an application, application performance can still be severely limited by the communication between processors. To minimize the communication overhead, a shared memory such as a scratchpad memory may be employed between the main processor and the accelerator coprocessor. However, this setup poses a significant challenge to the main processor, which now must manage data on the scratchpad explicitly, resulting in superfluous data copying due to the inflexibility of a scratchpad. In this article, we present an enhancement of a scratchpad, Configurable Range Memory (CRM), whose address range can be reprogrammed to minimize unnecessary data copying between processors and therefore promote data reuse on the accelerator, and also present a software management algorithm for the CRM. Our experimental results involving detailed simulation of full multimedia applications demonstrate that our CRM architecture can reduce the communication overhead quite effectively, reducing the kernel execution time by up to 28% and the application runtime by up to 12.8%, in addition to considerable system energy reduction, compared to the conventional architecture based on a scratchpad. Jongeun Lee, Seongseok Seo, Jong Kyung Paek, Kiyoung Choi |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2014 | Introduction to the Special Issue on the 11th International Conference on Field-Programmable Technology (FPT'12)abstractNo abstract available. Jason Helge Anderson, Kiyoung Choi |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2014 | LASIC: Loop-Aware Sleepy Instruction Caches Based on STT-RAM TechnologyabstractThis brief presents an approach to reduce static power consumption in peripheral circuits of spin-transfer torque RAM (STT-RAM) instruction caches. It is based on the key observation that only a small set of instructions is accessed inside a program loop. We propose to add a small static RAM cache called loop cache between the processor and the L1 instruction cache made of STT-RAM. When the loop cache has an entire loop cached, the L1 instruction cache can be turned off to save energy during the execution of the loop. Experimental results show that the proposed approach achieves 49% reduction in energy consumption over the STT-RAM baseline. Junwhan Ahn, Kiyoung Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Selectively protecting error-correcting code for area-efficient and reliable STT-RAM cachesabstractRecent researches on STT-RAM revealed that device scaling makes its write operations unreliable. To mitigate the impact of this problem, this paper proposes a low-cost, ECC-based solution for STT-RAM caches. In particular, it proposes to share storage for ECC among different blocks within a set and to use them only for unsuccessful write operations. Experimental results show that our scheme reduces 74% to 98% of area overhead incurred by the conventional per-block ECC while maintaining system performance and reliability. Junwhan Ahn, Sungjoo Yoo, Kiyoung Choi |
ASP-DAC | 3 |
| 2013 | Deflection routing in 3D Network-on-Chip with TSV serializationabstractThis paper proposes a deflection routing for 3D NoC with serialized TSVs. Bufferless deflection routing provides area- and power-efficient communication under low to medium traffic load. Under 3D circumstances, the bufferless deflection routing can yield even better performance than buffered routing when key aspects are properly taken into account. Evaluation of the proposed scheme shows its effectiveness in throughput, latency, and energy consumption. Jinho Lee 0001, Sunwook Kim, Kiyoung Choi |
ASP-DAC | 4 |
| 2013 | Compiling control-intensive loops for CGRAs with state-based full predicationabstractPredication is an essential technique to accelerate kernels with control flow on CGRAs. While state-based full predication (SFP) can remove wasteful power consumption on issuing/decoding instructions from conventional full predication, generating code for SFP is challenging for general CGRAs, especially when there are multiple conditionals to be handled due to exploiting data level parallelism. In this paper, we present a novel compiler framework addressing central issues such as how to express the parallelism between multiple conditionals, and how to allocate resources to them to maximize the parallelism. In particular, by separating the handling of control flow and data flow, our framework can be integrated with conventional mapping algorithms for mapping data flow. Experimental results demonstrate that our framework can find and exploit parallelism between multiple conditionals, thereby leading to 2.21 times higher performance on average than a naive approach. Kyuseung Han, Kiyoung Choi, Jongeun Lee |
DATE | 2 |
| 2013 | Write intensity prediction for energy-efficient non-volatile cachesabstractThis paper presents a novel concept called write intensity prediction for energy-efficient non-volatile caches as well as the architecture that implements the concept. The key idea is to correlate write intensity of cache blocks with addresses of memory access instructions that incur cache misses of those blocks. The predictor keeps track of instructions that tend to load write-intensive blocks and utilizes that information to predict write intensity of blocks. Based on this concept, we propose a block placement strategy driven by write intensity prediction for SRAM/STT-RAM hybrid caches. Experimental results show that the proposed approach reduces write energy consumption by 55% on average compared to the existing hybrid cache architecture. Junwhan Ahn, Sungjoo Yoo, Kiyoung Choi |
ISLPED | 3 |
| 2013 | A deadlock-free routing algorithm requiring no virtual channel on 3D-NoCs with partial vertical connectionsabstractElevator-first routing algorithm has been introduced for partially connected 3D network-on-chips, as a low-cost, distributed and deadlock-free routing algorithm using two virtual channels. This paper proposes Redelf, a modification of the elevator-first routing algorithm on a 3D mesh topology. The proposed algorithm requires no virtual channel to ensure deadlock-freedom. Jinho Lee 0001, Kiyoung Choi |
NOCS | 2 |
| 2013 | CPU-based speed acceleration techniques for shear warp volume rendering
Kiyoung Choi, Sung-Up Jo, Hwa-Min Lee, Changsung Jeong |
Multim. Tools Appl. | 1 |
| 2013 | Power-Efficient Predication Techniques for Acceleration of Control Flow Execution on CGRAabstractCoarse-grained reconfigurable architecture typically has an array of processing elements which are controlled by a centralized unit. This makes it difficult to execute programs having control divergence among PEs without predication. However, conventional predication techniques have a negative impact on both performance and power consumption due to longer instruction words and unnecessary instruction-fetching decoding nullifying steps. This article reveals performance and power issues in predicated execution which have not been well-addressed yet. Furthermore, it proposes fast and power-efficient predication mechanisms. Experiments conducted through gate-level simulation show that our mechanism improves energy-delay product by 11.9% to 23.8% on average. Kyuseung Han, Junwhan Ahn, Kiyoung Choi |
ACM Trans. Archit. Code Optim. | 3 |
| 2013 | Isomorphism-Aware Identification of Custom Instructions With I/O SerializationabstractExtensible processors have been widely used to achieve the conflicting demands for performance improvement, low power consumption, and flexibility. As extensible processors have become more popular, several algorithms have been proposed for automatically identifying instruction-set extensions in order to reduce the effort of manual design and verification. However, most of them focus on finding large and complex instructions that are used only once, rather than repeatedly used ones. Moreover, some other approaches that consider recurrence are limited to finding small instructions. This paper proposes a novel algorithm that considers the instruction reusability as well as input/output (I/O) serialization. In order to overcome the high complexity of the problem, we develop a canonical-form construction algorithm for fast isomorphism detection on directed acyclic graphs and an incremental template generation algorithm that identifies the best custom instruction in terms of a user-defined fitness function. Moreover, our algorithm serializes I/O operations so that the numbers of inputs and outputs of custom instructions are not limited by the microarchitecture. This paper also proposes an algorithm for multiple custom instructions utilizing a well-known iterative selection algorithm. Last, it presents a hybrid algorithm composed of our algorithm and the previous algorithm that does not consider reusability. Experimental results show that our isomorphism-aware algorithm achieves significant improvement over previous approaches in terms of algorithm runtime, as well as performance gain obtained by custom instructions. Junwhan Ahn, Kiyoung Choi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2013 | Mapping and Scheduling of Tasks and Communications on Many-Core SoC Under Local Memory ConstraintabstractThere has been extensive research on mapping and scheduling tasks on a many-core SoC. However, none considers the optimization of communication types, which can significantly affect performance, energy consumption, and local memory usage of the SoC. This paper presents an approach to automatic mapping and scheduling of tasks and communications on a many-core SoC. The key idea is to decide the type of each communication between message passing and shared memory when we do the mapping and scheduling. By assigning a proper type to each communication, we can optimize the energy consumption, performance, or energy-delay product. To solve the optimization problem, the approach adopts a probabilistic algorithm coupled with some heuristics. To enhance throughput of the system, it performs software pipelined scheduling of the tasks using a modified iterative modulo scheduling technique. Experiments show that our algorithm achieves on average 50.1% lower energy consumption, 21.0% higher throughput, and 64.9% lower energy- delay product, compared to shared memory only communication. Jinho Lee 0001, Moo-Kyoung Chung, Yeongon Cho, Soojung Ryu, Jung Ho Ahn, Kiyoung Choi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2013 | Deflection routing in 3D network-on-chip with limited vertical bandwidthabstractThis article proposes a deflection routing for 3D NoC with serialized TSVs for vertical links. Compared to buffered routing, deflection routing provides area- and power-efficient communication and little loss of performance under low to medium traffic load. Under 3D environments, the deflection routing can yield even better performance than buffered routing when key aspects are properly taken into account. However, the existing deflection routing technique cannot be directly applied because the serialized TSV links will take longer time to send data than ordinary planar links and cause many problems. A naive deflection through a TSV link can cause significantly longer latency and more energy consumption even for communications through planar links. This article proposes a method to mitigate the effect and also solve arising deadlock and livelock problems. Evaluation of the proposed scheme shows its effectiveness in throughput, latency, and energy consumption. Jinho Lee 0001, Sunwook Kim, Kiyoung Choi |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2012 | Memory-aware mapping and scheduling of tasks and communications on many-core SoCabstractThis paper presents an approach to automatic task mapping, scheduling, and communication routing on a many-core SoC, considering the trade-offs between two different communication types—message passing and shared memory—for the communication routing in order to optimize the energy consumption or performance. To solve the optimization problem, the approach uses the quantum-inspired evolutionary algorithm. For the scheduling of the tasks with backward dependencies, it uses the iterative modulo scheduling technique. Experiments with random task graphs as well as real applications show the effectiveness of the proposed approach. Jinho Lee 0001, Kiyoung Choi |
ASP-DAC | 2 |
| 2012 | State-based full predication for low power coarse-grained reconfigurable architectureabstractIt has been one of the most fundamental challenges in architecture design to achieve high performance with low power while maintaining flexibility. Parallel architectures such as coarse-grained reconfigurable architecture, where multiple PEs are tightly coupled with each other, can be a viable solution to the problem. However, the PEs are typically controlled by a centralized control unit, which makes it hard to parallelize programs requiring different control of each PE. To overcome this limitation, it is essential to convert control flows into data flows by adopting the predicated execution technique, but it may incur additional power consumption. This paper reveals power issues in the predicated execution and proposes a novel technique to mitigate power overhead of predicated execution. Contrary to the conventional approach, the proposed mechanism can decide whether to suppress instruction execution or not without decoding the instructions and does not require additional instruction bits, thereby resulting in energy savings. Experimental results show that energy consumed by the reconfigurable array and its configuration memory is reduced by up to 23.9%. Kyuseung Han, Kiyoung Choi |
DATE | 3 |
| 2012 | Lower-bits cache for low power STT-RAM cachesabstractAs power-efficient design becomes more important, spin-transfer torque RAM (STT-RAM) has drawn a lot of attention due to its ability to meet both high performance and low power consumption. However, its high write energy incurs an increase of dynamic power consumption and may offset power saving due to its low static power. This paper proposes a novel technique called lower-bits caches for reducing write activities of STT-RAM L2 caches. Based on the observation that upper bits of data are not changed as frequently as lower bits in most applications, the technique tries to hide frequent bit changes in lower bits from the L2 cache. Experimental results show that our architecture reduced 25 percent of energy consumed by the L2 cache and slightly improved performance at the same time compared to the STT-RAM baseline. Junwhan Ahn, Kiyoung Choi |
ISCAS | 2 |
| 2012 | An adaptive routing algorithm for 3D mesh NoC with limited vertical bandwidth
Mingyang Zhu, Jinho Lee 0001, Kiyoung Choi |
VLSI-SoC | 3 |
| 2012 | Active Memory Processor for Network-on-Chip-Based ArchitectureabstractMemory-intensive operations and their memory access latency are often the performance bottleneck in parallel applications. In this paper, we investigate the concept of active memory operation which is an active data processing operation performed on the memory side. Utilizing the active memory operation, we can replace multiple transactions of memory accesses over the on-chip network and related computations on the processor side with a smaller number of high-level transactions and computations on the memory side. To realize the concept, we have designed a special-purpose processor called active memory processor which is tightly coupled with the memory and executes the active memory operations. In our case studies, we have applied the concept to five real-world applications (parallelized JPEG, FFT, text indexing for data mining, histogram, and eikonal equation solver) running on a 36--tile architecture with 64 cores and four memory tiles and found that the proposed approach can improve performance by 20.5 \sim 259.3 percent. Jun-hee Yoo, Sungjoo Yoo, Kiyoung Choi |
IEEE Trans. Computers | 3 |
| 2011 | A polynomial-time custom instruction identification algorithm based on dynamic programmingabstractThis paper introduces an innovative algorithm for automatic instruction-set extension, which gives a pseudo-optimal solution within polynomial time to the size of a graph. The algorithm uses top-down dynamic programming strategy with the branch-and-bound algorithm in order to exploit overlapping of subproblems. Correctness of the algorithm is formally proved, and time complexity is analyzed from it. Also, it is verified that the algorithm gives an optimal solution for some type of merit functions, and has very small possibility of obtaining non-optimal solution in general. Furthermore, several experimental results are presented as evidence of the fact that the proposed algorithm has notable performance improvement. Junwhan Ahn, Imyong Lee, Kiyoung Choi |
ASP-DAC | 3 |
| 2011 | High-level synthesis with distributed controller for fast timing closureabstractCentralized controllers commonly used in high-level synthesis often cause long wires and high load capacitance and that is why critical paths typically occur on paths from controllers to data registers. However, conventional high level synthesis has focused on the delay of datapaths making it difficult to solve the timing closure problem during physical synthesis. This paper presents a hardware architecture with a distributed controller, which makes the timing closure problem much easier. It also presents a novel high-level synthesis flow for synthesizing such hardware through datapath partitioning and controller optimization. According to our experimental results, the proposed approach reduces the controller and interconnect delay by 20.3-27.4% and the entire critical path delay by 6.6~10.3% with 0.2~13.3% area overhead. Even without area overhead, it reduces the critical path delay by 5.8~10%. Seokhyun Lee, Kiyoung Choi |
ICCAD | 2 |
| 2011 | Mapping Multi-Domain Applications Onto Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-grained reconfigurable architectures (CGRAs) have drawn increasing attention due to their performance and flexibility. However, their applications have been restricted to domains based on integer arithmetic since typical CGRAs support only integer arithmetic or logical operations. This paper introduces approaches to mapping applications onto CGRAs supporting both integer and floating-point arithmetic. After presenting an optimal formulation using integer linear programming, we present a fast heuristic mapping algorithm. Our experiments on randomly generated examples generate optimal mapping results using our heuristic algorithm for 97% of the examples within a few seconds. We observe similar results for practical examples from multimedia and 3-D graphics benchmarks. The applications mapped on a CGRA show up to 120 times performance improvement compared to software implementations, demonstrating the potential for application acceleration on CGRAs supporting floating-point operations. Ganghee Lee, Kiyoung Choi, Nikil Dutt |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2010 | Acceleration of control flow on CGRA using advanced predicated executionabstractCoarse-grained reconfigurable array is a very attractive architecture from the viewpoint of performance and flexibility. However, because the performance improvement is achieved by exploiting parallelism, the architecture is typically poor at handling control flow, which is sequential in nature. There have been many attempts to overcome this problem by using predicated execution techniques; however, they do not support all types of control flow or suffer from performance degradation in doing so. In addition, predicated execution schemes in general require a longer execution time because both the if- and else-paths are always executed. This paper proposes advanced predicated execution techniques that can handle and accelerate all types of control flow with only 2% hardware overhead. These techniques can also be easily extended to general SIMD machines. We implemented these techniques on a coarse-grained reconfigurable array architecture and verified its functionality and effectiveness by accelerating an H.264 deblocking filter, a kernel which is both data- and control-intensive. The results show that the proposed approach achieves up to 43% improvement in execution time compared to speculation by sacrificing 76% code size, and 24% improvement in execution time compared to the previous full predication approach, with a smaller code size. Kyuseung Han, Jong Kyung Paek, Kiyoung Choi |
FPT | 3 |
| 2010 | Design Space Exploration for Efficient Resource Utilization in Coarse-Grained Reconfigurable ArchitectureabstractCoarse-grained reconfigurable architectures (CGRAs) aim to achieve both goals of high performance and flexibility. In addition, power consumption is significant for the reconfigurable architecture to be used as a competitive processing core in embedded systems. However, the existing reconfigurable architectures require too much area and power. In this paper, we propose a new design space exploration flow, optimizing CGRA to reduce area and power with enhancing performance for digital signal processing (DSP) application domain. It reduces the array size through efficient arrangement of array components and customization of their interconnection, exploiting input patterns belonging to the DSP application domain. Such a design flow is based on pipelining and sharing of area/delay-critical resources in the processing element array. Experimental results show that for DSP applications, the proposed approach reduces area by up to 36.75%, average execution time by 36.78%, and average power by 31.85% when compared with the existing CGRA architecture. Yoonjin Kim, Rabi N. Mahapatra, Kiyoung Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2009 | Code decomposition and recomposition for enhancing embedded software performanceabstractMultitasking of concurrent processes implements the concurrency inherited from applications, increasing the utilization of limited resources. It requires an operating system and imposes significant runtime overhead. Serializing multitasking codes removes the need of operating system and the overhead as well. In this paper, we propose a software synthesis method to transform multitasking codes into a single process code. For this, we decompose multitasking codes into a set of code fractions and then recompose the code fractions into a single process code, preserving the functionality of the original codes. We present two different techniques for the transformation - code partitioning and code covering - and propose a hybrid technique that combines the two techniques. Youngchul Cho, Kiyoung Choi |
ASP-DAC | 2 |
| 2009 | Multiprocessor System-on-Chip designs with active memory processors for higher memory efficiencyabstractMemory access latency and memory-related operations are often the performance bottleneck in parallel applications. In this paper, we present a concept of active memory operations which is an on-chip network transaction that operates based on the microcode provided by the software designer. Utilizing the active memory operation, we can replace multiple transactions of memory accesses over the on-chip network and related local processing element computation with a smaller number of high-level transactions and near-memory computation. We implemented a processor called active memory processor which is located near the memory and executes the active memory operations. In our case studies, we applied the concept to three real-world applications (parallelized JPEG, FFT, and text indexing for data mining) running on a 36-tile architecture with 32 cores and 4 memories and found that the programmable transaction approach can improve performance by 34.3% to 618% at the cost of additional design effort. Jun-hee Yoo, Sungjoo Yoo, Kiyoung Choi |
DAC | 3 |
| 2009 | Low Power Reconfiguration Technique for Coarse-Grained Reconfigurable ArchitectureabstractCoarse-grained reconfigurable architectures (CGRAs) require many processing elements (PEs) and a configuration memory unit (configuration cache) for reconfiguration of its PE array. Although this structure is meant for high performance and flexibility, it consumes significant power. Specially, power consumption by configuration cache is explicit overhead compared to other types of intellectual property (IP) cores. Reducing power is very crucial for CGRA to be more competitive and reliable processing core in embedded systems. In this paper, we propose a reusable context pipelining (RCP) architecture to reduce power-overhead caused by reconfiguration. It shows that the power reduction can be achieved by using the characteristics of loop pipelining, which is a multiple instruction stream, multiple data stream (MIMD)-style execution model. RCP efficiently reduces power consumption in configuration cache without performance degradation. Experimental results show that the proposed approach saves much power even with reduced configuration cache size. Power reduction ratio in the configuration cache and the entire architecture are up to 86.33% and 37.19%, respectively, compared to the base architecture. Yoonjin Kim, Rabi N. Mahapatra, Ilhyun Park, Kiyoung Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2009 | Topology/Floorplan/Pipeline Co-Design of Cascaded Crossbar BusabstractOn-chip bus design has a significant impact on the die area, power consumption, performance and design cycle of complex system-on-chips (SoCs). Especially, for high frequency systems having on-chip buses pipelined extensively to cope with long wire delay, a naive bus design may yield a significant area/power cost mostly due to bus pipeline cost. The topology, floorplan, and pipeline are the most important design factors that affect the cost and frequency of the on-chip bus. Since they are strongly correlated with each other, it is imperative to codesign all of the three. In this paper, we present an automated codesign method for cascaded crossbar bus design. We present CADBUS (CAscadeD crossbar BUS design tool), an automated tool for AXI-based cascaded crossbar bus architecture design. The primary objective of this study is to design a cascaded crossbar bus, including the topology/floorplan/bus pipelines, having minimum area/power cost while satisfying the given constraints of communication bandwidth/latency or frequency. Experimental results of the three industrial strength SoCs show that, compared to the existing approach, the proposed method gives as much as 11.6%-34.2% (9.9%-33.5%) savings in bus area (power consumption). Jun-hee Yoo, Sungjoo Yoo, Kiyoung Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2008 | Entry control in network-on-chip for memory power reductionabstractAs high-end mobile embedded systems become data-intensive, the off-chip memory is becoming a major contributor to the total energy consumption. Especially, high-end mobile chips accommodate dedicated hardware blocks, e.g., codec and 3D graphics IP's, required for both performance and power consumption reasons. Those IP's usually do not have a large shared memory on chip. Thus, they communicate with each other via the off-chip DDR memory increasing off-chip memory accesses, which increases memory energy consumption during read/write operations. In this paper, we present a method of reducing memory energy consumption during read/write operations. It aims at minimizing the number of row opens and closes, which are the major source of energy consumption during read/write operations. The basic idea is to apply network entry control to prioritize consecutive open row memory accesses. The experimental results show up to 35% reduction in memory energy consumption with an industrial strength multimedia mobile SoC. Sungjoo Yoo, Kiyoung Choi |
ISLPED | 3 |
| 2008 | Introduction to embedded systems week 2006 special issueabstractintroduction Share on Introduction to embedded systems week 2006 special issue Editors: Soonhoi Ha Seoul National University Seoul National UniversityView Profile , Kiyoung Choi Seoul National University Seoul National UniversityView Profile , Taewhan Kim Seoul National University Seoul National UniversityView Profile , Krisztian Flautner ARM Ltd. U.K. ARM Ltd. U.K.View Profile , Sanglyul Min Seoul National University Seoul National UniversityView Profile , Wang Yi Uppsala University Uppsala UniversityView Profile Authors Info & Claims ACM Transactions on Embedded Computing SystemsVolume 7Issue 2Article No.: 8pp 1–3https://doi.org/10.1145/1331331.1331332Published:29 January 2008Publication History 0citation445DownloadsMetricsTotal Citations0Total Downloads445Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Soonhoi Ha, Kiyoung Choi, Krisztián Flautner, Sang Lyul Min, Wang Yi 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2008 | SoCDAL: System-on-chip design AcceLeratorabstractTime-to-market pressure and the ever-growing design complexity of multiprocessor system-on-chips have demanded an efficient design environment that enables fast exploration of large design space. In this article, we introduce a new design environment, called SoCDAL, for accelerating multiprocessor system-on-chip design through fast design-space exploration targeting real-time multimedia systems. SoCDAL is a set of mostly automated tools covering system specification, hardware/software estimation, application-to-architecture mapping, simulation model generation, and system verification through simulation. For system specification, the process network model has been widely used for system specification because of its modeling capability. However, it is hard to use for real-time systems design, since its behavior cannot be estimated statically. We introduce a new approach which enables analyzing a process network model statically with some restrictions. For the hardware/software estimation, we analyze codes statically. Application-to-architecture mapping process implements a novel algorithm to support an arbitrary number of processors, with performance evaluation by static scheduling considering communication behavior. Mapping results are used to generate simulation models automatically at several transaction levels to be pipelined to a commercial tool. We show the effectiveness of our approaches by some experimental results with multimedia applications such as JPEG, H.263, and H.264 encoders, as well as an H.264 decoder. Yongjin Ahn, Keesung Han, Ganghee Lee, Hyunjik Song, Jun-hee Yoo, Kiyoung Choi, Xingguang Feng |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2007 | Memory Operation Inclusive Instruction-Set Extensions and Data Path GenerationabstractApplication-specific instruction-set extension of an extensible processor is an effective way of improving the performance of a system. However, identifying an optimum instruction-set extension for a target application is a difficult problem. In this paper, we propose a novel approach to automatic identification of instruction-set extensions by combining branch and bound and high-level synthesis techniques. The approach allows adding instructions containing memory operations, improving the system performance significantly. We also propose an alias aware multi-port data cache architecture, which helps memory operation inclusive instructions get better performance by making them respond to multiple requests at the same cycle. Experimental result shows that the proposed approach improves the performance up to 3 times compared to the previous approaches. Imyong Lee, Kiyoung Choi |
ASAP | 3 |
| 2007 | Communication Architecture Synthesis of Cascaded Bus MatrixabstractFor high frequency on-chip communication architecture design, we propose cascaded bus matrix-based solutions. Due to the huge design space in cascaded bus matrix design, it is crucial to perform an efficient design space exploration. In our work, we present a simulated annealing-based design space exploration method. For an efficient representation of bus topology, we propose an encoding method called traffic group encoding and apply it to AMBA3 AXI-based bus system design. In addition, we propose a method of two-step simulated annealing to improve the quality of results. Experimental results show that the proposed methods allow designing complex communication architectures (ones with up to 31 masters and 71 slaves) with high frequency constraints to which existing methods could not give solutions. Jun-hee Yoo, Sungjoo Yoo, Kiyoung Choi |
ASP-DAC | 4 |
| 2007 | Buffer Size Reduction through Control-Flow DecompositionabstractSoftware synthesis from a data-flow model has been a very promising technique, especially for multimedia applications with contradicting requirements of high design complexity and fast time-to-market. In a dataflow model, buffer size is pessimistically determined through static analysis, thus results in large memory overhead even with optimization techniques such as buffer sharing and scheduling. So, reducing buffer size is one of the key issues of data-flow models. In this work, we propose a novel software synthesis technique to reduce buffer size through control-flow decomposition. We first traverse the control-flow within each actor of a data-flow graph and decompose it into a set of multiple execution paths. Then we transform the actor such that only one of the paths is executed at one invocation of the actor. The new actor may have to be invoked many times to complete the behavior of the original actor. By proper decomposition, we can make the new actor consume/produce much smaller amount of input/output data for each invocation, thereby reducing the input/output buffer size drastically. We automate the process of transformation and show the efficiency of the proposed approach through experiments with image/video multimedia applications. Youngchul Cho, Nacer-Eddine Zergainoh, Ahmed Amine Jerraya, Kiyoung Choi |
RTCSA | 4 |
| 2007 | Instruction set synthesis with efficient instruction encoding for configurable processorsabstractApplication-specific instructions can significantly improve the performance, energy-efficiency, and code size of configurable processors. While generating new instructions from application-specific operation patterns has been a common way to improve the instruction set (IS) of a configurable processor, automating the design of ISs for given applications poses new challenges---how to create as well as utilize new instructions in a systematic manner, and how to choose the best set of application-specific instructions considering the various effects the new instructions may have on the data path and the compilation? To address these problems, we present a novel IS synthesis framework that optimizes the IS through an efficient instruction encoding for the given application as well as for the given data path architecture. We first build a library of new instructions created with various encoding alternatives taking into account the data path architecture constraints, and then select the best set of instructions while satisfying the instruction bitwidth constraint. We formulate the problem using integer linear programming and also present an effective heuristic algorithm. Experimental results using our technique generate ISs that show improvements of up to about 40% over the native IS for several application benchmarks running on typical embedded RISC processors. Jongeun Lee, Kiyoung Choi, Nikil Dutt |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2006 | Worst case execution time analysis for synthesized hardwareabstractWe propose a hardware performance estimation flow for fast design space exploration, based on worst-case execution time analysis algorithms for software analysis. Test cases on some real-world applications show that our flow provides a fight upper bound of the execution time, and many useful hints to the designer. Jun-hee Yoo, Xingguang Feng, Kiyoung Choi, Eui-Young Chung, Kyu-Myung Choi |
ASP-DAC | 3 |
| 2006 | A spatial mapping algorithm for heterogeneous coarse-grained reconfigurable architecturesabstractIn this work, we investigate the problem of automatically mapping applications onto a coarse-grained reconfigurable architecture and propose an efficient algorithm to solve the problem. We formalize the mapping problem and show that it is NP-complete. To solve the problem within a reasonable amount of time, we divide it into three subproblems: covering, partitioning and layout. Our empirical results demonstrate that our technique produces nearly as good performance as hand-optimized outputs for many kernels. Minwook Ahn, Jonghee W. Yoon, Yunheung Paek, Yoonjin Kim, Mary Kiemb, Kiyoung Choi |
DATE | 6 |
| 2006 | Power-conscious configuration cache structure and code mapping for coarse-grained reconfigurable architectureabstractCoarse-grained reconfigurable architecture aims to achieve both performance and flexibility. However, power consumption is no less important for the reconfigurable architecture to be used as a competitive processing core in embedded systems. In this paper, we show how power is consumed in a typical coarse-grained reconfigurable architecture. Based on the power breakdown data, we suggest a power-conscious configuration cache structure and code mapping technique, which reduce power consumption without performance degradation. Experimental results show that the proposed approach saves much power even with reduced configuration cache size. Yoonjin Kim, Ilhyun Park, Kiyoung Choi, Yunheung Paek |
ISLPED | 3 |
| 2005 | Scheduler implementation in MP SoC designabstractIn the design of a heterogeneous multiprocessor system on chip, we face a new design problem; scheduler implementation. In this paper, we present an approach to implementing a static scheduler, which controls all the task executions and communication transactions of a system according to a pre-determined schedule. For the scheduler implementation, we consider both intra-processor and inter-processor synchronization. We also consider scheduler overhead, which is often neglected. In particular, we address the issue of centralized implementation versus distributed implementation. We investigate the pros and cons of the two different scheduler implementations. Through experiments with synthetic examples and a real world multimedia application, we show the effectiveness of our approach. Youngchul Cho, Sungjoo Yoo, Kiyoung Choi, Nacer-Eddine Zergainoh, Ahmed Amine Jerraya |
ASP-DAC | 3 |
| 2005 | Resource Sharing and Pipelining in Coarse-Grained Reconfigurable Architecture for Domain-Specific OptimizationabstractCoarse-grained reconfigurable architectures aim to achieve goals of both high performance and flexibility. However, existing reconfigurable array architectures require many resources without considering the specific application domain. Functional resources that take long latency and/or large area can be pipelined and/or shared among the processing elements. Therefore, the hardware cost and the delay can be effectively reduced without any performance degradation for some application domains. We suggest such a reconfigurable array architecture template and a design space exploration flow for domain-specific optimization. Experimental results show that our approach is much more efficient, in both performance and area, compared to existing reconfigurable architectures. Yoonjin Kim, Mary Kiemb, Chulsoo Park, Jinyong Jung, Kiyoung Choi |
DATE | 5 |
| 2005 | Pipelining with common operands for power-efficient linear systemsabstractWe propose a systematic pipelining method for a linear system to minimize power and maximize throughput, given a constraint on the number of pipeline stages and a set of resource constraints. Unlike most existing pipelining approaches, our method takes the number of pipeline stages as one of the constraints and considers the pipelining as an aspect of power minimization. Operations are retimed so that as many operations as possible take common operands as their inputs, using a novel technique called force-directed retiming; operand sharing is then determined, based on list scheduling. Experimental results show that the proposed approach reduces the power consumption of functional units by 27.8% on average and by more than 50% in some cases, compared to the state-of-the-art pipelining and operand sharing techniques. Daehong Kim, Dongwan Shin, Kiyoung Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2004 | Memory and architecture exploration with thread shifting for multithreaded processors in embedded systemsabstractIn embedded multithreaded architectures, the performance enhancement relative to the base single-threaded architecture is highly dependent on the characteristics of the application and memory configuration. When the application is well parallelized, the multithreading performance may be good even with a small cache since the memory access latency can be hidden. However, if there are complicated dependencies between threads, they cause frequent cache conflicts, so the performance may not be improved. For that reason, not only processor architecture but also memory configuration should be customized to get an optimal solution of an embedded multithreaded system. We suggest a design space exploration algorithm, which considers both memory configuration and multithreaded architecture and a thread shifting technique, which shifts threads in compile time to minimize cache conflict. Mary Kiemb, Kiyoung Choi |
CASES | 2 |
| 2004 | A Mobility Analysis Method of Closed-chain Mechanisms with Over-constraints and Non-holonomic ConstraintsabstractMobility for a great portion of robot mechanisms having over-constraint and non-holonomic constraints has not been clearly identified. This work us to introduce a method of mobility analysis for such systems using the concept of representative screws and pseudo-joint. The pseudo-joint is employed to effectively represent the real motion trajectory due to the rolling contact of the wheel. To show the validity and effectiveness of the proposed method, mobility of various types of planar mobile robots having over-constraint and non-holonomic constrains are examined. Whee Kuk Kim, Kiyoung Choi, Byung-Ju Yi |
ICRA | 2 |
| 2003 | Evaluating Memory Architectures for Media Applications on Coarse-Grained Recon.gurable Architectures
Jongeun Lee, Kiyoung Choi, Nikil Dutt |
ASAP | 2 |
| 2003 | Scheduling and Timing Analysis of HW/SW On-Chip Communication in MP SoC DesignabstractOn-chip communication design includes designing software (SW) parts (operating system, device drivers, interrupt service routines, etc.) as well as hardware (HW) parts (on-chip communication network, communication interfaces of processor/IP/memory, on-chip memory, etc.). For an efficient exploration of its design space, we need fast scheduling and timing analysis. In this work, we tackle two problems (one for SW and the other for HW) in on-chip communication design. One is to incorporate the dynamic behavior of SW (interrupt processing and context switching) into on-chip communication scheduling. The other is to reduce on-chip data storage required for on-chip communication, by sharing physical communication buffers with different communication transactions. To solve the problems, we present both ILP (integer linear programming) formulation and heuristic algorithm, which enable the designer to perform efficient onchip communication scheduling and obtain accurate timing information. Experimental results show the effectiveness of our work. Youngchul Cho, Ganghee Lee, Sungjoo Yoo, Kiyoung Choi, Nacer-Eddine Zergainoh |
DATE | 4 |
| 2003 | Energy-efficient instruction set synthesis for application-specific processorsabstractSeveral techniques have been proposed to enhance the energy-efficiency of ASIPs (Application-Specific Instruction set Processors). While those techniques can reduce the energy consumption with a minimal change in the instruction set (IS), they fail to exploit the opportunity of designing the entire IS from the energy-efficiency perspective. In this paper, we present an energy-efficient IS synthesis technique that can comprehensively reduce the energy-delay product (EDP) of ASIPs through optimal instruction encoding, considering both the instruction bitwidth and the dynamic instruction count. Experimental results with a typical embedded RISC processor show that our technique can generate application-specific IS's that are up to 40% more energy-efficient over the native IS for several application benchmarks. Jongeun Lee, Kiyoung Choi, Nikil Dutt |
ISLPED | 2 |
| 2003 | An algorithm for mapping loops onto coarse-grained reconfigurable architecturesabstractWith the increasing demand for flexible yet highly efficient architecture platforms for media applications, there is a growing interest in the Coarse-grained Reconfigurable Architectures (CRAs). While many CRAs have demonstrated impressive performance improvement, the lack of compilation technology for such architectures causes a bottleneck in the current design process. In this paper, we present a novel mapping algorithm designed to support Reconfigurable ALU Array (RAA) architectures, that represent a significant class of CRAs. More specifically we present a core mapping algorithm that addresses the problem of placing and routing the operations of a loop body onto the ALU array, to be executed in a loop pipelined fashion. Experimental results using our mapping algorithm on a typical RAA show that our algorithm not only has very fast compilation time but can also generate quality mappings exhibiting high memory bandwidth utilization and low global interconnection requirements. Comparison with manual mapping also indicates that our algorithm can generate near-optimal mappings for several loops. Jongeun Lee, Kiyoung Choi, Nikil Dutt |
LCTES | 2 |
| 2002 | Efficient instruction encoding for automatic instruction set design of configurable ASIPsabstractApplication-specific instructions can significantly improve the performance, energy, and code size of configurable processors. A common approach used in the design of such instructions is to convert application-specific operation patterns into new complex instructions. However, processors with a fixed instruction bitwidth cannot accommodate all the potentially interesting operation patterns, due to the limited code space afforded by the fixed instruction bitwidth. We present a novel instruction set synthesis technique that employs an efficient instruction encoding method to achieve maximal performance improvement. We build a library of complex instructions with various encoding alternatives and select the best set of complex instructions while satisfying the instruction bitwidth constraint. We formulate the problem using integer linear programming and also present an effective heuristic algorithm. Experimental results using our technique generate instruction sets that show improvements of up to 38% over the native instruction set for several realistic benchmark applications running on a typical embedded RISC processor. Jongeun Lee, Kiyoung Choi, Nikil Dutt |
ICCAD | 2 |
| 2002 | An intra-task dynamic voltage scaling method for SoC design with hierarchical FSM and synchronous dataflow modelabstractThis paper presents a method of intra-task dynamic voltage scaling (DVS) for SoC design with hierarchical FSM and synchronous dataflow model (in short, HFSM-SDF model). To have an optimal intra-task DVS, exact execution paths need to be determined in compile time or runtime. In general programs, since determining exact execution paths in compile time or runtime is not possible, existing methods assume worst/average-case execution paths and take static voltage scaling approaches. In our work, we exploit a property of HFSM-SDF model to calculate exact execution paths in runtime. With the information of exact execution paths, our DVS method can calculate exact remaining workload. The exact workload enables to calculate optimal voltage level which gives optimal energy consumption while satisfying the given timing constraint. Experiments show the effectiveness of the presented method in low-power design of an MPEG4 decoder system. Sunghyun Lee 0002, Kiyoung Choi, Sungjoo Yoo |
ISLPED | 2 |
| 2001 | High-level synthesis under multi-cycle interconnect delayabstractAs process technology goes into deep submicron range, interconnect delay becomes dominant among overall system delay, occupying most of the system clock cycle time. Interconnect delay is now a crucial factor that needs to be considered even during high-level synthesis. In this paper, we propose a concurrent scheduling and binding algorithm that takes interconnect delay into account. We first define our distributed target architecture, which minimizes the effect of interconnect delay on clock cycle time. We no longer assume that interconnect delay between functional units is a part of one clock cycle. Interconnect delay can span over multiple clock cycles. We incorporate the concept of multi-cycle interconnect delay into scheduling and binding process, to reduce the critical path length and therefore the system latency. We show that by introducing interconnect delay, we can obtain latency improvement of up to 54 % and of 37% on the average. Jinhwan Jeon, Daehong Kim, Dongwan Shin, Kiyoung Choi |
ASP-DAC | 4 |
| 2001 | Performance improvement of multi-processor systems cosimulation based on SW analysisabstractWe propose a method for performance improvement of multi-processor systems cosimulation by reducing synchronization overhead between multiple simulators. To reduce the amount of simulator synchronization, we predict synchronization time points based on a static analysis of application software running on each processor. In experiments with real embedded systems, we obtained orders of magnitude higher performance in cosimulation runtimes. Jinyong Jung, Sungjoo Yoo, Kiyoung Choi |
DATE | 3 |
| 2001 | Behavior-to-Placed RTL Synthesis with Performance-Driven PlacementabstractInterconnect delay should be considered together with computation delay during architectural synthesis in order to achieve timing closure in deep submicrometer technology. In this paper, we propose an architectural synthesis technique for distributed-register architecture, which separates interconnect delay for data transfer from component delay for computation. The technique incorporates performance-driven placement into the architectural synthesis to minimize performance overhead due to interconnect delay. Experimental results show that our methodology achieves performance improvement of up to 60% and 22% on the average. Daehong Kim, Jinyong Jung, Sunghyun Lee 0002, Jinhwan Jeon, Kiyoung Choi |
ICCAD | 5 |
| 2001 | Low power pipelining of linear systems: a common operand centric approachabstractArticle Low power pipelining of linear systems: a common operand centric approach Share on Authors: Daehong Kim EECS, Seoul National University, Seoul, Korea EECS, Seoul National University, Seoul, KoreaView Profile , Dongwan Shin ICS, UC Irvine ICS, UC IrvineView Profile , Kiyoung Choi EECS, Seoul National University, Seoul, Korea EECS, Seoul National University, Seoul, KoreaView Profile Authors Info & Claims ISLPED '01: Proceedings of the 2001 international symposium on Low power electronics and designAugust 2001 Pages 225–230https://doi.org/10.1145/383082.383141Online:06 August 2001Publication History 9citation224DownloadsMetricsTotal Citations9Total Downloads224Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Daehong Kim, Dongwan Shin, Kiyoung Choi |
ISLPED | 3 |
| 2001 | Performance-driven high-level synthesis with bit-level chaining andclock selectionabstractThis paper presents a new scheme for scheduling and control synthesis in high-level circuit design. The scheduling algorithm tries to maximize the performance of a design under resource constraints by maximizing the utilization of resources and minimizing clock slack. It exploits the technique of bit-level chaining (BLC) to target high-speed design. It also exploits noninteger multicycling and chaining, which allows multiple cycle execution of a set of chained operations and even sharing of chained functional units to obtain further performance at the cost of a small increase in the complexity of the control unit. Experimental results on several datapath-intensive designs show significant improvement in throughput over the conventional scheduling algorithms. Sanghun Park, Kiyoung Choi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2001 | Partial bus-invert coding for power optimization of application-specific systemsabstractThis paper presents two bus coding schemes for power optimization of application-specific systems: partial pus-invert coding and its extension to multiway partial bus-invert coding. In the first scheme, only a selected subgroup of bus lines is encoded to avoid unnecessary inversion of relatively inactive and/or uncorrelated bus lines which are not included in the subgroup. In the extended scheme, we partition a bus into multiple subbuses by clustering highly correlated bus lines and then encode each subbus independently. We describe a heuristic algorithm of partitioning a bus into subbuses for each encoding scheme. Experimental results for various examples indicate that both encoding schemes are highly efficient for application-specific systems. Youngsoo Shin, Soo-Ik Chae, Kiyoung Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2001 | Narrow bus encoding for low-power DSP systemsabstractHigh levels of integration in integrated circuits often lead to the problem of running out of pins. Narrow data buses can be used to alleviate this problem provided that the degraded performance due to wait cycles can be tolerated. We address bus coding methods for low-power core-based systems incorporating narrow buses. We show that transition signaling combined with bus-invert coding, which we call BITS coding, is particularly suitable for the data patterns of typical DSP applications on narrow data buses. The application of BITS coding to real circuit design is limited by the extra bus line introduced, which changes the pinout of the chip. We propose a new coding method, which does not require the extra bus line but retains the advantage of BITS. Youngsoo Shin, Kiyoung Choi, Young-Hoon Chang |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | Hardware-software cosynthesis for run-time incrementally reconfigurable FPGAsabstractThis paper presents a method for hardware-softw are cosynthesis with run-time incrementally recon gurable FPGAs.T o reduce the run-time o v erhead of recon guring FPGAs, we present a concept called early partial recon guration (EPR) which minimizes the ov erhead by performing recon guration for an operation (or a task in our terms) mapped to an FPGA as early as possible so that the operation is ready to start when its execution is requested.F or further reduction of the ov erhead,w ein tegrate the incremental reconguration (IR) of FPGAs with the EPR concept.We present an ILP formulation and an eÆcient heuristic algorithm based on the EPR and IR concepts.Experiments on embedded system examples and syn thetic examples show the eÆciency of the proposed method Byungil Jeong, Sungjoo Yoo, Sunghyun Lee 0002, Kiyoung Choi |
ASP-DAC | 4 |
| 2000 | Narrow bus encoding for low power systemsabstractHigh integration in integrated circuits often leads to the problem of running out of pins.Narrow data buses can be used to alleviate this problem at the cost of performance degradation due to wait cycles.In this paper, we address bus coding methods for low power core-based systems incorporating narrow buses.Although the conventional Bus-Invert code performs well for completely random patterns, we show that transition signaling combined with Bus-Invert, which we call BITS coding, can achieve much more power saving for data patterns of typical DSP applications.The application of BITS coding to a real circuit design is limited by the overhead of the encoder and decoder circuits and the extra bus line introduced.We propose an approximate version of BITS coding, which do not require the extra bus line while retaining the advantage of BITS coding. Youngsoo Shin, Kiyoung Choi |
ASP-DAC | 2 |
| 2000 | Schedulability-driven performance analysis of multiple mode embedded real-time systemsabstractProviding multiple modes to support dynamically changing environments, standards, and new services is prevalent in embedded systems, especially in mobile radio systems. Because such a system frequently contains time-constrained tasks, it is important to analyze the temporal requirements as well as the functional correctness. This paper presents a method to analyze temporal requirements imposed on an embedded real-time system supporting multiple modes. While most performance analysis methods focus only on testing the feasibility of a task or a system, our method goes further by addressing the problem of locating hot spots of a system thereby helping the designer to choose among alternative designs or architectures. We formally define the analysis problem and show that it is very unlikely to be solved efficiently. We present a heuristic algorithm, which is accurate and fast enough to be used in iterative processes in system-level analysis and design. The analysis problem is extended to accommodate probabilistic behavior exhibited by soft real-time tasks. Youngsoo Shin, Daehong Kim, Kiyoung Choi |
DAC | 3 |
| 2000 | Fast Hardware-Software Coverification by Optimistic Execution of Real ProcessorabstractTo achieve fast verification of the software part of an embedded system, we propose to run the target processor optimistically, which effectively reduces the synchronization overhead with other simulators. For the optimistic processor execution, we present a processor execution platform and state saving/restoration methods. We performed optimistic execution of ARM710A processor in the coverification of an IS-95 CDMA cellular phone system and obtained up to orders of magnitude higher performance compared with the case that the processor runs conservatively. Sungjoo Yoo, Jongeun Lee, Jinyong Jung, Kyoungseok Rha, Youngchul Cho, Kiyoung Choi |
DATE | 6 |
| 2000 | Power Optimization of Real-Time Embedded Systems on Variable Speed ProcessorsabstractPower efficient design of real-time embedded systems based on programmable processors becomes more important as system functionality is increasingly realized through software. This paper presents a power optimization method for real-time embedded applications on a variable speed processor. The method combines off-line and on-line components. The off-line component determines the lowest possible maximum processor speed while guaranteeing deadlines of all tasks. The on-line component dynamically varies the processor speed or brings a processor into a power-down mode according to the status of task set in order to exploit execution time variations and idle intervals. Experimental results show that the proposed method obtains a significant power reduction across several kinds of applications. Youngsoo Shin, Kiyoung Choi, Takayasu Sakurai |
ICCAD | 2 |
| 2000 | A new cost model for high-level power optimization and its applicationabstractOptimizing power at high level needs less computational effort than that at lower levels and has significant effect as the designer con explore the design space from a global view point. Correct power cost model is essential to estimate during the optimizing process which design alternative consumes less power. We devise a new power cost model for each execution unit. It is based on the behavior of the cells contained in the execution unit. It is simple but more accurate than the conventional Hamming distance model. Applying it to operand interchange, we obtained significantly better results than Hamming distance model. We also propose an optimal algorithm for operand interchange. Taekyoon Ahn, Kiyoung Choi, Ki-Hyun Kim, Seong-Kwan Hon |
ISCAS | 2 |
| 2000 | Power minimization of functional units partially guarded computationabstractThis paper deals with power minimization problem for data-dominated applications based on a novel concept called partially guarded computation. We divide a functional unit into two parts - MSP (Most Significant Part) and LSP (Least Significant Part) - and allow the functional unit to perform only the LSP computation if the range of output data can be covered by LSP. We dynamically disable MSP computation to remove unnecessary transitions thereby reducing power consumption. We also propose a systematic approach for determining optimal location of the boundary between the two parts during high-level synthesis. Experimental results show about 10~44% power reduction with about 30~36% area overhead and less than 3% delay overhead in functional units. Junghwan Choi, Jinhwan Jeon, Kiyoung Choi |
ISLPED | 3 |
| 2000 | Low power self-timed Radix-2 division (poster session)abstractA self-timed radix-2 division scheme for low power consumption is proposed. By replacing dual-rail dynamic circuits in non-critical data paths with single-rail static circuits, power dissipation is decreased, yet performance is maintained by speculative remainder computation. SPICE simulation results show that the proposed design can achieve 33.8-ns latency for 56-bit mantissa division and 47% energy reduction compared to a fully dual-rail version. Jae-Hee Won, Kiyoung Choi |
ISLPED | 2 |
| 2000 | Performance improvement of geographically distributed cosimulation by hierarchically grouped messagesabstractTo improve the performance of geographically distributed cosimulation, we propose a concept called hierarchically grouped message. The concept improves cosimulation performance, preserving the cosimulation accuracy, by hierarchically grouping messages transferred between simulators in a short period of simulated time into a single physical message, thereby reducing the number of physical messages. Applying the concept to hybrid and optimistic cosimulation, we can reduce the number of rollbacks as well as the communication overhead accompanying the message transfer. Experimental results show the efficiency of the proposed method for practical examples in an internationally distributed cosimulation environment. Sungjoo Yoo, Kiyoung Choi, Dong Sam Ha |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1999 | Performance-Driven Scheduling with Bit-Level ChainingabstractThis paper presents a new scheduling algorithm that maximizes the performance of a design under resource constraints in high-level synthesis. The algorithm tries to achieve the maximal utilization of resources and the minimal waste of clock slack time. Moreover, it exploits the technique of bit-level chaining to target high-speed designs. The algorithm tries non-integer multiple-cycling and chaining, which allows multiple cycle execution of chained operations, to further increase the performance at the cost of small increase in the complexity of the control unit. Experimental results on several datapath-intensive designs show significant improvement in execution time, over the conventional scheduling algorithms. 1 Introduction Scheduling in high-level synthesis takes a crucial role of determining the performance of synthesized hardware. As a performance measure, we can use data introduction interval (sample period) or latency when pipelining [1, 2, 3] is not used. Given a data flow g... Sanghun Park, Kiyoung Choi |
DAC | 2 |
| 1999 | Power Conscious Fixed Priority Scheduling for Hard Real-Time SystemsabstractPower efficient design of real-time systems based on programmable processors becomes more important as system functionality is increasingly realized through software. This paper presents a powerefficient version of a widely used fixed priority scheduling method. The method yields a power reduction by exploiting slack times, both those inherent in the system schedule and those arising from variations of execution times. The proposed run-time mechanism is simple enough to be implemented in most kernels. Experimental results show that the proposed scheduling method obtains a significant power reduction across several kinds of applications. 1 Introduction Recently, power consumption has been a critical design constraint in the design of digital systems due to widely used portable systems such as cellular phones and PDAs, which require low power consumption with high speed and complex functionality. The design of such systems often involves reprogrammable processors such as microprocessors... Youngsoo Shin, Kiyoung Choi |
DAC | 2 |
| 1999 | Exploiting Early Partial Reconfiguration of Run-Time Reconfigurable FPGAs in Embedded Systems DesignabstractNo abstract available. Byungil Jeong, Sungjoo Yoo, Kiyoung Choi |
FPGA | 3 |
| 1998 | Loop Pipelining in Hardware-Software PartitioningabstractThis paper presents a hardware-software partitioning algorithm that exploits a loop pipelining technique. The partitioning algorithm is based on iterative improvement. The algorithm tries to minimize hardware cost through hardware sharing and hardware implementation selection without violating given performance constraint. The proposed loop pipelining technique, which is an adaptation of a compiler optimization technique for instruction level parallelism, increases parallelism within a loop by transforming the structure of an input system description. By combining this technique with our partitioning algorithm, we can further reduce the hardware cost and/or improve the performance of the partitioned system. Experiments show about 19% performance improvement and 44% reduced hardware for a JPEG encoder design, compared to the results without loop pipelining. Jinhwan Jeon, Kiyoung Choi |
ASP-DAC | 2 |
| 1998 | Partial bus-invert coding for power optimization of system level busabstractWe presen t a partial bus-in vertcoding scheme for po wer optim ization of system level bus. In the proposed sch eme, we select a su b-group of bus lines involved in b us encoding to a void unnecessary inversion of b us lines not in the sub-group thereby redu cing th e total number of bus transitions. We propose a heuristic algorithm that selects the sub-grou p of bus lines for b us encoding. Ex periments on benchmark examples in dicate that the partial bus-in vert coding reduces the tot al bus tran sitions b y 62.6% on the av erage, compared to that of the unencoded patterns. Youngsoo Shin, Soo-Ik Chae, Kiyoung Choi |
ISLPED | 3 |
| 1997 | Power-conscious High Level Synthesis Using Loop FoldingabstractIn this paper, a transformation technique, called powerconscious loop folding is proposed for high level synthesis of a low power system. Our work is focused on reducing the power consumed by functional units through the decrease of switching activity in a data path dominated circuit containing loops. The transformation algorithm has been implemented and integrated into a high level synthesis system for experiments. In our experiments, we could achieve power reduction of up to 50% for circuits dominated by functional units. 1. Introduction Until 1980's, one of the most important factors that determined the quality of a system was the speed and so much effort had been made to increase the speed at a minimal cost or silicon area. But, in 1990's, as the portable system market grows rapidly and reliability problems due to high power dissipation are becoming an issue for systems operating in high clock frequency, low power design is becoming more and more important and is now one of the m... Daehong Kim, Kiyoung Choi |
DAC | 2 |
| 1997 | Low power high level synthesis by increasing data correlationabstractWith the increasing performance and density of VLSI circuits as well as the popularity of portable devices such as personal digital assistance, power consumption has emerged as an important issue in the design of electronic systems. Low power design techniques have been pursued at all design levels. However, it is more effective to attempt to reduce power dissipation at higher levels of abstraction which allow wider view. In this paper, we propose a simultaneous scheduling and binding scheme which increases the correlation between consecutive inputs to an execution unit so that the switched capacitance of the execution unit is reduced. The proposed method is implemented and integrated into the scheduling and assignment part of the HYPER synthesis environment. Compared with the original HYPER synthesis system, average power saving of 23.0% in execution units and 14.2% in the whole circuit, is obtained for a set of benchmark examples. Dongwan Shin, Kiyoung Choi |
ISLPED | 2 |
| 1996 | Software synthesis through task decomposition by dependency analysisabstractLatency tolerance is one of main problems of software synthesis in the design of hardware-software mixed systems. This paper presents a methodology for speeding up systems through latency tolerance which is obtained by decomposition of tasks and generation of an efficient scheduler. The task decomposition process focuses on the dependency analysis of system i/o operations. Scheduling of the decomposed tasks is performed in a mixed static and dynamic fashion. Experimental results show the significance of our approach. Youngsoo Shin, Kiyoung Choi |
ICCAD | 2 |
| 1996 | An integrated hardware-software cosimulation environment with automated interface generationabstractWe present a hardware-software cosimulation environment for heterogeneous systems. To be an efficient system verification environment for the rapid prototyping of heterogeneous systems, the environment provides following features: interface transparency, smooth transition to cosynthesis, simulation acceleration, and integrated user interface and internal representation. Among them, the first two are more important than the others. To support these two features, we have developed automatic interface generation schemes. As demonstrating experiments, two heterogeneous systems performing same function with different target architectures were cosimulated and prototyped successfully in our environment. The experimental results show that our environment can be a useful heterogeneous system specification/verification environment for rapid prototyping. Kyuseok Kim, Youngsoo Shin, Kiyoung Choi |
RSP | 4 |
| 1996 | Self-timed divider based on RSD number systemabstractThe authors propose a divider structure that combines a novel self timed ring structure and a carry-propagation-free division algorithm. The self-timed ring structure enables the divider to compute at a speed comparable to that of previously designed dividers with less silicon area. By exploiting the carry-propagation-free division algorithm, an even better performance can be achieved. The authors designed a layout of 54 b divider using 1.2 /spl mu/m CMOS technology and measured the area and speed. A speed of 135 ns per worst case division was obtained on 5.7 mm/sup 2/ of silicon area. KiJong Lee, Kiyoung Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 1995 | An integrated hardware-software cosimulation environment for heterogeneous systems prototypingabstractNo abstract available. Kyuseok Kim, Youngsoo Shin, Taekyoon Ahn, Wonyong Sung, Kiyoung Choi, Soonhoi Ha |
ASP-DAC | 6 |
| 1995 | Efficient Prototyping System Based on Incremental Design and Module-by-Module VerificationabstractThis paper presents an efficient hardware prototyping methodologies of digital systems. A is based on a low-cost and flexible prototyping system which consists of a general-purpose CPU and a FPGA-based custom board. Using our prototyping methodologies such as incremental system design and module-by-module verification, we can map partial system specification into hardware prototype, which is implemented by programming FPGAs on custom board. This allows flexible and efficient system verification as well as reduction in cost and time of prototype building. Youngsoo Shin, Kyuseok Kim, Jae-Hee Won, Kiyoung Choi |
ISCAS | 5 |
| 1994 | A Self-Timed Divider Using RSD Number SystemabstractThe paper proposes a divider structure that combines a novel self-timed ring structure and a carry-propagation-free division algorithm. The self-timed ring structure enables the divider to compute at a speed comparable to that of previously designed dividers with less silicon area. By exploiting the carry-propagation-free division algorithm, we can achieve even better performance. A 54-bit divider employing the proposed structure and algorithm was designed with 1.2 /spl mu/m CMOS technology. We obtained a speed of 215 ns per worst case division on 4.2 mm/sup 2/ of silicon area.> Kiyoung Choi, KiJong Lee, Jun-Woo Kang |
ICCD | 1 |
| 1988 | Incremental-in-time Algorithm for Digital Simulation
Kiyoung Choi, Sun Young Hwang, Tom Blank |
DAC | 1 |
| 1988 | Fast functional simulation: an incremental approachabstractIn an effort to speed up simulation, a novel algorithm,, called incremental simulation evaluates the circuit components that can be affected directly or indirectly by design changes, utilizing the information generated during the previous simulation to reduce the number of component evaluations to a minimum. The authors describe the design and implementation of the incremental algorithm for logic or functional simulation, which substantially improves the run-time performance over existing simulators by using the incremental property of the hardware design process.> Sun Young Hwang, Tom Blank, Kiyoung Choi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |