VLDB 2026 Research / reviewers in the wild / expert
In-Cheol Park
dblp:47/4676
· DBLP profile ↗
106ranked-venue papers
9as first author
14since 2021 · last 2026
0000-0003-3524-2838ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 85 · 9 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13Computer networks · 2Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hardware-Efficient Unified Approximation for Implementing Diverse Smooth Activation Functions
Jeongmin Kim 0001, Kangjoon Choi, In-Cheol Park |
IEEE Trans. Computers | 3 |
| 2026 | General-Purpose QC-LDPC Processor for Error-Correction Performance EvaluationabstractThis paper introduces a general-purpose quasi-cyclic (QC)-low-density parity-check (LDPC) processor (GPLP) developed to evaluate LDPC decoding performance, which is a highly flexible and programmable field programmable gate array (FPGA)-based simulation platform. The GPLP enables to configure LDPC decoding processes, depending on the decoding algorithms and design choices to be simulated and verified. In contrast to traditional FPGA-based approaches, which typically rely on a fixed dataflow, the GPLP offers a reconfigurable, processor-like environment that can be programmed for various design configurations. This paper defines essential atomic operations required for QC-LDPC decoding and proposes the GPLP architecture supporting those operations with vector-based processing. The proposed GPLP significantly enhances simulation speed, providing a performance improvement of 220x to 574x over CPU-based simulations. The platform provides a unique combination of simulation speed, flexibility, and reconfigurability, making it an invaluable tool to rapidly evaluate and optimize LDPC decoding performance. Kangjoon Choi, Jongmin Baek, In-Cheol Park |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2026 | Enhanced-Throughput Antithetic Sample Generation Architecture for Monte Carlo SimulationabstractMonte Carlo simulation imposes significant computational and hardware costs due to extensive sample generation. While antithetic variates can reduce sample requirements, direct hardware application doubles transformation paths. Based on a point-symmetric inverse cumulative distribution, this brief introduces a hardware architecture that generates antithetic pairs using a single transformation path and a sign inversion, guaranteeing strong negative correlation and effective variance reduction. Implemented on a Xilinx ZCU104 field-programmable gate array (FPGA), the architecture doubles throughput with minimal hardware overhead. In Monte Carlo bit error rate (BER) simulation for an 802.11n LDPC code, the proposed architecture maintains accuracy while achieving a$2.14\times $to$3.62\times $speed gain in random-variable generation. Kangjoon Choi, In-Cheol Park |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2026 | A Hardware-Efficient Reed-Solomon Decoding Architecture for Single-Symbol Error Correcting in High-Bandwidth Memory
Jaehoon Kwon, Jeongmin Kim 0001, Jeonghyeon Yoon, Hansol Jeong, In-Cheol Park |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | Hardware-Efficient Architecture for Multiple Quantized Gaussian Noise GenerationabstractThis paper presents two novel architectures to generate a number of quantized Gaussian noises. The first architecture exploits inversion through uniform segmentation, enabling a uniform look up table (LUT) splitting technique to efficiently generate quantized Gaussian noise while maintaining reasonable tail quality in Gaussian noise generation. The second architecture utilizes inversion through hierarchical segmentation and a probability-based LUT selection, significantly reducing the total LUT size while preserving the tail quality of the generated Gaussian noise. Both designs generate multiple uniform random numbers by cascading combinational circuits, which improves Gaussian noise generation efficiency compared to the conventional linear feedback shift register-based method. Compared to the previous architecture based on inversion through hierarchical segmentation, the proposed uniform segmentation architecture achieves a 6.02x improvement, when implemented on a field-programmable gate array device, in terms of throughput per configurable logic block, and the proposed hierarchical segmentation architecture achieves a 2.71x improvement. Kangjoon Choi, In-Cheol Park |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2025 | Multiple-Resolution Decoding Architecture for QC-LDPC CodesabstractThis paper proposes a hardware-efficient multiple-resolution decoding architecture for quasi-cyclic low-density parity-check (QC-LDPC) codes, specifically designed to meet the stringent requirements of the 5G New Radio (NR) standard. The proposed architecture adopts a single-instruction multiple-data (SIMD) approach to dynamically adjust the bitwidth of log-likelihood ratio (LLR) values based on theEb/N0condition, significantly reducing hardware complexity and improving throughput area ratio. Unlike conventional single-resolution decoders, the architecture processes 2-bit LLR values in highEb/N0regions and scales up to 4-bit or 8-bit LLR values for moderate and lowEb/N0conditions, maintaining robust error-correcting performance. Key innovations include SIMD-based design for variable-node units (VNUs), check-node units (CNUs), and quasi-cyclic shifting networks (QSNs), as well as optimized memory access scheduling to support all 51 lifting sizes defined in the 5G NR standard. Designed in a 65-nm CMOS process, the decoder achieves a peak throughput of 27.24 Gbps under error-free conditions with a throughput area ratio improvement of 2.07× compared to the state-of-the-art designs. Furthermore, the proposed architecture demonstrates superior throughput-area ratio and flexibility, supporting all code rates and lifting sizes specified in the 5G NR standard. Simulation results confirm that the proposed decoder meets the peak throughput requirement under error-free conditions, while maintaining robust performance in challenging channel environments. In-Cheol Park, Kangjoon Choi, Hyejung Jang, Jongmin Baek |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | Energy-Efficient Syndrome Calculation Architecture for BCH Decoders
Jeongmin Kim 0001, Jaehoon Kwon, Hansol Jeong, In-Cheol Park |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | Area-Efficient QC-LDPC Decoding Architecture With Thermometer Code-Based Sorting and Relative Quasi-Cyclic ShiftingabstractThe 5G New-Radio (NR) communication standard requires high throughput and low latency, so low-density parity-check (LDPC) codes, which have higher inherent parallelism and lower decoding complexity than turbo codes, were adopted as the main coding method for data channels. In traditional LDPC min-sum decoders, the check node unit was realized using a sorting unit based on the min-tree structure. However, this structure resulted in high hardware complexity and long latency. To address this issue, we propose a new sorting method based on the thermometer code-based number system. Additionally, we introduce a new LDPC decoding architecture that reduces the number of QSN stages from two to one, significantly lowering the shifting logic complexity needed to support different lifting sizes. This is achieved by using relative shift amounts instead of absolute shift amounts specified in the parity check matrix. The proposed decoder implemented using a partially parallel structure in a 65nm CMOS technology satisfies the various operation modes and the throughput requirements of the 5G NR standard, and boasts a higher normalized throughput than state-of-the-art LDPC decoders. Boseon Jang, Hyejung Jang, Kangjoon Choi, In-Cheol Park |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | Hardware-Efficient SoftMax Architecture With Bit-Wise Exponentiation and Reciprocal CalculationabstractThe SoftMax function is one of the activation functions used in deep neural networks (DNN) to normalize input values to the range of (0,1). With the advent of DNN models including the Transformer, operations utilizing SoftMax have gained significant attention, and the efficient hardware implementation of such operations has become a prominent issue in hardware realization. Implementing SoftMax often involves exponential and division operations, which can be a significant bottleneck in terms of hardware cost and performance. Various efforts have been made to address this challenge, and this paper introduces a novel approach to efficiently implement SoftMax. In most previous works, the maximum input value is subtracted from all the input values to ensure numerical stability. In the proposed approach, the maximum value is replaced with a different value to reduce the hardware complexity with ensuring numerical stability. Additionally, in exponential operations, simple Look-Up Tables (LUTs) with only one entry each are used for bit-wise calculations, and the reciprocal of the total exponential sum is computed to replace division with multiplication. Applying the proposed methods reduces the computational complexity significantly compared to the previous log-sum-exp approach. As a result, the proposed 8-bit SoftMax accelerator achieves a high operating frequency of 3.12GHz and a high throughput of 25G inputs/s. It also improves area efficiency and power consumption by at least 2 times. From an accuracy perspective, furthermore, it is associated with similar or even better accuracy compared to previous works. Jeongmin Kim 0001, Kangjoon Choi, In-Cheol Park |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | High-Speed Counter With Novel LFSR State ExtensionabstractThis paper presents a high-speed counter architecture associated with novel LFSR state extension. By employing the proposed state extension, an$\emph {m}$-bit LFSR counter with$(\mathrm{2^{\mathit{m}}}-1)$states is modified to cover$\mathrm{2^{\mathit{m}}}$states without degrading the counting rate. Based on the property that only the low-order bits are frequently switched, the proposed counter consists of two sub-counters to achieve a high counting rate and reduce the hardware complexity needed to convert an LFSR state into a binary state. The low-order sub-counter is implemented with the proposed LFSR counter, and the high-order sub-counter is designed by employing the conventional synchronous binary counter. In addition, the implemented counter takes into account the speed degradation caused by the large fan-out of the high-order sub-counter. The proposed counter designed with standard cells operates at 2.08 GHz in a 65 nm CMOS technology, and its counting rate is almost independent of the counter size. Hyungjoon Bae 0001, Yujin Hyun, Suchang Kim, Sangsoo Park, Jaeyoung Lee 0004, Boseon Jang, Suyoung Choi, In-Cheol Park |
IEEE Trans. Computers | 8 |
| 2023 | A CNN Inference Accelerator on FPGA With Compression and Layer-Chaining Techniques for Style Transfer ApplicationsabstractRecently, convolutional neural networks (CNNs) have actively been applied to computer vision applications such as style transfer that changes the style of a content image into that of a style image. As the style transfer CNNs are based on encoder-decoder network architecture and should deal with high-resolution images that become mainstream these days, the computational complexity and the feature map size are very large, preventing the CNNs from being implemented on an FPGA. This paper proposes a CNN inference accelerator for the style transfer applications, which employs network compression and layer-chaining techniques. The network compression technique is to make a style transfer CNN have low computational complexity and a small amount of parameters, and an efficient data compression method is proposed to reduce the feature map size. In addition, the layer-chaining technique is proposed to reduce the off-chip memory traffic and thus to increase the throughput at the cost of small hardware resources. In the proposed hardware architecture, a neural processing unit is designed by taking into account the proposed data compression and layer-chaining techniques. A prototype accelerator implemented on a FPGA board achieves a throughput comparable to the state-of-the-art accelerators developed for encoder-decoder CNNs. Suchang Kim, Boseon Jang, Jaeyoung Lee 0004, Hyungjoon Bae 0001, Hyejung Jang, In-Cheol Park |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | Multi-Mode QC-LDPC Decoding Architecture With Novel Memory Access Scheduling for 5G New-Radio StandardabstractAs the low-density parity-check (LDPC) code has a powerful error-correcting performance and can achieve high throughput, it is being used in many application areas and recently adopted as a channel coding method in the 5G New-Radio communication standard. Unlike other LDPC codes, the 5G LDPC code has various irregular lifting sizes to support diverse message lengths. To meet the demanding requirements of the 5G standard, many solutions have been presented, but all of them are either impractical or fail to satisfy all the requirements. This paper, for the first time, proposes an area-efficient QC-LDPC decoder that satisfies the peak throughput requirements of the 5G standard and supports all the lifting sizes specified in the 5G standard. Instead of relying on full parallelism like in the previous works, this work tries partial parallelism to mitigate the hardware complexity, which leads to high efficiency in hardware complexity. In addition, a novel memory access scheduling method is proposed to solve the data access and alignment problems caused by the partially parallel structure, which is effective in supporting all the lifting sizes. A LDPC decoder realized in 65-nm CMOS technology demonstrates that its decoding throughput is greater than 20Gbps and its area is smaller than the existing decoders. Seongjin Lee, Sangsoo Park, Boseon Jang, In-Cheol Park |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | Hybrid Convolution Architecture for Energy-Efficient Deep Neural Network ProcessingabstractThis paper presents a convolution process and its hardware architecture for energy-efficient deep neural network (DNN) processing. A DNN in general consists of a number of convolutional layers, and the number of input features involved in the convolution of a shallow layer is larger than that of kernels. As the layer deepens, however, the number of input features decreases, while that of kernels increases. The previous convolution architectures developed for enhancing energy efficiency have tried to reduce the memory accesses by increasing the reuse of the data once accessed from the memory. However, redundant memory accesses are still required as the change in the numbers of data has not been considered. We propose a hybrid convolution process that selects either a kernel-stay or feature-stay process by taking into account the numbers of data, and a forwarding technique to further reduce the memory accesses needed to store and load partial sums. The proposed convolution process is effective in maximizing data reuse, leading to an energy-efficient hybrid convolution architecture. Compared to the state-of-the- art architectures, the proposed architecture enhances the energy efficiency by up to 2.38 times in a 65nm CMOS process. Suchang Kim, Jihyuck Jo, In-Cheol Park |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Real-Time SSDLite Object Detection on FPGAabstractDeep neural network (DNN)-based object detection has been investigated and applied to various real-time applications. However, it is hard to employ the DNNs in embedded systems due to their high computational complexity and deep-layered structure. Although several field-programmable gate array (FPGA) implementations have been presented recently for real-time object detection, they suffer from either low throughput or low detection accuracy. In this article, we propose an efficient computing system for real-time SSDLite object detection on FPGA devices, which includes novel hardware architecture and system optimization techniques. In the proposed hardware architecture, a neural processing unit (NPU) that consists of heterogeneous units, such as band processing, scaling, and accumulating, and data fetching and formatting units is designed to accelerate the DNNs efficiently. In addition, system optimization techniques are presented to improve the throughput further. A task control unit is employed to balance the workload and increase the utilization of heterogeneous units in the NPU, and the object detection algorithm is refined accordingly. The proposed architecture is realized on an Intel Arria 10 FPGA and enhances the throughput by up to 13.6× compared to the state-of-the-art FPGA implementation. Suchang Kim, Seungho Na, Byeong Yong Kong, Jaewoong Choi, In-Cheol Park |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2020 | A Low-Latency Multi-Touch Detector Based on Concurrent Processing of Redesigned Overlap Split and Connected Component AnalysisabstractA low-latency multi-touch detector architecture for locating numerous touches in large-panel devices is presented in this paper. Two respective processors for the overlap split and the connected component analysis (CCA) are the key components in a multi-touch detector. Exploiting the simplicity of typical intensity maps in practice, the two processors are first redesigned under a principle that every pixel is processed only once in a raster-scan order. More specifically, the concept of valley point is introduced, and a simple yet effective overlap-split scheme called valley-point division (VPD) is newly developed based on the concept. In addition, equivalent labels in the CCA are handled on the fly instead of being kept to be processed later, and the logics and memories for corner cases that never occur in practice are unloaded. Enabled by the redesign, subsequently, the two processors are integrated into one to concurrently conduct both the VPD and the CCA during the same raster scan. As a result, the proposed detector completes the whole detection procedure with a single raster scan of a map, and is exempted from a large memory required to hold an entire map. Implementation results in a 65-nm CMOS for a large panel of 400 × 250 sensors demonstrate that the proposed detector takes less than a half latency of the existing ones while occupying only 17% silicon area and consuming 52% power on average. Byeong Yong Kong, Jooseung Lee, In-Cheol Park |
ISCAS | 3 |
| 2020 | A 120-mW 0.16-ms-Latency Connectivity-Scalable Multiuser Detector for Interleave Division Multiple AccessabstractTo facilitate the massive connectivity of the 5G New Radio, this brief proposes a connectivity-scalable multiuser detector (MUD) architecture for interleave division multiple access (IDMA). A regular inter-MUD interface makes the proposed MUD not only operable alone but also scalable to serve a massive number of users by connecting multiple MUDs. Besides, the memory subsystem is simplified to mitigate the power consumption and the silicon area. A prototype 16-MUD in a 65-nm CMOS consumes 120 mW, occupies 8.21 mm2, and takes 0.16 ms to process one 8192-chip frame. Such numbers represent 29% less silicon area and 59% less power dissipation than those of the state-of-the-art MUD. Byeong Yong Kong, In-Cheol Park |
ISCAS | 2 |
| 2019 | Parallel IDMA Architecture Based on Interleaving with Replicated SubpatternsabstractThis paper presents a parallel multiuser detector architecture for low-latency interleave division multiple access. To enable P-parallel processing, an interleaving pattern is divided into P disjoint subpatterns, and all the subpatterns are designed to be identical without degrading error-rate performance noticeably. Since the subpatterns are all disjoint, they can be processed in parallel. Besides, by exploiting that they access the same address of separate memory banks at the same time, the banks are integrated into one to minimize the silicon area and the power consumption. As a result, the proposed architecture reduces the latency by a factor of P at the expense of a little hardware overhead. A prototype 2-parallel 16-user detector in a 65-nm CMOS completes the entire detection procedure two times earlier than the state-of-the-art nonparallel detector, while occupying only 12% more silicon area and dissipating 20% more power. Byeong Yong Kong, In-Cheol Park |
ICC | 2 |
| 2018 | A 2.4pJ/bit, 6.37Gb/s SPC-enhanced BC-BCH decoder in 65nm CMOS for NAND flash storage systemsabstractThis paper present an energy-efficient block-concatenated BCH (BC-BCH) decoder which can achieve superior decoding performance for NAND flash storage systems. To enhance the error-correcting capability, an additional decoding step with single parity-check (SPC) block is newly employed. A novel memory based syndrome updating method effectively improves the energy efficiency as well as the decoding latency. Using the proposed methods, a prototype chip is implemented to decode a (36443, 32768) BC-BCH code in 65nm CMOS process. The proposed decoder provides a decoding throughput of 6.37Gb/s and an efficiency of 2.4pJ/bit, being superior to the state-of-the-art hard-decision decoders for storages. In-Cheol Park, Youngjoo Lee 0002 |
ASP-DAC | 2 |
| 2016 | Low-complexity symbol detection for massive MIMO uplink based on Jacobi methodabstractIn this paper, we propose a low-complexity symbol detection algorithm for massive multiple-input multiple-output (MIMO) uplink. Grounded on the fact that a primary property of the massive MIMO systems guarantees the convergence of the Jacobi method, the method is exploited in the linear detection so that the estimate of transmitted symbols can be obtained without employing the computationally intensive matrix inversion. In addition, we propose a multiplication-free initial estimate for the Jacobi method in order to lessen the computational complexity further. Owing to the elimination of matrix inversion and the efficient initial estimate, the proposed algorithm achieves near-optimal error-rate performance with fewer computations than the state-of-the-art schemes. Byeong Yong Kong, In-Cheol Park |
PIMRC | 2 |
| 2016 | Energy-Efficient Floating-Point MFCC Extraction Architecture for Speech Recognition SystemsabstractThis brief presents an energy-efficient architecture to extract mel-frequency cepstrum coefficients (MFCCs) for real-time speech recognition systems. Based on the algorithmic property of MFCC feature extraction, the architecture is designed with floating-point arithmetic units to cover a wide dynamic range with a small bit-width. Moreover, various operations required in the MFCC extraction are examined to optimize operational bit-width and lookup tables needed to compute nonlinear functions, such as trigonometric and logarithmic functions. In addition, the dataflow of MFCC extraction is tailored to minimize the computation time. As a result, the energy consumption is considerably reduced compared with previous MFCC extraction systems. Jihyuck Jo, Hoyoung Yoo, In-Cheol Park |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Narrow-range frequency estimation based on comprehensive optimization of DFT and interpolationabstractAn efficient procedure for frequency estimation is proposed in this paper to alleviate the computational complexity. Grounded on the fact that the frequency of a target signal usually lies in a known range in practical applications, two fundamental steps in the frequency estimation, i.e., the discrete Fourier transform (DFT) and the interpolation of the DFT samples, are modified accordingly. Unlike the previous works focusing on either the DFT or the interpolation, this paper does not decouple the two steps but optimizes the whole procedure comprehensively by considering the interrelationship between the two steps. As a result, the number of operations required for the estimation is remarkably diminished while the performance remains competitive with the recent works. Byeong Yong Kong, In-Cheol Park |
ICASSP | 2 |
| 2014 | 7.3 Gb/s universal BCH encoder and decoder for SSD controllersabstractThis paper presents a universal BCH encoder and decoder that can support multiple error-correction capabilities. A novel encoding architecture and on-demand syndrome calculation technique is proposed to reduce both hardware complexity and power consumption. Based on the proposed methods, 32-parallel universal encoder and decoder are designed for BCH (8192+14t, 8192, t) codes, where the error-correction capability t is configurable to 8, 11, 16, 24, 32, and 64. The prototype chip achieves a throughput of 7.3 Gb/s and occupies 2.24 mm2in 0.13μπι CMOS technology. Hoyoung Yoo, Youngjoo Lee 0002, In-Cheol Park |
ASP-DAC | 3 |
| 2014 | Low-Complexity Low-Latency Architecture for Matching of Data Encoded With Hard Systematic Error-Correcting CodesabstractA new architecture for matching the data protected with an error-correcting code (ECC) is presented in this brief to reduce latency and complexity. Based on the fact that the codeword of an ECC is usually represented in a systematic form consisting of the raw data and the parity information generated by encoding, the proposed architecture parallelizes the comparison of the data and that of the parity information. To further reduce the latency and complexity, in addition, a new butterfly-formed weight accumulator (BWA) is proposed for the efficient computation of the Hamming distance. Grounded on the BWA, the proposed architecture examines whether the incoming data matches the stored data if a certain number of erroneous bits are corrected. For a (40, 33) code, the proposed architecture reduces the latency and the hardware complexity by ~32% and 9%, respectively, compared with the most recent implementation. Byeong Yong Kong, Jihyuck Jo, Hyewon Jeong, Mina Hwang, Soyoung Cha, Bongjin Kim, In-Cheol Park |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2014 | High-Throughput and Low-Complexity BCH Decoding Architecture for Solid-State DrivesabstractThis paper presents a high-throughput and low-complexity BCH decoder for NAND flash memory applications, which is developed to achieve a high data rate demanded in the recent serial interface standards. To reduce the decoding latency, a data sequence read from a flash memory channel is re-encoded by using the encoder that is idle at that time. In addition, several optimizing methods are proposed to relax the hardware complexity of a massive-parallel BCH decoder and increase the operating frequency. In a 130-nm CMOS process, a (8640, 8192, 32) BCH decoder designed as a prototype provides a decoding throughput of 6.4 Gb/s while occupying an area of 0.85${\rm mm}^{2}$. Youngjoo Lee 0002, Hoyoung Yoo, Injae Yoo, In-Cheol Park |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2013 | A 3Gb/s 2.08mm2 100b error-correcting BCH decoder in 0.13µm CMOS processabstractThis paper presents a high-throughput BCH decoder that can correct 100 bit-errors. Several optimization methods are proposed to reduce the hardware complexity caused by the large error-correction capability. Based on the proposed methods, an 8-parallel decoder is designed for the (9592, 8192, 100) BCH code, which achieves a decoding throughput of 3Gb/s and occupies 2.08mm2in 0.13µm CMOS process. Youngjoo Lee 0002, Hoyoung Yoo, In-Cheol Park |
ASP-DAC | 3 |
| 2013 | Adaptive Metric Calculation for Improving Detection Capability of MIMO DetectorsabstractA simple yet effective scheme is proposed to improve the detection capability of multiple-input multiple-output (MIMO) symbol detectors in wireless communication systems at low signal-to-noise ratio (SNR). The proposed scheme is to change the metric calculation adaptively to the channel SNR, grounded on the fundamental relationship between the detection capability and the bit-error rate (BER) performance with respect to the channel SNR. An efficient hardware architecture implementing the proposed scheme is also presented to show its applicability in a practical sense and it is proven that the architecture induces only a negligible hardware overhead compared to the state-of-the-art MIMO symbol detector. Experimental results show that the proposed scheme can indeed increase the detection capability effectively without degrading the BER performance noticeably. Byeong Yong Kong, In-Cheol Park |
VTC Spring | 2 |
| 2013 | Memory-Optimized Hybrid Decoding Method for Multi-Rate Turbo CodesabstractTo provide near-optimal error-correcting performance for multi-rate turbo codes and minimize the size of additional memory, a new decoding method is proposed in this paper. The proposed decoding is based on two methods: hybrid sliding window and dynamic metric encoding. The new window scheme combines dummy metric calculation and border metric storing methods to halve the border metric memory and improve the error-correcting performance for multi-rate codes. The dynamic encoding of metrics also significantly reduces the border metric memory without degrading the performance. Employing the two proposed methods reduces the conventional border metric memory by 83%. In addition, the superb error-correcting performance of the proposed method is verified for various code rates by conducting intensive simulations for 3GPP LTE codes. Injae Yoo, Bongjin Kim, In-Cheol Park |
VTC Spring | 3 |
| 2012 | Small-area parallel syndrome calculation for strong BCH decodingabstractThis paper presents a new optimization method to reduce the hardware complexity of syndrome calculation in strong BCH decoding. All the operations required in the parallel syndrome calculation are reformulated as a single matrix computation to enlarge the search area for common sub-expressions. The computational complexity of syndrome calculation is significantly reduced by finding and sharing common terms in the single matrix computation. Implementation results show that the proposed architecture saves 55% of area overheads compared to the conventional structure. Youngjoo Lee 0002, Hoyoung Yoo, In-Cheol Park |
ICASSP | 3 |
| 2012 | SNR-Adaptive Input Quantization for Turbo DecodingabstractThis paper presents how to control the input bit-width of a turbo decoder according to the signal-to-noise ratio (SNR) adaptively. It is crucial to minimize the input bit-width while maintaining the error-correcting performance. Several quantization schemes have been presented to properly decide the input resolution of turbo decoding, but they all assume a fixed bit-width irrespective of the SNR. This paper proposes a new method to adaptively change the quantization bit-width with respect to the SNR. In addition, this paper suggests a novel method to estimate the frame-error rate (FER) performance of a turbo decoder dealing with quantized channel outputs. As a result, the whole SNR range is divided into a number of regions each of which can be processed with a certain quantization bit-width. Injae Yoo, In-Cheol Park |
VTC Spring | 2 |
| 2012 | FIR Filter Synthesis Based on Interleaved Processing of Coefficient Generation and Multiplier-Block SynthesisabstractAn efficient filter synthesis algorithm is proposed to minimize the number of adders required in the design of finite-impulse response filters. Given a specification, a filter can be synthesized by conducting two main steps: coefficient generation and multiplier-block synthesis. While most of previous works have focused on only one of the steps, the proposed algorithm integrates the two steps in an interleaved manner so as to take into account the effect of multiplier-block synthesis in generating coefficients. In addition, the concept of sensitivity is developed to reduce the complexity of computing the variable ranges of coefficients. Experimental results show that the proposed algorithm outperforms previous algorithms in terms of adder cost and takes a relatively short computation time. Byeong Yong Kong, In-Cheol Park |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2012 | Low-Complexity Tone Reservation for PAPR Reduction in OFDM Communication SystemsabstractOrthogonal frequency division multiplexing communication systems have a drawback that some signal values can be much higher than the average signal value. Transmitting such a high signal increases symbol error rate (SER) significantly, as the signal is usually distorted by the nonlinearity of power amplifiers. Most of previous methods presented to reduce the peak-to-average-power ratio are based on iterative computations of the fast Fourier transform (FFT) and inverse FFT associated with large computational complexity. To lower the computational complexity of the tone reservation method, this paper proposes approximate algorithms and their implementation structures. Simulation results show that the SER performance of the proposed structure is similar to that of the conventional one. Compared to the conventional structure, the proposed structure reduces hardware complexity and power consumption by 62.2% and 58.4%, respectively. Kangwoo Park, In-Cheol Park |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2011 | QC-LDPC Decoding Architecture based on Stride SchedulingabstractIn this paper, an area-efficient decoder architecture is proposed for the quasi-cyclic low-density parity check (QC-LDPC) codes specified in the WiMAX 802.16e standard. In order to achieve low area and maximize hardware utilization, the decoder utilizes 4 decoding function units, which is the greatest common divisor of the expansion factors. Furthermore, the decoder adopts a novel scheduling scheme, named as stride scheduling, to remove the conventional flexible permutation network and also minimize the number of memory accesses. The synthesized decoder costs 49K of logic gates and 54,144 bits of memory, while maintaining the throughput over the requirement of the WiMAX. Bongjin Kim, In-Cheol Park |
ISCAS | 2 |
| 2010 | Dual-rail decoding of low-density parity-check codesabstractIn this paper, a new scheduling scheme is proposed to increase the throughput of a low-density parity-check decoder by maximizing resource utilization. The operations of check nodes and variable nodes are fully overlapped in the proposed scheduling to achieve maximized utilization of hardware resources, which in turn increases the throughput and reduces the overall decoding latency. Moreover, no restriction is posed on the formation of the parity check matrix. To verify the effectiveness of the proposed scheme, a series of simulations is performed for irregular random LDPC codes with considering additive white Gaussian noise channel. Bongjin Kim, Hasan Ahmed, In-Cheol Park |
ISCAS | 3 |
| 2010 | Capacitor array structure and switching control scheme to reduce capacitor mismatch effects for SAR analog-to-digital convertersabstractThis paper presents a new capacitor array structure developed for SAR analog-to-digital converters and its switching algorithm that can alleviate capacitor mismatch effects. The capacitor mismatch can induce many missing codes. The proposed capacitor array structure is based on the junction-splitting method as is efficient in terms of power consumption. To reduce the capacitor mismatch effects, two capacitor arrays are employed to enable redundant search. Simulation results show that the proposed method significantly reduces the number of missing codes caused by the capacitance mismatch, and reduces the energy consumption by more than 70% compared to the conventional charge redistribution method. Youngjoo Lee 0002, In-Cheol Park |
ISCAS | 2 |
| 2010 | Low-complexity tone reservation method for PAPR reduction of OFDM systemsabstractThe OFDM communication system has a serious drawback that the peak signal value can be much higher than the average signal value. High peak signals can be easily distorted by the nonlinearity of power amplifiers, and can increase the symbol error rate (SER) significantly. Though many techniques have been proposed to reduce the peak-to-average-power ratio (PAPR), they are usually based on iterative FFT computation, needing lots of computation time. Based on the tone reservation method, this paper proposes a low complexity DFT structure in order to lower the complexity of PAPR reduction techniques. In the proposed method, several approximations are employed to reduce the computational complexity and to substitute complex multiplications with simple shift operations. The performance of the proposed tone reservation method is compared with that of the conventional method based on the radix-2 FFT. Simulation results show that there is almost no performance degradation if the radix-2 FFT is replaced with the proposed approximate DFT. Kangwoo Park, In-Cheol Park |
ISCAS | 2 |
| 2010 | Small-area and low-energy K-best MIMO detector using relaxed tree expansion and early forwardingabstractThis paper proposes a new K-best detection method that can realize a small-area and low-energy MIMO detector. To reduce the complexity required in K-best operations, tree expansion is relaxed by early evicting inferior children. Additionally, an efficient pipeline scheduling called early forwarding is proposed to reduce the overall processing latency and the number registers. A 4×4 16-QAM MIMO detector integrating four K-best detection units is implemented to demonstrate the proposed method. In a 0.18-μm CMOS technology, the entire detector occupies 1.9 mm2 and shows a throughput of 584 Mbps. The energy consumption is 443 pJ per bit at 1.8 V. In-Cheol Park |
ISLPED | 2 |
| 2010 | Optimization of Arithmetic Coding for JPEG2000abstractEmbedded block coding with optimized truncation (EBCOT) employed in the JPEG2000 standard accounts for the majority of the processing time, because the EBCOT is full of bit operations that cannot be implemented efficiently in software. The block coder consists of a bit-plane coder (BPC) followed by a binary arithmetic coder (BAC), where the most up-to-date BPC architectures are capable of producing symbols at a much higher rate than the conventional BACs can handle. This letter proposes a novel pipelined BAC architecture that can encode input symbols at a much higher rate than the conventional BAC architectures. The proposed architecture can significantly reduce the critical path delay and can achieve a throughput of 400 Msymbols/s. The critical path delay synthesized with 0.18 ¿m CMOS technology is 2.42 ns, which is almost half of the delay taken in conventional BAC architectures. Minsoo Rhu, In-Cheol Park |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Design of a Scalable and Programmable Sound SynthesizerabstractSound synthesis employed in many multimedia systems is a useful method to generate the sound of musical instruments. Although it has a long history of development, a few researches have been devoted to deriving efficient VLSI architectures. In this paper, we analyze the inherent dataflow of sound synthesis methods and propose a programmable VLSI architecture suitable for a scalable sound synthesizer. The sound quality and the level of polyphony can be enhanced only by increasing the operating speed and enlarging the memory. A fully integrated sound synthesis system is implemented as a prototype to verify the proposed architecture. The prototype chip fabricated in a 0.18- m CMOS process occupies 1.5 mm 1.5 mm, and can synthesize a 64-polyphonic sound in real time. The power consumption ranges from 2.05 to 13.8 mW depending on the level of polyphony and the sound quality. Youngjoo Lee 0002, In-Cheol Park |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2009 | Implementation of a High-Throughput and Area-Efficient MIMO Detector Based on Modified Dijkstra's SearchabstractThis paper presents a VLSI implementation of a high-throughput and area-efficient MIMO detector. We propose a modified Dijkstra's algorithm and a pre-calculation technique to improve the throughput by allowing overlapped processing. In addition, we propose a simple approximation of L2-norm to reduce the computational complexity without degrading the error performance noticeably. A MIMO detector based on the proposed algorithm is implemented using a 0.18-¿m CMOS technology, which occupies 0.49 mm2with 25.IK equivalent gates and shows a throughput of over 300 Mbps. In-Cheol Park |
GLOBECOM | 2 |
| 2009 | Multiplier-less and table-less linear approximation for square and square-rootabstractSquare and square-root are widely used in digital signal processing and digital communication algorithms, and their efficient realizations are commonly required to reduce the hardware complexity. In the implementation point of view, approximate realizations are often desired if they do not degrade performance significantly. In this paper, we propose new linear approximations for the square and square-root functions. The traditional linear approximations need multipliers to calculate slope offsets and tables to store initial offset values and slope values, whereas the proposed approximations exploit the inherent properties of square-related functions to linearly interpolate with only simple operations, such as shift, concatenation and addition, which are usually supported in modern VLSI systems. Regardless of the bit-width of the number system, more importantly, the maximum relative errors of the proposed approximations are bounded to 6.25% and 3.13% for square and square-root functions, respectively. In-Cheol Park |
ICCD | 1 |
| 2009 | Architecture design of a high-performance dual-symbol binary arithmetic coder for JPEG2000abstractThe embedded-block coding with optimized truncation (EBCOT), which consists of a bit-plane coder (BPC) and a binary arithmetic coder (BAC), is the bottleneck in realizing a high-performance JPEG2000 encoding system due to its characteristics of bit-wise processing. Although efficient architectures for BPCs have been presented, the performance of EBCOT is mainly restricted by BAC because of its sequential processing nature. In this paper, we propose a novel architecture for BAC which is based on our optimization technique named as trace pipelining. It enables parallel processing of the usual byte-out cases in BAC, and can achieve a throughput of 534 M symbols/sec, which is the highest compared to those of the previous BAC architectures. Minsoo Rhu, In-Cheol Park |
ICIP | 2 |
| 2009 | Memory-less bit-plane coder architecture for JPEG2000 with concurrent column-stripe codingabstractIn implementing an efficient block coder for JPEG2000, the memories required for storing the state variables dominate the hardware cost of a block coder. In this paper, we propose a novel bit-plane coder (BPC) architecture that derives all the state variables on the fly, thereby eliminating the memory requirement. In addition, we present a concurrent column-stripe coding algorithm which merges the scanning of all three coding passes into a single context window to generate all relevant context outputs concurrently in a single clock cycle. Experimental results show that the memory requirement and overall hardware cost of the proposed BPC are much smaller than those of previous architectures. Furthermore, as the column-stripe can be encoded for all three passes in a single clock cycle, a minimum of four context outputs are generated per cycle. Therefore, the following arithmetic coder that encodes the BPC outputs can never be in the idle state, enabling fast computation of overall block coding. Minsoo Rhu, In-Cheol Park |
ICIP | 2 |
| 2009 | A Scalable and Programmable Sound SynthesizerabstractSound synthesis employed in many multimedia systems is a useful method to generate sounds of musical instruments. In this paper, we propose a new VLSI architecture suitable for a scalable sound synthesizer based on a programmable data-flow. The sound quality and the level of polyphony can be enhanced only by increasing operating speed and enlarging memory. A fully integrated sound synthesis system is implemented as a prototype to verify the proposed architecture. The prototype chip fabricated in a 0.18-mum CMOS process occupies 1.5 times 1.5 mm2, and can synthesize up to 64-polyphonic sound. The power consumption ranges from 2.05 mW to 13.8 mW depending on the quality of the synthesized sound. Youngjoo Lee 0002, In-Cheol Park |
ISCAS | 3 |
| 2009 | Fast Frequency Acquisition Phase Frequency Detectors with Prediction-based edge BlockingabstractThis paper presents a new phase frequency detector (PFD) to enable fast frequency acquisition in the phase-locked loop (PLL). The three-state PFD is conventionally employed because it is simple and almost immune to the dead-zone problem, but it can miss some rising edges when the edges come during the reset time. Eliminating or reducing the missing edges caused by the reset pulse is essential in achieving fast acquisition. To cope with the missing edge problem, the proposed PFD predicts the reset signal and blocks the corresponding input signal during the reset time. The blocked edge is regenerated after the reset signal is deactivated. Experimental results show that the proposed PFD works correctly for the entire phase difference and achieves 42.1% speed-up in the acquisition time when it is applied to the conventional charge pump PLL implemented in a 0.18 mum CMOS technology. Kangwoo Park, In-Cheol Park |
ISCAS | 2 |
| 2009 | Novel Pipelined DWT Architecture for Dual-line ScanabstractA new discrete wavelet transform (DWT) architecture is proposed in this paper to realize a memory-efficient 2D DWT unit. The proposed DWT architecture alternately processes two lines to remove the transpose buffer whose size is proportional to the image row size. As a result, the hardware complexity of 2D DWT is significantly reduced. To maintain the same critical path delay as that of the previous pipelined DWT, serially concatenated additions are optimized by changing computation topology and applying arithmetic optimization. Jinook Song, In-Cheol Park |
ISCAS | 2 |
| 2008 | Duo-binary circular turbo decoder based on border metric encoding for WiMAXabstractThis paper presents a duo-binary circular turbo decoder based on border metric encoding. With the proposed method, the memory size for branch memory is reduced by half and the dummy calculation is removed at the cost of the small-sized memory which holds the encoded border metrics. Based on the proposed SISO decoder and the dedicated hardware interleaver, a duo-binary circular turbo decoder is designed for the WiMAX standard using a 0.13 mum CMOS process, which can support 24.26 Mbps at 200 MHz. Ji-Hoon Kim 0003, In-Cheol Park |
ASP-DAC | 2 |
| 2008 | Area and power efficient design of coarse time synchronizer and frequency offset estimator for fixed WiMAX systemsabstractTargeting fixed WiMAX systems, this paper presents a new architecture for coarse time synchronization and carrier frequency offset (CFO) estimation. The proposed architecture is based on a two-step approach where the data-paths are decoupled to individually optimize performance and area. Implemented with 0.13μm CMOS technology, the results show that the proposed architecture has advantages of less silicon area and power consumption as well as better performance compared to the previous joint approach. In-Cheol Park |
ASP-DAC | 2 |
| 2008 | Digital filter synthesis considering multiple adder graphs for a coefficientabstractIn this paper, a new FIR digital filter synthesis algorithm is proposed to consider multiple adder graphs for a coefficient. The proposed algorithm selects an adder graph that can be maximally sharable with the remaining coefficients, while previous dependence-graph algorithms consider only one adder graph when implementing a coefficient. In addition, we propose an addition reordering technique to reduce the computational overhead of finding multiple adder graphs. By using the proposed technique, multiple adder graphs are efficiently generated from a seed adder graph obtained by using previous dependence-graph algorithms. Experimental results show that the proposed algorithm reduces the hardware cost of FIR filters by 23% and 3.4% on average compared to the Hartely and RAGn-hybrid algorithms. In-Cheol Park |
ICCD | 2 |
| 2008 | Fast frequency acquisition all-digital PLL using PVT calibrationabstractFast frequency acquisition is crucial for phase-locked loops (PLLs) used in portable devices, as on-chip clocks are frequently scaled down or up in order to manage power consumption. This paper describes a new frequency acquisition method that is effective in all-digital PLLs (ADPLLs). To achieve fast frequency acquisition, the codeword of the digitally controlled oscillator (DCO) is predicted by measuring the variations of process, supply voltage and temperature (PVT). A PVT sensor implemented with a ring oscillator is employed to monitor the variations. As the sensor frequency at the current operating condition is directly related to the PVT variations, the sensor frequency is taken into account to compensate such variations in predicting the DCO codeword. The proposed method enables one-cycle frequency acquisition, and the frequency error is less than 1.5%. The proposed ADPLL implemented in a 0.18μm CMOS process operates from 150MHz to 500MHz and occupies 0.075mm2. Hae-Soo Jeon, Duk-Hyun You, In-Cheol Park |
ISCAS | 3 |
| 2008 | Capacitor array structure and switch control for energy-efficient SAR analog-to-digital convertersabstractThis paper presents a new capacitor array structure and its switch control method for binary weighted SAR analog-to-digital converters, which can significantly lower the energy consumed in charge redistribution steps. The proposed method is analyzed theoretically and simulations are performed to verify the theoretical analysis. Simulation results show that the proposed capacitor array structure and switching method can reduce the average energy consumed in the capacitor array by 75% and 60% compared to the conventional method and the splitting capacitor method, respectively. Jeong-Sup Lee, In-Cheol Park |
ISCAS | 2 |
| 2008 | Prediction-based real-time CABAC decoder for high definition H.264/AVCabstractThis paper proposes a prediction scheme to decode in real-time H.264/AVC bitstream coded in Context-based Adaptive Binary Arithmetic Coding (CABAC). The proposed scheme predicts the subsequent syntax element type, leading to a significant reduction in cycles invoked by the syntax element switching overhead that degrades the decoding performance significantly. In addition, a new pipelined architecture combined with the proposed scheme is presented for hardware implementation. The simulation results show that the proposed scheme achieves a decoding performance of 1.2 cycles/bin, which releases the CABAC parsing from a bottleneck in decoding H.264 bitstream. WonHee Son, In-Cheol Park |
ISCAS | 2 |
| 2008 | Time-Domain Joint Estimation of Fine Symbol Timing Offset and Integer Carrier Frequency OffsetabstractIn this paper, we propose an efficient synchronization method to jointly estimate fine symbol timing offset (STO) and integer carrier frequency offset (CFO). The proposed method is to perform cross-correlation between received samples and pre-rotated training sequences. Experimental results on IEEE 802.16d systems show that the proposed method is significantly superior to the previous approaches in both estimations. Since the fine STO and the integer CFO are jointly estimated in the time domain in an on-the-fly manner, the proposed method requires no additional buffers. Moreover, the proposed joint estimation has an effect of eliminating the processing ordering dependency of the two estimations, being attractive for systems requiring tight synchronization. In-Cheol Park |
VTC Spring | 2 |
| 2008 | FIR Filter Synthesis Considering Multiple Adder Graphs for a CoefficientabstractTo reduce the hardware complexity of finite-impulse response (FIR) digital filters, this paper proposes a new filter synthesis algorithm. Considering multiple adder graphs for a coefficient, the proposed algorithm selects an adder graph that can be maximally sharable with the remaining coefficients, whereas previous dependence-graph algorithms consider only one adder graph when implementing a coefficient. In addition, an addition reordering technique is proposed to derive multiple adder graphs from a seed adder graph generated by using previous dependence-graph algorithms. Experimental results show that the proposed algorithm reduces the hardware cost of FIR filters by 22% and 3.4%, on average, compared to the Hartley and -dimensional reduced adder graph hybrid algorithms, respectively. In-Cheol Park |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2008 | Low-Power and High-Accurate Synchronization for IEEE 802.16d SystemsabstractOrthogonal frequency division multiplexing (OFDM) is a viable technology for high-speed data transmission by virtue of its spectral efficiency and robustness to multi-path fading. These advantages can be achieved only with good synchronization both in time and frequency. This paper proposes new efficient synchronization methods for an OFDM-based system, IEEE 802.16d. For the coarse time synchronization and the fractional carrier frequency offset (CFO) estimation, a disjoint architecture is proposed that performs auto-correlations separately to achieve more reliable frequency synchronization and to reduce overall hardware complexity and power consumption. In addition, for the fine symbol timing offset (STO) and the integer CFO, a new joint estimation method employing parallel cross-correlations between the received samples and the pre-rotated training sequences is proposed. Experimental results show significantly superior performance to the previous synchronization methods. A prototype synchronizer based on the proposed methods is designed with a 0.25-mum CMOS process, which reduces power consumption by more than 60% compared to a conventional synchronizer. In-Cheol Park |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2007 | Twiddle factor transformation for pipelined FFT processingabstractThis paper presents a novel transformation technique that can derive various fast Fourier transform (FFT) in a unified paradigm. The proposed algorithm is to find a common twiddle factor at the input side of a butterfly and migrate it to the output side. Starting from the radix-2 FFT algorithm, the proposed common factor migration technique can generate most of previous FFT algorithms without using mathematical manipulation. In addition, we propose new FFT algorithms derived by applying the proposed twiddle factor moving technique, which reduce the number of twiddle factors significantly compared with the previous algorithms being widely used for pipelined FFT processing. In-Cheol Park, WonHee Son, Ji-Hoon Kim 0003 |
ICCD | 1 |
| 2007 | High Speed Sphere Decoding Based on Vertically Incremental ComputationabstractSphere decoding enables maximum likelihood (ML) detection with lower complexity than other decoding algorithms, but it still suffers from large computational delay. This paper proposes a vertical partial Euclidean distance (PED) computation method to reduce the critical path delay and computational resources. Since the proposed method computes ahead the PED of lower levels using upper level symbols, a high speed PED computation unit can be implemented with less hardware resources Se-Hyeon Kang, In-Cheol Park |
ISCAS | 2 |
| 2007 | Energy-Efficient Double-Binary Tail-Biting Turbo Decoder Based on Border Metric EncodingabstractThis paper presents an energy-efficient turbo decoder based on border metric encoding, which is especially suitable for non-binary circular turbo codes. In the proposed method, the size of the branch memory is reduced to half and the dummy calculation is removed at the cost of a small-sized memory that holds encoded border metrics. Due to the small size and infrequent access to the border memory, power consumption for soft-input soft-output (SISO) decoding is reduced by 26.0%. Based on the proposed SISO decoder and the dedicated hardware interleaver, a double-binary tail-biting turbo decoder is designed for WiMAX standard using a 0.18 μm CMOS process and it can support 12.14Mbps at operating frequency of 100MHz. Ji-Hoon Kim 0003, In-Cheol Park |
ISCAS | 2 |
| 2007 | Tiled Interleaving for Multi-Level 2-D Discrete Wavelet TransformabstractThis paper presents a new architecture of 2D discrete wavelet transform (DWT) proposed for JPEG 2000. In the proposed architecture, the image is segmented into tiles each of which is sequentially processed to minimize the size of buffers required to process 2D DWT, and multi-level DWTs are interleaved to reduce the size of the repeat buffer drastically. Compared to the conventional architecture, the overall memory size is reduced by 85% and 92% for 256×256 and 512×512 images, respectively. The proposed DWT processor needs only 5kB memory for 256×256 images, and operates at 250MHz in 0.25μ technology. Jung-Wook Kim, Jinook Song, Seokho Lee, In-Cheol Park |
ISCAS | 4 |
| 2007 | Fast and Area-Efficient Sphere Decoding Using Look-Ahead SearchabstractSphere decoding enables maximum likelihood (ML) detection with fairly low complexity in the MIMO wireless systems, but it takes hundreds cycles at low SNR environment. This paper proposes a fast decoding algorithm to reduce the decoding cycles using look-ahead search. Since the proposed decoding algorithm utilizes hardware resources to add other sub-trees into the search space successively, it helps not to go down into the sub-tree that has a small value at the root node but has large values at the child nodes. Scaling and enumeration techniques are also presented, which are effective in implementing the proposed sphere decoder. As a result, the proposed decoder saves about 30% decoding cycles at the cost of small hardware overhead compared to the conventional decoder. Se-Hyeon Kang, In-Cheol Park |
VTC Spring | 2 |
| 2007 | High-Speed H.264/AVC CABAC DecodingabstractThe decoding of context-based adaptive binary arithmetic coding (CABAC) imposes a heavy performance requirement on H.264/AVC decoding systems particularly for large-scale video sequences. As a simple approach of elevating the operating frequency is not sufficient to meet the performance requirement, this paper proposes an efficient approach to accelerate the decoding, which is effective under relatively low operating frequency. Since the CABAC decoding procedure is highly sequential and has strong data dependencies, it is difficult to exploit parallelism and pipeline schemes. The proposed approach resolves the difficulties by modifying the operation chain based on a thorough analysis, eventually enabling both parallel operations and pipelining. More specifically, 1) several context models are simultaneously loaded from memory while context selection is performed in parallel and 2) bin-level pipelining is enabled by employing a small storage to remove structural hazards and data dependencies. Experimental results show that the proposed approach leads to the real-time decoding of HD sequences Y. Yi, In-Cheol Park |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | SIMD Processor-Based Turbo Decoder Supporting Multiple Third-Generation Wireless StandardsabstractA programmable turbo decoder is designed to support multiple third-generation wireless communication standards. We propose a hybrid architecture of hardware and software, which has small size, low power, and high performance like hardware implementations, as well as the flexibility and programmability of software. It mainly consists of a configurable hardware soft-input-soft-output (SISO) decoder and a 16-b single-instruction multiple-data processor, which is equipped with five processing elements and special instructions customized for interleaving in order to provide interleaved data at the speed of the hardware SISO. A fast and flexible software implementation of the block interleaving algorithm is also proposed. The interleaver generation is split into two parts, preprocessing and on-the-fly generation, to reduce the timing overhead of changing the interleaver structure. We present detailed descriptions of the interleaving implementation applied to the W-CDMA and cdma2000 standard turbo codes. The decoder occupies 8.90$~$mm$^{2}$in a 0.25-$\mu$m CMOS with five metal layers and exhibits the maximum decoding rate of 5.48$~$Mb/s. Myoung-Cheol Shin, In-Cheol Park |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | Low-power hybrid turbo decoding based on reverse calculationabstractAs turbo decoding is a highly memory-intensive algorithm consuming large power, a major issue to be solved in practical implementation is to reduce power consumption. This paper presents an efficient reverse calculation method to lower the power consumption by reducing the number of memory accesses required in turbo decoding. The reverse calculation method is proposed for the max-log-MAP algorithm, and it is combined with a scaling technique to achieve a new decoding algorithm, called hybrid log-MAP, that results in a similar BER performance to the log-MAP algorithm. For the W-CDMA standard, experimental results show that 80% of memory accesses are reduced through the proposed reverse calculation method. A hybrid log-MAP turbo decoder based on the proposed reverse calculation reduces power consumption and memory size by 34.4% and 39.2%, respectively. Hye-Mi Choi, Ji-Hoon Kim 0003, In-Cheol Park |
ISCAS | 3 |
| 2006 | High speed decoding of context-based adaptive binary arithmetic codes using most probable symbol predictionabstractContext-based adaptive binary arithmetic coding (CABAC) is the major entropy-coding algorithm employed in H.264/AVC. Although the performance gain of H.264/AVC is mostly resulted from CABAC, it is difficult to achieve a fast decoder because the decoding algorithm is basically sequential. In this paper, a prediction scheme is proposed to enhance overall decoding performance by decoding two binary symbols at a time. A CABAC decoder based on the proposed prediction scheme improves the decoding performance by 24% compared to conventional decoders. Chung-Hyo Kim, In-Cheol Park |
ISCAS | 2 |
| 2006 | Combined image signal processing for CMOS image sensorsabstractThis paper presents an efficient image signal processing structure for CMOS image sensors to achieve low area and power consumption. Although CMOS image sensors (CISs) have various benefits compared with charge-coupled devices (CCDs), the images obtained from CISs have much lower quality than those from CCDs. To improve the quality of CIS images, it is required to do reproducing and enhancing processings such as color interpolation, white balancing, color correction, gamma correction and color conversion. They are implemented individually in most conventional designs though they have similar functional characteristics. In this proposed structure, the gamma correction block is moved to the front in order to combine several image signal processings into one block. An efficient compensation scheme is also proposed to reduce the errors caused by the moving of the nonlinear gamma correction. A prototype CIS image signal processor is implemented in Verilog-HDL and synthesized with 0.18/spl mu/m standard cell library. Experimental results show that the proposed structure reduces area and power consumption by 23.8% and 31.1%, respectively. Kimo Kim, In-Cheol Park |
ISCAS | 2 |
| 2005 | SAT-based unbounded symbolic model checkingabstractThis paper describes a Boolean satisfiability checking (SAT)-based unbounded symbolic model-checking algorithm. The conjunctive normal form is used to represent sets of states and transition relation. A logical operation on state sets is implemented as an operation on conjunctive normal form formulas. A satisfy-all procedure is proposed to compute the existential quantification required in obtaining the preimage and fix point. The proposed satisfy-all procedure is implemented by modifying a SAT procedure to generate all the satisfying assignments of the input formula, which is based on new efficient techniques such as line justification to make an assignment covering more search space, excluding clause management, and two-level logic minimization to compress the set of found assignments. In addition, a cache table is introduced into the satisfy-all procedure. It is a difficult problem for a satisfy-all procedure to detect the case that a previous result can be reused. This paper shows that the case can be detected by comparing sets of undetermined variables and clauses. Experimental results show that the proposed algorithm can check more circuits than binary decision diagram-based and previous SAT-based model-checking algorithms. Hyeong-Ju Kang, In-Cheol Park |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2003 | SAT-based unbounded symbolic model checkingabstractThis paper describes a SAT-based unbounded symbolic model checking algorithm. BDDs have been widely used for symbolic model checking, but the approach suffers from memory overflow. The SAT procedure was exploited to overcome the problem, but it verified only the states reachable through a bounded number of transitions. The proposed algorithm deals with unbounded symbolic model checking. The conjunctive normal form is used to represent sets of states and the transition relation, and a SAT procedure is modified to compute the existential quantification required in obtaining a pre-image. Some optimization techniques are exploited, and the depth first search method is used for efficient safety-property checking. Experimental results show the proposed algorithm can check more circuits than BDD-based symbolic model checking tools. Hyeong-Ju Kang, In-Cheol Park |
DAC | 2 |
| 2003 | Low-power hybrid structure of digital matched filters for direct sequence spread spectrum systemsabstractThe paper presents a low-power structure of digital matched filters (DMFs), which is proposed for direct sequence spread spectrum systems. Traditionally, low-power approaches for DMFs are based on either the transposed-form structure or the direct-form one. A new hybrid structure that employs the direct-form structure for local addition and the transposed-form structure for global addition is used to take advantage of both structures. For a 128-tap DMF, the proposed DMF that processes 32 addends a cycle consumes 46% less power at the expense of 6% area overhead as compared to the state-of-the-art low-power DMF (Liou, M. and Chiueh, T., IEEE J. Solid-State Circuits, vol.31, p.933-43, 2001). Sungwon Lee 0005, In-Cheol Park |
ICASSP (2) | 2 |
| 2003 | Fast Cycle-accurate Behavioral Simulation for Pipelined Processors Using Early Pipeline Evaluation
In-Cheol Park, Se-Hyeon Kang, Yongseok Yi |
ICCAD | 1 |
| 2003 | Low-power hybrid structure of digital matched filters for direct sequence spread spectrum systemsabstractThis paper presents a low-power structure of digital matched filters (DMFs), which is proposed for direct sequence spread spectrum systems. Traditionally, low-power approaches for DMFs are based on either the transposed-form structure or the direct-form one. A new hybrid structure that employs the direct-form structure for local addition and the transposed-form structure for global addition is used to take advantages of both structures. For a 128-tap DMF, the proposed DMF that processes 32 addends a cycle consumes 46 % less power at the expense of 6 % area overhead as compared to the state-of-the-art low-power DMF [M. Liou et al., 2001]. Sungwon Lee 0005, In-Cheol Park |
ICME | 2 |
| 2003 | Timed compiled-code functional simulation of embedded software for performance analysis of SOC designabstractA new timing generation method is proposed for the performance analysis of embedded software. The time stamp generation of input/output (I/O) accesses is crucial to performance estimation and architecture exploration in the timed functional simulation that simulates the whole design at a functional level with timing. A portable compiler is modified to generate time deltas which are the estimated cycle counts between two adjacent I/O accesses by counting the cycles of the intermediate representation (IR) operations and using a machine description that contains information on a target processor. Since the proposed method is based on the machine-independent IR of a compiler, the method can be applied to various processors by changing the machine description. The experimental results show that the proposed method is effective in that the average estimation error is about 2% and the maximum speed-up over the corresponding instruction-set simulators is about 300 times. The proposed method is also verified in a timed functional simulation environment. Jong-Yeol Lee, In-Cheol Park |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2003 | Address code generation for DSP instruction-set architecturesabstractThis paper presents a new DSP-oriented code optimization method to enhance performance by exploiting the specific architectural features of digital signal processors. In the proposed method, a source code is translated into the static single assignment form while preserving the high-level information related to the address computation of array accesses. The information is used in generating auto-modification addressing operations provided by most digital signal processors. In addition to the conventional control-data flow graph, a new graph is employed to find auto-modification addressing modes efficiently. Experimental results on benchmark programs show that the proposed method is effective in improving performance and reducing code size. Jong-Yeol Lee, In-Cheol Park |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2002 | Timed compiled-code simulation of embedded software for performance analysis of SOC designabstractIn this paper, a new timing generation method is proposed for the performance analysis of embedded software. The time stamp generation of I/O accesses is crucial to performance estimation and architecture exploration in the timed functional simulation, which simulates the whole design at a functional level with timing. A portable compiler is modified to generate time-deltas, which are the estimated cycle counts between two adjacent I/O accesses, by counting the cycles of the intermediate representation (IR) operations and using a machine description that contains information on a target processor. Since the proposed method is based on the machine-independent IR of a compiler, the method can be applied to various processors by changing the machine description. The experimental results show that the proposed method is effective in that the average estimation error is about 2% and the maximum speed-up over the corresponding instruction-set simulators is about 300 times. The proposed method is also verified in a timed functional simulation environment. Jong-Yeol Lee, In-Cheol Park |
DAC | 2 |
| 2002 | A high-speed and low-latency Reed-Solomon decoder based on a dual-line structureabstractThis paper presents a new decoding structure of Reed-Solomon ( RS) codes that are widely used for channel coding. Although many decoding structures have been developed, the serial structures have long latency and the parallel structures are not fast enough to deal with the demands of high-speed decoding. To achieve both short latency and fast ope,ration, the summation of the products of syndromes is eliminated and the difference used to calculate the error locator polynomial is incrementally updated. The proposed structure called a dual-line structure can operate as fast as the serial structure and has as short latency as the parallel structure. In addition, the dual-line structure is regular and easy to implement. Experimental results confirm these advantages at the cost of a small hardware increase. Hyeong-Ju Kang, In-Cheol Park |
ICASSP | 2 |
| 2002 | Interface synthesis between software chip model and target board
Seungjong Lee, Ando Ki, In-Cheol Park, Chong-Min Kyung |
J. Syst. Archit. | 3 |
| 2002 | Digital filter synthesis based on an algorithm to generate all minimal signed digit representationsabstractIn this paper, the authors propose an algorithm to find all the minimal signed digit (MSD) representations of a constant and present an algorithm to synthesize digital filters based on the MSD representation. The hardware complexity of a digital signal processing system is dependent on the number system used for the implementation. Although the canonical signed digit (CSD) representation is widely employed, as it is unique and guarantees the minimal number of nonzero digits for a constant, the MSD representation provides multiple representations that have the same number of nonzero digits as the CSD representation. The proposed filter synthesis algorithm utilizes this redundancy of the MSD representation to make common subexpressions, as many as possible, leading to smaller filters. By applying the proposed algorithm to the hardware synthesis of finite impulse response filters, the authors obtained multiplier blocks that are 7% smaller than those generated from the CSD representation. In-Cheol Park, Hyeong-Ju Kang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2001 | Low-power high-level synthesis using latchesabstractHigh-level synthesis using latches has many merits in power, area and even in speed. But latches cannot be read and written at the same time and usually requires two-phase non-overlapping clock that is unpleasant choice for short-term design. In this paper we propose a storage allocation method that makes it possible to use latches as storage elements in single clocking scheme. The proposed method modifies the lifetime of variables slightly so that it can be applied to any high-level synthesis systems with small modification. The experimental results show 39 ~ 65% reduction in power consumption within almost same area compared to the conventional power management scheme using clock gating. Woo-Seung Yang, In-Cheol Park, Chong-Min Kyung |
ASP-DAC | 2 |
| 2001 | Digital Filter Synthesis Based on Minimal Signed Digit RepresentationabstractAs the complexity of digital filters is dominated by the number of multiplications, many works have focused on minimizing the complexity of multiplier blocks that compute the constant coefficient multiplications required in filters. The complexity of multiplier blocks can be significantly reduced by using an efficient number system. Although the canonical signed digit representation is commonly used as it guarantees the minimal number of additions for a constant multiplication, we propose in this paper a digital filter synthesis algorithm that is based on the minimal signed digit (MSD) representation. The MSD representation is attractive because it provides a number of forms that have the minimal number of non-zero digits for a constant. This redundancy can lead to efficient filters if a proper MSD representation is selected for each constant. In experimental results, the proposed algorithm resulted in superior filters to those generated from the CSD representation. In-Cheol Park, Hyeong-Ju Kang |
DAC | 1 |
| 2001 | An Area-Efficient Iterative Modified-Booth Multiplier Based on Self-Timed ClockingabstractA new iterative multiplier based on a self-timed clocking scheme is presented. To reduce the area required for the multiplier, only two CSA rows are iteratively used to complete a multiplication. The partial CSA array is controlled by a fast internal clock generated using a self-timed technique. Compared with the array implementation, the proposed multiplier yields an 86.6% area reduction at the expense of 18.8% slow down for 64/spl times/64-bit multiplication. Myoung-Cheol Shin, Se-Hyeon Kang, In-Cheol Park |
ICCD | 3 |
| 2001 | High-performance and low-power memory-interface architecture for video processing applicationsabstractTo improve memory bandwidth and power consumption in video applications, a new memory-interface architecture is proposed. The architecture adopts an array address-translation technique to utilize the fact that video processing algorithms have regular memory-access patterns. Since the translation can minimize the number of overhead cycles needed for row-activations in synchronous DRAM (SDRAM), we can improve the memory bandwidth and energy consumption significantly. The features of SDRAM and memory-access patterns of video processing applications are considered to find a suitable address translation. Compared to the conventional linear translation, experimental results show that the proposed architecture reduces about 89% of row-activations and increases the memory bandwidth by 50%. In addition, the proposed architecture reduces the energy consumption by 30% on the average. In-Cheol Park |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2000 | A hardware accelerator for the specular intensity of phong illumination model in 3-dimensional graphicsabstractAbstract — This paper presents a special hardware implementation developed for the computation of the specular term which is the most time consuming part in the Phong's illumination. In the Phong shading, the exponentiation operation of two floating-point numbers is necessary for each point inside a polygon. An approximation algorithm is developed to speed up the exponentiation operation, and it is supported by simple hardware that can be easily merged into a floating-point multiplier. The exponentiation operation takes just 4 cycles in the proposed hardware while it takes about 100-200 cycles in conventional floating-point units. Although an approximation algorithm is employed for the exponentiation operation, the amount of error is so minimal that the difference is virtually indistinguishable. I. Young-Su Kwon, In-Cheol Park, Chong-Min Kyung |
ASP-DAC | 2 |
| 2000 | FIR Filter Synthesis Algorithms for Minimizing the Delay and the Number of AddersabstractAs the complexity of digital filters is dominated by the number of multiplications, many works have focused on minimizing the complexity of multiplier blocks that compute the constant coefficient multiplications required in filters. Although the complexity of multiplier blocks is significantly reduced by using efficient techniques such as decomposing multiplications into simple operations and sharing common subexpressions, previous works have not considered the delay of multiplier blocks which is a critical factor in the design of complex filters. In this paper, we present new algorithms to minimize the complexity of multiplier blocks under the given delay constraints. By analyzing multiplier blocks in view of delay, three delay reduction methods are proposed and combined into previous algorithms. Since the proposed algorithms can generate multiplier blocks that meet the specified delay, a trade-off between delay and hardware complexity is enabled by changing the delay constraints. Experimental results show that the proposed algorithms can reduce the delay of multiplier blocks at the cost of a little increase of complexity. Hyeong-Ju Kang, In-Cheol Park |
ICCAD | 3 |
| 2000 | Synthesis and Optimization of Interface Hardware between IP's Operating at Different Clock FrequenciesabstractIn system-on-a-chip design, interfacing of Intellectual Property (IP) blocks is one of the most important issues. Since most IPs are provided by different vendors, they have different interface schemes and different operating frequencies. In this paper, we propose a new interface synthesis method that enables one not only to handle the interface between IPs with different operating frequencies but also to minimize the hardware resource required for the interface. We have demonstrated the proposed algorithm by applying it to a real design example, MP3 decoder, and verified the IIS-to-PCI protocol converter on a real hardware system. Bong-Il Park, Hoon Choi, In-Cheol Park, Chong-Min Kyung |
ICCD | 3 |
| 2000 | Pyramid Texture Compression and Decompression Using Interpolative Vector QuantizationabstractTexture mapping is a common technique used to increase the visual quality of 3D scenes. As texture mapping requires a large amount of memory to deal with large textures generally required in the current visual systems, we propose an algorithm for compressing a pyramid texture used for mipmapping. Vector quantization is used to compress all levels of the pyramid texture to one representative value databook, one residual codebook and one index map. The proposed compression scheme uses interpolative texel difference vector quantization that compresses the difference between the interpolated surfaces generated by the representative value and the correct uncompressed texels of the texture at each level. The compressed pyramid texture can be accessed randomly and decompressed without loss of visual quality. We also propose a hardware architecture that performs the trilinear filtering with the compressed pyramid texture. Young-Su Kwon, In-Cheol Park, Chong-Min Kyung |
ICIP | 2 |
| 2000 | Array address translation for SDRAM-based video processing applications
In-Cheol Park |
VCIP | 2 |
| 2000 | Optimal down-conversion in compressed DCT domain with minimal operations
Myoung-Cheol Shin, In-Cheol Park |
VCIP | 2 |
| 2000 | MetaCore: an application-specific programmable DSP development systemabstractThis paper describes the MetaCore system which is an application-specific instruction-set processor (ASIP) development system targeted for digital signal processor (DSP) applications. The goal of the MetaCore system is to offer an efficient design methodology meeting specifications given as a combination of performance, cost, and design turnaround time. The MetaCore system consists of two major design stages: design exploration and design generation. In the design exploration stage, MetaCore system accepts a set of benchmark programs and structural/behavioral specifications for the target processor and estimates the hardware cost and performance for each hardware configuration being explored. Once a hardware configuration and instruction set are chosen, the system helps generate the target processor design in the form of hardware description language (HDL) along with the application program development tools such as C compiler, assembler, and instruction set simulator. The effectiveness of the MetaCore system was verified with a successful design of MDSP-II, a programmable DSP processor targeted for mobile communication. Jin-Hyuk Yang, Byoung-Woon Kim, Sang-Joon Nam, Young-Su Kwon, Dae-Hyun Lee, Jong-Yeol Lee, Chan-Soo Hwang, Yong Hoon Lee, Seung Ho Hwang, In-Cheol Park, Chong-Min Kyung |
IEEE Trans. Very Large Scale Integr. Syst. | 10 |
| 1999 | Node Sampling Technique to Speed Up Probability-Based Power Estimation MethodsabstractWe propose a new technique called node sampling to speed up the probability-based power estimation methods. It samples and processes only a small portion of total nodes to estimate the power consumption of a circuit. It is different from the previous speed-up techniques for probability-based methods in that the previous techniques reduce the processing time for each node while our method reduces the number of nodes actually processed. In addition, it is also different from the previous statistical sampling simulation techniques for simulation-based methods in that the previous methods sample the input vectors while our method samples the nodes in the network. The experimental results are very encouraging. The proposed method shows on the average more than 80% and 60% reductions of simulation run time under 20% and 5% error bounds, respectively. Hoon Choi, In-Cheol Park, Seung Ho Hwang, Chong-Min Kyung |
ASP-DAC | 3 |
| 1999 | A New Single-Clock Flip-Clop for Half-Swing ClockingabstractWe propose a new flip-flop configuration which saves about 60% of total clocking power using a half-swing clock. To use the half-swing clock, level converters or special clock drivers are traditionally required and the power consumptions of this logic cannot be ignored. In the proposed scheme, only NMOS devices are clocked with a half-swing clock in order to make it operate without the level converter or any other additional logics, and the random logic circuits except for the clock and flip-flops are supplied by V/sub cc/ while the clock network is supplied by V/sub cc//2. Compared to the conventional scheme, a great amount of power consumed in clocking which is responsible for a large portion of total chip power can be saved with the proposed new flip-flop configuration. Young-Su Kwon, Bong-Il Park, In-Cheol Park, Chong-Min Kyung |
ASP-DAC | 3 |
| 1999 | Verification of a Microprocessor Using Real World ApplicationsabstractIn this paper, we describe a fast and convenient verification method-ology for microprocessor using large-size, real application pro-grams as test vectors. The verification environment is based on automatic consistency checking between the golden behavioral ref-erence model and the target HDL model, which are run in an hand-shaking fashion. In conjunction with the automatic comparison facility, a new HDL saver is proposed to accelerate the verifica-tion process. The proposed saver allows 'restart ' from the nearest checkpoint before the point of inconsistency detection regardless of whether any modification on the source code is made or not. It is to be contrasted with conventional saver that does not allow restart when some design change, or debugging is made. We have proved the effectiveness of the environment through applying it to a real-world example, i.e., Pentium-compatible processor design process. It was shown that the HDL verification with the proposed saver can be faster and more flexible than the hardware emulation approach. In short, it was demonstrated that restartability with source code modification capability is very important in obtaining the short de-bugging turnaround time by eliminating a large number of redun-dant simulations. 1 You-Sung Chang, Seungjong Lee, In-Cheol Park, Chong-Min Kyung |
DAC | 3 |
| 1999 | Exploiting Intellectual Properties in ASIP Designs for Embedded DSP SoftwareabstractThe growing requirements on the correct design of a highperformance system in a short time force us to use IP's in many designs. In this paper, we propose a new approach to select the optimal set of IP's and interfaces to make the application program meet the performance constraints in ASIP designs. The proposed approach selects IP's with considering interfaces and supports concurrent execution of parts of task in kernel as software code with others in IP's, while the previous state-of-the-art approaches do not consider IP's and interfaces simultaneously and cannot support the concurrent execution. The experimental results on real applications show that the proposed approach is effective in making application programs meet the performance constraints using IP's. Hoon Choi, Ju Hwan Yi, Jong-Yeol Lee, In-Cheol Park, Chong-Min Kyung |
DAC | 4 |
| 1999 | Customization of a CISC Processor Core for Low-Power ApplicationsabstractThis paper describes a core-customization process of a CISC processor core for a given application program. It aims at the power reduction in the CISC processor core by fully utilizing the microcode-based control scheme, that is one of the most characterizing features of a CISC processor The optimization process includes two key techniques, generation of application-specific complex instructions (ASCI) and low-power-oriented microcode-ROM compilation, which independently operate at the two different levels of optimization. As a means of architectural level of optimization, application-specific complex instructions are generated so as to reduce the activities of fetch and decode units, and in the point of physical level of optimization, the microcode-ROM is compiled to reduce the bit-line toggling for each microcode-ROM access. Our experimental results based on transistor-level simulation show the proposed techniques can jointly reduce the total power consumption of the CISC processor core by up to 41%. You-Sung Chang, Bong-Il Park, In-Cheol Park, Chong-Min Kyung |
ICCD | 3 |
| 1999 | A Regular Layout Structured Multiplier Based on Weighted Carry-Save AddersabstractA new parallel array multiplier based on a new circuit called a weighted carry-save adder (WCSA) is presented in this paper. Each row of the array consists of a (n+3) bit carry-save adder and one WCSA. Since the proposed WCSA enables the multiplier to be very regular as well as to have less operation complexity at the final addition stage than that of conventional implementations, the proposed WCSA is better suited for hardware implementation. Compared with the previous implementations, the proposed multiplier yields an area reduction of 21% for 64/spl times/64 multiplication. A 16/spl times/16 multiplier implemented in 0.8 /spl mu/m CMOS DLM technology functions at more than 60 MHz. The chip is 1.04/spl times/1.15 mm/sup 2/ with 7877 transistors. Bong-Il Park, In-Cheol Park, Chong-Min Kyung |
ICCD | 2 |
| 1999 | Synthesis of Application Specific Instructions for Embedded DSP SoftwareabstractApplication specific instructions play an important role in reducing the required code size and increasing performance in embedded DSP systems. This paper describes a new approach to generate application specific instructions for DSP applications. The proposed approach is based on a modified subset-sum problem and supports multicycle complex instructions, as well as single-cycle instructions, while the previous state-of-the-art approaches generate only the single-cycle instructions or just select instructions from the fixed super-set of possible instructions. In addition, the proposed approach can also be applied to the case that instructions are predefined. Experimental results on real applications show that Various given constraints can be met by the generated set of application specific instructions without attaching special hardware accelerators. Hoon Choi, Jong-Sun Kim, Chi-Won Yoon, In-Cheol Park, Seung Ho Hwang, Chong-Min Kyung |
IEEE Trans. Computers | 4 |
| 1998 | Virtual Chip: Making Functional Models Work on Real Target SystemsabstractAs design complexity increases, functional verification becomes a crucial issue to ensure design correctness at an early design stage. Traditional methods for verifying functional designs are based on the HDL simulation, which is becoming the bottleneck of the design cycle because of the increasing design complexity. The accurate verification ability at the architectural level through a large set of the test programs and real world applications is a foundation for the next design step. In this paper, we describe how to verify a functional model on a real target system. The proposed methodology called virtual chip makes it possible not only to check the functional correctness on real systems, but also to explore design space by measuring the performance effectiveness of various architecture parameters under real applications. Experimental results show that functional models can be verified on real systems using complicated application programs. The proposed functional verification method is faster than HDL simulation and even comparable to emulation. Namseung Kim, Hoon Choi, Seungjong Lee, Seungwang Lee, In-Cheol Park, Chong-Min Kyung |
DAC | 5 |
| 1998 | MetaCore: An Application Specific DSP Development SystemabstractThis paper describes the MetaCore system which is an ASIP (Application-Specific Instruction set Processor) development system targeted for DSP applications. The goal of MetaCore system is to offer an efficient design methodology meeting specifications given as a combination of performance, cost and design turnaround time. Jin-Hyuk Yang, Byoung-Woon Kim, Sang-Jun Nam, Jang-Ho Cho, Chang-Ho Ryu, Young-Su Kwon, Dae-Hyun Lee, Jong-Yeol Lee, Jong-Sun Kim, Hyun-Dhong Yoon, Jae-Yeol Kim, Kun-Moo Lee, Chan-Soo Hwang, In-Hyung Kim, Jun Sung Kim, Kwang-Il Park, Yong Hoon Lee, Seung Ho Hwang, In-Cheol Park, Chong-Min Kyung |
DAC | 21 |
| 1998 | Multiple Behavior Module Synthesis Based on Selective GroupingsabstractIn this paper, we present an approach to synthesize multiple behavior modules. Given n DFGs to be implemented the previous methods scheduled each of them sequentially, and implemented them as a single module. Though the method Is appropriate for sharing the functional units, it ignored the following two aspects: (1) different interconnection patterns among DFGs can increase the interconnection area and delay of the critical path, (2) the sequential scheduling of DFGs has a difficulty in considering the effects on the other DFGs not scheduled yet. We show an efficient way to solve the problems using a selective grouping method and the extensions of the traditional scheduling methods. The experimentation reveals that the result obtained by the proposed method is better to reduce interconnection area and to meet the timing constraints than those obtained by the previous methods. Ju Hwan Yi, Hoon Choi, In-Cheol Park, Seung Ho Hwang, Chong-Min Kyung |
DATE | 3 |
| 1998 | Synthesis of application specific instructions for embedded DSP softwareabstractArticle Free Access Share on Synthesis of application specific instructions for embedded DSP software Authors: Hoon Choi Department of Electrical Engineering, Korea Advanced Institute of Science and Technology, Taejon, Korea Department of Electrical Engineering, Korea Advanced Institute of Science and Technology, Taejon, KoreaView Profile , Seung Ho Hwang Department of Electrical Engineering, Korea Advanced Institute of Science and Technology, Taejon, Korea Department of Electrical Engineering, Korea Advanced Institute of Science and Technology, Taejon, KoreaView Profile , Chong-Min Kyung Department of Electrical Engineering, Korea Advanced Institute of Science and Technology, Taejon, Korea Department of Electrical Engineering, Korea Advanced Institute of Science and Technology, Taejon, KoreaView Profile , In-Cheol Park Department of Electrical Engineering, Korea Advanced Institute of Science and Technology, Taejon, Korea Department of Electrical Engineering, Korea Advanced Institute of Science and Technology, Taejon, KoreaView Profile Authors Info & Claims ICCAD '98: Proceedings of the 1998 IEEE/ACM international conference on Computer-aided designNovember 1998 Pages 665–671https://doi.org/10.1145/288548.289109Published:01 November 1998Publication History 18citation223DownloadsMetricsTotal Citations18Total Downloads223Last 12 Months16Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Hoon Choi, Seung Ho Hwang, Chong-Min Kyung, In-Cheol Park |
ICCAD | 4 |
| 1997 | HK386: an x86-compatible 32-bit CISC microprocessorabstractThe authors describe the implementation and design methodology of a microprocessor, called HK386. The microprocessor is compatible with Intel 80386 with respect to the behavior of each instruction set. As the extraction of the exact behavior of each instruction set is the single most important step in compatible chip design, they focused their effort on establishing the reliable verification strategy ensuring the complete instruction level compatibility. The HK386 was successfully designed and fabricated using 0.8 /spl mu/m CMOS technology. Chong-Min Kyung, In-Cheol Park, Se-Kyoung Hong, K. S. Seong, B. S. Kong, Seungjong Lee, Hoon Choi, S. R. Maeng, D. T. Kim, Jong-Sun Kim, S. H. Park, Y. J. Kang |
ASP-DAC | 2 |
| 1997 | Multi-project chip activities in Korea-IDEC perspectiveabstractThis paper describes the current status of multi-project chip (MPC) services in Korea to promote full-custom and semi-custom IC design activities in universities. Although MPC foundry services for IC designs were started in a lesser scale more than 10 years ago, it is only recently that systematic and effective education has developed. The MPC foundry services program called IDEC (IC design education center) was launched with the planned support of the government and three major semiconductor companies in Korea. The paper introduces the activities of IDEC and other MPC foundry services. Chong-Min Kyung, In-Cheol Park, Ho-Jun Song |
ASP-DAC | 2 |
| 1997 | Single cycle access cache for the misaligned data and instruction prefetchabstractIn microprocessors, reducing the cache access time and the pipeline stall is critical to improve the system performance. To overcome the pipeline stall caused by the misaligned multi-words data or multi cycle accesses of prefetch codes which are placed over two cache lines, we proposed the Separated Word-line Decoding (SEWD) cache. SEWD cache makes it possible to access misaligned multiple words as well as aligned words in one clock cycle. This feature is invaluable in most microprocessors because the branch target address is usually misaligned, and many of data accesses are misaligned. 8K-byte SEWD cache chip consists of 489,000 transistors on a die size of 0.853/spl times/0.827 cm/sup 2/ and is implemented in 0.8 /spl mu/m DLM CMOS process operating at 60 MHz. Joon-Seo Yim, Hee-Choul Lee, Bong-Il Park, Chang-Jae Park, In-Cheol Park, Chong-Min Kyung |
ASP-DAC | 6 |
| 1997 | Verification methodology of compatible microprocessorsabstractAs the complexity of high-performance microprocessor increases, functional verification becomes more difficult and emerges as the bottleneck of the design cycle. In this paper, we suggest a functional verification methodology, especially for the compatible microprocessor design. To guarantee the perfect compatibility with previous microprocessors, we developed three C models in different representation levels, i.e., Polaris, MCV(Micro-Code Verifier) and StreC. C models are co-simulated with consistency checking between two different models. The simulation speed of C models makes it possible to test the "real-world" application programs on the RTL design with a software board model. To increase the confidence level of verifications, Profiler reports the verification coverage of the test vector, which is fed back to the automatic test program generator. Restartability feature also helps significantly reduce the total simulation time. Using the proposed verification methodology, we designed and verified an Intel 486-compatible microprocessor successfully. Joon-Seo Yim, Chang-Jae Park, Woo-Seung Yang, Hun-Seung Oh, Hee-Choul Lee, Hoon Choi, Seungjong Lee, Nara Won, Yung-Hei Lee, In-Cheol Park, Chong-Min Kyung |
ASP-DAC | 11 |
| 1997 | A C-Based RTL Design Verification Methodology for Complex MicroprocessorabstractAs the complexity of high-performance microprocessor increases,functional verification becomes more and more difficultand RTL simulation emerges as the bottleneck of thedesign cycle.In this paper, we suggest C language-based designand verification methodology to enhance the simulationspeed instead of the conventional HDL-based methodologies.RTL C model (StreC) describes the cycle-based behaviors ofsynchronous circuits and is followed by model refining andoptimization using LifeTime Analyzer (LTA) and Cleaner.The simulation speed of cycle-based C model makes it possibleto test the RTL design with the "real-world" applicationprograms in the order-of-magnitude faster speed thanthe commercial event-driven simulators.Using the proposedfunctional verification methodology, HK486, an intel 80486 - compatiblemicroprocessor was successfully designed and verified. Joon-Seo Yim, Yoon-Ho Hwang, Chang-Jae Park, Hoon Choi, Woo-Seung Yang, Hun-Seung Oh, In-Cheol Park, Chong-Min Kyung |
DAC | 7 |
| 1994 | Two Complementary Approaches for Microcode Bit OptimizationabstractIn the design of microprogrammed processors, the minimization of microcode width is very crucial to reduce the required microcode ROM area. The paper suggests two different procedures which are complementary in nature: first an integer linear programming formulation which guarantees an optimal solution for small or medium size problems; and second, a heuristic algorithm based on the graph bipartitioning to deal with large size problems. Experimental results show that the proposed heuristic algorithm yields near-optimal solutions with polynomial time complexity.> In-Cheol Park, Se-Kyoung Hong, Chong-Min Kyung |
IEEE Trans. Computers | 1 |
| 1993 | FAMOS: an efficient scheduling algorithm for high-level synthesisabstractFAMOS, an iterative improvement scheduling algorithm for the high-level synthesis of digital systems, is described. The algorithm is based on a move acceptance strategy and various selection functions defined to represent the cost of hardware resources such as functional units and registers. A main feature of the algorithm is that it can escape from local minima. The algorithm can deal with diverse design styles such as multi-cycle operations, chained operations, pipelined datapaths, pipelined functional units and conditional branches. Register costs and maximal time constraints are also considered. To efficiently represent information on the design styles, a graph model called weighted precedence graph is proposed as a general model on which the scheduling algorithm is based. Despite the iterative nature, the proposed algorithm has a polynomial time complexity. Although the optimality of the algorithm is not guaranteed, optimal solutions were obtained for several examples available from the literature.> In-Cheol Park, Chong-Min Kyung |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1991 | Fast and Near Optimal Scheduling in Automatic Data Path AynthesisabstractA new heuristic scheduling algorithm which has a feature of escaping from local minima is presented.The algorithm has a polynomial time complexity in spite of its iterative nature.Although there is no guarantee for the optimality, the algorithm produced optimal results for the experimental examples of earlier works.A graph model which contains information on the real world constraints such as multi-cycle operations, chained operations and pipelined data paths is also proposed as a general model on which our scheduling algorithm is based. In-Cheol Park, Chong-Min Kyung |
DAC | 1 |
| 1990 | An O(n3logn)-Heuristic for Microcode Bit OptimizationabstractThe authors address the problem of minimizing the control ROM width, which is important in the design of microprogrammed processors, because it directly reduces the silicon area of the control unit. A heuristic algorithm is proposed which is based on graph partitioning. This algorithm results in nearly optimal solutions with the time complexity of O(n/sup 3/log n), where n denotes the number of distinct microoperations. Comparison of the results with earlier works shows that the proposed heuristic performs better in terms of cost and/or CPU time.> Se-Kyoung Hong, In-Cheol Park, Chong-Min Kyung |
ICCAD | 2 |