Ki-Seok Chung

dblp:27/5623 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0002-2908-8443ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorComputer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 TiMM-ViT: Token-Importance-Aware Token Merging with Multi-Exit Training and an FPGA-based Vision Transformer Accelerator for Image Classification
abstract
Token merging (ToMe) reduces the computational cost of Vision Transformers (ViTs) by merging similar tokens. However, ToMe may inadvertently merge critical tokens, leading to accuracy degradation. Furthermore, the sequential execution of ViT and ToMe operations incurs increased latency on edge GPUs. To address these challenges, we propose TiMM, a token-importance-aware merging method that preserves essential tokens by considering both token importance and similarity. We also present a dedicated FPGA-based accelerator designed to exploit the concurrent execution between ViT and TiMM. Experimental results across DeiT-Tiny, Small, and Base demonstrate that TiMM achieves accuracy improvements of 0.23%-0.73% over ToMe, respectively, within a 1% accuracy drop margin relative to the baseline. When implemented on the Xilinx ZCU102, it demonstrates speedups of 1.45×-10×, along with energy efficiency gains of 1.03×-3.1× compared to the NVIDIA Jetson TX2.
Soomin Rho, Sangki Park, Kwangrae Kim, Min-Gwon Song, Ki-Seok Chung
ISLPED6
2024 FloatMax: An Efficient Accelerator for Transformer-Based Models Exploiting Tensor-Wise Adaptive Floating-Point Quantization
abstract
The rapid growth of the Transformer model size results in significant computational resources and memory requirements. To mitigate this complexity, quantization to reduce the bit width to represent numbers is being actively studied. However, prior quantization methods that use integers of 8 bits or less suffer from accuracy loss due to lower resolution because tensors of the Transformer model often have a non-uniform distribution containing outliers. In this paper, we propose a novel quantization method that utilizes tensor-wise adaptive floating-point quantization. Two key strategies address the issue of accuracy degradation. First, we leverage the characteristics of the floating-point data type, which provides higher precision for normal values and lower precision for outlier values. Second, since the degree of non-uniformity and outliers in each tensor vary from layer to layer, we adaptively assign different bit widths to the exponent and the mantissa of the floating-point representation for each tensor, quantizing them with different levels of precision. This approach allows high accuracy across the tensor distribution while efficiently managing outliers. We design the FloatMax processing element and decoders to efficiently carry out these floating-point computations. In addition, FloatMax is integrated into a systolic array to accelerate the linear layer and attention mechanism. Evaluation results show that FloatMax achieves 0.75×, 0.85×, and 0.79× area reduction and 1.70×, 1.84×, and 1.89× average performance improvement over Olive, ANT, and AdaFloat, respectively, without accuracy loss.
Seoho Chung, Kwangrae Kim, Soomin Rho, Chanhoon Kim, Ki-Seok Chung
ICCD5
2023 Token Merging with Class Importance Score
abstract
Vision Transformers have achieved high performance in computer vision tasks, but their high computational cost and low throughput are weaknesses. Therefore, much research has been done to reduce the size of Vision Transformers. Among them, studies on pruning unnecessary tokens are being actively conducted to reduce the number of tokens used for self-attention computation inside the Vision Transformer. Recently, token merging has been proposed as a new alternative approach. These studies aim to increase throughput with a small accuracy drop by merging similar tokens instead of pruning them. A previous study finds similar tokens using cosine similarity and merges them with a weighted average. However, merging a large number of tokens at once may lead to an accuracy drop because of the underestimating of important information. In this paper, we propose ToMeCIS, a method that merges similar tokens through a weighted average using the class importance score of tokens to reduce the accuracy drop. When ToMeCIS is applied to a pretrained DeiT-S and evaluated on the ImageNet-1k dataset, the throughput is increased by about 50% with an accuracy drop of less than 1% without additional training. In addition, importance scores were evaluated with different metrics to find the best accuracy versus throughput trade-off.
Kwang-Soo Seol, Si-Dong Roh, Ki-Seok Chung
IECON3
2023 Dynamic Partitioning Method for Near-Memory Parallel Processing of Sparse Matrix-Vector Multiplication
abstract
Near-memory processing (NMP), which places lightweight processing units near the DRAM memory, has been actively studied to speed up the execution of memory-intensive applications by reducing the amount of data traffic between the DRAM and the CPU. Sparse matrix-vector multiplication (SpMV) is a representative memory-bound kernel used in various applications such as graph analytics, scientific computing, and machine learning. There are prior works to accelerate SpMV by NMP employing a fixed partitioning scheme that hides random access of SpMV using a parallel NMP core. However, due to the various distributions of the matrix, the fixed partitioning of prior works causes a load imbalance in which a sparse matrix is unevenly allocated to processing units for NMP. To resolve this, dynamic partitioning methods to distribute matrices and vectors to the NMP processing units can be effective. In this paper, we propose a dynamic partitioning algorithm (DPA) that analyzes the distribution of non-zero elements in a sparse matrix to classify it into three types (even distribution, skewed distribution, and power-law distribution) and partitions the matrix according to each distribution. Our proposed distribution scheme alleviates load imbalance by up to 73% when compared to static distribution schemes, and such improvement achieves an average speed-up 1.37x (up to 1.84x) over the NMP architecture with static distribution schemes.
Dae-Eun Wi, Kwangrae Kim, Ki-Seok Chung
IECON3
2021 HammerFilter: Robust Protection and Low Hardware Overhead Method for RowHammer
abstract
The continuous scaling-down of the dynamic random access memory (DRAM) manufacturing process has made it possible to improve DRAM density. However, it makes small DRAM cells susceptible to electromagnetic interference between nearby cells. Unless DRAM cells are adequately isolated from each other, the frequent switching access of some cells may lead to unintended bit flips in adjacent cells. This phenomenon is commonly referred to as RowHammer. It is often considered a security issue because unusually frequent accesses to a small set of rows generated by malicious attacks can cause bit flips. Such bit flips may also be caused by general applications. Although several solutions have been proposed, most approaches either incur excessive area overhead or exhibit limited prevention capabilities against maliciously crafted attack patterns. Therefore, the goals of this study are (1) to mitigate RowHammer, even when the number of aggressor rows increases and attack patterns become complicated, and (2) to implement the method with a low area overhead.We propose a robust hardware-based protection method for RowHammer attacks with a low hardware cost called HammerFilter, which employs a modified version of the counting bloom filter. It tracks all attacking rows efficiently by leveraging the fact that the counting bloom filter is a space-efficient data structure, and we add an operation, HALF-DELETE, to mitigate the energy overhead. According to our experimental results, the proposed method can completely prevent bit flips when facing artificially crafted attack patterns (five patterns in our experiments), whereas state-of-the-art probabilistic solutions can only mitigate less than 56% of bit flips on average. Furthermore, the proposed method has a much lower area cost compared to existing counter-based solutions (40.6× better than TWiCe and 2.3× better than Graphene).
Kwangrae Kim, Jeonghyun Woo, Junsu Kim 0004, Ki-Seok Chung
ICCD4
2020 A Confidence-Calibrated MOBA Game Winner Predictor
abstract
In this paper, we propose a confidence-calibration method for predicting the winner of a famous multiplayer online battle arena (MOBA) game, League of Legends. In MOBA games, the dataset may contain a large amount of input-dependent noise; not all of such noise is observable. Hence, it is desirable to attempt a confidence-calibrated prediction. Unfortunately, most existing confidence calibration methods are pertaining to image and document classification tasks where consideration on uncertainty is not crucial. In this paper, we propose a novel calibration method that takes data uncertainty into consideration. The proposed method achieves an outstanding expected calibration error (ECE) (0.57%) mainly owing to data uncertainty consideration, compared to a conventional temperature scaling method of which ECE value is 1.11%.
Dong-Hee Kim, Changwoo Lee 0001, Ki-Seok Chung
CoG3
2020 Direct Conversion: Accelerating Convolutional Neural Networks Utilizing Sparse Input Activation
abstract
The amount of computation and the number of parameters of neural networks are increasing rapidly as the depth of convolutional neural networks (CNNs) is increasing. Therefore, it is very crucial to reduce both the amount of computation and that of memory usage. The pruning method, which compresses a neural network, has been actively studied. Depending on the layer characteristics, the sparsity level of each layer varies significantly after the pruning is conducted. If weights are sparse, most results of convolution operations will be zeroes. Although several studies have proposed methods to utilize the weight sparsity to avoid carrying out meaningless operations, those studies lack consideration that input activations may also have a high sparsity level. The Rectified Linear Unit (ReLU) function is one of the most popular activation functions because it is simple and yet pretty effective. Due to properties of the ReLU function, it is often observed that the input activation sparsity level is high (up to 85%). Therefore, it is important to consider both the input activation sparsity and the weight one to accelerate CNN to minimize carrying out meaningless computation. In this paper, we propose a new acceleration method called Direct Conversion that considers the weight sparsity under the sparse input activation condition. The Direct Conversion method converts a 3D input tensor directly into a compressed format. This method selectively applies one of two different methods: a method called image to Compressed Sparse Row (im2CSR) when input activations are sparse and weights are dense; the other method called image to Compressed Sparse Overlapped Activations (im2CSOA) when both input activations and weights are sparse. Our experimental results show that Direct Conversion improves the inference speed up to 2.82× compared to the conventional method.
Wonhyuk Lee, Si-Dong Roh, Sangki Park, Ki-Seok Chung
IECON4
2019 GRAM: Gradient Rescaling Attention Model for Data Uncertainty Estimation in Single Image Super Resolution
abstract
In this paper, a new learning method to quantify data uncertainty without suffering from performance degradation in Single Image Super Resolution (SISR) is proposed. Our work is motivated by the fact that the idea of loss design for capturing uncertainty and that for solving SISR are contradictory. As to capturing data uncertainty, we often model the output of a network as a Euclidian distance divided by a predictive variance, negative log-likelihood (NLL) for the Gaussian distribution, so that images with high variance have less impact on training. On the other hand, in the SISR domain, recent works give more weights to the loss of challenging images to improve the performance by using attention models. Nonetheless, the conflict should be handled to make neural networks capable of predicting the uncertainty of a super-resolved image, without suffering from performance degradation. Therefore, we propose a method called Gradient Rescaling Attention Model (GRAM) that combines both attempts effectively. Since variance may reflect the difficulty of an image, we rescale the gradient of NLL by the degree of variance. Hence, the neural network can focus on the challenging images, similarly to attention models. We conduct performance evaluation using standard SISR benchmarks in terms of peak signal-noise ratio (PSNR) and structural similarity (SSIM). The experimental results show that the proposed gradient rescaling method generates negligible performance degradation compared to SISR outputs with the Euclidian loss, whereas NLL without attention degrades the SR quality.
Ki-Seok Chung, Changwoo Lee 0001
ICMLA1
2016 User-Centric Power Management for Embedded CPUs Using CPU Bandwidth Control
abstract
Dynamic power management for mobile processors has become very important due to the increased clock speed and number of cores. There have been various power management governors using dynamic voltage and frequency scaling (DVFS). Among them, a user-centric power management has received a lot of attention as a method to save power while maintaining the quality of user experience (UX) referring to the perceived quality of system services to end users. Most user-centric governors have employed DVFS as a method to reduce the power consumption. However, DVFS may not be adequate enough to guarantee UX qualities for all task because the CPU clock speed changed by DVFS can affect all tasks running at the same processor. In order to minimize such inter-task interferences by DVFS, it is necessary to employ task-specific power management methods. This paper shows that CPU bandwidth control developed for CPU resource management within Linux kernel can be employed as a task-specific power management method, and a novel CPU power management scheme employing both DVFS and CPU bandwidth control is proposed. Experimental results show that the proposed governor can reduce the power consumption more than the Ondemand governor can achieve while maintaining the quality of UX.
Youngho Ahn, Ki-Seok Chung
IEEE Trans. Mob. Comput.2
2015 Parallel LDPC decoding on a GPU using OpenCL and global memory for accelerators
abstract
This paper introduces a parallel software decoder of Low Density Parity Check (LDPC) codes with an Open Computing Language (OpenCL) framework including Global Memory for ACcelerators (GMAC). The LDPC code is one of the most popular and strongest error correcting codes for mobile communication systems. OpenCL is an open standard programming framework that supports programming languages and application programming interfaces (APIs) for heterogeneous platforms. GMAC is a software implementation of Asymmetric Distributed Shared Memory (ADSM) that maintains a shared logical memory space for the host to access memory objects in the physical memory of an OpenCL device. In this paper, we parallelize the iterative LDPC decoding steps on a graphics processing unit (GPU) using OpenCL. To improve the performance of the proposed decoder, data transfer optimization techniques between the host and the GPU including pre-pinned OpenCL memory objects for GMAC are applied. In terms of the entire decoding time, the speedup of the proposed LDPC decoder over a conventional OpenCL implementation is 1.28.
Jung-Hyun Hong, Ki-Seok Chung
NAS2
2015 Analysis of various DRAM devices from power consumption's perspective
abstract
DRAM has become a crucial component in terms of system power consumption as the size of main memory increases. To improve power efficiency of DRAM devices, we need to analyze characteristics of DRAM behavior from power consumption's perspective. In this paper, we analyze the characteristics of various DRAM devices from major vendors under real system operating environment. As the size of the DRAM increases, power consumption due to activate, precharge and burst operations remains about the same, but that due to background and refresh operations increases steadily. Especially, power consumption due to the refresh operation for 3D stacked DRAMs increases by 91% at high temperatures, which strongly implies that the refresh operation will become more crucial for DRAMs in the future.
Dong-Ik Jeon, Min-Kyu Lee, Ki-Seok Chung
NAS3
2013 Dynamic Power Management Technique for Multicore Based Embedded Mobile Devices
abstract
As the proliferation of ubiquitous computing environments becomes a reality, the need for high speed data processing and intelligent system management increases rapidly. In particular, the need for low-power designs and power-aware system management is getting stronger. While multicore systems are deployed in many embedded system areas, an effective power management technique for multicores is not available yet. In this paper, we propose a novel power management technique based on a parallel programming model. OpenMP is a well-known programming paradigm for shared memory multicore systems. OpenMP is based on library routines for parallel processing. By identifying the invoked library routines, how many cores will be adequate for a certain application can be determined, and the number of necessary cores for a given task can be determined during run-time. By turning off unnecessary cores, we can reduce power consumption. We implemented this method by adding capabilities in an OpenMP-compliant compiler and conducted experiments with various benchmarks. We were able to reduce the power consumption by 18% on average compared to other conventional power management methods.
Young-Si Hwang, Ki-Seok Chung
IEEE Trans. Ind. Informatics2
2011 LDPC decoding for CMMB utilizing OpenMP and CUDA parallelization
abstract
As the 4G mobile communication systems require high transmission rate with reliability, the demand for efficient error correcting code increases. In this paper, a novel LDPC (Low Density Parity Check) decoding method is introduced. We address a parallel software implementation of LDPC decoding for CMMB (China Multimedia Mobile Broadcasting) standard. LDPC codes for CMMB employ a regular H-matrix which has the fixed row and column weights. While effectively utilizing the regularity of the H-matrix, we process information on H-matrices for multiple code rates using OpenMP pragmas on a multi-core processor and execute the decoding algorithm in parallel using CUDA (Compute Unified Device Architecture) on a GPU (Graphics Processing Unit). We evaluated the performance of the proposed implementation with respect to two different code rates, and verified that the proposed implementation satisfies the bandwidth requirement for CMMB.
Joo-Yul Park, Ki-Seok Chung
APCC2
2009 Thermal sensor allocation and placement for reconfigurable systems
abstract
A dynamic monitoring of thermal behavior of hardware resources using thermal sensors is very important to maintain the operation of systems safe and reliable. This article addresses the problem of thermal sensor allocation and placement for reconfigurable systems. For programmable logic arrays, the degree of the use of hardware resources in the systems highly depends on the target application to be implemented, making the allocation of thermal sensors at the manufacturing stage inadequate (or too costly if implemented) due to the unpredictable thermal profile. This means that the thermal sensor allocation could be processed at the time when the reconfigurable logic is implemented (i.e., at the post manufacturing stage). This work proposes an effective solution to the problem of thermal sensor allocation and placement at the post-manufacturing stage. Specifically, we define the Sensor Allocation and Placement Problem (SAPP), and propose a solution which formulates SAPP into the Unate-Covering Problem (UCP) and solves it optimally. Also we combine SAPP with temperature correlation to reduce required sensors more aggressively and propose a solution by applying UCP again. We then provide an extended solution to handle a practical design issue where the hardware resources for the sensor implementation on specific array locations have already been used up by the application logic. Experimental results using MCNC benchmarks show that our proposed technique uses 62.4% and 19.7% less number of sensors to monitor hotspots on the average than that used by the grid-based and the bisection-based approaches while the overhead of auxiliary circuitry is minimized, respectively.
Ki-Seok Chung, Bontae Koo, Nak-Woong Eum, Taewhan Kim 0001
ACM Trans. Design Autom. Electr. Syst.2
2008 Predictive power aware management for embedded mobile devices
abstract
Intelligent power management of mobile devices is getting more important as ubiquitous computing is coming true in daily life. Power aware system management relies on techniques of collecting and analyzing information on the status of I/O devices or processors while the system is running applications. However, the overhead of collecting information using software while the system is running is so huge that performance of the system may be severely deteriorated. Therefore, it is very crucial to design a PMU (power management unit) which collects information in hardware so that the performance of the system is not degraded. In this paper, we propose a novel PMU design which collects access patterns to I/O devices while an application is being executed. And a predictive power aware management is carried out based on the collected information. Experiments with various applications have been conducted to show the effectiveness of our approach.
Young-Si Hwang, Sung-Kwan Ku, Chan-Min Jung, Ki-Seok Chung
ASP-DAC4
2007 Design of Low Power MAC Operator with Dual Precision Mode
abstract
MAC (multiply and accumulate) operations are heavily involved in many DSP applications. To improve the performance of MAC operations, it is very crucial to make multiplication fast. In this paper, we propose a set of novel MAC operator designs with constant coefficients. It is well-known that shifting can replace a constant multiplication if the constant is a power of 2. We extend this idea in such a way that by employing more than 2 barrel shifters and a supplementary multiplier we can design a very efficient constant multiplier for general constants. To enhance the applicability of our design for a wide range of constants, a support for variable precision computations has been implemented. Experimental results show that our designs achieve consistent enhancement in terms of power consumption when our design is com pared with other existing optimized designs.
Young-Geun Lee, Joo-Yul Park, Ki-Seok Chung
RTCSA3
2004 Profile-based optimal intra-task voltage scheduling for hard real-time applications
abstract
This paper presents a set of comprehensive techniques for the intratask voltage scheduling problem to reduce energy consumption in hard real-time tasks of embedded systems. Based on the execution profile of the task, a voltage scheduling technique that optimally determines the operating voltages to individual basic blocks in the task is proposed. The obtained voltage schedule guarantees minimum average energy consumption. The proposed technique is then extended to solve practical issues regarding transition overheads, which are totally or partially ignored in the existing approaches. Finally, a technique involving a novel extension of our optimal scheduler is proposed to solve the scheduling problem in a discretely variable voltage environment. In summary, it is confirmed from experiments that the proposed optimal scheduling technique reduces energy consumption by 20.2 % over that of one of the state-of-the-art schedulers [11] and, further, the extended technique in a discrete voltage environment reduces energy consumption by 45.3 % on average.
Jaewon Seo, Taewhan Kim 0001, Ki-Seok Chung
DAC3
2001 A Static Estimation Technique of Power Sensitivity in Logic Circuits
abstract
In this paper, we study a new problem of statically estimating the power sensitivity of a given logic circuit with respect to the primary inputs. The power sensitivity defines the characteristics of power dissipation due to changes in state of primary inputs, Consequently, estimating the power sensitivity among the inputs is essential not only to measure the power consumption of the circuit efficiently but also to provide potential opportunities of redesigning the circuit for low power, In this context, we propose a fast and reliable static estimation technique for power sensitivity based on a new concept called power equations, which are then collectively transformed into a table called power table. Experimental data on MCNC benchmark examples show that the proposed technique is useful and effective in estimating power consumption. In summary, the relative error for the estimation of maximum power consumption is 9.4% with a huge speed-up in simulation.
Taewhan Kim 0001, Ki-Seok Chung, Chien-Liang Liu
DAC2
2000 Behavioral-level partitioning for low power design in control-dominated application
abstract
In this paper, we study the problem of behavioral-level partitioning for low power design. By behavioral-level partitioning, we mean a partitioning which is done at the behavioral description where scheduling and allocation have not been carried out. The motivation is that turning on/off individual operations cycle-by-cycle is very expensive, thereby we provide a partitioning solution so that all operations in the same partition can be controlled by the same gated clock signal. Our partitioning algorithm is specifically focused on the applications which contain many nested conditional branches and loops.
Ki-Seok Chung, Taewhan Kim 0001, Chien-Liang Liu
ACM Great Lakes Symposium on VLSI1
1998 Local transformation techniques for multi-level logiccircuits utilizing circuit symmetries for power reduction
abstract
In this pap er, we present sever al optimization techniques for power reduction utilizing circuit symmetries. There are four kinds of symmetries that we dete ct in a given circuit implementation. First, we pr op ose an algorithm for dete cting the four different typ es ofsymmetries in a given circuit implementation of a Boole an function. Sever alre-synthesis techniques utilizing such symmetries are prop ose d. These techniques enable us to optimize power consumption and delay with no (or very little) ar ea overhead. We have carrie dout experiments on MCNC benchmark circuits to demonstrate the efficiency of the prop ose dtechniques. The aver age power reduction is 14% with little or none ar ea and/or delay overhead.
Ki-Seok Chung, C. L. Liu 0001
ISLPED1
1997 Low power multiplexer decomposition
abstract
The advent of poTtable digital devices such as laptop peTsona1 computers has made low poweT ciTcuit design an incTeasingly impoTtant TeseaTch area.Recently, low power decomposition foT simple logic gates such as AND and OR has been extensively Teseazhed.HoweveT, the pToblem of MUX decomposition to minimize poweT dissipation has not been addTessed.In this papeT, we study the pToblem of low power multiplexer (MUX) decomposition.MUX decomposition is the procedure of tTansfo?ming an n-to-one MUX into an equivalent tTee of two-to-one MUXes.We propose a formulation for the minimum power MUX decomposition problem based on the common CMOS pass tTan-sistoT implementation of a MUX.Given the occuTTence pTobabilities of the data signals and theiT on probabilities, we analyze the poweT dissipation of OUT MUX implementation and give a geneTa1 method for computing the poweT dissipation of a MUX tTee decomposition.We then present seveTa1 algorithms which eficientiy generate minimum power MUX decompositiotas.We demon&ate the effectiveness of OUT algoTi&ns wi& experimental Tesults.
Unni Narayanan, Hon Wai Leong, Ki-Seok Chung, Chien-Liang Liu
ISLPED3
1996 An algorithm for synthesis of system-level interface circuits
abstract
We describe an algorithm for the synthesis and optimization of interface circuits for embedded system components such as microprocessors, memory ASIC, and network subsystems with fixed interfaces. The algorithm accepts the timing characteristics of two system components as input, and generates a combinational interface (glue logic) circuit. The algorithm consists of two parts. In the first part, we determine the direct pin-to-pin connections in the interface circuit employing a 0/1 ILP formulation to minimize wiring area and dynamic power consumption. In the second part, we determine logic subcircuits in the interface circuit, utilizing the timing diagrams of the system components. The proposed algorithm has been implemented in a software package SYNTERFACE. Experimental results are presented to demonstrate the effectiveness of the algorithm.
Ki-Seok Chung, Rajesh K. Gupta 0001, C. L. Liu 0001
ICCAD1