Ik Joon Chang

dblp:92/6854 · also Ik-Joon Chang · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-8871-8695ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
YearPublicationVenuePosition
2026 L2Mersit: A Scaling-Free Sub-8-bit Data Format for On-Device Reliable Large Language Model Serving
abstract
On-device large language model (LLM) serving drives low-precision computing to address memory and compute limits. This paper presents L2Mersit, a scaling-free, range-adjustable exponent-encoded data format tailored for sub-8-bit LLM quantization. Building upon the Mersit framework, L2Mersit employs dual mode operation, comprising range-expanded and precision-enhanced modes that dynamically adapt to activation distributions with minimal control overhead. The proposed design eliminates on-the-fly scaling and auxiliary computations while effectively preserving range and precision, thereby achieving both superior perplexity and hardware efficiency. Experimental results demonstrate that L2Mersit achieves the highest accuracy among all 6-bit exponent-encoded formats while reducing the hardware complexity of auxiliary units for low-precision computing, resulting in a 62.7% area reduction.
Myeongjin Kim, Hyeonseong Kim, Ik Joon Chang, Seungkyu Choi
ISLPED3
2024 MERSIT: A Hardware-Efficient 8-bit Data Format with Enhanced Post-Training Quantization DNN Accuracy
abstract
Post-training quantization (PTQ) models utilizing conventional 8-bit Integer or floating-point formats still exhibit significant accuracy drops in modern deep neural networks (DNNs), rendering them unreliable. This paper presents MERSIT, a novel 8-bit PTQ data format designed for various DNNs. While leveraging the dynamic configuration of exponent and fraction bits derived from Posit data format, MERSIT demonstrates enhanced hardware efficiency through the proposed merged decoding scheme. Our evaluation indicates that MERSIT yields more reliable 8-bit PTQ models, exhibiting superior accuracy across various DNNs compared to conventional floating-point formats. Furthermore, the proposed processing unit saves 26.6% in area and 22.2% in power consumption compared to the Posit-based unit, while maintaining comparable efficiency to the floating-point-based unit.
Nguyen-Dong Ho, Gyujun Jeong, Cheol-Min Kang, Seungkyu Choi, Ik Joon Chang
DAC5
2024 Skew-CIM: Process-Variation-Resilient and Energy-Efficient Computation-in-Memory Design Technique With Skewed Weights
abstract
In analog-mixed-signal (AMS) compute-in-memory (CIM) systems, the two’s-complement (2SC) format provides better area efficiency than the sign-and-magnitude (SNM) one. However, the 2SC format exacerbates the challenges of AMS-CIM systems, suffering from significant DNN accuracy drop under process variations and high computation currents from activating multiple WLs. In the 2SC format, ‘0’ and ‘1’ are nearly balanced for all logical-order bits, unlike ‘0’-skewed higher-order bits in the SNM format. Consequently, the 2SC-based AMS-CIM systems have much more on-cells than the SNM-based counterpart, deteriorating the above challenges. We propose Skew-CIM, a software-hardware co-design technique to relax these challenges. Our proposed weight skewing (WESK) breaks the ‘0’ and ‘1’ balance at the software level. The potential accuracy drops resulting from WESK are successfully compensated by retraining DNNs. The offsets caused by WESK can be easily corrected using online hardware-level processing. Our Skew-CIM technique can be applied to most AMS-CIM systems with memories showing large on-off cell current ratios. As an example, we use it in a custom-designed 8T-SRAM-based CIM device, demonstrating a significant reduction in the DNN classification error by 7.6 times compared to the 2SC-based AMS-CIM without our Skew-CIM technique. Furthermore, our Skew-CIM markedly enhances energy efficiency by up to 39.9%, outperforming conventional SNM-based AMS-CIM systems.
Donghyeon Yi, Injun Choi, Gichan Yun, Edward Choi 0001, Jonghee Park, Jonghoon Kwak, Sung-Joon Jang, Sohmyung Ha, Ik Joon Chang, Minkyu Je
IEEE Trans. Circuits Syst. I Regul. Pap.10
2021 TCL: an ANN-to-SNN Conversion with Trainable Clipping Layers
abstract
Spiking-neural-networks (SNNs) are promising at edge devices since the event-driven operations of SNNs provides significantly lower power compared to analog-neural-networks (ANNs). Although it is difficult to efficiently train SNNs, many techniques to convert trained ANNs to SNNs have been developed. However, after the conversion, a trade-off relation between accuracy and latency exists in SNNs, causing considerable latency in large size datasets such as ImageNet. We present a technique, named as TCL, to alleviate the trade-off problem, enabling the accuracy of 73.87% (VGG-16) and 70.37% (ResNet-34) for ImageNet with the moderate latency of 250 cycles in SNNs.
Nguyen-Dong Ho, Ik Joon Chang
DAC2
2021 Posit Arithmetic for the Training and Deployment of Generative Adversarial Networks
abstract
This paper proposes a set of methods that enables low precision posit ™ arithmetic to be successfully used for the training of generative adversarial networks (GANs) with minimal quality loss. We show that ultra low precision posits, as small as 6 bits, can achieve high quality output for the generation phase after training. We also evaluate floating-point (float) formats and compare them to 8-bit posits in the context of GAN training. Our scaling and adaptive calibration techniques are capable of producing superior training quality for 8-bit posits that surpasses 8-bit floats and matches the results of 16-bit floats. Hardware simulation results indicate that our methods have higher energy efficiency compared to both 16- and 8-bit float training systems.
Nhut-Minh Ho, Duy Thanh Nguyen, Himeshi De Silva, John L. Gustafson, Weng-Fai Wong, Ik Joon Chang
DATE6
2021 STT-MRAM Architecture with Parallel Accumulator for In-Memory Binary Neural Networks
abstract
In this paper, a row-wise XNOR accumulator architecture for STT-MRAM arrays is proposed for parallel and efficient multiply-and-accumulate (MAC) operation. The proposed accumulator supports in-memory computing and binary neural network (BNN) applications. In the proposed architecture, inputs are fed from the complementary bitlines, whereas readout is performed through a time-based sense amplifier (TBS). The proposed architecture that does not require any ADC can exhibit an average error rate of 0.085 for XNOR vector size (i.e., accumulate capacity) of 128 bits, which translates into 98.45% classification accuracy of a multi-layer perceptron (MLP) on the MNIST dataset.
Thi-Nhan Pham, Quang-Kien Trinh, Ik Joon Chang, Massimo Alioto
ISCAS3
2020 O-2A: Low Overhead DNN Compression with Outlier-Aware Approximation
abstract
We present a low-latency DNN compression technique to reduce DRAM energy, significant in DNN inferences, namely Outlier-Aware Approximation (O-2A) coding. This technique compresses 8-bit integer, de-facto standard of DNN inferences, to 6-bit without degrading the accuracies of DNNs. The hardware for the O-2A coding can be easily embedded to DRAM controllers due to small overhead. In an Eyeriss platform, the O-2A coding improves both DRAM energy and system performance by 18~20%. The O-2A coding enables us to implement an error-correction scheme without additional parity overhead, opening the possibility of an approximate DRAM to simultaneously reduce DRAM accessing and refresh energy.
Nguyen-Dong Ho, Minh-Son Le, Ik Joon Chang
DAC3
2020 PCM: Precision-Controlled Memory System for Energy Efficient Deep Neural Network Training
abstract
Deep neural network (DNN) training suffers from the significant energy consumption in memory system, and most existing energy reduction techniques for memory system have focused on introducing low precision that is compatible with computing unit (e.g., FP16, FP8). These researches have shown that even in learning the networks with FP16 data precision, it is possible to provide training accuracy as good as FP32, de facto standard of the DNN training. However, our extensive experiments show that we can further reduce the data precision while maintaining the training accuracy of DNNs, which can be obtained by truncating some least significant bits (LSBs) of FP16, named as hard approximation. Nevertheless, the existing hard-ware structures for DNN training cannot efficiently support such low precision. In this work, we propose a novel memory system architecture for GPUs, named as precision-controlled memory system (PCM), which allows for flexible management at the level of hard approximation. PCM provides high DRAM bandwidth by distributing each precision to different channels with as transposed data mapping on DRAM. In addition, PCM supports fine-grained hard approximation in the L1 data cache using software-controlled registers, which can reduce data movement and thereby improve energy saving and system performance. Furthermore, PCM facilitates the reduction of data maintenance energy, which accounts for a considerable portion of memory energy consumption, by controlling refresh period of DRAM. The experimental results show that in training CIFAR-100 dataset on Resnet-20 with precision tuning, PCM achieves energy saving and performance enhancement by 66% and 20%, respectively, without loss of accuracy.
Boyeal Kim, Hyun Kim 0001, Duy Thanh Nguyen, Minh-Son Le, Ik Joon Chang, Dohun Kwon, Jin Hyeok Yoo
DATE6
2020 DRAMA: An Approximate DRAM Architecture for High-performance and Energy-efficient Deep Training System
abstract
As the density of DRAM becomes larger, the refresh overhead becomes more significant. This becomes more problematic in the systems that require large DRAM capacity, such as the training of deep neural networks (DNNs). To solve this problem, we present DRAMA, a novel architecture which employs the approximate characteristic of DNNs. We make that non-critical bits are not refreshed while critical bits are normally refreshed. The refresh time of the critical bits are concealed by employing per-bank refreshes, significantly improving the training system performance. Furthermore, the potential racing hazard of bank-refresh technique is simply prevented by our novel command scheduler in DRAM controllers. Our experiments on various recent DNNs show that DRAMA can improve the training system performance by 10.4% and save 23.77% DRAM energy compared to the conventional architecture.
Duy Thanh Nguyen, Changhong Min, Nhut-Minh Ho, Ik Joon Chang
ICCAD4
2019 St-DRC: Stretchable DRAM Refresh Controller with No Parity-overhead Error Correction Scheme for Energy-efficient DNNs
abstract
We present a stretchable DRAM refresh control for energy-efficient processing of DNNs, namely St-DRC. We exploit the characteristic that the recognition accuracy of DNNs is insensitive to errors of insignificant bits. By replacing some insignificant bits with parity bits for the error-correction of significant bits, the St-DRC can protect the significant bits under stretched refresh periods. This significantly improves DRAM refresh energy without performance degradation of DNNs, applicable to both training and inference operations. Our simulation shows that in training, the St-DRC obtains 23%/12% DRAM energy savings for graphic/main memories, respectively. Further, the St-DRC accelerates the training speed by 0.43 ~ 4.12%.
Duy Thanh Nguyen, Nhut-Minh Ho, Ik Joon Chang
DAC3
2019 Segmented Tag Cache: A Novel Cache Organization for Reducing Dynamic Read Energy
abstract
A set-associative cache organization is widely used to achieve a high hit-rate in modern caches, leading to a considerable enhancement in the system performance. In a conventional set-associative cache, tag and data arrays are simultaneously accessed to achieve fast access, but doing so causes a large amount of energy consumption. Previous attempts, such as way prediction and sequential tag access for energy reduction, are still associated with substantial energy consumption due to the tag access. This paper presents a new cache organization termed Segmented Tag Cache (STC), which reduces the amount of energy consumed during the tag access. The proposed organization initially implements a new tag organization that supports partial tag access. More specifically, the tag array is segmented into two parts, and partial tag access only reads low-order part of the tag. A cache access scheme suitable for the proposed tag organization then are developed. Under the scheme, the delay of the tag organization is hidden by the data access delay, avoiding an increase in the overall cache access time to be maintained. Simulation results show that the STC reduces the energy-delay product by approximately 58 percent compared to the conventional cache organizations with a negligible performance penalty.
Moonsoo Kim, Ik Joon Chang
IEEE Trans. Computers2
2018 An Approximate Memory Architecture for a Reduction of Refresh Power Consumption in Deep Learning Applications
abstract
A DRAM device requires periodic refresh operations to preserve data integrity, which incurs significant power consumption. This paper proposes a new memory architecture to reduce the power consumption by refresh operations by slowing down the refresh rate. Slow refresh may cause a loss of data stored in a DRAM cell, which affects the correctness of the computation using the lost data. The proposed memory architecture attempts to avoid the problem caused by lost data by taking advantage of the error-tolerant property of deep learning applications that are tolerant to presence of a small amount of errors. For data storage in deep learning applications, the approximate DRAM architecture stores the data in a transposed manner so that data are sorted according to their significance. DRAM organization is modified to support the control of the refresh period according to the significance of stored data. Simulation results with GoogLeNet and VGG-16 show that the power consumption is reduced by 69.68% with a negligible drop of the classification accuracy for both GoogLeNet and VGG-16.
Duy Thanh Nguyen, Hyun Kim 0001, Ik Joon Chang
ISCAS4
2014 a-SAD: power efficient SAD calculator for real time H.264 video encoder using MSB-approximation technique
abstract
We propose a power efficient SAD calculator, namely a-SAD. We use MSB-approximation where some highest-order MSB's are approximated to single MSB. Our theoretical analysis shows that this technique simultaneously improves performance and power of SAD circuit. We obtain optimal number of approximated MSB's from video experiments, which is the largest number not to affect video compression rate. In our simulations, our a-SAD circuit delivers higher performance compared to previous SAD circuits. We compare power dissipation under iso-performance scenario, where our a-SAD circuit shows 27% power saving compared to a previous design.
Trang Le Dinh Dang, Ik Joon Chang
ISLPED2
2013 Low-complexity decision directed method for carrier frequency offset estimation of IEEE 802.11ad
abstract
In this paper, a new carrier frequency offset (CFO) estimation algorithm is proposed for MIMO-OFDM based wireless local area network (WLAN) systems, IEEE 802.11ad. The existing hardware-efficient auto-correlation scheme which subsamples the short train symbol can be used for low-complexity. However this method has a limitation of performance degradation. Therefore, we combine the subsampled auto-correlation scheme in a decision-directed (DD) method for the performance improvement. We show that the proposed scheme improves BER performance about 3dB with reasonable complexity increase.
Junghyun Ha, Janghyuk Yoon, Ik Joon Chang
ISCAS3
2013 Low complexity image correction using color and focus matching for stereo video coding
abstract
In three-dimensional video (3DV), two cameras capture the same scene from different viewpoints. Color and focus variations between the camera views may deteriorate the 3DV quality and performance of 3DV coding. Therefore, we need to correct the color and focus discrepancy between the camera views. In this paper, we propose algorithms that color and focus correction are combined in a preprocessing step of the stereo video coding. For both color and focus matching, we calculate only disparity vector (DV) once during disparity estimation (DE), since the proposed color and focus correction algorithm can share the disparity vector. The experimental results show that the proposed color and focus correction algorithm provides better image quality and coding performance.
Wooseok Kim, Joohan Kim, Minsu Choi, Ik Joon Chang
ISCAS4
2013 High Performance and Hardware Efficient Multiview Video Coding Frame Scheduling Algorithms and Architectures
abstract
Multiview video coding (MVC) provides more realistic 3-D scenes adding depth information derived from multiple cameras than single-or stereo-view video coding. In MVC, video frames obtained from each view are simply scheduled to corresponding encoding channels. However, under such a conventional scheduling technique the encoding times of each channel may not be identical, degrading encoding performance. To address this problem, this paper proposes two MVC frame scheduling schemes and their architectures: a hardware resource aware scheduling and a frame waiting time aware scheduling (WTaS). Here, WTaS considers the waiting time of each frame stored in on-chip SRAM during frame scheduling, thereby reducing SRAM size significantly. Experimental results show that the proposed frame scheduling schemes provide 29.4% faster processing time, compared to the conventional counterpart. In addition, we can improve the core area, on-chip SRAM area, and the power dissipation by 26.7%, 23.2%, and 26.6%, respectively.
Minsu Choi, Ik Joon Chang
IEEE Trans. Circuits Syst. Video Technol.2
2011 A Priority-Based 6T/8T Hybrid SRAM Architecture for Aggressive Voltage Scaling in Video Applications
abstract
We present a voltage-scalable and process-variation resilient, hybrid memory architecture, suitable for use in MPEG-4 video processors such that power dissipation can be traded for graceful degradation in “quality.” The key innovation in our proposed work is a hybrid memory array, which is a mixture of conventional 6T and 8T SRAM bit-cells. The fundamental premise of our approach lies in the fact that the human visual system is mostly sensitive to higher order bits of luminance pixels in video data. We implemented a preferential storage policy in which the higher order luma bits are stored in robust 8T bit-cells while the lower order bits are stored in conventional 6T bit-cells. This facilitates aggressive scaling of supply voltage in memory as the important luma bits, stored in 8T bit-cells, remain relatively unaffected by voltage scaling. The not-so-important lower order luma bits, stored in 6T bit-cells, if affected, contribute insignificantly to the overall degradation in output video quality. Simulation results show that under iso-area condition, we can obtain at least 32% power savings in the hybrid memory array compared to the conventional 6T SRAM array.
Ik Joon Chang, Debabrata Mohapatra, Kaushik Roy 0001
IEEE Trans. Circuits Syst. Video Technol.1
2011 Robust Level Converter for Sub-Threshold/Super-Threshold Operation: 100 mV to 2.5 V
abstract
For ultra low power application, digital sub-threshold logic design has been explored. Extremely low power supply (VDD) of sub-threshold logic results in significant power reduction. However, it is difficult to convert signals from core logic to input/output (I/O) circuits since core VDD is vastly different from high I/O supply voltage. In this work, we propose a level converter based on dynamic logic style for sub-threshold I/O part, having a large dynamic range of conversion. For the level converter, high voltage clock signal needs to be delivered through separate clock path from core logic, leading to clock synchronization problem between high voltage and low voltage clocks. To overcome this issue, we employed a Clock Synchronizer. A test chip is fabricated in 130-nm CMOS technology in order to verify the proposed technique. Hardware measurement results show that the level converter successfully converts 0.3 V 8 MHz pulse to 2.5 V signal.
Ik Joon Chang, Jae-Joon Kim, Keejong Kim, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2009 A voltage-scalable & process variation resilient hybrid SRAM architecture for MPEG-4 video processors
abstract
We present a voltage-scalable and process-variation resilient memory architecture, suitable for MPEG-4 video processors such that power dissipation can be traded for graceful degradation in "quality". The key innovation in our proposed work is a hybrid memory array, which is mixture of conventional 6T and 8T SRAM bit-cells. The fundamental premise of our approach lies in the fact that human visual system (HVS) is mostly sensitive to higher order bits of luminance pixels in video data. We implemented a preferential storage policy in which the higher order luma bits are stored in robust 8T bit-cells while the lower order bits are stored in conventional 6T bit-cells. This facilitates aggressive scaling of supply voltage in memory as the important luma bits, stored in 8T bit-cells, remain relatively unaffected by voltage scaling. The not-so-important lower order luma bits, stored in 6T bit-cells, if affected, contribute insignificantly to the overall degradation in output video quality. Simulation results show average power savings of up to 56%, in the hybrid memory array compared to the conventional 6T SRAM array implemented in 65nm CMOS. The area overhead and maximum output quality degradation (PSNR) incurred were 11.5% and 0.56 dB, respectively.
Ik Joon Chang, Debabrata Mohapatra, Kaushik Roy 0001
DAC1
2006 Robust level converter design for sub-threshold logic
abstract
The large supply voltage difference between sub-threshlold core logic and I/O makes it extremely challenging to convert signals from core circuit to I/O circuit. In this paper, we propose two novel circuits, Clock Synchronizer and Reduced Swing Inverter to design dynamic and static level converters for sub-threshold logic. Circuit simulations shows that our level converters work at frequency > 500Khz between 20®C and 40®C with a supply voltage of 0.25V.
Ik Joon Chang, Jae-Joon Kim, Kaushik Roy 0001
ISLPED1