VLDB 2026 Research / reviewers in the wild / expert
Yongtae Kim 0001
dblp:77/2860-1
· DBLP profile ↗
13ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0001-7039-973XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | quEStab: Towards Scalable Quantum Circuit Simulation on Multi-GPU using an Extended Stabilizer Formalism
Hyunjoon Shin, Seokhyeon Lee, Myeongjin Kwak, Yongtae Kim 0001 |
ICS | 4 |
| 2025 | QQ: Is 2-bit Enough? Exploiting Quantization to Enhance Computation and Memory Efficiency in Quantum SimulationabstractQuantum computer simulations are essential for validating quantum algorithms but face significant computational and memory challenges due to massive floating-point arithmetic with complex numbers. In this paper, we propose a Quantization for Quantum computer simulation (QQ), an integer-based quantization scheme that considerably reduces these overheads while maintaining simulation accuracy. By exploiting the structure of quantum gate matrices, the proposed method separates real and imaginary components, eliminating zero-value operations. Across 25 quantum benchmark circuits, our quantization scheme reduces matrix-vector multiplications and additions by up to 50% and 66.7%, respectively, leading to energy savings of up to 400,834× compared to conventional quantum simulations. It also maintains simulation fidelity above 0.99 for most benchmark circuits, with over half preserving accuracy even at 2-bit precision. Additionally, it reduces memory usage by up to 93.75%, enabling scalable and efficient quantum computer simulations. Hyoju Seo, Seokhyeon Lee, Yongtae Kim 0001 |
ICCAD | 3 |
| 2024 | Enabling Quantum Computer Simulation under Minimal Precision Floating-Point using Irrational Value DecompositionabstractQuantum computing holds immense promise for solving complex problems beyond the reach of classical computers. However, the realization of practical quantum computers faces significant challenges, including its inherent noise and error rates. To address these challenges and explore the capabilities of quantum computers, accurate and efficient quantum computer simulations are essential. This paper presents a novel quantum computer simulation framework designed to accommodate a wide range of custom reduced precision floating-point formats. Leveraging our proposed quantum computation algorithm, our framework enables the quantum computer simulations using exceptionally low precision floating-point numbers, while maintaining high accuracy levels. Through a systematic analysis, we explore the impact of floating-point precision on the accuracy of the quantum computer simulation, providing optimized bit-widths tailored for floating-point numbers. Particularly, using our proposed quantum computation algorithm, our framework achieves quantum state vector outcomes comparable to full-precision quantum simulations, even with extremely low precision floating-point formats ranging from 4 to 6 bits. Overall, these findings highlight the viability of precision reduction techniques in enhancing the efficiency of quantum computer simulations. Hyoju Seo, Yongtae Kim 0001 |
MASCOTS | 2 |
| 2024 | A comprehensive exploration of approximate DNN models with a novel floating-point simulation framework
Myeongjin Kwak, Jeonggeun Kim, Yongtae Kim 0001 |
Perform. Evaluation | 3 |
| 2023 | Towards Quantized Stochastic Computing by Leveraging Reduced Precision Binary Numbers through Bit TruncationabstractStochastic computing (SC) offers high hardware efficiency and error tolerance but faces challenges, such as the overhead of converting between binary and stochastic forms. This paper introduces a novel quantized SC architecture, significantly reducing stochastic number generator (SNG) hardware complexity. We achieve this by quantizing binary numbers to lower precision using various bit truncation schemes, thereby reducing SNG overhead. Implemented in a 65-nm CMOS process, our proposed quantized SNG reduces area and power by up to 65.5% and 73.0%, respectively, compared to the conventional full-precision SNG. We also demonstrate that our SC schemes have minimal impact on processing quality while greatly improving hardware efficiency, as seen in a digital image processing application. Donghui Lee, Yongtae Kim 0001 |
ICCD | 2 |
| 2023 | TorchAxf: Enabling Rapid Simulation of Approximate DNN Models Using GPU-Based Floating-Point Computing FrameworkabstractThis paper presents an approximate floating-point computing framework TorctiAxf1 that enables fast simulation of various approximate deep neural network (DNN) models, including spiking neural networks (SNNs), using various types of approximate adders and multipliers. Additionally, it supports the standard reduced precision floating-point formats, such as bfloat16, and any user-customized precision representation. TorchAxf leverages GPU acceleration to expedite approximate DNN training and inference running on the PyTorch framework. Any arbitrary approximate arithmetic algorithm with C/C++ behavioral models can be readily integrated with TorchAxf to emulate approximate DNN accelerators. Through extensive experiments, we reveal an appropriate degree of the floating-point arithmetic that can be approximated for DNN models without any significant accuracy loss. We also show that approximate-aware re-training can recover errors and refine pre-trained DNN models under reduced precision formats. Besides, TorchAxf running on GPU enables the simulation time of complex DNN models using approximate arithmetic to reduce up to$43.17\times$compared to the baseline optimized CPU implementation. Myeongjin Kwak, Jeonggeun Kim, Yongtae Kim 0001 |
MASCOTS | 3 |
| 2022 | Do Not Forget: Exploiting Stability-Plasticity Dilemma to Expedite Unsupervised SNN Training for Neuromorphic ProcessorsabstractThis paper presents a novel early training termination technique that significantly improves the training speed and energy efficiency of unsupervised learning-based spiking neural networks (SNNs) by skipping redundant training samples. To achieve early termination, we leveraged the key observation that unsupervised SNNs tend to stably maintain previously learned information and systematically analyze the spike firing activity of the network during training. To make a training termination decision, we exploit the difference between the number of spikes generated by the previous and current input training samples. Our termination algorithm is adopted in an SNN using the spike-timing-dependent plasticity (STDP) learning rule for a pattern classification application. The proposed scheme made an early termination decision with insignificant accuracy performance loss by adequately ignoring redundant training samples. Specifically, it enhances the training speedup and energy efficiency by up to 5.07 × and 5.14 × , respectively, with less than 1 percent points (pp) accuracy loss compared to the baseline counterparts by skipping up to 80% of the training samples. Additionally, when employed in a VLSI-based neuromorphic chip environment, it exhibits up to 4.95 × better energy efficiency than the baseline. Myeongjin Kwak, Yongtae Kim 0001 |
ICCD | 2 |
| 2016 | Neuromorphic Processors with Memristive Synapses: Synaptic Interface and Architectural ExplorationabstractDue to their nonvolatile nature, excellent scalability, and high density, memristive nanodevices provide a promising solution for low-cost on-chip storage. Integrating memristor-based synaptic crossbars into digital neuromorphic processors (DNPs) may facilitate efficient realization of brain-inspired computing. This article investigates architectural design exploration of DNPs with memristive synapses by proposing two synapse readout schemes. The key design tradeoffs involving different analog-to-digital conversions and memory accessing styles are thoroughly investigated. A novel storage strategy optimized for feedforward neural networks is proposed in this work, which greatly reduces the energy and area cost of the memristor array and its peripherals. Qian Wang 0003, Yongtae Kim 0001, Peng Li 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2015 | A Reconfigurable Digital Neuromorphic Processor with Memristive Synaptic Crossbar for Cognitive ComputingabstractThis article presents a brain-inspired reconfigurable digital neuromorphic processor (DNP) architecture for large-scale spiking neural networks. The proposed architecture integrates an arbitrary number of N digital leaky integrate-and-fire (LIF) silicon neurons to mimic their biological counterparts and on-chip learning circuits to realize spike-timing-dependent plasticity (STDP) learning rules. We leverage memristor nanodevices to build an N × N crossbar array to store not only multibit synaptic weight values but also network configuration data with significantly reduced area overhead. Additionally, the crossbar array is designed to be accessible both column- and row-wise to expedite the synaptic weight update process for learning. The proposed digital pulse width modulator (PWM) produces binary pulses with various durations for reading and writing the multilevel memristive crossbar. The proposed column based analog-to-digital conversion (ADC) scheme efficiently accumulates the presynaptic weights of each neuron and reduces silicon area overhead by using a shared arithmetic unit to process the LIF operations of all N neurons. With 256 silicon neurons, learning circuits and 64K synapses, the power dissipation and area of our DNP are 6.45 mW and 1.86 mm 2 , respectively, when implemented in a 90-nm CMOS technology. The functionality of the proposed DNP architecture is demonstrated by realizing an unsupervised-learning based character recognition system. Yongtae Kim 0001, Yong Zhang 0049, Peng Li 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2015 | Energy Efficient Approximate Arithmetic for Error Resilient Neuromorphic ComputingabstractThis brief proposes a novel design scheme for approximate adders and comparators to significantly reduce energy consumption while maintaining a very low error rate. The considerably improved error rate and critical path delay stem from the employed carry prediction technique that leverages the information from less significant input bits in a parallel manner. The proposed designs have been adopted in a VLSI-based neuromorphic character recognition chip with unsupervised learning implemented on chip. The approximation errors of the proposed arithmetic units have been shown to have negligible impact on the training process while archiving good energy efficiency. Yongtae Kim 0001, Yong Zhang 0049, Peng Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | A Parallel Digital VLSI Architecture for Integrated Support Vector Machine Training and ClassificationabstractThis paper presents a parallel digital VLSI architecture for combined support vector machine (SVM) training and classification. For the first time, cascade SVM, a powerful training algorithm, is leveraged to significantly improve the scalability of hardware-based SVM training and develop an efficient parallel VLSI architecture. The presented architecture achieves excellent scalability by spreading the training workload of a given data set over multiple SVM processing units with minimal communication overhead. Hardware-friendly implementation of the cascade algorithm is employed to achieve low hardware overhead and allow for training over data sets of variable size. In the proposed parallel cascade architecture, a multilayer system bus and multiple distributed memories are used to fully exploit parallelism. In addition, the proposed architecture is rather flexible and can be tailored to realize hybrid use of hardware parallel processing and temporal reuse of processing resources, leading to good tradeoffs between throughput, silicon overhead and power dissipation. Several parallel cascade SVM processors have been designed with a commercial 90-nm CMOS technology, which provide up to a 561× training time speedup and a significant estimated 21 859× energy reduction compared with the software SVM algorithm running on a 45-nm commercial general-purpose CPU. Qian Wang 0003, Peng Li 0001, Yongtae Kim 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | An energy efficient approximate adder with carry skip for error resilient neuromorphic VLSI systemsabstractWe propose a novel approximate adder design to significantly reduce energy consumption with a very moderate error rate. The significantly improved error rate and critical path delay stem from the employed carry prediction technique that leverages the information from less significant input bits in a parallel manner. An error magnitude reduction scheme is proposed to further reduce amount of error once detected with low cost. Implemented in a commercial 90 nm CMOS process, it is shown that the proposed adder is up to 2.4× faster and 43% more energy efficient over traditional adders while having an error rate of only 0.18%. The proposed adder has been adopted in a VLSI-based neuromorphic character recognition chip using unsupervised learning. The approximation errors of the proposed adder have been shown to have negligible impact on the training process. Moreover, the energy savings of up to 48.5% over traditional adders is achieved for the neuromorphic circuit with scaled supply level. Finally, we achieve error-free operations by including a low-overhead error correction logic. Yongtae Kim 0001, Yong Zhang 0049, Peng Li 0001 |
ICCAD | 1 |
| 2011 | High effective-resolution built-in jitter characterization with quantization noise shapingabstractA novel built-in jitter characterization architecture combining quantization noise shaping and a partial Vernier delay structure is proposed for high resolution jitter measurement. The effective resolution is optimized at the system level as well as the circuit level. Using 90nm CMOS technology, an area of 0.008mm2 is occupied. The power consumption is 1.85mW. An effective resolution of 1.5ps is achieved. Leyi Yin, Yongtae Kim 0001, Peng Li 0001 |
DAC | 2 |