EDBT 2026 Demo / reviewers in the wild / expert
Kyeongho Lee
dblp:26/393
· DBLP profile ↗
14ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-2550-2272ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Design Framework of Heterogeneous Approximate DCIM-Based Accelerator for Energy-Efficient NN ProcessingabstractStatic random-access memory (SRAM) based digital compute-in-memory (DCIM) provides error-resilient computation at the expense of considerable power overhead of adder tree. In recent works, DCIM macro based on approximate computing mitigates the adder tree overheads, however, it faces a trade-off between power and neural network (NN) accuracy. The trade-off becomes more complicated in array-level CIM architecture since output channels of NN model have different sensitivities to approximation errors. In this paper, we propose a heterogeneous approximate DCIM-based accelerator design framework that achieves a good energy-accuracy trade-off for a specific NN model. The framework includes three key features: 1) Evolutionary algorithm-based search finds cost-efficient approximation points by pruning the design space. 2) Genetic algorithm-based channel-wise mapping creates heterogeneous approximation methods that effectively reduce DCIM energy consumption while maintaining high accuracy. 3) A hardware generation strategy decides the number of DCIM macros and their sizes, resulting in an energy-efficient DCIM-based accelerator tailored for the given NN model. Experimental results show that employing the proposed heterogeneous channel-wise mapping significantly enhances the energy efficiency compared to a homogeneous mapping. Moreover, the proposed framework can produce heterogeneous DCIM-based accelerators that consume less energy than state-of-the-art approximate DCIM approaches. Kyeongho Lee, Hyeyeong Lee, Jongsun Park 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | TP-DCIM: Transposable Digital SRAM CIM Architecture for Energy-Efficient and High Throughput Transformer AccelerationabstractTo accelerate the execution of transformer models, compute-in-memory (CIM) has been widely adopted. However, the CIM architecture has the drawback of fixed one-way computing structure supporting only horizontal input sharing vertical accumulation (HIVA). So, it faces two major obstacles: 1) Matrix transposition processing needs large hardware overheads, 2) A fixed dataflow incurs CIM underutilization problem, degrading throughput. In this paper, we present a digital SRAM CIM (DCIM)-based transformer accelerator that supports two-way computing: HIVA and vertical input sharing horizontal accumulation (VIHA). The proposed two-way computing DCIM macro features a novel 9T SRAM bitcell and transposable adder tree, which efficiently reduces matrix transposition costs. We also present a novel dataflow to improve overall self-attention latency by resolving CIM underutilization and hiding dynamic input generation latency. In addition, CIM-friendly computation skipping scheme is exploited to enhance energy efficiency with negligible accuracy loss. The simulation results show that the proposed DCIM-based accelerator achieves up to 2.5x latency improvement and 24% energy savings compared to previous CIM-based accelerators. Junwoo Park, Kyeongho Lee, Jongsun Park 0001 |
ICCAD | 2 |
| 2023 | A 2941-TOPS/W Charge-Domain 10T SRAM Compute-in-Memory for Ternary Neural NetworkabstractIn this paper, we present a 10T SRAM compute-in memory (CiM) macro to process the multiplication-accumulation (MAC) operations between ternary-inputs and binary-weights. In the proposed 10T SRAM bitcell, the charge-domain analog computations are employed to improve the noise tolerance of bit-line (BL) signals where the MAC results are represented in CiM. Parallel processing of 3 different analog levels for ternary input activations is also performed in the proposed single 10T bitcell. To reduce the analog-to-digital converter (ADC) bit-resolutions without sacrificing deep neural network (DNN) accuracies, a confined-slope non-uniform integration (CS-NUI) ADC is proposed, which can provide layer-wise adaptive quantization for multiple different layers with different MAC distributions. In addition, by sharing the ADC reference voltage generator in every single column of SRAM array, the ADC area is effectively reduced with improved energy efficiencies of CiM. The$256\times 64.10\text{T}$SRAM CiM macro with the proposed charge-sharing scheme and CS-NUI ADCs has been implemented using 28nm CMOS process. The silicon measurement results show that the proposed CiM shows the accuracies of 98.66% and 88.48% with MNIST dataset on MLP, and CIFAR-10 dataset on VGGNet-7, respectively, with the energy efficiency of 2941-TOPS/W and the area efficiency of 59.584-TOPS/mm2. Sungsoo Cheon, Kyeongho Lee, Jongsun Park 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Low-Cost 7T-SRAM Compute-in-Memory Design Based on Bit-Line Charge-Sharing Based Analog-to-Digital ConversionabstractAlthough compute-in-memory (CIM) is considered as one of the promising solutions to overcome memory wall problem, the variations in analog voltage computation and analog-to-digital-converter (ADC) cost still remain as design challenges. In this paper, we present a 7T SRAM CIM that seamlessly supports multiply-accumulation (MAC) operation between 4-bit inputs and 8-bit weights. In the proposed CIM, highly parallel and robust MAC operations are enabled by exploiting the bit-line charge-sharing scheme to simultaneously process multiple inputs. For the readout of analog MAC values, instead of adopting the conventional ADC structure, the bit-line charge-sharing is efficiently used to reduce the implementation cost of the reference voltage generations. Based on the in-SRAM reference voltage generation and the parallel analog readout in all columns, the proposed CIM efficiently reduces ADC power and area cost. In addition, the variation models from Monte-Carlo simulations are also used during training to reduce the accuracy drop due to process variations. The implementation of 256×64 7T SRAM CIM using 28nm CMOS process shows that it operates in the wide voltage range from 0.6V to 1.2V with energy efficiency of 45.8-TOPS/W at 0.6V. Kyeongho Lee, Joonhyung Kim, Jongsun Park 0001 |
ICCAD | 1 |
| 2022 | A Charge Domain P-8T SRAM Compute-In-Memory with Low-Cost DAC/ADC Operation for 4-bit Input ProcessingabstractThis paper presents a low cost PMOS-based 8T (P-8T) SRAM Compute-In-Memory (CIM) architecture that efficiently per-forms the multiply-accumulate (MAC) operations between 4-bit input activations and 8-bit weights. First, bit-line (BL) charge-sharing technique is employed to design the low-cost and reliable digital-to-analog conversion of 4-bit input activations in the pro-posed SRAM CIM, where the charge domain analog computing provides variation tolerant and linear MAC outputs. The 16 local arrays are also effectively exploited to implement the analog mul-tiplication unit (AMU) that simultaneously produces 16 multipli-cation results between 4-bit input activations and 1-bit weights. For the hardware cost reduction of analog-to-digital converter (ADC) without sacrificing DNN accuracy, hardware aware system simulations are performed to decide the ADC bit-resolutions and the number of activated rows in the proposed CIM macro. In addition, for the ADC operation, the AMU-based reference col-umns are utilized for generating ADC reference voltages, with which low-cost 4-bit coarse-fine flash ADC has been designed. The 256×80 P-8T SRAM CIM macro implementation using 28nm CMOS process shows that the proposed CIM shows the accuracies of 91.46% and 66.67% with CIFAR-10 and CIFAR-100 dataset, respectively, with the energy efficiency of 50.07-TOPS/W. Joonhyung Kim, Kyeongho Lee, Jongsun Park 0001 |
ISLPED | 2 |
| 2021 | A Charge-Sharing based 8T SRAM In-Memory Computing for Edge DNN AccelerationabstractThis paper presents a charge-sharing based customized 8T SRAM in-memory computing (IMC) architecture. In the proposed IMC approach, the multiply-accumulate (MAC) operation of multi-bit activations and weights is supported using the charge sharing between bit-line (BL) parasitic capacitances. The area-efficient customized 8T SRAM macro can achieve robust and voltage-scalable MAC operations due to the charge-domain computation. We also propose a split capacitor structure-based 5/6-bit reconfigurable successive approximation register analog-to-digital converter (SAR-ADC) to reduce the hardware cost of an analog readout circuit while supporting higher precision MAC operations. The proposed reconfigurable SAR-ADC has been exploited to implement layer-by-layer mixed bit-precisions in convolution layer for increasing energy efficiency with negligible accuracy loss. The 256×64 8T SRAM IMC macro has been implemented using 28nm CMOS process technology. The proposed SRAM macro achieves 11. 20-TOPS/W with a maximum clock frequency of 125MHz at 1. 0V. It also supports supply voltage scaling from 0.5V to 1.1V with the energy efficiency ranging from 8.3-TOPS/W to 35.4-TOPS/W within 1 % accuracy loss. Kyeongho Lee, Sungsoo Cheon, Joongho Jo, Woong Choi, Jongsun Park 0001 |
DAC | 1 |
| 2020 | Bit Parallel 6T SRAM In-memory Computing with Reconfigurable Bit-PrecisionabstractThis paper presents 6T SRAM cell-based bit-parallel in-memory computing (IMC) architecture to support various computations with reconfigurable bit-precision. In the proposed technique, bit-line computation is performed with a short WL followed by BL boosting circuits, which can reduce BL computing delays. By per-forming carry-propagation between each near-memory circuit, bit-parallel complex computations are also enabled by iterating operations with low latency. In addition, reconfigurable bit-precision is also supported based on carry-propagation size. Our 128KB in/near memory computing architecture has been implemented using a 28nm CMOS process, and it can achieve 2.25GHz clock frequency at 0.9V with 5.2% of area overhead. The proposed architecture also achieves 0.68, 8.09 TOPS/W for the parallel addition and multiplication, respectively. In addition, the proposed work also supports a wide range of supply voltage, from 0.6V to 1.1V. Kyeongho Lee, Jinho Jeong 0002, Sungsoo Cheon, Woong Choi, Jongsun Park 0001 |
DAC | 1 |
| 2019 | Low Cost Ternary Content Addressable Memory Based on Early Termination Precharge SchemeabstractIn this paper, we present early termination match-line (ML) precharge scheme for low power and high speed ternary content addressable memory (TCAM). In the proposed TCAM, by employing the pre-decision based early termination, unnecessary ML precharging has been effectively eliminated while improving the search speed and achieving error-free operation. The reference voltage generator used to implement the proposed early termination approach can be simply designed using dummy row without large area overhead. According to the post-layout simulations with the 65nm CMOS process, the proposed early termination ML precharge scheme shows up to 30.4% of sensing delay improvement and 65.9% of ML power savings compared to the conventional approach. It also shows 8% of FOM (energy/bit/search) improvement compared to state-of-the-art works. Kyeongho Lee, Geon Ko, Jongsun Park 0001 |
ISCAS | 1 |
| 2018 | Content addressable memory based binarized neural network accelerator using time-domain signal processingabstractBinarized neural network (BNN) is one of the most promising solution for low-cost convolutional neural network acceleration. Since BNN is based on binarized bit-level operations, there exist great opportunities to reduce power-hungry data transfers and complex arithmetic operations. In this paper, we propose a content addressable memory (CAM) based BNN accelerator. By using time-domain signal processing, the huge convolution operations of BNN can be effectively replaced to the CAM search operation. In addition, thanks to fully parallel search of CAM, the parallel convolution operations for non-overlapped filtering window is enabled for high throughput data processing. To verify the effectiveness of the proposed CAM based BNN accelerator, the convolutional layer of LeNet-5 model has been implemented using 65nm CMOS technology. The implementation results show that the proposed BNN accelerator achieves 9.4% and 38.5% of area and energy savings, respectively. The parallel convolution operation of the proposed approach also shows 2.4x improved processing time. Woong Choi, Kwanghyo Jeong, Kyungrak Choi, Kyeongho Lee, Jongsun Park 0001 |
DAC | 4 |
| 2018 | Low Cost Ternary Content Addressable Memory Using Adaptive Matchline Discharging SchemeabstractThis paper presents an adaptive match-line (ML) discharging scheme for low power and high speed ternary content addressable memory (TCAM). In the proposed TCAM, by employing the gated ML pulldown path and ML boosting scheme, the redundant ML discharging and SL switching are eliminated while improving the search speed. By considering the number of mismatch and ML discharging speed, the ML discharging is adaptively controlled in the proposed TCAM. The simulation results with the 65nm CMOS technology show that the proposed adaptive ML discharging scheme improves up to 19% of sensing delay and saves 81% of ML power compared to the conventional approach. When compared with the state-of-the-art work, the post-layout simulations show 10% improvement of FOM (energy/bit/search). Woong Choi, Kyeongho Lee, Jongsun Park 0001 |
ISCAS | 2 |
| 2015 | A Preliminary Study on Human Trust Measurements by EEG for Human-Machine InteractionsabstractWe propose a novel experiment paradigm to measure human trust on machine during a collaborative and egoistic theory-of-mind game. To show a different level of human trust on machine partners, we control the technical capability and humanlike cues of the autonomous agent in the cognitive experiments while recording participant's electroencephalography (EEG). The measured human trust values at various situations will be used to develop a dynamic trust model for efficient human-machine systems. Suh-Yeon Dong, Bo-Kyeong Kim, Kyeongho Lee, Soo-Young Lee |
HAI | 3 |
| 2015 | Implicit Shopping Intention Recognition with Eye Tracking Data and Response TimeabstractImplicit intention is the intention that is not expressed externally but having in one's mind. Implicit intention is difficult to be recognized, but it can be significant information if it is recognized with some measures. When people buy something, they also have implicit intention in their mind, whether I buy this or not. We proposed an experimental paradigm to recognize shopper's implicit intention, and the result of experiment was analyzed in this paper. On the experiment, subjects were instructed to select items to buy from the candidates, and eye-tracking and speech data were recorded during the selection. On data analysis, measures discriminating the existence of implicit shopping intention were selected and compared. From the result, fixation duration, fixation count, multiplication of first fixation duration, and visit count showed different tendency between two cases: when people have intention to buy it and when people do not have. By using this standards, implicit shopping intention of people can be recognized. Dong-Gun Lee, Kyeongho Lee, Soo-Young Lee |
HAI | 2 |
| 2008 | Wire-speed application flow generation in hardware platform for multi-gigabit traffic monitoringabstractFlow-based traffic monitoring has received strong attention from the traffic measurement research communities. Despite the value provided by flow measurement, its usage has limited for relatively lower speed traffic mainly due to the performance impact it introduces. Recent traffic measurement research has tried to overcome such a limitation of simple flow-based monitoring by utilizing payload inspection for applications signatures or/and by identifying target application group based on common traffic characteristics. Such improvement, however, requires much higher analysis performance, particularly when it comes to the monitoring of high-speed links like 2.5 Gbps or higher. Thus, the traditional approach which consists of traffic capture in a dedicated hardware and traffic analysis on a server with database may be unable to meet such harsh requirements. Although a dedicated hardware is used for traffic measurement, the generation of application flows is normally done in the software. In this paper, we propose our novel traffic measurement methodology which pushes application flow generation functionality into the network processor (NP) based hardware to meet such requirements. We have embedded packet capturing capability without loss, deep packet inspection, and flow generation into the hardware. We describe major features, design concepts, implementation, and performance evaluation result. Taesang Choi, Sangsik Yoon, Dongwon Kang, Joonkyung Lee, Kyeongho Lee |
NOMS | 6 |
| 2006 | Novel Traffic Measurement Methodology for High Precision Applications Awareness in Multi-gigabit Networks
Taesang Choi, Sangsik Yoon, Dongwon Kang, Joonkyung Lee, Kyeongho Lee |
APNOMS | 6 |