Kyungchul Lee

dblp:169/2968 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0000-0003-1479-8810ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 C-STEP: Compute-Efficient Spiking Transformers with Temporal Exit and Early-Guided Pruning
abstract
Spiking Transformers (STs) have emerged as an efficient alternative to artificial neural networks, yet they still incur substantial computation, motivating computational reduction techniques. In this paper, we present C-STEP, a unified technique for reducing computation in ST inference. First, we introduce LightSoftmax, a novel lightweight scoring technique that enables confidence-based temporal exit with negligible overhead, allowing inputs with confident predictions to terminate at early timesteps. Second, for the inputs that do not exit, we apply early-time-guided dynamic channel pruning to remove low-contribution channels in later timesteps. Third, we devise a synaptic computation scheme that decomposes spikes into a locally common component and token-specific residuals. The common component is computed once and reused across the tokens, preserving functional equivalence. We have designed an end-to-end SNN architecture that seamlessly executes the proposed low complexity schemes. C-STEP reduces synaptic operations by up to 65.4% relative to the original ST backbones.
Kyungchul Lee
DATE1
2026 A Structured-Sparsity-Based Design Approach for Energy-Efficient Spiking Transformer Processing
abstract
Spiking transformers (STs) have emerged as promising architectures that achieve competitive accuracy with artificial neural network (ANN) on large-scale datasets. Despite this progress, the efforts to reduce the computational complexity of spiking self-attention (SSA) have seldom been explored. In this work, we present sparsity-based design approaches to reduce the computational complexity of SSA execution by systematically exploiting the structured sparsity in SSA. The proposed approaches are based on the observation that SSA naturally exhibits structured sparsity, which can be exploited to identify and skip redundant computations in SSA blocks. First, by employing a novel fully spiking SSA operator (FSSA) incorporating additional leaky integrate-and-fire (LIF) neurons, the average structured sparsity of SSA has been increased by$1.8\times $with less than 1% loss in accuracy. In addition, by adopting a filtering strategy that ignores neurons with low spike rate when detecting structured sparsity, additional$1.5\times $structured sparsity has been achieved compared to FSSA with negligible accuracy loss. Then, a sparsity-aware dataflow and hardware design convert the proposed structured sparsity patterns into runtime skipping of computations and weight transfers in SSA block execution. In postsynthesis simulations across the evaluated models, the proposed SSA operator with filtering techniques improve SSA block effective throughput and energy efficiency to 234.44 GOP/s and 319.12 GOP/J while keeping model accuracy loss below 1%, corresponding to 24.3% and 28.2% gains over the baseline spiking neural network (SNN) accelerator.
Hyunseok Jung, Kyungchul Lee, Jongsun Park 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2022 A time-to-first-spike coding and conversion aware training for energy-efficient deep spiking neural network processor design
abstract
In this paper, we present an energy-efficient SNN architecture, which can seamlessly run deep spiking neural networks (SNNs) with improved accuracy. First, we propose a conversion aware training (CAT) to reduce ANN-to-SNN conversion loss without hardware implementation overhead. In the proposed CAT, the activation function developed for simulating SNN during ANN training, is efficiently exploited to reduce the data representation error after conversion. Based on the CAT technique, we also present a time-to-first-spike coding that allows lightweight logarithmic computation by utilizing spike time information. The SNN processor design that supports the proposed techniques has been implemented using 28nm CMOS process. The processor achieves the top-1 accuracies of 91.7%, 67.9% and 57.4% with inference energy of 486.7uJ, 503.6uJ, and 1426uJ to process CIFAR-10, CIFAR-100, and Tiny-ImageNet, respectively, when running VGG-16 with 5bit logarithmic weights.
Dongwoo Lew, Kyungchul Lee, Jongsun Park 0001
DAC2