Changmin Ye

dblp:283/1272 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0009-0004-2610-7794ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 IterL2Norm: Fast Iterative L2-Normalization
abstract
Transformer-based large language models are a memory-bound model whose operation is based on a large amount of data that are marginally reused. Thus, the data movement between a host and accelerator likely dictates the total wall-clock time. Layer normalization is one of the key workloads in the transformer model, following each of multi-head attention and feed-forward network blocks. To reduce data movement, layer normalization needs to be performed on the same chip as the matrix-matrix multiplication engine. To this end, we introduce an iterative L2-normalization method for 1D input (IterL2Norm), ensuring fast convergence to the steady-state solution within five iteration steps and high precision, outperforming the fast inverse square root algorithm in six out of nine cases for FP32 and five out of nine for BFloat16 across the embedding lengths used in the OPT models. Implemented in 32/28nm CMOS, the IterL2Norm macro normalizes d-dimensional vectors, where 64 ≤$d$≤ 1024, with a latency of 116–227 cycles at 100MHz/l.05V.
Changmin Ye, Yonguk Sim, Youngchae Kim, SeongMin Jin, Doo Seok Jeong
DATE1
2025 Optimal strategy for mapping spiking neural networks onto manycore neuromorphic processors
abstract
Manycore digital neuromorphic event processors execute ad hoc event routing between spiking neurons distributed across multiple cores. Due to limited hardware resources, such as on-chip memory capacity, only a limited number of neurons and their fan-in weights can be accommodated per core. This challenge is particularly significant for convolutional layers, where placing neurons from the same layer in different cores hinders weight reuse, as the same weight must be duplicated across cores. To address this, we propose an optimal mapping method for spiking units across multiple cores, considering hardware resource constraints. This method is based on the discrete Lagrange Multiplier Method, which uses a total memory usage as an objective function alongside constraint functions (memory usage per core). Our results show that this method achieves optimal spiking unit distributions with high core memory utilization (> 70%) for the reduced ResNet models.
Changmin Ye, Doo Seok Jeong
ISCAS1
2023 LaCERA: Layer-centric event-routing architecture
Changmin Ye, Vladimir Kornijcuk, Donghyung Yoo, Jeeson Kim, Doo Seok Jeong
Neurocomputing1
2021 Hardware-Efficient Emulation of Leaky Integrate-and-Fire Model Using Template-Scaling-Based Exponential Function Approximation
abstract
We present a method to emulate a leaky integrate-and-fire (LIF) model in a field-programmable gate array (FPGA) in a hardware-efficient manner. The simplified spike-response model (SRM0) is chosen as an LIF model. For the hardware-efficient implementation of SRM0, we adopt the template-scaling-based exponential function approximation (TS-EFA). This method allows high precision and low latency exponential function approximations with the efficient use of hardware resources. We subsequently propose an algorithm for SRM0, which leverages the advantage of TS-EFA. An implementation of 512 neurons conforming to SRM0in an FPGA highlights (i) high precision of SRM0emulation (mean squared error of membrane potential approximation: 4×10-12- 1×10-10), (ii) low latency (eight clock cycles), and (iii) high efficiency in hardware usage (only 125b memory per neuron).
Jeeson Kim, Vladimir Kornijcuk, Changmin Ye, Doo Seok Jeong
IEEE Trans. Circuits Syst. I Regul. Pap.3