EDBT 2026 Demo / reviewers in the wild / expert
Chang Eun Song
dblp:377/0261
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0002-4235-6243ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NOVA-PIM: Noise-Aware Hyperdimensional Processing in Memory with Optimized Vector Allocation and Minimal ADCsabstractHyperdimensional computing (HDC) is an emerging brain-inspired paradigm that enables highly efficient and robust inference and learning. Analog processing in memory (PIM) has become a promising solution to accelerate HDC by processing lengthy hypervectors (HVs) directly in memory, thereby reducing costly data movement and leveraging massive parallelism. Despite its efficiency, analog PIM suffers from non-idealities that reduce reliability and accuracy. Although the similarity search stage in HDC is inherently error-tolerant given the high dimensionality of HVs, the encoding stage, which transforms raw input data into HVs, remains sensitive to analog noise. Moreover, encoding accounts for a dominant portion of energy consumption, creating a long-standing bottleneck that limits the overall efficiency of analog PIM-based HDC systems. To overcome this challenge, we propose a noise-aware partitioning scheme that improves HDC inference accuracy by processing a critical subset of HV dimensions digitally, while offloading most of the non-critical dimensions to analog PIM. To further synergize the PIM operations across the two consecutive stages, we eliminate the analog-to-digital converters (ADCs) overhead for encoding by employing pulse width modulation (PWM), allowing direct interfacing with the subsequent similarity search stage. The proposed system achieves a 2.6 × reduction in area, 1.5 × –10.3 × lower energy consumption, and 4.2 × –6.5 × speedup compared with state-of-the-art, while maintaining inference accuracy. Keming Fan, Chang Eun Song, Xuan Wang 0040, Tajana Rosing, Mingu Kang |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | Clo-HDnn: Continual On-Device Learning Accelerator with Hyperdimensional Computing via Progressive SearchabstractClo-HDnn is an on-device learning (ODL) accelerator designed for emerging continual learning (CL) tasks. Clo-HDnn integrates hyperdimensional computing (HDC) along with low-cost Kronecker HD Encoder and weight clustering feature extraction (WCFE) to optimize accuracy and efficiency. Clo-HDnn adopts gradient-free CL to efficiently update and store the learned knowledge in the form of class hypervectors. Its dual-mode operation enables bypassing costly feature ex- traction for simpler datasets, while progressive search reduces complexity by up to $61 \%$ by encoding and comparing only partial query hypervectors. Achieving 4.66 TFLOPS/W (FE) and 3.78 TOPS/W (classifier), Clo-HDnn delivers $7.77 \times$ and $4.85 \times$ higher energy efficiency compared to SOTA ODL accelerators. Chang Eun Song, Keming Fan, Soumil Jain, Gopabandhu Hota, Haichao Yang, Leo Liu, Meng-Fan Chang, Carlos H. Diaz, Gert Cauwenberghs, Tajana Rosing, Mingu Kang |
HCS | 1 |
| 2025 | Hybrid SLC-MLC RRAM Mixed-Signal Processing-in-Memory Architecture for Transformer Acceleration via Gradient RedistributionabstractTransformers, while revolutionary, face challenges due to their demanding computational cost and large data movement.To address this, we propose HyFlexPIM, a novel mixed-signal processingin-memory (PIM) accelerator for inference that flexibly utilizes both single-level cell (SLC) and multi-level cell (MLC) RRAM technologies to trade-off accuracy and efficiency.HyFlexPIM achieves efficient dual-mode operation by utilizing digital PIM for highprecision and write-intensive operations while analog PIM for high parallel and low-precision computations.The analog PIM further distributes tasks between SLC and MLC PIM operations, where a single analog PIM module can be reconfigured to switch between two operations (SLC/MLC) with minimal overhead (<1% for area & energy).Critical weights are allocated to SLC RRAM for high accuracy, while less critical weights are assigned to MLC RRAM to maximize capacity, power, and latency efficiency.However, despite employing such a hybrid mechanism, brute-force mapping on hardware fails to deliver significant benefits due to the limited proportion of weights accelerated by the MLC and the noticeable degradation in accuracy.To maximize the potential of our hybrid hardware architecture, we propose an algorithm co-optimization technique, called gradient redistribution, which uses Singular Value Decomposition (SVD) to decompose and truncate matrices based on their importance, then fine-tune them to concentrate significance into a small subset of weights.By doing so, only 5-10% of the weights have dominantly large gradients, making it favorable for HyFlexPIM by minimizing the use of expensive SLC RRAM while maximizing the efficient MLC RRAM.Our evaluation shows that HyFlexPIM significantly enhances computational throughput and energy efficiency, achieving maximum 1.86× and 1.45× higher than state-of-the-art methods. Chang Eun Song, Priyansh Bhatnagar, Zihan Xia 0002, Nam Sung Kim, Tajana Rosing, Mingu Kang |
ISCA | 1 |
| 2024 | Efficient Transformer Acceleration via Reconfiguration for Encoder and Decoder Models and Sparsity-Aware Algorithm MappingabstractTwo essential computing blocks of Transformers, encoder and decoder, used for summarization and generation stages, respectively, present distinct data flow and computation requirements. This paper proposes an architecture to efficiently support both stages, maximizing the parallelism and hardware utilization. We re-purpose the widely deployed 2D systolic array to inherit its efficiency in processing matrix multiplications and to maintain the compatibility with other models with a minor hardware addition (4.9%/3.9% overhead for area/energy) for the reconfigurability between two modes. The design also incorporates token pruning and bit precision reconfigurability without altering the 2D processing array. We also introduce a tailored data mapping for the attention, dubbed score stationary, which leverages the unique sparsity pattern from token pruning to further reduce power consumption. The proposed architecture achieves 27.4X energy savings and 10.7X performance benefits for decoder processing, while obtaining 2.65X energy reduction for the encoder at iso-throughput, presenting a promising unified solution for these two distinct key tasks. Chang Eun Song, Ashkan Moradifirouzabadi, Tajana Rosing, Mingu Kang |
ISLPED | 1 |