Hunjun Lee

dblp:250/8881 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-3209-9274ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 TierX: A Simulation Framework for Multi-tier BCI System Design Evaluation and Exploration
abstract
Brain-computer interfaces (BCIs) have made remarkable progress in recent years, driven by advances in neuroscience and clinical applications. For practical use, underlying processing systems must meet strict latency and power budgets. However, existing BCI systems typically rely on a single processing node to handle the entire workload, making it difficult to satisfy these budgets across diverse applications. In this work, we present TierX, the first simulation framework for design space exploration of multi-tier BCI systems. TierX models heterogeneous processing nodes across tiers, including implanted processors, body-attached devices, and external servers, together with diverse communication and powering methods. It navigates the extensive design space to identify optimal (1) workload partitioning options and (2) system configurations that leverage the strengths of each tier. We validate TierX on representative system configurations and demonstrate its effectiveness across diverse use cases.
Seunghyun Song, Yeongwoo Jang, Daye Jung, Kyungsoo Park, Gwangjin Kim, Hunjun Lee, Jerald Yoo, Jangwoo Kim
ASPLOS (2)7
2026 FLUX: Frequency Scaling with Layer-wise Utilization for Energy-Efficient NPU Execution (WIP)
abstract
With the widespread adoption of Deep Neural Networks (DNNs), Neural Processing Units (NPUs) are emerging as energy-efficient alternatives to GPUs through parallel processing and high data reuse. However, since diverse deep learning kernels have different memory and computation resource requirements, a utilization imbalance between memory and computation resources often occurs.
Inho Lee 0002, Ky Yeop Lim, Hyejun Kim, Beomseok Kim, Dongsuk Jeon, Hunjun Lee, Yongjun Park 0001
LCTES6
2025 InfiniMind: A Learning-Optimized Large-Scale Brain-Computer Interface
abstract
Brain-computer interfaces (BCIs) provide an interactive closed-loop connection between the brain and a computer.By employing signal processors implanted within the brain, BCIs are driving innovations across various fields in neuroscience and medicine.Recent studies highlight the need to integrate non-volatile memories (NVMs) into the implanted system for large-scale applications.At the same time, they emphasize the importance of continual learning within the system to address non-stationarities in the recorded signals.This work is the first to address the performance and lifetime issues of deploying learning on NVM-assisted BCI systems.To reduce excessive write overhead associated with learning support, we propose four optimization schemes tailored for BCI workloads.First, update filtering minimizes unnecessary writes by leveraging the sparse and recurring nature of BCI signals.Second, delta buffering exploits temporal locality inherent in BCI signals to minimize NVM writes.Third, out-of-place flushing reduces write amplification by packing multiple sub-page updates into a single page write.Fourth, waveform compression decreases the volume of written data by exploiting the structural characteristics of neural signals.We implement these optimizations in a memory controller and integrate it into the state-of-the-art NVM-assisted BCI system, realizing an endto-end learning-optimized system.Evaluation results show that our system improves performance and lifetime by 5.39× and 23.52×, respectively, on representative continual learning algorithms.
Yeongwoo Jang, Daye Jung, Seunghyun Song, Hunjun Lee, Jangwoo Kim
ISCA4
2024 Rearchitecting a Neuromorphic Processor for Spike-Driven Brain-Computer Interfacing
abstract
Brain-computer interfaces (BCIs) are electrophysiological devices (e.g., electrode arrays) that connect the brain to a computer. They offer neuroscientific and neurological innovations by utilizing a dedicated processor for continuous BCI signal processing. Recent studies propose a scaled-up BCI that adopts an order of magnitude larger number of electrodes to more precisely interface with the brain. As the BCI scales, utilizing a spike-driven processor emerges as an alternative processing method, where the BCI offloads computations to the processor upon detecting spikes. However, the processor design for spike-driven processing has been relatively unexplored compared to that of the continuous processor. In this work, we propose NeuroLobe, a flexible and efficient processor design for spike-driven processing. The key idea is to utilize a neuromorphic processor to take advantage of its event-driven computing nature. We carefully rearchitect the existing neuromorphic system for the purpose of flexibly and efficiently deploying the BCI algorithms. First, we extend the instruction set architecture of the existing neuromorphic processor to flexibly deploy representative spike-driven BCI algorithms. Second, we redesign the connection controller and execution path to improve the performance. Third, we design a custom synchronization unit for scalable processing. Fourth, we implement a custom software stack to minimize load imbalance among the cores. Lastly, we design a multitask controller to simultaneously process multiple algorithms. We evaluate NeuroLobe on four representative BCI algorithms with 11 configurations. Evaluation results show that NeuroLobe surpasses CPU and GPU in terms of speed and energy efficiency.
Hunjun Lee, Yeongwoo Jang, Daye Jung, Seunghyun Song, Jangwoo Kim
MICRO1
2022 NeuroSync: A Scalable and Accurate Brain Simulator Using Safe and Efficient Speculation
abstract
To understand and mimic the working mechanism of the brain, neuroscientists rely on brain simulations that operate in a time-driven manner. The simulation involves evaluating how the neurons change their states over time and transferring spikes to the connected neurons through synapses. It also simulates learning by evaluating how the synapses change their weights according to the spiking activity of the neurons. To explore various behaviors of the brain and thus make great advances, neuroscientists need a methodology to support large-scale simulations in both an accurate and efficient manner. For accurate simulations, existing simulators adopt a time-precise simulation methodology where the simulator computes all the neuronal and the synaptic state changes in time order. Unfortunately, they suffer from significant underutilization and energy inefficiency as the simulator scales.In this paper, we present NeuroSync, a fast, energy-efficient, and scalable hardware-based accelerator for accurate brain simulations. The key idea is to adopt a speculative simulation methodology at a minimum overhead along with architectural support. NeuroSync achieves high efficiency using an optimal dataflow for the speculative simulations. At the same time, it ensures simulation accuracy by carefully designing a rollback and recovery mechanism to handle mis-speculations. To implement the methodology at a low cost, NeuroSync further proposes a speculation-optimal learning simulation method. Our evaluations show that 64-chip NeuroSync achieves 3.37× speedup and 3.81× higher energy efficiency with only 10.96% area overhead. The evaluations also show that NeuroSync is extremely scalable with higher speedup as the system scales.
Hunjun Lee, Chanmyeong Kim, Minseop Kim, Yujin Chung, Jangwoo Kim
HPCA1
2022 3D-FPIM: An Extreme Energy-Efficient DNN Acceleration System Using 3D NAND Flash-Based In-Situ PIM Unit
abstract
The crossbar structure of the nonvolatile memory enables highly parallel and energy-efficient analog matrix-vector-multiply (MVM) operations. To exploit its efficiency, existing works design a mixed-signal deep neural network (DNN) accelerator, which offloads low-precision MVM operations to the memory array. However, they fail to accurately and efficiently support the low-precision networks due to their naive ADC designs. In addition, they cannot be applied to the latest technology nodes due to their premature RRAM-based memory array.In this work, we present 3D-FPIM, an energy-efficient and robust mixed-signal DNN acceleration system. 3D-FPIM is a full-stack 3D NAND flash-based architecture to accurately deploy low-precision networks. We design the hardware stack by carefully architecting a specialized analog-to-digital conversion method and utilizing the three-dimensional structure to achieve high accuracy, energy efficiency, and robustness. To accurately and efficiently deploy the networks, we provide a DNN retraining framework and a customized compiler. For evaluation, we implement an industry-validated circuit-level simulator. The result shows that 3D-FPIM achieves an average of 2.09x higher performance per area and 13.18x higher energy efficiency compared to the baseline 2D RRAM-based accelerator.
Hunjun Lee, Minseop Kim, Dongmoon Min, Joonsung Kim 0001, Jongwon Back, Honam Yoo, Jong-Ho Lee 0002, Jangwoo Kim
MICRO1
2021 NeuroEngine: a hardware-based event-driven simulation system for advanced brain-inspired computing
abstract
Brain-inspired computing aims to understand the cognitive mechanisms of a brain and apply them to advance various areas in computer science. Deep learning is an example to greatly improve the field of pattern recognition and classification by utilizing an artificial neural network (ANN). To exploit advanced mechanisms of a brain and thus make more great advances, researchers need a methodology that can simulate neural networks with higher computational capabilities such as advanced spiking neural networks (SNNs) with two-stage neurons and synaptic delays. However, existing SNN simulation methodologies are too slow and energy-inefficient due to their software-based simulation or hardware-based but time-driven execution mechanisms.
Hunjun Lee, Chanmyeong Kim, Yujin Chung, Jangwoo Kim
ASPLOS1
2021 UC-Check: Characterizing Micro-operation Caches in x86 Processors and Implications in Security and Performance
abstract
The modern x86 processor (e.g., Intel, AMD) translates CISC-style x86 instructions to RISC-style micro operations (uops) as RISC pipelines are more efficient than CISC pipelines. However, this x86 decoding process requires complex hardware logic (i.e., x86 decoder) to identify variable-length x86 instructions, which incurs high translation overhead. To avoid this overhead, the x86 processors adopt a micro-operation cache (uop cache) to bypass the expensive x86 decoder by caching the decoded uops.
Joonsung Kim 0001, Hamin Jang, Hunjun Lee, Jangwoo Kim
MICRO3
2021 An accurate and fair evaluation methodology for SNN-based inferencing with full-stack hardware design space explorations
Hunjun Lee, Chanmyeong Kim, Eunjin Baek, Jangwoo Kim
Neurocomputing1
2019 FlexLearn: Fast and Highly Efficient Brain Simulations Using Flexible On-Chip Learning
abstract
To understand how the human brain works, neuroscientists heavily rely on brain simulations which incorporate the concept of time to their operating model. In the simulations, neurons transmit their signals through synapses whose weights change over time and by the activity of the associated neurons. Such changes in synaptic weights, known as learning, are thought to contribute to memory, and various learning rules exist to model different behaviors of the human brain. Due to the diverse neurons and learning rules, neuroscientists perform the simulations using highly programmable general-purpose processors. Unfortunately, the processors greatly suffer from the high computational overheads of the learning rules. As an alternative, brain simulation accelerators achieve orders of magnitude higher performance; however, they have limited flexibility and cannot support the diverse neurons and learning rules.
Eunjin Baek, Hunjun Lee, Youngsok Kim, Jangwoo Kim
MICRO2