Kyeongwon Lee

dblp:228/0798 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
0009-0002-0505-3753ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 LoRA-Edge: Tensor-Train-Assisted LoRA for Practical CNN Fine-Tuning on Edge Devices
abstract
On-device fine-tuning of CNNs is essential to with-stand domain shift in edge applications such as Human Activity Recognition (HAR), yet full fine-tuning is infeasible under strict memory, compute, and energy budgets. We present LoRA-Edge, a parameter-efficient fine-tuning (PEFT) method that builds on Low-Rank Adaptation (LoRA) with tensor-train assistance. LoRA-Edge (i) applies Tensor-Train Singular Value Decomposition (TT-SVD) to pre-trained convolutional layers, (ii) selectively updates only the output-side core with zero-initialization to keep the auxiliary path inactive at the start, and (iii) fuses the update back into dense kernels, leaving inference cost unchanged. This design preserves convolutional structure and reduces the number of trainable parameters by up to two orders of magnitude compared to full fine-tuning. Across diverse HAR datasets and CNN backbones, LoRA-Edge achieves accuracy within 4.7% of full fine-tuning while updating at most 1.49% of parameters, consistently outperforming prior parameter-efficient baselines under similar budgets. On a Jetson Orin Nano, TT-SVD initialization and selective-core training yield 1.4–3.8× faster convergence to target F1. LoRA-Edge thus makes structure-aligned, parameter-efficient on-device CNN adaptation practical for edge platforms.
Hyunseok Kwak, Kyeongwon Lee, Jae-Jin Lee
DATE2
2026 TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
abstract
The growing demands of distributed learning on resource-constrained edge devices underscore the importance of efficient on-device model compression. Tensor-Train Decomposition (TTD) offers high compression ratios with minimal accuracy loss, yet repeated singular value decompositions (SVDs) and matrix multiplications can impose significant latency and energy costs on low-power processors. In this work, we present TT-Edge, a hardware–software co-designed framework aimed at overcoming these challenges. By splitting SVD into two phases—bidiagonalization and diagonalization, TT-Edge offloads the most compute-intensive tasks to a specialized TTD-Engine. This engine integrates tightly with an existing GEMM accelerator, thereby curtailing the frequent matrix–vector transfers that often undermine system performance and energy efficiency. Implemented on a RISC-V-based edge AI processor, TT-Edge achieves a 1.7× speedup compared to a GEMM-only baseline when compressing a ResNet-32 model via TTD, all while reducing overall energy usage by 40.2%. Notably, these gains come with only a 4% increase in total power and minimal hardware overhead—enabled by a lightweight design that reuses GEMM resources and employs a shared floating-point unit. Our experimental results on both FPGA prototypes and post-synthesis power analysis at 45nm demonstrate that TT-Edge effectively addresses the latency/energy bottlenecks of TTD-based compression in edge environments.
Hyunseok Kwak, Kyeongwon Lee, Kyeongpil Min, Chaebin Jung
DATE2
2026 Demo Abstract: An Energy-Efficient Crossbar-Free Oscillatory Neural Network Accelerator for Near-Sensor Edge AI
Jeongmin Jin, Mundo Jeong, Kyeongwon Lee, Chaebin Jung, Eun-Su Cho
ISLPED3
2026 SA-Kura: An Energy-Efficient Systolic Array Accelerator for Locally-Coupled Kuramoto Drift in Diffusion Sampling
Jeongmin Jin, Kyeongwon Lee, Mundo Jeong
ISLPED2
2026 CMAX-CAMEL: A Coarse-to-Fine Adaptive, Memory-Efficient, and Low-Power Edge Processor for Contrast Maximization
Kyeongpil Min, Kyeongwon Lee
ISLPED3
2025 HH-PIM: Dynamic Optimization of Power and Performance with Heterogeneous-Hybrid PIM for Edge AI Devices
abstract
Processing-in-Memory (PIM) architectures offer promising solutions for efficiently handling AI applications in energy-constrained edge environments. While traditional PIM designs enhance performance and energy efficiency by reducing data movement between memory and processing units, they are limited in edge devices due to continuous power demands and the storage requirements of large neural network weights in SRAM and DRAM. Hybrid PIM architectures, incorporating nonvolatile memories like MRAM and ReRAM, mitigate these limitations but struggle with a mismatch between fixed computing resources and dynamically changing inference workloads. To address these challenges, this study introduces a Heterogeneous-Hybrid PIM (HH-PIM) architecture, comprising high-performance MRAM-SRAM PIM modules and low-power MRAM-SRAM PIM modules. We further propose a data placement optimization algorithm that dynamically allocates data based on computational demand, maximizing energy efficiency. FPGA prototyping and power simulations with processors featuring HH-PIM and other PIM types demonstrate that the proposed HH-PIM achieves up to 60.43% average energy savings over conventional PIMs while meeting application latency requirements. These results confirm HH-PIM’s suitability for adaptive, energy-efficient AI processing in edge devices.
Sangmin Jeon, Kangju Lee, Kyeongwon Lee
DAC3
2025 Demo Abstract: Radar-PIM-Lite: Ultra-Low-Power PIM Processor for Real-Time UWB Radar Respiration Detection on UAVs
abstract
We recently proposed Radar-PIM, a Processing-in-Memory (PIM) solution for real-time, low-power UWB radar respiration detection. To meet stringent energy constraints for UAV-based rescue operations, we further developed Radar-PIM-Lite, significantly reducing resource usage and power consumption. We implemented a processor based on our proposed technology and validated its superior ultra-low-power performance and reliable real-time detection capability through FPGA prototyping and application demonstrations. We will showcase this FPGA-based Radar-PIM-Lite prototype through a live demonstration at ISLPED 2025.
Kyeongwon Lee, Hyunseok Kwak, Kyeongpil Min, Chaebin Jung, Sangmin Jeon, Jina Park, Massoud Pedram
ISLPED1
2025 Posterior Contraction for Sparse Neural Networks in Besov Spaces with Intrinsic Dimensionality
abstract
This work establishes that sparse Bayesian neural networks achieve optimal posterior contraction rates over anisotropic Besov spaces and their hierarchical compositions. These structures reflect the intrinsic dimensionality of the underlying function, thereby mitigating the curse of dimensionality. Our analysis shows that Bayesian neural networks equipped with either sparse or continuous shrinkage priors attain the optimal rates which are dependent on the intrinsic dimension of the true structures. Moreover, we show that these priors enable rate adaptation, allowing the posterior to contract at the optimal rate even when the smoothness level of the true function is unknown. The proposed framework accommodates a broad class of functions, including additive and multiplicative Besov functions as special cases. These results advance the theoretical foundations of Bayesian neural networks and provide rigorous justification for their practical effectiveness in high-dimensional, structured estimation problems.
Kyeongwon Lee, Lizhen Lin, Jaewoo Park 0002, Seonghyun Jeong
NeurIPS1
2025 Radar-PIM: Developing IoT Processors Utilizing Processing-in-Memory Architecture for Ultrawideband-Radar-Based Respiration Detection
abstract
The adoption of ultrawideband (UWB) radar technology in IoT and healthcare applications for respiration detection is rapidly expanding, opening up a wide array of potential use cases. Despite its burgeoning utility, the integration of UWB radar-based respiration detection in IoT end node devices faces significant challenges due to the memory-intensive nature of these tasks, which strain the capabilities of IoT processors. This article introduces a streamlined UWB radar-based respiration detection application designed for operation on IoT processors, emphasizing that when executed on conventional IoT processors, the limited processing power still results in significant data loss, underscoring the need for enhanced processing solutions. To address these challenges, we propose the adoption of processing-in-memory (PIM) technology and unveil the novel Radar-PIM architecture. This architecture is meticulously engineered to boost the efficiency of respiration detection while ensuring seamless integration with existing embedded processor frameworks. This article extensively describes the Radar-PIM architecture and its operational mechanisms. We further demonstrate its superior performance by implementing and empirically testing a Radar-PIM processor prototype. Next, we present an optimization strategy tailored for designing energy-efficient Radar-PIM processors, specifically adapted for diverse UWB radar-based respiration detection applications. For instance, a Radar-PIM processor prototype, optimized for a particular application, achieved approximately 42% energy savings compared to its unoptimized counterpart and delivered performance nearly three times greater than that of a multicore processor with equivalent power consumption. This demonstrates the transformative potential of our proposed solution in enhancing the capabilities of radar-based respiration detection systems for IoT end nodes.
Kyeongwon Lee, Sangmin Jeon, Kangju Lee, Massoud Pedram
IEEE Internet Things J.1
2024 Day-Night architecture: Development of an ultra-low power RISC-V processor for wearable anomaly detection
abstract
In healthcare, anomaly detection has emerged as a central application. This study presents an ultra-low power processor tailored for wearable devices dedicated to anomaly detection. Introducing a unique Day-Night architecture, the processor is bifurcated into two distinct segments: The Day segment and the Night segment, both of which function autonomously. The Day segment, catering to generic wearable applications, is designed to remain largely inactive, awakening only for specific tasks. This approach leads to considerable power savings by incorporating the Main-CPU and system interconnect, both major power consumers. Conversely, the Night segment is dedicated to real-time anomaly detection using sensor data analytics. It comprises a Sub-CPU and a minimal set of IPs, operating continuously but with minimized power consumption. To further enhance this architecture, the paper presents an ultra-lightweight RISC-V core, All-Night core, specialized for anomaly detection applications, replacing the traditional Sub-CPU. To validate the Day-Night architecture, we developed a prototype processor and implemented it on an FPGA board. An anomaly detection application, optimized for this prototype, was also developed to showcase its functional prowess. Finally, when we synthesized the processor prototype using 45 nm process technology, it affirmed our assertion of achieving an energy reduction of up to 57%.
Eunjin Choi, Jina Park, Kyeongwon Lee, Jae-Jin Lee, Kyuseung Han
J. Syst. Archit.3
2024 Designing Low-Power RISC-V Multicore Processors With a Shared Lightweight Floating Point Unit for IoT Endnodes
abstract
The increasing interest in RISC-V from both academia and industry has motivated the development and release of a number of free, open-source cores based on the RISC-V instruction set architecture. Specifically, the use of lightweight RISC-V cores in processors tailored for IoT endnode devices is on the rise. As the range and complexity of these applications grow, there is an increasing demand for multicore processors that can handle floating-point operations. This poses a significant challenge because most lightweight RISC-V cores are integer cores lacking a floating-point unit (FPU). This limitation makes it difficult to design processors optimized for applications that require floating-point operations concurrently with integer operations. While it is inefficient to have a dedicated FPU per core in a multicore processor (because it would give rise to unnecessary power consumption), it is crucial to find a solution that balances performance and energy efficiency. To address this challenge, we propose to utilize an external lightweight FPU that can be added to any RISC-V integer core, along with a low-power multicore architecture that shares the said FPU. We have applied this concept to design a RISC-V processor that integrates these technologies, implemented it on an FPGA device, and completed the fabrication of a System-on-Chip for functional verification. Our experiments, which involved testing various applications on different processor prototypes, demonstrated significant energy savings of up to 79.6% in a quad-core processor prototype, highlighting the potential energy efficiency of our proposed technology.
Jina Park, Kyuseung Han, Eunjin Choi, Jae-Jin Lee, Kyeongwon Lee, Massoud Pedram
IEEE Trans. Circuits Syst. I Regul. Pap.5
2022 Asymptotic Properties for Bayesian Neural Network in Besov Space
abstract
Neural networks have shown great predictive power when applied to unstructured data such as images and natural languages. The Bayesian neural network captures the uncertainty of prediction by computing the posterior distribution of the model parameters. In this paper, we show that the Bayesian neural network with spikeand-slab prior has posterior consistency with a near minimax optimal convergence rate when the true regression function belongs to the Besov space. The spikeand-slab prior is adaptive to the smoothness of the regression function and the posterior convergence rate does not change even when the smoothness of the regression function is unknown. We also consider the shrinkage prior, which is computationally more feasible than the spike-and-slab prior, and show that it has the same posterior convergence rate as the spike-and-slab prior.
Kyeongwon Lee
NeurIPS1