EDBT 2026 Demo / reviewers in the wild / expert
Hyunseok Kwak
dblp:386/9843
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0000-1862-0230ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LoRA-Edge: Tensor-Train-Assisted LoRA for Practical CNN Fine-Tuning on Edge DevicesabstractOn-device fine-tuning of CNNs is essential to with-stand domain shift in edge applications such as Human Activity Recognition (HAR), yet full fine-tuning is infeasible under strict memory, compute, and energy budgets. We present LoRA-Edge, a parameter-efficient fine-tuning (PEFT) method that builds on Low-Rank Adaptation (LoRA) with tensor-train assistance. LoRA-Edge (i) applies Tensor-Train Singular Value Decomposition (TT-SVD) to pre-trained convolutional layers, (ii) selectively updates only the output-side core with zero-initialization to keep the auxiliary path inactive at the start, and (iii) fuses the update back into dense kernels, leaving inference cost unchanged. This design preserves convolutional structure and reduces the number of trainable parameters by up to two orders of magnitude compared to full fine-tuning. Across diverse HAR datasets and CNN backbones, LoRA-Edge achieves accuracy within 4.7% of full fine-tuning while updating at most 1.49% of parameters, consistently outperforming prior parameter-efficient baselines under similar budgets. On a Jetson Orin Nano, TT-SVD initialization and selective-core training yield 1.4–3.8× faster convergence to target F1. LoRA-Edge thus makes structure-aligned, parameter-efficient on-device CNN adaptation practical for edge platforms. Hyunseok Kwak, Kyeongwon Lee, Jae-Jin Lee |
DATE | 1 |
| 2026 | TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AIabstractThe growing demands of distributed learning on resource-constrained edge devices underscore the importance of efficient on-device model compression. Tensor-Train Decomposition (TTD) offers high compression ratios with minimal accuracy loss, yet repeated singular value decompositions (SVDs) and matrix multiplications can impose significant latency and energy costs on low-power processors. In this work, we present TT-Edge, a hardware–software co-designed framework aimed at overcoming these challenges. By splitting SVD into two phases—bidiagonalization and diagonalization, TT-Edge offloads the most compute-intensive tasks to a specialized TTD-Engine. This engine integrates tightly with an existing GEMM accelerator, thereby curtailing the frequent matrix–vector transfers that often undermine system performance and energy efficiency. Implemented on a RISC-V-based edge AI processor, TT-Edge achieves a 1.7× speedup compared to a GEMM-only baseline when compressing a ResNet-32 model via TTD, all while reducing overall energy usage by 40.2%. Notably, these gains come with only a 4% increase in total power and minimal hardware overhead—enabled by a lightweight design that reuses GEMM resources and employs a shared floating-point unit. Our experimental results on both FPGA prototypes and post-synthesis power analysis at 45nm demonstrate that TT-Edge effectively addresses the latency/energy bottlenecks of TTD-based compression in edge environments. Hyunseok Kwak, Kyeongwon Lee, Kyeongpil Min, Chaebin Jung |
DATE | 1 |
| 2025 | Demo Abstract: Radar-PIM-Lite: Ultra-Low-Power PIM Processor for Real-Time UWB Radar Respiration Detection on UAVsabstractWe recently proposed Radar-PIM, a Processing-in-Memory (PIM) solution for real-time, low-power UWB radar respiration detection. To meet stringent energy constraints for UAV-based rescue operations, we further developed Radar-PIM-Lite, significantly reducing resource usage and power consumption. We implemented a processor based on our proposed technology and validated its superior ultra-low-power performance and reliable real-time detection capability through FPGA prototyping and application demonstrations. We will showcase this FPGA-based Radar-PIM-Lite prototype through a live demonstration at ISLPED 2025. Kyeongwon Lee, Hyunseok Kwak, Kyeongpil Min, Chaebin Jung, Sangmin Jeon, Jina Park, Massoud Pedram |
ISLPED | 2 |
| 2024 | STARC: Crafting Low-Power Mixed-Signal Neuromorphic Processors by Bridging SNN Frameworks and Analog DesignsabstractDeveloping low-power neuromorphic processors capable of inferring outcomes from SNN Frameworks presents significant challenges, largely due to the gap between frameworks and analog circuit-based SNNs. This paper analyzes the root of this gap as stemming from over/underflow issues and proposes mixed-signal neurons as a solution, further developing a neural core composed of these neurons. In the development of the neural core, we incorporate a design methodology for application-specific neural core optimization. We advance to develop a neural engine as an independent IP, ultimately introducing the snnTorch Architecture (STARC), an integrated mixed-signal neuromorphic processor architecture. The STARC processor, developed as a prototype, demonstrates operational correctness and exceptional low-power performance. Kyuseung Han, Hyunseok Kwak, Kwang-Il Oh, Sukho Lee, Hyeonguk Jang, Jae-Jin Lee |
ISLPED | 2 |