VLDB 2026 Research / reviewers in the wild / expert
Kyuseung Han
dblp:56/7539
· DBLP profile ↗
20ranked-venue papers
11as first author
9since 2021 · last 2026
0000-0002-9151-3447ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 9 first-author · 8 since 2021Software engineering, systems software and programming languages · 4 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Neural Thresholding on Mixed-Signal Neuromorphic Processors Enabled by Integrated Learning and Hardware DesignabstractSpiking neural networks (SNNs) can improve inference accuracy through joint optimization of synaptic weights and neuronal thresholds. However, mixed-signal neuromorphic processors, which are designed for energy efficiency using analog circuits, face practical limitations. In particular, digital to analog converters (DACs) often lack sufficient resolution to represent the large threshold values required by joint optimization. To address this issue, we propose a mixed-signal neuromorphic processor architecture that shifts threshold control to digital logic. This approach removes the need for high-resolution DACs and allows dynamic threshold adjustment without modifying the analog neural core. We also propose a learning method tailored to this architecture. We evaluate the proposed design on five image classification benchmarks, measuring accuracy, latency, and energy consumption. The results show that our architecture consistently improves accuracy across benchmarks while incurring only minimal latency and energy overhead. This demonstrates that the proven benefits of joint weight and threshold learning can be realized in energy efficient analog hardware. Kyuseung Han, Kwang-Il Oh, Sukho Lee, Hyeonguk Jang, Jae-Jin Lee, Sooyoung Jang |
DATE | 1 |
| 2025 | NPX: Automating Neuromorphic Processor Design from Spike-Based Learning to FPGA PrototypingabstractNeuromorphic Processor eXpress (NPX) is a framework designed to facilitate the development of lightweight neuromorphic processors. To evaluate its efficacy, we conducted a case study involving the FPGA-based implementation of a traffic sign recognition system using an NPX-generated neuromorphic processor. The prototype integrates camera input and OLED output, successfully demonstrating full functionality. This case study confirms that NPX substantially streamlines the design and deployment of efficient neuromorphic processors for embedded artificial intelligence applications. Kyuseung Han, Hyeonguk Jang, Sukho Lee, Sung-Eun Kim, Kyudong Hwang, Jae-Jin Lee |
FPL | 1 |
| 2025 | NeuGEMM: A Reordering-Free Unified GEMM-Conv2D Accelerator for Lightweight Neuromorphic ProcessorsabstractNeuromorphic inference applications primarily rely on general matrix-matrix multiplication (GEMM) and twodimensional convolution (Conv2D) operations. When conventional artificial neural network (ANN) acceleration techniques are employed, these computations often necessitate extensive data reordering, which imposes significant overheads, especially in lightweight embedded systems with limited CPU and memory bandwidth. To address this challenge, we propose a unified accelerator architecture executing GEMM and Conv2D operations without data reordering. The accelerator is co-designed with a neuromorphic software framework tailored for lightweight embedded systems. To validate effectiveness, we implement a neuromorphic processor incorporating the proposed accelerator on an FPGA. Evaluation results across four representative neuromorphic applications demonstrate that the proposed design reduces execution time and energy consumption by 69% and 89%, respectively, compared to conventional ANN accelerators. Hyeonguk Jang, Sukho Lee, Jae-Jin Lee, Kyuseung Han |
FPL | 4 |
| 2024 | STARC: Crafting Low-Power Mixed-Signal Neuromorphic Processors by Bridging SNN Frameworks and Analog DesignsabstractDeveloping low-power neuromorphic processors capable of inferring outcomes from SNN Frameworks presents significant challenges, largely due to the gap between frameworks and analog circuit-based SNNs. This paper analyzes the root of this gap as stemming from over/underflow issues and proposes mixed-signal neurons as a solution, further developing a neural core composed of these neurons. In the development of the neural core, we incorporate a design methodology for application-specific neural core optimization. We advance to develop a neural engine as an independent IP, ultimately introducing the snnTorch Architecture (STARC), an integrated mixed-signal neuromorphic processor architecture. The STARC processor, developed as a prototype, demonstrates operational correctness and exceptional low-power performance. Kyuseung Han, Hyunseok Kwak, Kwang-Il Oh, Sukho Lee, Hyeonguk Jang, Jae-Jin Lee |
ISLPED | 1 |
| 2024 | Day-Night architecture: Development of an ultra-low power RISC-V processor for wearable anomaly detectionabstractIn healthcare, anomaly detection has emerged as a central application. This study presents an ultra-low power processor tailored for wearable devices dedicated to anomaly detection. Introducing a unique Day-Night architecture, the processor is bifurcated into two distinct segments: The Day segment and the Night segment, both of which function autonomously. The Day segment, catering to generic wearable applications, is designed to remain largely inactive, awakening only for specific tasks. This approach leads to considerable power savings by incorporating the Main-CPU and system interconnect, both major power consumers. Conversely, the Night segment is dedicated to real-time anomaly detection using sensor data analytics. It comprises a Sub-CPU and a minimal set of IPs, operating continuously but with minimized power consumption. To further enhance this architecture, the paper presents an ultra-lightweight RISC-V core, All-Night core, specialized for anomaly detection applications, replacing the traditional Sub-CPU. To validate the Day-Night architecture, we developed a prototype processor and implemented it on an FPGA board. An anomaly detection application, optimized for this prototype, was also developed to showcase its functional prowess. Finally, when we synthesized the processor prototype using 45 nm process technology, it affirmed our assertion of achieving an energy reduction of up to 57%. Eunjin Choi, Jina Park, Kyeongwon Lee, Jae-Jin Lee, Kyuseung Han |
J. Syst. Archit. | 5 |
| 2024 | Designing Low-Power RISC-V Multicore Processors With a Shared Lightweight Floating Point Unit for IoT EndnodesabstractThe increasing interest in RISC-V from both academia and industry has motivated the development and release of a number of free, open-source cores based on the RISC-V instruction set architecture. Specifically, the use of lightweight RISC-V cores in processors tailored for IoT endnode devices is on the rise. As the range and complexity of these applications grow, there is an increasing demand for multicore processors that can handle floating-point operations. This poses a significant challenge because most lightweight RISC-V cores are integer cores lacking a floating-point unit (FPU). This limitation makes it difficult to design processors optimized for applications that require floating-point operations concurrently with integer operations. While it is inefficient to have a dedicated FPU per core in a multicore processor (because it would give rise to unnecessary power consumption), it is crucial to find a solution that balances performance and energy efficiency. To address this challenge, we propose to utilize an external lightweight FPU that can be added to any RISC-V integer core, along with a low-power multicore architecture that shares the said FPU. We have applied this concept to design a RISC-V processor that integrates these technologies, implemented it on an FPGA device, and completed the fabrication of a System-on-Chip for functional verification. Our experiments, which involved testing various applications on different processor prototypes, demonstrated significant energy savings of up to 79.6% in a quad-core processor prototype, highlighting the potential energy efficiency of our proposed technology. Jina Park, Kyuseung Han, Eunjin Choi, Jae-Jin Lee, Kyeongwon Lee, Massoud Pedram |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | Developing an Ultra-low Power RISC-V Processor for Anomaly DetectionabstractThis paper aims to develop an ultra-low power processor for wearable devices for anomaly detection. To this end, this paper proposes a processor architecture that divides the architecture into a part for general applications running on wearable devices (day part) and a part that performs anomaly detection by analyzing sensor data (night parts), and each part operates completely independently. This day-night architecture allows the day part, which contains the power-hungry main-CPU and system interconnect, to be turned off most of the time except for intermittent work, and the night part, which consists only of the sub-CPU and minimal IPs, can run all the time with low power. By developing a processor based on the proposed processor architecture, the design verification of the proposed technology and the superiority of power saving are demonstrated. Jina Park, Eunjin Choi, Kyungwon Lee, Jae-Jin Lee, Kyuseung Han |
DATE | 5 |
| 2023 | Florian: Developing a Low-Power RISC-V Multicore Processor with a Shared Lightweight FPUabstractAs applications running on lightweight RISC-V processors become increasingly diverse and complex, the need for multicore processors supporting floating-point units (FPUs) is riseing, making processor designs using existing open-source RISC-V cores challenging. With the exception of a very few, most open lightweight RISC-V cores are integer cores without FPUs, which greatly reduces the design exploration space, making it impossible to design a processor optimized for each application. For example, most of these applications mainly perform integer operations, but occasionally perform floating-point operations. For them, a multicore processor with FPU per core is overkill and wastes power, which is a critical problem for processors where low-power design is paramount. To address the problem, we propose an external lightweight FPU that can be attached to any RISC-V integer core and a low-power multicore architecture using the designed FPU. For verification, we designed a RISC-V processor that implements all the proposed technologies, prototyped it on an FPGA device, and finally fabricated it as a System-on-Chip. Through experiments, it was confirmed that the proposed technology can cut energy consumption energy by up to 23%. Jina Park, Kyuseung Han, Eunjin Choi, Sukho Lee, Jae-Jin Lee, Massoud Pedram |
ISLPED | 2 |
| 2021 | Developing TEI-Aware Ultralow-Power SoC Platforms for IoT End NodesabstractRanging from circuit-level characterization to designing a platform architecture, developing a design automation tool, and fabricating a System on Chip (SoC), this article deals with the entire development process for ultralow-power (ULP) SoCs for Internet-of-Things (IoT) end nodes. More precisely, this article first focuses on the unique characteristics of the ULP circuits, the temperature effect inversion (TEI), i.e., the delay of the ULP circuits decreases with increasing temperature. Existing TEI-aware low-power (TEI-LP) techniques have incredible potential to further reduce the power consumption of conventional ULP SoCs, but there is a critical limitation to be widely adopted in real SoCs. To address this limitation and realize the ULP SoCs that can fully benefit from the TEI-LP techniques, this article proposes a new TEI-inspired SoC platform (TIP) architecture. On top of that, taking into account that the highly complex, time consuming, and labor-intensive development process of these ULP SoCs may hinder their widespread use for IoT end nodes, this article presents a new electronic design automation tool to accelerate ULP SoC development, RISC-V express (RVX). Finally, by using the RVX, this article introduces a TIP prototyping chip fabricated in 28-nm FD-SOI technology. This chip demonstrates that power savings of up to 35% can be achieved by lowering the supply voltage from 0.54 to 0.48 V at 25 °C and 0.44 V at 80 °C while continuing to operate at a target 50-MHz clock frequency. Kyuseung Han, Sukho Lee, Kwang-Il Oh, Younghwan Bae, Hyeonguk Jang, Jae-Jin Lee, Massoud Pedram |
IEEE Internet Things J. | 1 |
| 2019 | TIP: A Temperature Effect Inversion-Aware Ultra-Low Power System-on-Chip PlatformabstractResearchers have been trying to exploit the temperature effect inversion (TEI) phenomenon to improve energy efficiency of system-on-chip (SoC) designs without sacrificing its performance. However, TEI-aware low power methods have a critical limitation in that they can only be applied to components within the SoC that do not contain long (global) wires. This is because wire delays continue to increase with rising temperatures irrespective of the operating supply voltage level, which tends to cancel out positive effects of the TEI phenomenon in SoCs. To tackle this limitation and thoroughly utilize the TEI-aware methods, this paper presents new TEI-inspired SoC platform (called TIP), which relies on network-on-chip architecture (called μNoC) to realize system interconnects. The μNoC successfully reduces the total number and length of global wires. By fabricating a TIP prototyping chip in Samsung 28nm FD-SOI technology, we verify the effectiveness of TIP. Extensive post-fabrication measurements demonstrate that the chip while continuing to operate at a target 50MHz clock frequency can lower its supply voltage from 0.54V to 0.48V at 25°C and to 0.44V at 80°C, which results in up to 35% power saving. Kyuseung Han, Sukho Lee, Jae-Jin Lee, Massoud Pedram |
ISLPED | 1 |
| 2019 | TEI-ULP: Exploiting Body Biasing to Improve the TEI-Aware Ultralow Power MethodsabstractTemperature effect inversion (TEI) phenomenon in ultralow power (ULP) very large scale integration circuits has been identified as an important effect by both academia and industry. Although a number of ULP methods that attempt to exploit the TEI phenomenon have been proposed, the small size of the design exploration space when applying these methods to ULP circuits hinders them from achieving their full potential. This is mainly due to the limited granularity of the supply voltage level control. Starting with an intuition that the body biasing (BB) technique is a key to overcome this limitation, this paper exploits the BB technique along with the TEI-aware voltage scaling (TEI-VS) method and TEI-aware frequency scaling (TEI-FS) method, so as to substantially increase the design spaces of these methods. Techniques for optimally combining the BB technique with TEI-VS and TEI-FS are introduced. Simulation results with the latest commercial CMOS process technologies for ULP designs demonstrate the effectiveness of the proposed methodology. Jae-Jin Lee, Kyuseung Han, Joongheon Kim, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | TEI-NoC: Optimizing Ultralow Power NoCs Exploiting the Temperature Effect InversionabstractThe era of the Internet of Things (IoT) is upon us. In this era, minimizing power consumption becomes a primary concern of system-on-chip designers. Ultralow power (ULP) very large-scale integration circuits have been receiving considerable interest from both academia and industry as the best-suited techniques for IoT devices, which can take full advantage of power-saving that voltage scaling potentially achieves. Consequently, research on ULP designs has begun to yield tangible outcomes, namely ULP circuits. However, little attention has been paid to ULP network-on-chip (NoC), although the NoC is an essential of the ULP chips, and its power consumption accounts for a significant portion of the total power. This paper focuses on ULP NoCs, and presents a new power management method that exploits delay versus temperature characteristics of ULP circuits. Recent studies on ULP circuits show that delay versus temperature characteristics are fundamentally different from normal circuits, i.e., the delay of the ULP circuits implemented in state-of-the-art bulk CMOS operating at low supply voltages or in FinFET technologies decreases with increasing temperature, a phenomenon known as the temperature effect inversion (TEI). Starting with an intuition that at a certain temperature point, power savings without performance penalty can be achieved by increasing the router frequency to create the opportunity to turn off some routers in ULP NoCs, or by decreasing the NoC supply voltage level, an optimization method is presented to maximize the power savings with minor performance penalty. To validate the proposed method, a concrete ULP NoC simulator, TEI-Noxim, has been developed. Experimental results demonstrate that TEI-aware NoC achieves an average of 36.0% power reduction over 21 applications. Kyuseung Han, Jae-Jin Lee, Jinho Lee 0001, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2017 | TEI-power: Temperature Effect Inversion-Aware Dynamic Thermal ManagementabstractFinFETs have emerged as a promising replacement for planar CMOS devices in sub-20nm technology nodes. However, based on the temperature effect inversion (TEI) phenomenon observed in FinFET devices, the delay characteristics of FinFET circuits in sub-, near-, and superthreshold voltage regimes may be fundamentally different from those of CMOS circuits with nominal voltage operation. For example, FinFET circuits may run faster in higher temperatures. Therefore, the existing CMOS-based and TEI-unaware dynamic power and thermal management techniques would not be applicable. In this article, we present TEI-power, a dynamic voltage and frequency scaling--based dynamic thermal management technique that considers the TEI phenomenon and also the superlinear dependencies of power consumption components on the temperature and outlines a real-time trade-off between delay and power consumption as a function of the chip temperature to provide significant energy savings, with no performance penalty—namely, up to 42% energy savings for small circuits where the logic cell delay is dominant and up to 36% energy savings for larger circuits where the interconnect delay is considerable. Kyuseung Han, Yanzhi Wang 0001, Tiansong Cui, Shahin Nazarian, Massoud Pedram |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2014 | Leveraging parallelism in the presence of control flow on CGRAsabstractCoarse-Grained Reconfigurable Architectures (CGRAs) are suitable for accelerating data-intensive applications in embedded systems due to high performance and power efficiency. However, as application programs become complex having more control flows in them, it becomes harder to accelerate such programs on CGRAs. Previous researches on this issue have focused on correct execution of control flows rather than their acceleration. This paper reveals how control flows degrade the performance of programs and proposes a software approaches to accelerating control flows by exploiting parallelism residing in each conditionals as well as among conditionals. Experiments show that our proposed techniques improve performance by 2.51 times on average. Jihyun Ryoo, Kyuseung Han, Kiyoung Choi |
ASP-DAC | 2 |
| 2014 | Design of a coarse-grained reconfigurable architecture with floating-point support and comparative study
Manhwee Jo, Kyuseung Han, Kiyoung Choi |
Integr. | 3 |
| 2014 | Software-Level Approaches for Tolerating Transient Faults in a Coarse-GrainedReconfigurable ArchitectureabstractCoarse-grained reconfigurable architectures have drawn increasing attention due to their merits in performance and flexibility. Typically, they have many processing elements in the form of an array, which is suitable for implementing spatial redundancy used for fault-tolerant systems design. This paper presents a purely software-level approach to implementing transient-fault-tolerance on an existing processing element array without any modification to the architecture. It includes automated design flow to construct a fault-tolerant system and mathematical modeling for analyzing system reliability. Experiments with real-world applications show the effectiveness of the proposed approaches in terms of yield enhancement and system reliability. Kyuseung Han, Ganghee Lee, Kiyoung Choi |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2013 | Compiling control-intensive loops for CGRAs with state-based full predicationabstractPredication is an essential technique to accelerate kernels with control flow on CGRAs. While state-based full predication (SFP) can remove wasteful power consumption on issuing/decoding instructions from conventional full predication, generating code for SFP is challenging for general CGRAs, especially when there are multiple conditionals to be handled due to exploiting data level parallelism. In this paper, we present a novel compiler framework addressing central issues such as how to express the parallelism between multiple conditionals, and how to allocate resources to them to maximize the parallelism. In particular, by separating the handling of control flow and data flow, our framework can be integrated with conventional mapping algorithms for mapping data flow. Experimental results demonstrate that our framework can find and exploit parallelism between multiple conditionals, thereby leading to 2.21 times higher performance on average than a naive approach. Kyuseung Han, Kiyoung Choi, Jongeun Lee |
DATE | 1 |
| 2013 | Power-Efficient Predication Techniques for Acceleration of Control Flow Execution on CGRAabstractCoarse-grained reconfigurable architecture typically has an array of processing elements which are controlled by a centralized unit. This makes it difficult to execute programs having control divergence among PEs without predication. However, conventional predication techniques have a negative impact on both performance and power consumption due to longer instruction words and unnecessary instruction-fetching decoding nullifying steps. This article reveals performance and power issues in predicated execution which have not been well-addressed yet. Furthermore, it proposes fast and power-efficient predication mechanisms. Experiments conducted through gate-level simulation show that our mechanism improves energy-delay product by 11.9% to 23.8% on average. Kyuseung Han, Junwhan Ahn, Kiyoung Choi |
ACM Trans. Archit. Code Optim. | 1 |
| 2012 | State-based full predication for low power coarse-grained reconfigurable architectureabstractIt has been one of the most fundamental challenges in architecture design to achieve high performance with low power while maintaining flexibility. Parallel architectures such as coarse-grained reconfigurable architecture, where multiple PEs are tightly coupled with each other, can be a viable solution to the problem. However, the PEs are typically controlled by a centralized control unit, which makes it hard to parallelize programs requiring different control of each PE. To overcome this limitation, it is essential to convert control flows into data flows by adopting the predicated execution technique, but it may incur additional power consumption. This paper reveals power issues in the predicated execution and proposes a novel technique to mitigate power overhead of predicated execution. Contrary to the conventional approach, the proposed mechanism can decide whether to suppress instruction execution or not without decoding the instructions and does not require additional instruction bits, thereby resulting in energy savings. Experimental results show that energy consumed by the reconfigurable array and its configuration memory is reduced by up to 23.9%. Kyuseung Han, Kiyoung Choi |
DATE | 1 |
| 2010 | Acceleration of control flow on CGRA using advanced predicated executionabstractCoarse-grained reconfigurable array is a very attractive architecture from the viewpoint of performance and flexibility. However, because the performance improvement is achieved by exploiting parallelism, the architecture is typically poor at handling control flow, which is sequential in nature. There have been many attempts to overcome this problem by using predicated execution techniques; however, they do not support all types of control flow or suffer from performance degradation in doing so. In addition, predicated execution schemes in general require a longer execution time because both the if- and else-paths are always executed. This paper proposes advanced predicated execution techniques that can handle and accelerate all types of control flow with only 2% hardware overhead. These techniques can also be easily extended to general SIMD machines. We implemented these techniques on a coarse-grained reconfigurable array architecture and verified its functionality and effectiveness by accelerating an H.264 deblocking filter, a kernel which is both data- and control-intensive. The results show that the proposed approach achieves up to 43% improvement in execution time compared to speculation by sacrificing 76% code size, and 24% improvement in execution time compared to the previous full predication approach, with a smaller code size. Kyuseung Han, Jong Kyung Paek, Kiyoung Choi |
FPT | 1 |