EDBT 2026 Demo / reviewers in the wild / expert
Hyuk-Jun Lee
dblp:75/702 · also Hyukjun Lee
· DBLP profile ↗
14ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0003-2981-0800ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Energy-efficient computing · 27% Memory systems · 23% Storage systems · 20% | |
| Software engineering, system software, and programming languages
1 paper |
Operating systems · 100% |
Topics — the 21 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Reconfigurable computing and FPGAs
FPGA architecture |
0.6 | 1 | 2022 | STT-MRAM-Based Multicontext FPGA for Multithreading Computing Environment · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022 |
Reconfigurable computing and FPGAs › dynamic reconfiguration
multicontext FPGA |
0.6 | 1 | 2022 | STT-MRAM-Based Multicontext FPGA for Multithreading Computing Environment · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022 |
Memory systems
non-volatile memory |
0.6 | 1 | 2022 | STT-MRAM-Based Multicontext FPGA for Multithreading Computing Environment · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022 |
Memory systems › non-volatile memory › magnetic random access memory
STT-MRAM |
0.6 | 1 | 2022 | STT-MRAM-Based Multicontext FPGA for Multithreading Computing Environment · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022 |
Energy-efficient computing › power management › memory power management
cache energy reduction |
0.4 | 1 | 2020 | Per-Operation Reusability Based Allocation and Migration Policy for Hybrid Cache · IEEE Trans. Computers 2020 |
Energy-efficient computing
energy-aware scheduling |
0.4 | 1 | 2020 | EANeM: Energy-Aware Network Stack Management for Mobile Devices · DAC 2020 |
Storage systems
flash and SSD |
0.4 | 1 | 2020 | Multitoken-Based Power Management for NAND Flash Storage Devices · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Memory systems › cache › cache technology
hybrid cache |
0.4 | 1 | 2020 | Per-Operation Reusability Based Allocation and Migration Policy for Hybrid Cache · IEEE Trans. Computers 2020 |
Embedded and real-time systems
mobile computing |
0.4 | 1 | 2020 | EANeM: Energy-Aware Network Stack Management for Mobile Devices · DAC 2020 |
Storage systems › flash and SSD › flash memory › NAND flash
NAND flash storage |
0.4 | 1 | 2020 | Multitoken-Based Power Management for NAND Flash Storage Devices · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Energy-efficient computing
power management |
0.4 | 1 | 2020 | Multitoken-Based Power Management for NAND Flash Storage Devices · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Energy-efficient computing
storage power management |
0.4 | 1 | 2020 | Multitoken-Based Power Management for NAND Flash Storage Devices · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Parallel and multicore computing › parallel scheduling
thread scheduling |
0.4 | 1 | 2020 | EANeM: Energy-Aware Network Stack Management for Mobile Devices · DAC 2020 |
Storage systems › flash and SSD › flash memory
flash storage |
0.2 | 1 | 2014 | An Adaptive Idle-Time Exploiting Method for Low Latency NAND Flash-Based Storage Devices · IEEE Trans. Computers 2014 |
Storage systems › flash and SSD › flash memory management
flash translation layer |
0.2 | 1 | 2014 | An Adaptive Idle-Time Exploiting Method for Low Latency NAND Flash-Based Storage Devices · IEEE Trans. Computers 2014 |
Energy-efficient computing › low-power design
power optimization |
0.2 | 1 | 2022 | STT-MRAM-Based Multicontext FPGA for Multithreading Computing Environment · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022 |
Operating systems › resource management › process management
CPU scheduling |
0.1 | 1 | 2020 | EANeM: Energy-Aware Network Stack Management for Mobile Devices · DAC 2020 |
Storage systems
storage reliability |
0.1 | 1 | 2020 | Multitoken-Based Power Management for NAND Flash Storage Devices · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2014 | An Adaptive Idle-Time Exploiting Method for Low Latency NAND Flash-Based Storage Devices · IEEE Trans. Computers 2014 |
Reconfigurable computing and FPGAs › FPGA architecture
FPGA logic block architecture |
0.0 | 1 | 2000 | Coarse-grained carry architecture for FPGA (poster abstract) · FPGA 2000 |
Integrated circuit design › digital circuit design
arithmetic circuit design |
0.0 | 1 | 2000 | Coarse-grained carry architecture for FPGA (poster abstract) · FPGA 2000 |
Methods — techniques the papers use, named apart from their topics
bandwidth control · 0.9CPU load estimation · 0.9multicontext-aware CAD flow · 0.6token-based power management · 0.4deadlock resolution · 0.4cost function · 0.4LRU replacement · 0.4scheduling · 0.2hidden markov model · 0.2adaptive time-out · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Token Pruning in Audio Transformers: Optimizing Performance and Decoding Patch ImportanceabstractVision Transformers (ViTs) have achieved state-of-the-art performance across various computer vision tasks, but their high computational cost remains a challenge. Token pruning has been proposed to reduce this cost by selectively removing less important tokens. While effective in vision tasks by discarding non-object regions, applying this technique to audio tasks presents unique challenges, as distinguishing relevant from irrelevant regions in time-frequency representations is less straightforward. In this study, for the first time, we applied token pruning to ViT-based audio classification models using Mel-spectrograms and analyzed the trade-offs between model performance and computational cost: TopK token pruning can reduce MAC operations of AudioMAE and AST by 30-40%, with less than a 1% drop in accuracy. Our analysis reveals that while high-intensity or high-variation tokens contribute significantly to model accuracy, low-intensity or low-variation tokens also remain important when token pruning is applied; pruning solely based on the intensity or variation of signals in a patch leads to a noticeable drop in accuracy. We support our claim by measuring high correlation between attention scores and these statistical features and by showing retained tokens consistently receive distinct attention compared to pruned ones. We also show that AudioMAE retains more low-intensity tokens than AST. This can be explained by AudioMAE’s self-supervised reconstruction objective, which encourages attention to all patches, whereas AST’s supervised training focuses on label-relevant tokens. Taehan Lee, Hyuk-Jun Lee |
ECAI | 2 |
| 2022 | STT-MRAM-Based Multicontext FPGA for Multithreading Computing EnvironmentabstractThe demand for high-performance computing and rapidly increasing power consumption has increased the necessity for application-specific accelerators. In the datacenter and mobile system, more applications are increasingly relying on accelerators. Field-programmable gate arrays (FPGAs) emerge as a good candidate because they have high programmability and power efficiency. As the number of applications requiring acceleration increases, there is huge demand for FPGAs that support multiple contexts. Previous FPGA designs that support multicontext have various shortcomings such as volatility, poor power efficiency, large performance, area, and reconfiguration overhead. In this article, we propose a spin-transfer torque magnetic RAM (STT-MRAM)-based nonvolatile multicontext FPGA (NVMC-FPGA) that overcomes these shortcomings. We introduce the NVMC-FPGA architecture and operation modes that take advantage of nonvolatility and support multicontext. We also develop the multicontext-aware FPGA computer aided design flow to make the most of the NVMC-FPGA. Compared to the conventional SRAM-based FPGA, when eight identical circuits are mapped, the NVMC-FPGA improves the performance by 15.3% on average and reduces the power consumption by 11.2%–80.7%, depending on the number of simultaneously activated circuits. Moreover, when eight different circuits are mapped, the NVMC-FPGA improves the performance by 58.5% on average and reduces the power consumption by 6.2%–63.3%, depending on the number of simultaneously activated circuits. Jeongbin Kim 0001, Yongwoon Song, Kyungseon Cho, Hyuk-Jun Lee, Hongil Yoon, Eui-Young Chung |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Highly Available Packet Buffer Design With Hybrid Nonvolatile MemoryabstractInternet routers/switches are vulnerable to system failures which require a power reset. Tremendous efforts are made to guarantee the high availability of the systems. A recent work shows that a phase change memory (PCM)-based routing lookup table can achieve high availability in the destination lookup of routers/switches. However, a packet buffer in routers/switches cannot benefit from the lookup table approach because it requires a much larger and higher-bandwidth memory system and its memory traffic is equally divided into reads and writes, while routing table accesses are mostly read-dominant. PCM can provide high availability even when the system undergoes a power reset but exhibits unacceptable write bandwidth. In this work, we propose a magnetic RAM (MRAM)/PCM-based hybrid memory packet buffer and a packet mapping method. A small MRAM combined with a large PCM can outperform the dynamic random access memory (DRAM)-based packet buffer by 28.5% or 22.4% on average for internet-mix packet traffic when optimizing only bandwidth or both bandwidth and lifetime. The proposed adaptive packet mapping method maps small packets less than a predetermined packet size threshold to MRAM for maximizing PCM bandwidth as small packets degrade row buffer locality in PCM. In addition, the mapping method dynamically changes the packet size threshold to capture the working set of packet buffering, which significantly improves PCM lifetime. Yongwoon Song, Jooyoung Hwang, Insoon Jo, Hyuk-Jun Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2020 | EANeM: Energy-Aware Network Stack Management for Mobile DevicesabstractIn mobile computing, various energy-efficient thread scheduling schemes for heterogeneous multiprocessing architectures are proposed. For network applications, however, inaccurate prediction of the CPU load and high-priority network packet processing overdrive CPU cores, leading to large energy consumption. We present a framework including a network stack monitor, a bandwidth controller, and an energy-efficient thread scheduling scheme, which accurately estimates the CPU load for packet processing and optimally schedules CPU resource. It improves performance/watt by 4.79 times over the baseline Linux scheduler for FTP applications. In multi-threaded environments, it improves performance/watt by 2.35–3.11 times for PARSEC benchmark and network applications. Chungseop Lee, Keonhyuk Lee, Mingoo Kang, Hyuk-Jun Lee |
DAC | 4 |
| 2020 | Per-Operation Reusability Based Allocation and Migration Policy for Hybrid CacheabstractRecently, a hybrid cache consisting of SRAM and STT-RAM has attracted much attention as a future memory by complementing each other with different memory characteristics. Prior works focused on developing data allocation and migration techniques considering write-intensity to reduce write energy at STT-RAM. However, these works often neglect the impact of operation-specific reusability of a cache line. In this paper, we propose an energy-efficient per-operation reusability-based allocation and migration policy (ORAM) with a unified LRU replacement policy. First, to select an adequate memory type for allocation, we propose a cost function based on per-operation reusability - gain from an allocated cache line and loss from an evicted cache line for different memory types - which exploits the temporal locality. Besides, we present a migration policy, victim and target cache line selection scheme, to resolve memory type inconsistency between replacement policy and the allocation policy, with further energy reduction. Experiment results show an average energy reduction in the LLC and the main memory by 12.3 and 21.2 percent, and the improvement of latency and execution time by 21.2 and 8.8 percent, respectively, compared with a baseline hybrid cache management. In addition, the Energy-Delay Product (EDP) is improved by 36.9 percent over the baseline. Minsik Oh, Kwangsu Kim, Duheon Choi, Hyuk-Jun Lee, Eui-Young Chung |
IEEE Trans. Computers | 4 |
| 2020 | Multitoken-Based Power Management for NAND Flash Storage DevicesabstractNAND flash-based storage devices (NFSDs) have been widely employed in various systems, including cloud servers as well as mobile devices. The core component of NFSDs is NAND flash memory (NFM) which has several advantages over the conventional hard disk drives (HDDs). An NFSD typically adopts a bunch of NFMs which are operated in parallel for maximizing the I/O throughput. However, optimizing for performance may not be desirable from the power budget (PB) perspective. In other words, concurrent operations of NFMs often drain inordinate current, which leads to the violation of the PB allocated for a storage device. In this article, we propose a novel power management scheme which maximizes concurrent operations of NFMs under the given power constraint. The proposed method quantizes the given power constraint of an NFSD. A quantum also called token is the basic unit of power management. The proposed power management scheme allocates tokens to NFMs and only the NFMs having enough tokens can perform their operations. We call this method multitoken-based power management (MTPM). The critical issue of MTPM is a deadlock which is resolved with the key allocation scheme. Furthermore, we enhance MTPM to improve performance. The extended method called keyless MTPM (KMTPM) improves the overall performance by relaxing the key acquisition requirement and allowing subatomic operations. In the experimental results, we confirm that the proposed methods always meet the given power constraint. The proposed KMTPM improves throughput by 22.85% compared to state of the art technique. In addition, KMTPM only incurs 3.8% of performance overhead and 0.015% of area overhead. Tae-Hee You, Sangwoo Han, Young Min Park, Hyuk-Jun Lee, Eui-Young Chung |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | A Novel NAND Flash Memory Architecture for Maximally Exploiting Plane-Level ParallelismabstractSolid-state drive (SSD) has become one of the most dominant storage devices and is rapidly replacing conventional storage devices. The core component of SSD is NAND flash memory (NFM), where the actual data are stored. Cost pressure is the most critical factor limiting the further deployment of SSDs and past researches have focused on developing cost-effective high-density NFM. Although the cost-driven technology development increases per-chip capacity, it reduces channel-/way-level parallelisms for the given device capacity, resulting in the performance degradation. Such observation directs us to focus on a novel NFM architecture exploiting plane-level parallelism. The distinct features of this architecture are: 1) enabling a decoupled word-line (WL) selection for the mated planes and 2) segmenting each plane into subplanes for further maximizing the plane-level parallelism. The experimental results show that decoupled WL selection improves the throughput by up to 21.3% with a marginal overhead of less than 1%, compared to the conventional NFM architecture. In addition, adopting the plane segmentation improves the throughput by up to 43.9% with an additional overhead of 14%. Considering the tradeoff between performance and overhead, the proposed NFM architecture is a cost-efficient method to secure high performance under decreasing channel-/way-level parallelisms in high-density NFM. Myeongjin Kim, Wontaeck Jung, Hyuk-Jun Lee, Eui-Young Chung |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Parameterised codebook design based on channel statistics for efficient multi-rank MIMO transmissionabstractAs the number of antenna elements increases in multiple‐input multiple‐output (MIMO) systems, an efficient feedback and beamforming strategy becomes more important to achieve the desired bandwidth efficiency required by the next‐generation wireless systems. The authors investigate the statistical characteristics of the MIMO channel generated by the three‐dimensional spatial channel model to recognise some of the limitations inherent in the existing codebooks adopted by the current third generation partnership project specifications. Such limitations include unnecessary co‐phasing combinations for cross‐polarised (X‐pol) antenna arrays and, in particular, long‐term codevector selections, which are not best suited to the channel distributions. In this study, the observed shortcomings are improved by proposing a new multi‐rank codebook applicable to the X‐pol array of antenna elements supporting up to the quadruple rank. The proposed beamforming codevectors are presented using variable parameters so that they can be adapted to changing channel characteristics. The efficiency of the proposed codebook is further improved by eliminating redundant co‐phasing values. The performance of the proposed codebook is evaluated using the TR 36.873 urban macro environment and its effectiveness is demonstrated by quantifying its gain over the existing codebooks. Sungin Shin, Hyuk-Jun Lee, Wonjin Sung |
IET Commun. | 2 |
| 2017 | Scalable Bandwidth Shaping Scheme via Adaptively Managed Parallel Heaps in Manycore-Based Network ProcessorsabstractScalability of network processor-based routers heavily depends on limitations imposed by memory accesses and associated power consumption. Bandwidth shaping of a flow is a key function, which requires a token bucket per output queue and abuses memory bandwidth. As the number of output queues increases, managing token buckets becomes prohibitively expensive and limits scalability. In this work, we propose a scalable software-based token bucket management scheme that can reduce memory accesses and power consumption significantly. To satisfy real-time and low-cost constraints, we propose novel parallel heap data structures running on a manycore-based network processor. By using cache locking, the performance of heap processing is enhanced significantly and is more predictable. In addition, we quantitatively analyze the performance and memory footprint of the proposed software scheme using stochastic modeling and the Lyapunov central limit theorem. Finally, the proposed scheme provides an adaptive method to limit the size of heaps in the case of oversubscribed queues, which can successfully isolate the queues showing unideal behavior. The proposed scheme reduces memory accesses by up to three orders of magnitude for one million queues sharing a 100Gbps interface of the router while maintaining stability under stressful scenarios. Jongbum Lim, Jinku Kim, Woo-Cheol Cho, Eui-Young Chung, Hyuk-Jun Lee |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2016 | Implementation of a large-scale language model adaptation in a cloud environment
Kwang-Ho Kim, Dae-Young Jung, Donghyun Lee 0001, Hyuk-Jun Lee, Sungyong Park, Myoung-Wan Koo, Jeong-Sik Park, Hyung-Bae Jeon |
Multim. Tools Appl. | 4 |
| 2014 | An Adaptive Idle-Time Exploiting Method for Low Latency NAND Flash-Based Storage DevicesabstractThe market share of NAND flash-based storage devices (NFSDs) has rapidly grown in recent years since many characteristics, such as non-volatility, low latency, and high reliability, meet the requirements for various types of storage devices. However, the unique characteristic of NAND flash memories (NFMs), erase-before-write, causes problems for NFSDs from a performance perspective. Specifically, performance degradation is incurred by extra operations that serve to hide the bad characteristics of NFMs. In order to resolve this problem, many attractive methods have been proposed. Various algorithms for flash translation layers (FTLs) are representative methods that provide space redundancy to NFSDs for better performance. However, the amount of space redundancy is limited by the capacity of NFMs and thus, space redundancy is still insufficient for improving the performance of NFSDs. Consequently, a new type of redundancy, termed temporal redundancy, has recently been introduced for NFSDs. More precisely, the idleness of NFSDs is exploited so as to precede extra operations for NFSDs while minimizing the overhead of extra operations. In this paper, we propose an adaptive time-out method based on the Hidden-Markov Model (HMM) to efficiently utilize idle periods. In addition, we also suggest a simple scheduling scheme for extra operations that can be customized for general FTLs. The experimental results demonstrate that the proposed method yields performance improvements in terms of average write latency and peak latency, 74% and 76% better than the existing method, respectively, and approaching within average 9% and 5% of the optimal case, respectively. Sang-Hoon Park, Donggun Kim 0005, Kwanhu Bang, Hyuk-Jun Lee, Sungjoo Yoo, Eui-Young Chung |
IEEE Trans. Computers | 4 |
| 2008 | Scalable QoS-Aware Memory Controller for High-Bandwidth Packet MemoryabstractThis paper proposes a high-performance scalable quality-of-service (QoS)-aware memory controller for the packet memory where packet data are stored in network routers. A major challenge in the packet memory controller design is to make the design scalable. As the input and output bandwidth requirement and the number of output queues for routers increase, the memory system becomes a bottleneck that limits the performance and scalability. Existing schemes require an input and output buffer that store packet data temporarily before they are written into or read from the memory. With the buffer size proportional to the number of output queues, the buffer becomes a limiting factor for scalability. Our scheme consists of a hashing logic and a reorder buffer whose size is not proportional to the number of output queues and is scalable with the increasing number of output queues. Another major challenge in the packet memory controller design is supporting QoS. As an increasing number of Internet packets become latency sensitive, it is critical that the memory controller is capable of providing different QoS to packets belonging to different classes. To the best of our knowledge, no published work on the packet memory controller supports QoS. In this paper, we show our scheme reduces the SRAM buffer size of the existing schemes by an order of magnitude whereas guaranteeing a packet loss probability as low as 10-20. Our QoS-aware scheduler shows that it meets the latency requirements assigned to multiple service classes under dynamically changing input loads for multiple classes using a feedback control loop. Hyuk-Jun Lee, Eui-Young Chung |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2007 | Slack-based Bus Arbitration Scheme for Soft Real-time Constrained Embedded SystemsabstractWe present a bus arbitration scheme for soft real-time constrained embedded systems. Some masters in such systems are required to complete their work for given timing constraints, resulting in the satisfaction of system-level timing constraints. The computation time of each master is predictable, but it is not easy to predict its data transfer time since the communication architecture is mostly shared by several masters. Previous works solved this issue by minimizing the latencies of several latency-critical masters, but the side effect of these methods is that it can increase the latencies of other masters, hence they may violate the given timing constraints. Unlike previous works, our method uses the concept of "slack" in order to make the latency as close as its given constraint, resulting in the reduction of the side effect. The proposed arbitration scheme consists of bandwidth-conscious arbiter and scheduler. The arbiter can be any existing bandwidth-conscious arbiter and the scheduler implements the latency-awareness proposed in this paper. The scheduler is involved in the arbitration only when it observes a request whose slack is not sufficient for the given timing constraint. The experimental results show that our method outperforms the conventional round-robin arbiter by more than 100% in the best case in terms of the longest violated cycles. Minje Jun, Kwanhu Bang, Hyuk-Jun Lee, Naehyuck Chang, Eui-Young Chung |
ASP-DAC | 3 |
| 2000 | Coarse-grained carry architecture for FPGA (poster abstract)abstractThe fine grain size of current FPGA has been a major performance bottleneck. In this paper, we introduce a coarse-grained carry architecture that increases the grain size from a two-bit addition/subtraction per logic block to an m-bit addition/subtraction. The m-bit addition is implemented by increasing the number of read-ports for a look-up table from 1 to m-2. In addition, we use a dedicated selection logic to implement an m-bit conditional addition/subtraction. The proposed architecture improves the performance of applications containing intensive arithmetic operations. We use throughout density as a cost-performance metric to justify the benefit of the new architecture and find the optimal grain size. We could achieve roughly up to 5 times larger throughput density for selected applications at the cost of 5-10% area penalty. Hyuk-Jun Lee, Michael J. Flynn |
FPGA | 1 |