EDBT 2026 Demo / reviewers in the wild / expert
Hwansoo Han
dblp:74/4345
· DBLP profile ↗
28ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0001-7182-8452ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Compile-Time QoS Scheme for Deep Learning InferencesabstractWith the proliferation of deep learning technologies across various service domains, the sharing of accelerators such as GPUs, TPUs, and NPUs for inference processing has become increasingly common. These accelerators must efficiently handle multiple deep learning services operating concurrently. However, inference requests, characterized by sequences of short-duration kernels, create significant challenges for online schedulers attempting to maintain Quality of Service (QoS) guarantees. This paper presents QoSlicer, a novel compile-time QoS management framework that employs kernel slicing to relieve the burden on schedulers. By generating multiple pre-determined slicing plans, QoSlicer enables more efficient, lightweight QoS scheduling while ensuring target latency requirements are met. Our approach incorporates a heuristic search algorithm to identify optimal slicing plans and implements robust performance estimation models to validate these plans. Our experimental evaluation across 75 diverse workload combinations demonstrates that QoSlicer improves throughput by an average of 20.2% compared to state-of-the-art scheduling techniques. Sungin Hong, Hwansoo Han |
SC | 3 |
| 2024 | GPU thread throttling for page-level thrashing reduction via static analysis
Hwansoo Han |
J. Supercomput. | 2 |
| 2023 | Accelerating Deep Neural Networks on Mobile Multicore NPUsabstractNeural processing units (NPUs) have become indispensable parts of mobile SoCs. Furthermore, integrating multiple NPU cores into a single chip becomes a promising solution for ever-increasing computing power demands in mobile devices. This paper addresses techniques to maximize the utilization of NPU cores and reduce the latency of on-device inference. Mobile NPUs typically have a small amount of local memory (or scratch pad memory, SPM) that provides space only enough for input/output tensors and weights of one layer operation in deep neural networks (DNNs). Even in multicore NPUs, such local memories are distributed across the cores. In such systems, executing network layer operations in parallel is the primary vehicle to achieve performance. By partitioning a layer of DNNs into multiple sub-layers, we can execute them in parallel on multicore NPUs. Within a core, we can also employ pipelined execution to reduce the execution time of a sub-layer. In this execution model, synchronizing parallel execution and loading/storing intermediate tensors in global memory are the main bottlenecks. To alleviate these problems, we propose novel optimization techniques which carefully consider partitioning direction, execution order, synchronization, and global memory access. Using six popular convolutional neural networks (CNNs), we evaluate our optimization techniques in a flagship mobile SoC with three cores. Compared to the highest-performing partitioning approach, our techniques improve performance by 23%, achieving a speedup of 2.1x over single-core systems. Hanwoong Jung, Hexiang Ji, Alexey Pushchin, Maxim Ostapenko, Wenlong Niu, Ilya Palachev, Yutian Qu, Pavel Fedin, Yuri Gribov, Heewoo Nam, Dongguen Lim, Joonho Song, Hwansoo Han |
CGO | 15 |
| 2023 | Performance Analysis of ZNS-aware File Systems on Distributed ApplicationsabstractRecently, ZNS SSDs have been actively researched to handle the functions of the FTL directly on the host system. ZNS SSDs can improve the I/O performance and spatial efficiency of SSDs by eliminating GC and eliminating overprovisioning. In this paper, we analyze the performance by running distributed applications on a file system which supports ZNS SSDs. We found that much of the time difference occurs in the process of executing open and unlink operations rather than the performance of read and write operations. It has been confirmed that the performance difference of read/write operations is not significant for the applications. If structural optimization is made on file system, it has been found that ZNS SSDs are better than CNS SSDs of the same capacity in terms of price. Yongwoo Cho 0004, Seongsoo Park 0001, Hwansoo Han |
ICDCS | 3 |
| 2021 | Libpubl: exploiting persistent user buffers as logs for write atomicityabstractWith recent advancement in NVM technologies, non-volatile main memory(NVMM) has been highlighted as promising systems for large scale data centers. By utilizing memory-mapped IO, researchers have attempted to provide fast IOs without kernel mode switches and complex kernel IO layers. Even though they improve the performance by taking advantage of NVMM characteristics, memory-mapped IOs still require additional user-level logging techniques to guarantee write atomicity and file system integrity. We propose a user-level library file system, Libpubl, which is designed to minimize the write amplification overhead of traditional logging by exploiting persistent user buffers as logs. Libpubl ensures atomic updates of data and improves the performance and the scalability. Compared to the state-of-the-art NVM file systems, Libpubl increases the performance by 50~120% in Fio benchmark. Jungsik Choi 0001, Hwansoo Han |
HotStorage | 3 |
| 2021 | Idempotence-Based Preemptive GPU Kernel Scheduling for Embedded SystemsabstractMission-critical embedded systems simultaneously run multiple graphics-processing-unit (GPU) computing tasks with different criticality and timeliness requirements. Considerable research effort has been dedicated to supporting the preemptive priority scheduling of GPU kernels. However, hardware-supported preemption leads to lengthy scheduling delays and complicated designs, and most software approaches depend on the voluntary yielding of GPU resources from restructured kernels. We propose a preemptive GPU kernel scheduling scheme that harnesses the idempotence property of kernels. The proposed scheme distinguishes idempotent kernels through static source code analysis. If a kernel is not idempotent, then GPU kernels are transactionized at the operating system (OS) level. Both idempotent and transactionized kernels can be aborted at any point during their execution and rolled back to their initial state for reexecution. Therefore, low-priority kernel instances can be preempted for high-priority kernel instances and reexecuted after the GPU becomes available again. Our evaluation using the Rodinia benchmark suite showed that the proposed approach limits the preemption delay to 18 ms in the 99.9th percentile, with an average delay in execution time of less than 10 percent for high-priority tasks under a heavy load in most cases. Cheolgi Kim, Hwansoo Han, Euiseong Seo |
IEEE Trans. Computers | 4 |
| 2020 | Effective Profiling for Data-Intensive GPU Programs: Work-in-ProgressabstractGPU profilers have been successfully used to analyze bottlenecks and slowdowns of GPU programs. Several instrumentation tools for profiling GPU binaries are introduced, but these tools take little consideration into GPU architectures. In this paper, we investigate common factors of performance degradation on existing GPU profilers and provide design guide to improve the performance. Hwiwon Kim, Hwansoo Han |
CASES | 3 |
| 2020 | Page Reuse in Cyclic Thrashing of GPU Under Oversubscription: Work-in-ProgressabstractNvidia GPUs can oversubscribe CPU-side RAM, elevating GPU programmability and enabling GPU to use data beyond the size of GPU memory. However, oversubsrciption accompanies overhead of constant exchange of pages between the GPU and the CPU. To this problem, while previous works implemented and experimented solutions overhead in GPU simulators, we propose a new method implemented in the driver, which runs actual GPU hardware. We developed algorithms that detect repeated page replacement of the pages and pin pages in GPU memory to be reused in the future without replacement. Dojin Park, Hwansoo Han |
CASES | 3 |
| 2020 | Libnvmmio: Reconstructing Software IO Path with Failure-Atomic Memory-Mapped Interface
Jungsik Choi 0001, Jaewan Hong, Youngjin Kwon, Hwansoo Han |
USENIX ATC | 4 |
| 2020 | Static code transformations for thread-dense memory accesses in GPU computingabstractSummary Due to the GPU's complex memory system and massive thread‐level parallelism, application programmers often have difficulty optimizing GPU programs. An essential approach to memory optimization is to utilize low‐latency on‐chip memory to avoid high latency of off‐chip memory accesses. Shared memory is an on‐chip memory, which is explicitly managed by programmers. Shared memory has a read/write latency similar to that of the L1 cache, but poor data management can degrade performance. In this paper, we present a static code transformation that preloads dataset in GPU's shared memory. Our static analysis primarily targets global memory requests with high thread‐density for preloading in shared memory. The thread‐dense memory access pattern is a pattern in which many threads efficiently manage the address space of shared memory, as well as reuse the same data in a thread block. We limit the usage of shared memory so that thread‐level parallelism remains at the same level when selecting datasets for preloading. Finally, our source‐to‐source compiler allows to preload selected datasets in shared memory by transforming non‐optimized GPU kernel code. Our methods achieve 1.26× and 1.62× speedups on average (geometric mean), respectively with GTX980 and P100 GPUs. Sungin Hong, Hwansoo Han |
Concurr. Comput. Pract. Exp. | 4 |
| 2019 | Compiler-Assisted GPU Thread Throttling for Reduced Cache ContentionabstractModern GPUs concurrently deploy thousands of threads to maximize thread level parallelism (TLP) for performance. For some applications, however, maximized TLP leads to significant performance degradation, as many concurrent threads compete for the limited amount of the data cache. In this paper, we propose a compiler-assisted thread throttling scheme, which limits the number of active thread groups to reduce cache contention and consequently improve the performance. A few dynamic thread throttling schemes have been proposed to alleviate cache contention by monitoring the cache behavior, but they often fail to provide timely responses to the dynamic changes in the cache behavior, as they adjust the parallelism afterwards in response to the monitored behavior. Our thread throttling scheme relies on compile-time adjustment of active thread groups to fit their memory footprints to the L1D capacity. We evaluated the proposed scheme with GPU programs that suffer from cache contention. Our approach improved the performance of original programs by 42.96% on average, and this is 8.97% performance boost in comparison to the static thread throttling schemes. Sungin Hong, Euiseong Seo, Hwansoo Han |
ICPP | 5 |
| 2017 | Balanced cache bypassing for critical warp reduction: work-in-progressabstractWarp-level cache bypassing has been proposed to resolve GPU memory resource contention on GPU computing. However, the proposed cache bypassing scheme has sub-optimal performance due to warp criticality problem in balanced workload. In this paper, we show that warp-level cache bypassing is a sub-optimal solution and propose a balanced cache bypassing scheme to solve this problem. Sungin Hong, Hwansoo Han |
CASES | 3 |
| 2017 | Efficient Memory Mapped File I/O for In-Memory File Systems
Jungsik Choi 0001, Jiwon Kim 0001, Hwansoo Han |
HotStorage | 3 |
| 2012 | Efficient SIMD code generation for irregular kernelsabstractArray indirection causes several challenges for compilers to utilize single instruction, multiple data (SIMD) instructions. Disjoint memory references, arbitrarily misaligned memory references, and dependence cycles in loops are main challenges to handle for SIMD compilers. Due to those challenges, existing SIMD compilers have excluded loops with array indirection from their candidate loops for SIMD vectorization. However, addressing those challenges is inevitable, since many important compute-intensive applications extensively use array indirection to reduce memory and computation requirements. In this work, we propose a method to generate efficient SIMD code for loops containing indirected memory references. We extract both inter- and intra-iteration parallelism, taking data reorganization overhead into consideration. We also optimally place data reorganization code in order to amortize the reorganization overhead through the performance gain of SIMD vectorization. Experiments on four array indirection kernels, which are extracted from real-world scientific applications, show that our proposed method effectively generates SIMD code for irregular kernels with array indirection. Compared to the existing SIMD vectorization methods, our proposed method significantly improves the performance of irregular kernels by 91%, on average. Seonggun Kim, Hwansoo Han |
PPoPP | 2 |
| 2011 | Region-based parallelization of irregular reductions on explicitly managed memory hierarchies
Seonggun Kim, Hwansoo Han, Kwang-Moo Choe |
J. Supercomput. | 2 |
| 2010 | Filtering false alarms of buffer overflow analysis using SMT solvers
Youil Kim, Jooyong Yi, Hwansoo Han, Kwang-Moo Choe |
Inf. Softw. Technol. | 3 |
| 2010 | Composition-based Cache simulation for structure reorganization
Keoncheol Shin, Hwansoo Han, Kwang-Moo Choe |
J. Syst. Archit. | 2 |
| 2009 | Abstracting access patterns of dynamic memory using regular expressionsabstractUnless the speed gap between CPU and memory disappears, efficient memory usage remains a decisive factor for performance. To optimize data usage of programs in the presence of the memory hierarchy, we are particularly interested in two compiler techniques: pool allocation and field layout restructuring . Since foreseeing runtime behaviors of programs at compile time is difficult, most of the previous work relied on profiling. On the contrary, our goal is to develop a fully automatic compiler that statically transforms input codes to use memory efficiently. Noticing that regular expressions , which denote repetition explicitly, are sufficient for memory access patterns, we describe how to extract memory access patterns as regular expressions in detail. Based on static patterns presented in regular expressions, we apply pool allocation to repeatedly accessed structures and exploit field layout restructuring according to field affinity relations of chosen structures. To make a scalable framework, we devise and apply new abstraction techniques, which build and interpret access patterns for the whole programs in a bottom-up fashion. We implement our analyses and transformations with the CIL compiler. To verify the effect and scalability of our scheme, we examine 17 benchmarks including 2 SPECINT 2000 benchmarks whose source lines of code are larger than 10,000. Our experiments demonstrate that the static layout transformations for dynamic memory can reduce L1D cache misses by 16% and execution times by 14% on average. Jinseong Jeon, Keoncheol Shin, Hwansoo Han |
ACM Trans. Archit. Code Optim. | 3 |
| 2008 | Shared heap management for memory-limited java virtual machinesabstractOne scarce resource in embedded systems is memory. Multitasking makes the lack of memory problem even worse. Most current embedded systems, which do not provide virtual memory, simply divide physical memory and evenly assign contiguous memory chunks to multiple applications. Such simple memory management can frequently cause the lack of available memory for some applications, while others are not using the full amount of assigned memory. To overcome inefficiency in current memory management, we present an efficient heap management scheme that allows multiple applications to share heap space. To reduce overall heap memory usage, applications adaptively acquire subheaps out of shared pool of memory and release surplus subheaps to shared pool. As a result, applications see noncontiguous multiple subheaps as a heap in their address space. We target Java applications to implement our heap-sharing scheme in the KVM from Sun Microsystems. To protect fragmented heap space with a limited number of regions in memory protection unit (MPU), we maintain only a limited number of subheaps. We experimentally evaluate our heap management scheme with J2ME MIDP applications. Our static and dynamic schemes reduce heap memory usage, on average, by 30 and 27%, respectively. For both schemes, overheads are kept low. The execution times in our schemes are increased only by 0.01% for static scheme and 0.35% for dynamic scheme, on average. Yoonseo Choi, Hwansoo Han |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2007 | Layout Transformations for Heap Objects Using Static Access Patterns
Jinseong Jeon, Keoncheol Shin, Hwansoo Han |
CC | 3 |
| 2006 | Protected heap sharing for memory-constrained java environmentsabstractMultitasking is one of capabilities we often want to have in memoryconstrained embedded systems. To support multiple address spaces within a small physical memory, a simple memory management frequently encounters the lack of available memory. Our paper presents an efficient heap memory management scheme that reduces memory footprints by adaptively sharing heaps among multiple tasks in JVM environments. We modified KVM from Sun Microsystems so that Java applications acquire or release heaps in a shared pool on an as-needed basis. To protect address spaces among tasks in the absence of virtual memory capabilities, we use memory protection units (MPUs) by incorporating them into our heap sharing scheme. Our experiments with J2ME MIDP applications show significant reductions by 33 % on average, ranging from 6 % to 50 % in memory usage over the execution. The overheads of our scheme in garbage collection are kept low. The execution times in our scheme increase only by 0.2 % on average. Yoonseo Choi, Hwansoo Han |
CASES | 2 |
| 2006 | Restructuring field layouts for embedded memory systemsabstractIn many computer systems with large data computations, the delay of memory access is one of the major performance bottlenecks. In this paper, we propose an enhanced field remapping scheme for dynamically allocated structures in order to provide better locality than conventional field lay outs. Our proposed scheme reduces cache miss rates drastically by aggregating and grouping fields from multiple instances of the same structure, which implies the performance improvement and power reduction. Our methodology will become more important in the design space exploration, especially as the embedded systems for data oriented application become prevalent. Experimental results show that average L1 and L2 data cache misses are reduced by 23% and 17%, respectively. Due to the enhanced localities, our remapping achieves 13% faster execution time on average than original programs. It also reduces power consumption by 18% for data cache. Keoncheol Shin, Jungeun Kim, Seonggun Kim, Hwansoo Han |
DATE | 4 |
| 2006 | Optimal register reassignment for register stack overflow minimizationabstractArchitectures with a register stack can implement efficient calling conventions. Using the overlapping of callers' and callees' registers, callers are able to pass parameters to callees without a memory stack. The most recent instance of a register stack can be found in the Intel Itanium architecture. A hardware component called the register stack engine (RSE) provides an illusion of an infinite-length register stack using a memory-backed process to handle overflow and underflow for a physically limited number of registers. Despite such hardware support, some applications suffer from the overhead required to handle register stack overflow and underflow. The memory latency associated with the overflow and underflow of a register stack can be reduced by generating multiple register allocation instructions within a procedure [Settle et al. 2003]. Live analysis is utilized to find a set of registers that are not required to keep their values across procedure boundaries. However, among those dead registers, only the registers that are consecutively located in a certain part of the register stack frame can be removed. We propose a compiler-supported register reassignment technique that reduces RSE overflow/underflow further. By reassigning registers based on live analysis, our technique forces as many dead registers to be removed as possible. We define the problem of optimal register reassignment, which minimizes interprocedural register stack heights considering multiple call sites within a procedure. We present how this problem is related to a path-finding problem in a graph called a sequence graph . We also propose an efficient heuristic algorithm for the problem. Finally, we present the measurement of effects of the proposed techniques on SPEC CINT2000 benchmark suite and the analysis of the results. The result shows that our approach reduces the RSE cycles by 6.4% and total cpu cycles by 1.7% on average. Yoonseo Choi, Hwansoo Han |
ACM Trans. Archit. Code Optim. | 2 |
| 2006 | Exploiting Locality for Irregular Scientific CodesabstractIrregular scientific codes experience poor cache performance due to their irregular memory access patterns. In this paper, we present two new locality improving techniques for irregular scientific codes. Our techniques exploit geometric structures hidden in data access patterns and computation structures. Our new data reordering (GPART) finds the graph structure within data accesses and applies hierarchical clustering. Quality partitions are constructed quickly by clustering multiple neighbor nodes with priority on nodes with high degree and repeating a few passes. Overhead is kept low by clustering multiple nodes in each pass and considering only edges between partitions. Our new computation reordering (Z-SORT) treats the values of index arrays as coordinates and reorders corresponding computations in Z-curve order. Applied to dense inputs, Z-SORT achieves performance close to data reordering combined with other computation reordering but without the overhead involved in data reordering. Experiments on irregular scientific codes for a variety of meshes show locality optimization techniques are effective for both sequential and parallelized codes, improving performance by 60-87 percent. GPART achieved within 1-2 percent of the performance of more sophisticated partitioning algorithms, but with one third of the overhead. Z-SORT also yields the performance improvement of 64 percent for dense inputs, which is comparable with data reordering combined with computation reordering. Hwansoo Han, Chau-Wen Tseng |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2005 | A Path Sensitive Type System for Resource Usage Verification of C Like Languages
Hyun-Goo Kang, Youil Kim, Taisook Han, Hwansoo Han |
APLAS | 4 |
| 2005 | Memory layout techniques for variables utilizing efficient DRAM access modes in embedded system designabstractThe delay of memory access is one of the major bottlenecks in embedded systems' performance. In software compilation, it is known that there are high variations in memory access delay depending on the ways of storing/retrieving the variables in code to/from the memories. In this paper, we propose effective storage assignment techniques for variables to maximize the use of memory bandwidth. Specifically, we study the problem of DRAM memory layout for storing the nonarray variables in code to achieve a maximum utilization of page and/or burst modes for the memory accesses. The contributions of our work are, for each page and burst modes: 1) we prove that the problem is NP-hard and 2) we propose an exact formulation of the problem and efficient memory layout algorithms, called Solve-MLP for the page mode and Solve-MLB for the burst mode. From experiments with a set of benchmark programs, we confirm that our proposed techniques use on average 28.2% and 10.1% more page accesses and 82.9% and 107% more burst accesses than those by the order of first use and the technique of Panda et al. in Proc. Int. Conf. Computer-Aided Design, 1997, and Panda et al. in ACM Trans. Design Automation Electron. Syst., 1997, respectively. Yoonseo Choi, Taewhan Kim 0001, Hwansoo Han |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2001 | A Comparison of Parallelization Techniques for Irregular ReductionsabstractA large class of scientific applications are comprised of irregular reductions on large data sets. On shared-memory multiprocessors these reductions are typically parallelized by computing partial results into replicated buffers, then combining the values into shared data using synchronization. Recently, a number of alternative techniques have been developed based on selective privatization, local writes, and synchronized writes. In this paper, we present a more efficient version of the local write algorithm which is 56% faster on average. We then experimentally compare the performance of each technique using a number of representative kernels. Results show speedups vary greatly depending on application characteristics such as connectivity, locality, and adaptivity. In general, we find the local write technique provides the best performance, particularly when applications display good locality. Hwansoo Han, Chau-Wen Tseng |
IPDPS | 1 |
| 2000 | Efficient compiler and run-time support for parallel irregular reductions
Hwansoo Han, Chau-Wen Tseng |
Parallel Comput. | 1 |