VLDB 2026 Research / reviewers in the wild / expert
Kiung Jung
dblp:17/2961
· DBLP profile ↗
4ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0009-5720-4657ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Compiler and System Optimizations for Gem5 SimulatorabstractArchitectural simulators are indispensable for modern computer architecture research, but they remain notoriously slow due to their event-driven, cycle-level execution model. In this work, we present a set of software- and system-level optimizations to accelerate large-scale design-space exploration with gem5. First, we reduce per-instance simulation time via compiler-level optimization. We demonstrate that although gem5 suffers severe frontend stalls on modern CPUs stemming from its large instruction footprints, naïve Profile-Guided Optimization (PGO) is impractical in this setting because it requires frequent reprofiling and recompilation. To address this, we challenge the conventional reliance on self-profiling and instead construct a universal, performance-driven profile that generalizes across simulation inputs. Second, we improve aggregate simulation throughput by strengthening performance isolation using Sub-NUMA clustering (SNC). Finally, we show that a simple co-scheduling heuristic has great potential for reducing resource stranding and boosting multi-instance efficiency. Together, these techniques improve single simulation speed by 17 % and aggregate throughput by $27 \%$, making large-scale design-space exploration more practical and efficient. Haneul Park, Siddharth Agarwal, Pradyun Narkadamilli, Kiung Jung, Yongjun Park 0001, Ipoom Jeong, Nam Sung Kim |
ISPASS | 4 |
| 2025 | SortingHat: System Topology-aware Scheduling of Deep Neural Network Models on Multi-GPU SystemsabstractThe advent of cutting-edge AI applications has emphasized the importance of reducing inference latency.Consequently, efficient model-parallel execution on multiple GPUs represents a key challenge in achieving high performance through the partitioning of the target neural network.Nevertheless, in recent complex deep learning models, as the size of parameters continues to increase and overall inference latency is no longer solely dominated by kernel execution, performance improvements using multiple GPUs cannot be achieved by simply exploiting model parallelism without considering data transfer parallelism and the system topology.To address this challenge, this paper proposes SortingHat, which generates an efficient schedule of target neural network models on multi-GPU systems to minimize inference latency.Initially, SortingHat partitions a target model into multiple submodels based on dominator analysis to find the best solution within a reasonable time.Subsequently, SortingHat finds the best schedule for each submodel using Mixed Integer Linear Programming, taking system topology into account to exploit both model parallelism and data transfer parallelism.Once the schedules of all submodels are found, they are merged and executed on the ready queue-based executor.Evaluations on diverse multi-GPU environments with various large language models show that SortingHat achieves an average speedup of 2.28× and up to 2.96× over the single GPU on TVM baseline. Seok Namkoong, Taehyeong Park 0001, Kiung Jung, Yongjun Park 0001 |
ICS | 3 |
| 2024 | Orchestrating Multiple Mixed Precision Models on a Shared Precision-Scalable NPUabstractMixed-precision quantization can reduce the computational requirements of Deep Neural Network (DNN) models with minimal loss of accuracy. As executing mixed-precision DNN models on Neural Processing Units (NPUs) incurs significant under-utilization of computational resources, Precision-Scalable NPUs (PSNPUs) which can process multiple low-precision layers simultaneously have been proposed. However, the under-utilization still remains significant due to the lack of adequate scheduling algorithms to support multiple mixed-precision models on PSNPUs. Therefore, in this paper, we propose a dynamic programming-based scheduling algorithm for the operations of multiple mixed-precision models. Our scheduling algorithm finds the optimal execution plan that exploits the precision-scalable MACs to improve the end-to-end inference latency of mixed-precision models. We evaluate the performance of this algorithm in terms of hardware utilization, inference latency, and schedule search time compared to baseline scheduling algorithms. The experimental results show 1.23 inference latency improvements over the baseline algorithms within the allowed minutes. Kiung Jung, Seok Namkoong, Hongjun Um, Hyejun Kim, Youngsok Kim, Yongjun Park 0001 |
LCTES | 1 |
| 2005 | Simulcast packet transmission in ad hoc networksabstractIn previous work, unequal error-protection techniques have been applied to improve the throughput of a wireless communication system in which a transmission is received by several radios with different capabilities. For instance, these capabilities may correspond to differences in path loss, fading, or interference. By taking advantage of the broadcast nature of the channel, additional messages for the more-capable receivers can be included on transmissions to the less-capable receivers at very little cost (in terms of required energy at the transmitter or error probabilities at the receivers). This technique has been termed simulcasting or multicast signaling. In this paper, we consider the use of these techniques in an ad hoc network. These techniques impact the link throughput, end-to-end throughput, and network connectivity. We investigate how the choice of parameters for the simulcasting technique affects these network performance metrics. The results indicate that a properly chosen simulcasting technique can improve the link and end-to-end throughput in wireless ad hoc networks with only a slight degradation in other metrics, such as network connectivity. Kiung Jung, John M. Shea |
IEEE J. Sel. Areas Commun. | 1 |