Cheol-Ho Hong

dblp:05/7433 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
4since 2021 · last 2024
0000-0003-4730-950XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 3 since 2021Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
abstract
In cloud environments, GPU-based deep neural network (DNN) inference servers are required to meet the Service Level Objective (SLO) latency for each workload under a specified request rate, while also minimizing GPU resource consumption. However, previous studies have not fully achieved this objective. In this paper, we propose ParvaGPU, a technology that facilitates spatial GPU sharing for large-scale DNN inference in cloud computing. ParvaGPU integrates NVIDIA’s Multi-Instance GPU (MIG) and Multi-Process Service (MPS) technologies to enhance GPU utilization, with the goal of meeting the diverse SLOs of each workload and reducing overall GPU usage. Specifically, ParvaGPU addresses the challenges of minimizing underutilization within allocated GPU space partitions and external fragmentation in combined MIG and MPS environments. We conducted our assessment on multiple A100 GPUs, evaluating 11 diverse DNN workloads with varying SLOs. Our evaluation revealed no SLO violations and a significant reduction in GPU usage compared to state-of-the-art frameworks.
Munkyu Lee, Sihoon Seong, Minki Kang, Jihyuk Lee, Gap-Joo Na, In-Geol Chun, Dimitrios S. Nikolopoulos, Cheol-Ho Hong
SC8
2024 Batch-MOT: Batch-Enabled Real-Time Scheduling for Multiobject Tracking Tasks
abstract
Targeting a multiobject tracking (MOT) system with multiple MOT tasks, this article develops Batch-MOT, the first system design that achieves both (G1) timing guarantee and (G2) accuracy maximization, by utilizing batch execution that allows multiple deep neural network (DNN) executions to perform simultaneously in a single DNN inference resulting in significantly decreased execution time without accuracy loss. To this end, we propose an adaptable scheduling framework that allows run-time execution behaviors deviated from our base scheduling algorithm (i.e., nonpreemptive fixed-priority scheduling) without compromising G1. Based on the adaptable framework, we then develop 1) a run-time batching mechanism that finds and executes a batch set of MOT tasks and 2) a run-time idling mechanism that waits for the future releases of MOT tasks for batch execution. Both run-time mechanisms can achieve G1 and G2 without incurring high run-time overhead, as they systematically exploit the run-time execution behaviors allowed by the adaptive framework. Our evaluation conducted with a real-world data set demonstrates the effectiveness of Batch-MOT in improving tracking accuracy while providing a timing guarantee compared to the state-of-the-art real-time MOT system for multiple MOT tasks.
Donghwa Kang, Seunghoon Lee 0002, Cheol-Ho Hong, Jinkyu Lee 0001, Hyeongboo Baek
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 ScissionLite: Accelerating Distributed Deep Learning With Lightweight Data Compression for IIoT
abstract
Industrial Internet of Things (IIoT) applications can greatly benefit from leveraging edge computing. For instance, applications relying on deep neural network (DNN) models can be sliced and distributed across IIoT devices and the network edge to reduce inference latency. However, low network performance between IIoT devices and the edge often becomes a bottleneck. In this study, we propose ScissionLite, a holistic framework designed to accelerate distributed DNN inference using lightweight data compression. Our compression method features a novel lightweight down/upsampling network tailored for performance-limited IIoT devices, which is inserted at the slicing point of a DNN model to reduce outbound network traffic without causing a significant drop in accuracy. In addition, we have developed a benchmarking tool to accurately identify the optimal slicing point of the DNN for the best inference latency. ScissionLite improves inference latency by up to 15.7× with minimal accuracy degradation.
Hyunho Ahn, Munkyu Lee, Sihoon Seong, Gap-Joo Na, In-Geol Chun, Blesson Varghese, Cheol-Ho Hong
IEEE Trans. Ind. Informatics7
2022 gShare: A centralized GPU memory management framework to enable GPU memory sharing for containers
Munkyu Lee, Hyunho Ahn, Cheol-Ho Hong, Dimitrios S. Nikolopoulos
Future Gener. Comput. Syst.3
2018 Kafe: Can OS Kernels Forward Packets Fast Enough for Software Routers?
abstract
It is widely believed that software routers based on commodity operating systems cannot deliver high-speed packet processing, and a number of alternative approaches (including user-space network stacks) have been proposed. This paper revisits the inefficiency of kernel-level packet processing inside modern OS-based software routers and explores whether a redesign of kernel network stacks can improve the incompetence. We present a case contrary to the belief through a redesign: Kafe-a kernel-based advanced forwarding engine that can process packets as fast as user-space network stacks. The Kafe neither adds any new API nor depends on proprietary hardware features, but the Kafe outperforms Linux by seven times and RouteBricks by three times. The current implementation of the Kafe can forward 64-byte IPv4 packets at 28.2 Gbps using eight cores running at 2.6 GHz. Our evaluation results show that the Kafe achieves similar packet forwarding performance to Intel DPDK while consuming much less CPU and memory resources.
Cheol-Ho Hong, Kyungwoon Lee, Jae-Hyun Hwang, Hyunchan Park, Chuck Yoo
IEEE/ACM Trans. Netw.1
2017 FairGV: Fair and Fast GPU Virtualization
abstract
Increasingly high performance computing (HPC) application developers are opting to use cloud resources due to higher availability. Virtualized GPUs would be an obvious and attractive option for HPC application developers using cloud hosting services. Unfortunately, existing GPU virtualization software is not ready to address fairness, utilization, and performance limitations associated with consolidating mixed HPC workloads. This paper presents FairGV, a radically redesigned GPU virtualization system that achieves system-wide weighted fair sharing and strong performance isolation in mixed workloads that use GPUs with variable degrees of intensity. To achieve its objectives, FairGV introduces a trap-less GPU processing architecture, a new fair queuing method integrated with work-conserving and GPU-centric coscheduling polices, and a collaborative scheduling method for non-preemptive GPUs. Our prototype implementation achieves near ideal fairness (≥ 0.97 Min-Max Ratio) with little performance degradation (≤ 1.02 aggregated overhead) in a range of mixed HPC workloads that leverage GPUs.
Cheol-Ho Hong, Ivor T. A. Spence, Dimitrios S. Nikolopoulos
IEEE Trans. Parallel Distributed Syst.1
2016 Synchronization support for parallel applications in virtualized clouds
Cheol-Ho Hong, Young-Pil Kim, Hyunchan Park, Chuck Yoo
J. Supercomput.1
2016 Storage SLA Guarantee with Novel SSD I/O Scheduler in Virtualized Data Centers
abstract
Service level agreements (SLAs) for storage performance in virtualized systems are difficult to guarantee, because different consolidated virtual machines have their own performance requirements. Moreover, hard disk drives (HDDs) in virtualized systems are being replaced by solid-state drives (SSDs). SSDs have higher throughput and lower latency than HDDs; however, they pose new challenges in terms of SLAs. In this paper, we determine that existing I/O schedulers working with SSDs fail to guarantee SLAs among virtualmachines, and do not effectively utilize the high performance of SSDs. To address this issue, we propose the opportunistic I/O scheduler (OIOS), a novel I/O scheduler for SSDs. OIOS guarantees SLAs and fully utilizes the high performance of SSDs. To support realistic SLAs, OIOS provides diverse SLA support functions, including reservations, limitations, and proportional sharing. In addition, OIOS accepts SLAs that are specified in four measurement types: bandwidth, I/Os per second (IOPS), latency, and utilization. Experimental results show that OIOS increases the aggregated bandwidth of VMs by 80 percent compared to mClock, while achieving a similar level of fairness. In addition, we evaluate the proposed scheduler with realistic benchmarks, such as Filebench and the Yahoo CloudServing Benchmark. OIOS successfully guarantees the requirements of diverse SLAs with different metrics.
Hyunchan Park, See-hwan Yoo, Cheol-Ho Hong, Chuck Yoo
IEEE Trans. Parallel Distributed Syst.3
2015 SSD-Tailor: Automated Customization System for Solid-State Drives
abstract
Enterprise servers require customized solid-state drives (SSDs) to satisfy their specialized I/O performance and reliability requirements. For effective use of SSDs for enterprise purposes, SSDs must be designed considering requirements such as those related to performance, lifetime, and cost constraints. However, SSDs have numerous hardware and software design options, such as flash memory types and block allocation methods, which have not been well analyzed yet, but on which the SSD performance depends. Furthermore, there is no methodology for determining the optimal design for a particular I/O workload. This paper proposes SSD-Tailor, a customization tool for SSDs. SSD-Tailor determines a near-optimal set of design options for a given workload. SSD designers can use SSD-Tailor to customize SSDs in the early design stage to meet the customer requirements. We evaluate SSD-Tailor with nine I/O workload traces collected from real-world enterprise servers. We observe that SSD-Tailor finds near-optimal SSD designs for these workloads by exploring only about 1% of the entire set of design candidates. We also show that the near-optimal designs increase the average I/O operations per second by up to 17% and decrease the average response time by up to 163% as compared to an SSD with a general design.
Hyunchan Park, Hanchan Jo, Cheol-Ho Hong, Young-Pil Kim, See-hwan Yoo, Chuck Yoo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2014 Performance Prediction and Evaluation of Parallel Applications in KVM, Xen, and VMware
Cheol-Ho Hong, Young-Pil Kim, Hyunchan Park, Chuck Yoo
Euro-Par1