VLDB 2026 Research / reviewers in the wild / expert
Hwijoon Lim
dblp:187/0217
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0002-9872-6234ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SAND: A New Programming Abstraction for Video-based Deep LearningabstractVideo-based deep learning (VDL) is increasingly used across diverse applications and has become highly popular, but it faces significant challenges in preprocessing highly compressed video data. Preprocessing pipelines are complex, requiring extensive engineering effort, and introduce computational bottlenecks, with latency exceeding GPU training time. Existing solutions partially mitigate these issues but remain inefficient and resource-constrained. Juncheol Ye, Seungkook Lee, Hwijoon Lim, Jihyuk Lee, Uitaek Hong, Youngjin Kwon, Dongsu Han |
SOSP | 3 |
| 2024 | Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model TrainingabstractMixture-of-Experts (MoE) is a powerful technique for enhancing the performance of neural networks while decoupling computational complexity from the number of parameters. However, despite this, scaling the number of experts requires adding more GPUs. In addition, the load imbalance in token load across experts causes unnecessary computation or straggler problems. We present ES-MoE, a novel method for efficient scaling MoE training. It offloads expert parameters to host memory and leverages pipelined expert processing to overlap GPU-CPU communication with GPU computation. It dynamically balances token loads across GPUs, improving computational efficiency. ES-MoE accelerates MoE training on a limited number of GPUs without degradation in model performance. We validate our approach on GPT-based MoE models, demonstrating 67$\times$ better scalability and up to 17.5$\times$ better throughput over existing frameworks. Yechan Kim, Hwijoon Lim, Dongsu Han |
ICML | 2 |
| 2024 | Accelerating Model Training in Multi-cluster Environments with Consumer-grade GPUsabstractRapid advances in machine learning necessitate significant computing power and memory for training, which is accessible only to large corporations today. Small-scale players like academics often only have consumer-grade GPU clusters locally and can afford cloud GPU instances to a limited extent. However, training performance significantly degrades in this multi-cluster setting. In this paper, we identify unique opportunities to accelerate training and propose StellaTrain, a holistic framework that achieves near-optimal training speeds in multi-cloud environments. StellaTrain dynamically adapts a combination of acceleration techniques to minimize time-to-accuracy in model training. StellaTrain introduces novel acceleration techniques such as cache-aware gradient compression and a CPU-based sparse optimizer to maximize GPU utilization and optimize the training pipeline. With the optimized pipeline, StellaTrain holistically determines the training configurations to optimize the total training time. We show that StellaTrain achieves up to 104× speedup over PyTorch DDP in inter-cluster settings by adapting training configurations to fluctuating dynamic network bandwidth. StellaTrain demonstrates that we can cope with the scarce network bandwidth through systematic optimization, achieving up to 257.3× and 78.1× speed-ups on the network bandwidths of 100 Mbps and 500 Mbps, respectively. Finally, StellaTrain enables efficient co-training using on-premises and cloud clusters to reduce costs by 64.5% in conjunction with a reduced training time of 28.9%. Hwijoon Lim, Juncheol Ye, Sangeetha Abdu Jyothi, Dongsu Han |
SIGCOMM | 1 |
| 2024 | TopFull: An Adaptive Top-Down Overload Control for SLO-Oriented MicroservicesabstractMicroservice has become a de facto standard for building large-scale cloud applications. Overload control is essential in preventing microservice failures and maintaining system performance under overloads. Although several approaches have been proposed, they are limited to mitigating the overload of individual microservices, lacking assessments of interdependent microservices and APIs. Youngmok Jung, Hwijoon Lim, Hyunho Yeo, Dongsu Han |
SIGCOMM | 4 |
| 2023 | FlexPass: A Case for Flexible Credit-based Transport for Datacenter NetworksabstractProactive transports explicitly allocate bandwidth to each sender with credits which schedule packet transmission. While promising, existing proactive solutions share a stringent deployment requirement; they assume the perfect control of every link and packet in the network. However, the assumption breaks in practice because new transports are usually deployed gradually over time and legacy traffic always coexists. In this paper, we present FlexPass, a credit-based transport that takes deployment flexibility as a first-class citizen. FlexPass uses a novel combination of network and end-host designs to solve the problem of co-existence and gradual deployment. FlexPass leverages a proactive control loop to send credit-scheduled packets and a complementary reactive control loop to send unscheduled packets to utilize the spare bandwidth. Finally, FlexPass prevents queue buildups of both scheduled and unscheduled packets, and recovers lost packets efficiently. Our evaluation on the testbed shows that FlexPass maintains co-existence with legacy transports (DCTCP), while preserving the high-performance properties of the proactive transport. In large-scale simulations, we show that FlexPass delivers the best incremental benefits during the gradual deployment. We find traffic upgraded to FlexPass benefits from the bounded queue and reduced flow completion time by up to 44% compared to the legacy traffic, while minimizing the side-effect on the legacy flows. Hwijoon Lim, Jaehong Kim 0002, Inho Cho, Keon Jang, Wei Bai 0001, Dongsu Han |
EuroSys | 1 |
| 2023 | SAND: A Storage Abstraction for Video-based Deep LearningabstractDeep learning has gained significant success in video applications such as classification, analytics, and self-supervised learning. However, when scaling out to a large volume of videos, existing approaches suffer from a fundamental limitation; they cannot efficiently utilize GPUs for training deep neural networks (DNNs). This is because video decoding in data preparation incurs a prohibitive amount of computing overhead, making GPU idle for the majority of training time. Otherwise, caching raw videos in memory or storage to bypass decoding is not scalable as they account for from tens to hundreds of terabytes. Uitaek Hong, Hwijoon Lim, Hyunho Yeo, Dongsu Han |
HotStorage | 2 |
| 2023 | Neural Cloud Storage: Innovative Cloud Storage Solution for Cold VideoabstractCloud storage providers offer different pricing tiers based on the access frequency of stored data. This pricing plan offers cost benefits for videos that are accessed less than once per month. However, the stringent requirement falls short in addressing the large number of "cold" videos stored today. This paper proposes Neural Cloud Storage (NCS), a pioneering approach to address the problem by applying neural enhancement, specifically content-aware super-resolution (SR). According to our preliminary cost-benefit analysis, NCS can further save an annual 14% total cost of ownership (TCO) compared to the cheapest AWS storage service for cold video. By reducing the cost, it expands the cold video coverage (from 25% to 38%) that can benefit from the multi-tiered service. As deep learning and computational resources continue to advance, we believe that neural enhancement will revolutionize the field of cloud storage. Jinyeong Lim, Juncheol Ye, Jaehong Kim 0002, Hwijoon Lim, Hyunho Yeo, Dongsu Han |
HotStorage | 4 |
| 2022 | OutRAN: co-optimizing for flow completion time in radio access networkabstractTraffic from interactive applications demanding low latency has become dominant in cellular networks. However, existing schedulers of cellular network base stations fall short in delivering low latency when prior information (i.e., dedicated Quality of Service (QoS)) is unavailable; they become service agnostic and perform towards maximizing the radio resource utilization or user fairness. We identify a new opportunity of providing a better latency for those latency-sensitive traffic flows by additionally taking the Flow Completion Time (FCT) into account in downlink scheduling at the base stations. However, the key challenges are 1) it can bring a severe cost in optimization metrics of the existing scheduler and 2) it should work without prior knowledge of the traffic. Jaehong Kim 0002, Yunheon Lee, Hwijoon Lim, Youngmok Jung, Song Min Kim, Dongsu Han |
CoNEXT | 3 |
| 2022 | TSPipe: Learn from Teacher Faster with PipelinesabstractThe teacher-student (TS) framework, training a (student) network by utilizing an auxiliary superior (teacher) network, has been adopted as a popular training paradigm in many machine learning schemes, since the seminal work—Knowledge distillation (KD) for model compression and transfer learning. Many recent self-supervised learning (SSL) schemes also adopt the TS framework, where teacher networks are maintained as the moving average of student networks, called the momentum networks. This paper presents TSPipe, a pipelined approach to accelerate the training process of any TS frameworks including KD and SSL. Under the observation that the teacher network does not need a backward pass, our main idea is to schedule the computation of the teacher and student network separately, and fully utilize the GPU during training by interleaving the computations of the two networks and relaxing their dependencies. In case the teacher network requires a momentum update, we use delayed parameter updates only on the teacher network to attain high model accuracy. Compared to existing pipeline parallelism schemes, which sacrifice either training throughput or model accuracy, TSPipe provides better performance trade-offs, achieving up to 12.15x higher throughput. Hwijoon Lim, Yechan Kim, Sukmin Yun, Jinwoo Shin, Dongsu Han |
ICML | 1 |
| 2022 | NeuroScaler: neural video enhancement at scaleabstractHigh-definition live streaming has experienced tremendous growth. However, the video quality of live video is often limited by the streamer's uplink bandwidth. Recently, neural-enhanced live streaming has shown great promise in enhancing the video quality by running neural super-resolution at the ingest server. Despite its benefit, it is too expensive to be deployed at scale. To overcome the limitation, we present NeuroScaler, a framework that delivers efficient and scalable neural enhancement for live streams. First, to accelerate end-to-end neural enhancement, we propose novel algorithms that significantly reduce the overhead of video super-resolution, encoding, and GPU context switching. Second, to maximize the overall quality gain, we devise a resource scheduler that considers the unique characteristics of the neural-enhancing workload. Our evaluation on a public cloud shows NeuroScaler reduces the overall cost by 22.3× and 3.0--11.1× compared to the latest per-frame and selective neural-enhancing systems, respectively. Hyunho Yeo, Hwijoon Lim, Jaehong Kim 0002, Youngmok Jung, Juncheol Ye, Dongsu Han |
SIGCOMM | 2 |
| 2021 | Towards timeout-less transport in commodity datacenter networksabstractDespite recent advances in datacenter networks, timeouts caused by congestion packet losses still remain a major cause of high tail latency. Priority-based Flow Control (PFC) was introduced to make the network lossless, but its Head-of-Line blocking nature causes various performance and management problems. In this paper, we ask if it is possible to design a network that achieves (near) zero timeout only using commodity hardware in datacenters. Hwijoon Lim, Wei Bai 0001, Yibo Zhu 0001, Youngmok Jung, Dongsu Han |
EuroSys | 1 |