VLDB 2026 Research / reviewers in the wild / expert
Dongsu Han
dblp:12/5388
· DBLP profile ↗
87ranked-venue papers
5as first author
31since 2021 · last 2026
0000-0001-6922-7244ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 61 · 4 first-author · 18 since 2021Systems, architecture and hardware · 10 · 5 since 2021Security and privacy · 6 · 1 first-authorArtificial intelligence and machine learning · 5 · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Trinity: Three-Dimensional Tensor Program Optimization via Tile-level Equality SaturationabstractModern tensor program optimizers operate at two separate levels: graph-level optimizations (operator fusion, algebraic rewrites) and operator-level scheduling (tiling, parallelization). This separation prevents them from discovering cross-operator, tile-level optimizations that make hand-tuned kernels like FlashAttention effective. We present Trinity, the first tensor program optimizer that achieves scalable joint optimization through tile-level equality saturation. Our key insight is that optimal performance requires simultaneously optimizing three interdependent dimensions -- algebraic equivalence, memory I/O, and compute orchestration. To enable this, Trinity introduces a novel fine-grained IR that exposes all three axes as first-class, rewritable entities and applies equality saturation to perform scalable joint optimization. As a result, Trinity automatically discovers complex optimizations that require coordinated reasoning across all three dimensions. Across diverse Transformer variants, Trinity achieves up to 2.09× speedup over TensorRT and 2.35× over TorchInductor, both state-of-the-art production compilers. Haechan An, Gieun Jeong, Jeehoon Kang, Dongsu Han |
ASPLOS (2) | 6 |
| 2026 | BlenDR: Bandwidth-efficient RGB-D Representation and Delivery for Live 3D Video StreamingabstractLive volumetric streaming is experiencing rapid growth due to the availability of depth sensors and 3D cameras. The RGB-D format that utilizes 2D video codecs has emerged as a promising solution for streaming. However, it falls short in delivering high-quality volumetric capture scenes at Internet-friendly bitrates due to inefficient compression, stemming from the unique challenges of adapting the live RGB-D data format to video codecs. Jaehong Kim 0002, Joon Ha Kim, Yunheon Lee, Dongsu Han |
MobiSys | 4 |
| 2026 | NerVast: Compression-Efficient Scaling of Implicit Neural Video Representations via Scene-based Parameter-sharingabstractImplicit neural representation (INR) has emerged as a new data representation for compressing videos and now shows on-par performance with the conventional codecs. The next challenge in the field is to make INR scalable for its practical use. Existing works realize this by utilizing small INR models to scale for long and high-resolution video, which achieves better encoding and decoding speeds. However, they fail to fully exploit the temporal nature of video data when encoding it into multiple separate INRs across time, which leads to sub-optimal compression efficiency. In this work, we propose NerVast, a new encoding scheme for video INR, that improves compression efficiency while still enjoying the low computation and transfer costs of small INR models. When a video is represented in separate INR segments, NerVast effectively reduces the total volume required for representation by sharing the parameters between models during encoding. Without expensive training, NerVast selects the most efficient parameters to share. Then it jointly trains both shared and non-shared parameters in a way that minimizes the quality drop imposed by sharing. While maintaining real-time decoding speed (> 30 fps), NerVast provides better compression (39.9 % reduction in parameters on average) compared to the compute-efficient INR models. In other words, NerVast is better in encoding quality (1.57 dB higher in PSNR) with the same bitrate. Yunheon Lee, Juncheol Ye, Jaehong Kim 0002, Dongsu Han |
WACV | 4 |
| 2025 | Towards an Agentic Workflow for Internet Measurement ResearchabstractInternet measurement research faces an accessibility crisis: complex analyses require custom integration of multiple specialized tools that demands specialized domain expertise. When network disruptions occur, operators need rapid diagnostic workflows spanning infrastructure mapping, routing analysis, and dependency modeling. However, developing these workflows requires specialized knowledge and significant manual effort. Alagappan Ramanathan, Eunju Kang, Dongsu Han, Sangeetha Abdu Jyothi |
HotNets | 3 |
| 2025 | Presto: Hybrid CPU-GPU Preprocessing Framework for Video-based AI Inference SystemabstractThe growing adoption of video-based AI models has created a pressing demand for high throughput, low latency inference systems. However, existing preprocessing frameworks—whether CPU or GPU based—struggle to keep up with the computational burdens of video decoding and data augmentation, resulting in suboptimal GPU utilization and degraded inference system performance. Jihyuk Lee, Dongsu Han, Jaehong Kim 0002 |
MobiSys | 2 |
| 2025 | SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMsabstractLarge language models (LLMs) power many modern applications, but serving them at scale remains costly and resource-intensive. Current server-centric systems overlook consumer-grade GPUs at the edge. We introduce SpecEdge, an edge-assisted inference framework that splits LLM workloads between edge and server GPUs using a speculative decoding scheme, exchanging only token outputs over the network. SpecEdge employs proactive edge drafting to overlap edge token creation with server verification and pipeline-aware scheduling that interleaves multiple user requests to increase server-side throughput. Experiments show SpecEdge enhances overall cost efficiency by **1.91×** through achieving **2.22×** server throughput, and reduces inter token latency by **11.24\%** compared to a server-only baseline, introducing a scalable, cost-effective paradigm for LLM serving. The code is available at https://github.com/kaist-ina/specedge Seunggeun Cho, Dongsu Han |
NeurIPS | 3 |
| 2025 | Agua: A Concept-Based Explainer for Learning-Enabled SystemsabstractWhile deep learning offers superior performance in systems and networking, adoption is often hindered by difficulties in understanding and debugging. Explainability aims to bridge this gap by providing insight into the model's decisions. However, existing methods primarily identify the most influential input features, forcing operators to perform extensive manual analysis of low-level signals (e.g., buffer t - 1). Dongsu Han, Nina Narodytska, Sangeetha Abdu Jyothi |
SIGCOMM | 2 |
| 2025 | SAND: A New Programming Abstraction for Video-based Deep LearningabstractVideo-based deep learning (VDL) is increasingly used across diverse applications and has become highly popular, but it faces significant challenges in preprocessing highly compressed video data. Preprocessing pipelines are complex, requiring extensive engineering effort, and introduce computational bottlenecks, with latency exceeding GPU training time. Existing solutions partially mitigate these issues but remain inefficient and resource-constrained. Juncheol Ye, Seungkook Lee, Hwijoon Lim, Jihyuk Lee, Uitaek Hong, Youngjin Kwon, Dongsu Han |
SOSP | 7 |
| 2025 | NarrAD: Automatic Generation of Audio Descriptions for Movies with Rich Narrative ContextabstractAudio Description (AD) is a narration designed to enhance accessibility for visually impaired individuals by conveying the key visual elements of a video. Thus, automating AD generation for long-form videos, such as movies and dramas, provides high social value but is a challenging task. First, AD must reflect the narrative context of the entire movie, including the storyline, names of characters and places, and the cultural setting. Second, to avoid disrupting the immersive experience of the movie, AD must not overlap with the characters' dialogues, requiring the delivery of numerous visual elements in concise sentences. This paper presents NarrAD, a training-free AD generation framework that satisfies both of the requirements by leveraging rich narrative context in movie scripts and curating information across narration slots. Experiments on the MAD dataset demonstrate that our approach outperforms prior works in both captioning and LLM-based metrics. In the user study with 600 subjects, NarrAD achieves the highest user experience and movie comprehension. NarrAD's AD samples are available at https://bit.ly/4aSwOTr. Juncheol Ye, Seungkook Lee, Hyun W. Ka, Dongsu Han |
WACV | 5 |
| 2025 | Low-Overhead Intra-Host Container Communication With Hardware OffloadingabstractContainers are widely embraced for their deployment and performance benefits over virtual machines. Yet, for many data-intensive applications in containerized clouds, bulky data transfers may impose performance issues. In particular, communication across co-located containers on the same host incurs large overheads in memory copy and the kernel’s TCP stack. Existing solutions such as shared-memory networking and RDMA have their own limitations, including insufficient memory isolation and limited scalability. This paper presents PipeDevice, a new system for low overhead intra-host container communication. PipeDevice follows a hardware-software co-design approach — it offloads data forwarding entirely onto hardware, which accesses application data in hugepages on the host, thereby eliminating CPU overhead from memory copy and TCP processing. PipeDevice preserves memory isolation and scales well to connections, making it deployable in public clouds. Isolation is achieved by allocating dedicated memory to each connection from hugepages. To achieve high scalability, PipeDevice stores the connection states entirely in host DRAM and manages them in software. Evaluation with a prototype implementation on commodity FPGA shows that for delivering 80Gbps across containers PipeDevice saves 63.2% CPU compared to kernel TCP stack, and 40.5% over FreeFlow. PipeDevice provides salient benefits to applications. For example, we port baidu-allreduce to PipeDevice and obtain$\sim 2.2\times $gains in allreduce throughput. Zhixiong Niu, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong, Dongsu Han, Chun Jason Xue, Hong Xu 0001 |
IEEE Trans. Netw. | 6 |
| 2024 | Toward Trustworthy Learning-Enabled Systems with Concept-Based ExplanationsabstractDespite the superior performance of deep learning-based controllers in network applications, their practical adoption is limited due to the difficulty in understanding and trusting them. Existing explainability solutions largely focus on interpreting these controllers by providing insights into the top features used by the model. Although these insights can help reveal an important aspect of the controller, they require operators to deal with low-level features, requiring extensive manual analysis and interpretation. Dongsu Han, Nina Narodytska, Sangeetha Abdu Jyothi |
HotNets | 2 |
| 2024 | Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model TrainingabstractMixture-of-Experts (MoE) is a powerful technique for enhancing the performance of neural networks while decoupling computational complexity from the number of parameters. However, despite this, scaling the number of experts requires adding more GPUs. In addition, the load imbalance in token load across experts causes unnecessary computation or straggler problems. We present ES-MoE, a novel method for efficient scaling MoE training. It offloads expert parameters to host memory and leverages pipelined expert processing to overlap GPU-CPU communication with GPU computation. It dynamically balances token loads across GPUs, improving computational efficiency. ES-MoE accelerates MoE training on a limited number of GPUs without degradation in model performance. We validate our approach on GPT-based MoE models, demonstrating 67$\times$ better scalability and up to 17.5$\times$ better throughput over existing frameworks. Yechan Kim, Hwijoon Lim, Dongsu Han |
ICML | 3 |
| 2024 | Accelerating Model Training in Multi-cluster Environments with Consumer-grade GPUsabstractRapid advances in machine learning necessitate significant computing power and memory for training, which is accessible only to large corporations today. Small-scale players like academics often only have consumer-grade GPU clusters locally and can afford cloud GPU instances to a limited extent. However, training performance significantly degrades in this multi-cluster setting. In this paper, we identify unique opportunities to accelerate training and propose StellaTrain, a holistic framework that achieves near-optimal training speeds in multi-cloud environments. StellaTrain dynamically adapts a combination of acceleration techniques to minimize time-to-accuracy in model training. StellaTrain introduces novel acceleration techniques such as cache-aware gradient compression and a CPU-based sparse optimizer to maximize GPU utilization and optimize the training pipeline. With the optimized pipeline, StellaTrain holistically determines the training configurations to optimize the total training time. We show that StellaTrain achieves up to 104× speedup over PyTorch DDP in inter-cluster settings by adapting training configurations to fluctuating dynamic network bandwidth. StellaTrain demonstrates that we can cope with the scarce network bandwidth through systematic optimization, achieving up to 257.3× and 78.1× speed-ups on the network bandwidths of 100 Mbps and 500 Mbps, respectively. Finally, StellaTrain enables efficient co-training using on-premises and cloud clusters to reduce costs by 64.5% in conjunction with a reduced training time of 28.9%. Hwijoon Lim, Juncheol Ye, Sangeetha Abdu Jyothi, Dongsu Han |
SIGCOMM | 4 |
| 2024 | TopFull: An Adaptive Top-Down Overload Control for SLO-Oriented MicroservicesabstractMicroservice has become a de facto standard for building large-scale cloud applications. Overload control is essential in preventing microservice failures and maintaining system performance under overloads. Although several approaches have been proposed, they are limited to mitigating the overload of individual microservices, lacking assessments of interdependent microservices and APIs. Youngmok Jung, Hwijoon Lim, Hyunho Yeo, Dongsu Han |
SIGCOMM | 6 |
| 2024 | Graph Neural Network-Based SLO-Aware Proactive Resource Autoscaling Framework for MicroservicesabstractMicroservice is an architectural style widely adopted in various latency-sensitive cloud applications. Similar to the monolith, autoscaling has attracted the attention of operators for managing the resource utilization of microservices. However, it is still challenging to optimize resources in terms of latency service-level-objective (SLO) without human intervention. In this paper, we present GRAF, a graph neural network-based SLO-aware proactive resource autoscaling framework for minimizing total CPU resources while satisfying latency SLO. GRAF leverages front-end workload, distributed tracing data, and machine learning approaches to (a) observe/estimate the impact of traffic change (b) find optimal resource combinations (c) make proactive resource allocation. Experiments using various open-source benchmarks demonstrate that GRAF successfully targets latency SLO while saving up to 19% of total CPU resources compared to the fine-tuned autoscaler. GRAF also handles a traffic surge with 36% fewer resources while achieving up to 2.6x faster tail latency convergence compared to the Kubernetes autoscaler. Moreover, we verify the scalability of GRAF on large-scale deployments, where GRAF saves 21.6% and 25.4% for CPU resources and memory resources, respectively. Byungkwon Choi, Chunghan Lee, Dongsu Han |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | AccelIR: Task-aware Image Compression for Accelerating Neural RestorationabstractRecently, deep neural networks have been successfully applied for image restoration (IR) (e.g., super-resolution, de-noising, de-blurring). Despite their promising performance, running IR networks requires heavy computation. A large body of work has been devoted to addressing this issue by designing novel neural networks or pruning their parameters. However, the common limitation is that while images are saved in a compressed format before being enhanced by IR, prior work does not consider the impact of compression on the IR quality. In this paper, we present AccelIR, a framework that optimizes image compression considering the end-to-end pipeline of IR tasks. AccelIR encodes an image through IR-aware compression that optimizes compression levels across image blocks within an image according to the impact on the IR quality. Then, it runs a lightweight IR network on the compressed image, effectively reducing IR computation, while maintaining the same IR quality and image size. Our extensive evaluation using nine IR networks shows that AccelIR can reduce the computing overhead of super-resolution, de-nosing, and de-blurring by 49%, 29%, and 32% on average, respectively. Juncheol Ye, Hyunho Yeo, Dongsu Han |
CVPR | 4 |
| 2023 | FlexPass: A Case for Flexible Credit-based Transport for Datacenter NetworksabstractProactive transports explicitly allocate bandwidth to each sender with credits which schedule packet transmission. While promising, existing proactive solutions share a stringent deployment requirement; they assume the perfect control of every link and packet in the network. However, the assumption breaks in practice because new transports are usually deployed gradually over time and legacy traffic always coexists. In this paper, we present FlexPass, a credit-based transport that takes deployment flexibility as a first-class citizen. FlexPass uses a novel combination of network and end-host designs to solve the problem of co-existence and gradual deployment. FlexPass leverages a proactive control loop to send credit-scheduled packets and a complementary reactive control loop to send unscheduled packets to utilize the spare bandwidth. Finally, FlexPass prevents queue buildups of both scheduled and unscheduled packets, and recovers lost packets efficiently. Our evaluation on the testbed shows that FlexPass maintains co-existence with legacy transports (DCTCP), while preserving the high-performance properties of the proactive transport. In large-scale simulations, we show that FlexPass delivers the best incremental benefits during the gradual deployment. We find traffic upgraded to FlexPass benefits from the bounded queue and reduced flow completion time by up to 44% compared to the legacy traffic, while minimizing the side-effect on the legacy flows. Hwijoon Lim, Jaehong Kim 0002, Inho Cho, Keon Jang, Wei Bai 0001, Dongsu Han |
EuroSys | 6 |
| 2023 | SAND: A Storage Abstraction for Video-based Deep LearningabstractDeep learning has gained significant success in video applications such as classification, analytics, and self-supervised learning. However, when scaling out to a large volume of videos, existing approaches suffer from a fundamental limitation; they cannot efficiently utilize GPUs for training deep neural networks (DNNs). This is because video decoding in data preparation incurs a prohibitive amount of computing overhead, making GPU idle for the majority of training time. Otherwise, caching raw videos in memory or storage to bypass decoding is not scalable as they account for from tens to hundreds of terabytes. Uitaek Hong, Hwijoon Lim, Hyunho Yeo, Dongsu Han |
HotStorage | 5 |
| 2023 | Neural Cloud Storage: Innovative Cloud Storage Solution for Cold VideoabstractCloud storage providers offer different pricing tiers based on the access frequency of stored data. This pricing plan offers cost benefits for videos that are accessed less than once per month. However, the stringent requirement falls short in addressing the large number of "cold" videos stored today. This paper proposes Neural Cloud Storage (NCS), a pioneering approach to address the problem by applying neural enhancement, specifically content-aware super-resolution (SR). According to our preliminary cost-benefit analysis, NCS can further save an annual 14% total cost of ownership (TCO) compared to the cheapest AWS storage service for cold video. By reducing the cost, it expands the cold video coverage (from 25% to 38%) that can benefit from the multi-tiered service. As deep learning and computational resources continue to advance, we believe that neural enhancement will revolutionize the field of cloud storage. Jinyeong Lim, Juncheol Ye, Jaehong Kim 0002, Hwijoon Lim, Hyunho Yeo, Dongsu Han |
HotStorage | 6 |
| 2023 | Low-earth Orbit Satellite Network Optimization and Statistical Qubit Freezing on Quantum AnnealerabstractRecent studies have shown promising results indicating the potential of Noisy Intermediate Scale Quantum (NISQ) devices. In this work, Low-earth Orbit (LEO) satellite network design problem is formulated as a Quadratic Unconstrained Binary Optimization (QUBO) problem, and Quantum Annealing (QA) is utilized to solve the resulting problem with real world setting. Compare to classical approaches, the experimental results indicate improvements in terms of network path length and number of satellite hops in network path that amount to network latency. Further, an iterative post-processing method, Statistical Qubit Freezing (SQF), which freezes initial states of qubits and reduces the size of the problem in each annealing cycle, is proposed and evaluated. Solution found with SQF indicates that SQF in fact allows the system to reach lower energy state. Jeung Rac Lee, Yunheon Lee, Dongsu Han, Changjun Kim, Bo Hyun Choi, June-Koo Kevin Rhee |
ICC | 3 |
| 2023 | Scalable and Secure Virtualization of HSM With ScaleTrustabstractHardware security modules (HSMs) have been utilized as a trustworthy foundation for cloud services. Unfortunately, existing systems using HSMs fail to meet multi-tenant scalability arising from the emerging trends such as microservices, which utilize frequent cryptographic operations. As an alternative, cloud vendors provide HSMs as a service. However, such cloud-managed HSM usage models raise security concerns due to their untrusted and shared operating environment. We propose ScaleTrust, a scalable and secure system for key management. ScaleTrust allows us to scale the number of virtual HSM partitions, each of which is isolated with respect to each other and is robust against cloud insider attacks, while preserving physical isolation of the root of trust. To enable this, ScaleTrust uses Intel SGX and multiple HSM features, such as restricting key usage by controlling key attributes of in-HSM keys and establishing a secure channel using only HSM commands. Finally, we apply ScaleTrust to four real-world systems: Keyless SSL for TLS private key offloading, JSON Web Token authentication for microservices, key provisioning, and encryption in database systems. Our evaluation shows that ScaleTrust achieves multi-tenancy in a scalable way by providing multiple virtual HSMs with legacy HSM devices that are designed to support a single tenant. ScaleTrust provides security against insider threats while incurring 11.9% and 39.0% of end-to-end throughput and latency overhead for Keyless SSL compared to stand-alone HSMs. Juhyeng Han, Insu Yun, Taesoo Kim, Sooel Son, Dongsu Han |
IEEE/ACM Trans. Netw. | 6 |
| 2022 | OutRAN: co-optimizing for flow completion time in radio access networkabstractTraffic from interactive applications demanding low latency has become dominant in cellular networks. However, existing schedulers of cellular network base stations fall short in delivering low latency when prior information (i.e., dedicated Quality of Service (QoS)) is unavailable; they become service agnostic and perform towards maximizing the radio resource utilization or user fairness. We identify a new opportunity of providing a better latency for those latency-sensitive traffic flows by additionally taking the Flow Completion Time (FCT) into account in downlink scheduling at the base stations. However, the key challenges are 1) it can bring a severe cost in optimization metrics of the existing scheduler and 2) it should work without prior knowledge of the traffic. Jaehong Kim 0002, Yunheon Lee, Hwijoon Lim, Youngmok Jung, Song Min Kim, Dongsu Han |
CoNEXT | 6 |
| 2022 | PipeDevice: a hardware-software co-design approach to intra-host container communicationabstractContainers are prevalently adopted due to the deployment and performance advantages over virtual machines. For many containerized data-intensive applications, however, bulky data transfers may pose performance issues. In particular, communication across co-located containers on the same host incurs large overheads in memory copy and the kernel's TCP stack. Existing solutions such as shared-memory networking and RDMA have their own limitations, including insufficient memory isolation and limited scalability. Chuanwen Wang, Zhixiong Niu, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong, Dongsu Han, Chun Jason Xue, Hong Xu 0001 |
CoNEXT | 7 |
| 2022 | TSPipe: Learn from Teacher Faster with PipelinesabstractThe teacher-student (TS) framework, training a (student) network by utilizing an auxiliary superior (teacher) network, has been adopted as a popular training paradigm in many machine learning schemes, since the seminal work—Knowledge distillation (KD) for model compression and transfer learning. Many recent self-supervised learning (SSL) schemes also adopt the TS framework, where teacher networks are maintained as the moving average of student networks, called the momentum networks. This paper presents TSPipe, a pipelined approach to accelerate the training process of any TS frameworks including KD and SSL. Under the observation that the teacher network does not need a backward pass, our main idea is to schedule the computation of the teacher and student network separately, and fully utilize the GPU during training by interleaving the computations of the two networks and relaxing their dependencies. In case the teacher network requires a momentum update, we use delayed parameter updates only on the teacher network to attain high model accuracy. Compared to existing pipeline parallelism schemes, which sacrifice either training throughput or model accuracy, TSPipe provides better performance trade-offs, achieving up to 12.15x higher throughput. Hwijoon Lim, Yechan Kim, Sukmin Yun, Jinwoo Shin, Dongsu Han |
ICML | 5 |
| 2022 | NeuroScaler: neural video enhancement at scaleabstractHigh-definition live streaming has experienced tremendous growth. However, the video quality of live video is often limited by the streamer's uplink bandwidth. Recently, neural-enhanced live streaming has shown great promise in enhancing the video quality by running neural super-resolution at the ingest server. Despite its benefit, it is too expensive to be deployed at scale. To overcome the limitation, we present NeuroScaler, a framework that delivers efficient and scalable neural enhancement for live streams. First, to accelerate end-to-end neural enhancement, we propose novel algorithms that significantly reduce the overhead of video super-resolution, encoding, and GPU context switching. Second, to maximize the overall quality gain, we devise a resource scheduler that considers the unique characteristics of the neural-enhancing workload. Our evaluation on a public cloud shows NeuroScaler reduces the overall cost by 22.3× and 3.0--11.1× compared to the latest per-frame and selective neural-enhancing systems, respectively. Hyunho Yeo, Hwijoon Lim, Jaehong Kim 0002, Youngmok Jung, Juncheol Ye, Dongsu Han |
SIGCOMM | 6 |
| 2022 | BWA-MEME: BWA-MEM emulated with a machine learning approachabstractMOTIVATION: The growing use of next-generation sequencing and enlarged sequencing throughput require efficient short-read alignment, where seeding is one of the major performance bottlenecks. The key challenge in the seeding phase is searching for exact matches of substrings of short reads in the reference DNA sequence. Existing algorithms, however, present limitations in performance due to their frequent memory accesses. RESULTS: This article presents BWA-MEME, the first full-fledged short read alignment software that leverages learned indices for solving the exact match search problem for efficient seeding. BWA-MEME is a practical and efficient seeding algorithm based on a suffix array search algorithm that solves the challenges in utilizing learned indices for SMEM search which is extensively used in the seeding phase. Our evaluation shows that BWA-MEME achieves up to 3.45× speedup in seeding throughput over BWA-MEM2 by reducing the number of instructions by 4.60×, memory accesses by 8.77× and LLC misses by 2.21×, while ensuring the identical SAM output to BWA-MEM2. AVAILABILITY AND IMPLEMENTATION: The source code and test scripts are available for academic use at https://github.com/kaist-ina/BWA-MEME/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Youngmok Jung, Dongsu Han |
Bioinform. | 2 |
| 2022 | NetKernel: Making Network Stack Part of the Virtualized InfrastructureabstractThis paper presents a system called NetKernel that decouples the network stack from the guest virtual machine and offers it as an independent module. NetKernel represents a new paradigm where network stack can be managed as part of the virtualized infrastructure. It provides important efficiency benefits: By gaining control and visibility of the network stack, operators can perform network management more directly and flexibly, such as multiplexing VMs running different applications to the same network stack module to save CPU cores, and enforcing fair bandwidth sharing. Users also benefit from the simplified stack deployment and better performance: For example mTCP can be deployed without API change to support nginx natively, and shared memory networking can be readily enabled to improve performance of colocated VMs. Testbed evaluation using 100G NICs shows that NetKernel preserves the performance and scalability of both kernel and userspace network stacks, and provides the same isolation as the current architecture. Zhixiong Niu, Peng Cheng 0005, Yongqiang Xiong, Dongsu Han, Keith Winstein, Chun Jason Xue, Hong Xu 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2022 | Congestion Control for Cross-Datacenter NetworksabstractGeographically distributed applications hosted on cloud are becoming prevalent. They run oncross-datacenter networkthat consists of multiple data center networks (DCNs) connected by a wide area network (WAN). Such a cross-DC network poses significant challenges in transport design because the DCN and WAN segments have vastly distinct characteristics (e.g., buffer depths, RTTs). In this paper, we find that existing DCN or WAN transport reacting to ECN or delay alone do not (and cannot be extended to) work well for such an environment. The key reason is that neither of the signals, by itself only, can simultaneously capture the location and degree of congestion, mainly due to the discrepancies between DCN and WAN. Motivated by this, we present the design and implementation of GEMINI that strategically integrates both ECN and delay signals for cross-DC congestion control. To achieve low latency, GEMINI bounds the inter-DC latency with delay signal and prevents the intra-DC packet loss with ECN. To maintain high throughput, GEMINI modulates the window dynamics and maintains low buffer occupancy utilizing both congestion signals. GEMINI is implemented in Linux kernel and evaluated by extensive testbed experiments. Results show that GEMINI achieves up to 53%, 31%, 76% and 2% reduction of small flow average completion times, and up to 34%, 39%, 9% and 58% reduction of large flow average completion times compared to TCP Cubic, DCTCP, BBR and TCP Vegas. Gaoxiong Zeng, Wei Bai 0001, Kai Chen 0005, Dongsu Han, Yibo Zhu 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2021 | pHPA: A Proactive Autoscaling Framework for Microservice ChainabstractMicroservice is an architectural style that breaks down monolithic applications into smaller microservices and has been widely adopted by a variety of enterprises. Like the monolith, autoscaling has attracted the attention of operators in scaling microservices. However, most existing approaches of autoscaling do not consider microservice chain and severely degrade the performance of microservices when traffic surges. In this paper, we present pHPA, an autoscaling framework for the microservice chain. pHPA proactively allocates resources to the microservice chains and effectively handles traffic surges. Our evaluation using various open-source benchmarks shows that pHPA reduces 99%-tile latency and resource usage by up to 70% and 58% respectively compared to the most widely used autoscaler when traffic surges. Byungkwon Choi, Chunghan Lee, Dongsu Han |
APNet | 4 |
| 2021 | GRAF: a graph neural network based proactive resource allocation framework for SLO-oriented microservicesabstractMicroservice is an architectural style that has been widely adopted in various latency-sensitive applications. Similar to the monolith, autoscaling has attracted the attention of operators for managing resource utilization of microservices. However, it is still challenging to optimize resources in terms of latency service-level-objective (SLO) without human intervention. In this paper, we present GRAF, a graph neural network-based proactive resource allocation framework for minimizing total CPU resources while satisfying latency SLO. GRAF leverages front-end workload, distributed tracing data, and machine learning approaches to (a) observe/estimate impact of traffic change (b) find optimal resource combinations (c) make proactive resource allocation. Experiments using various open-source benchmarks demonstrate that GRAF successfully targets latency SLO while saving up to 19% of total CPU resources compared to the fine-tuned autoscaler. Moreover, GRAF handles traffic surge with 36% fewer resources while achieving up to 2.6x faster tail latency convergence compared to the Kubernetes autoscaler. Byungkwon Choi, Chunghan Lee, Dongsu Han |
CoNEXT | 4 |
| 2021 | Towards timeout-less transport in commodity datacenter networksabstractDespite recent advances in datacenter networks, timeouts caused by congestion packet losses still remain a major cause of high tail latency. Priority-based Flow Control (PFC) was introduced to make the network lossless, but its Head-of-Line blocking nature causes various performance and management problems. In this paper, we ask if it is possible to design a network that achieves (near) zero timeout only using commodity hardware in datacenters. Hwijoon Lim, Wei Bai 0001, Yibo Zhu 0001, Youngmok Jung, Dongsu Han |
EuroSys | 5 |
| 2020 | Leveraging SIMD parallelism for accelerating network applicationsabstractSoftware packet processing frameworks act as critical components in modern network architecture, as their performance has a vital impact on the quality of the network services. Motivated by the increasing number and capability for advanced vector instructions in recent mainstream CPUs, this paper explores a new parallel processing design and implementation of data structures and algorithms that are frequently used for building network applications. In particular, we propose effective SIMD optimization techniques for the bloom filter and Open vSwitch megaflow cache. Our design reduces memory access latency via careful prefetching and a new design that meets the needs of fast data consuming instructions. Our evaluation shows performance improvements up to 162% in bloom filter and 48% in Open vSwitch compared to their scalar version. Hejing Li, Juhyeng Han, Dongsu Han |
APNet | 3 |
| 2020 | Lumos: Improving Smart Home IoT Visibility and Interoperability Through Analyzing Mobile AppsabstractThe era of Smart Homes and the Internet of Things (IoT) calsl for integrating diverse "smart" devices, including sensors, actuators, and home appliances. However, enabling interoperation across heterogeneous IoT devices is a challenging task because vendors use their own control and communication protocols. Prior approaches have attempted to solve this problem by asking for vendor support, or even fundamentally re-designing the architecture of IoT devices. These approaches face limitations as they require disruptive changes.This paper explores a new approach to improving IoT interoperability without requiring architectural changes or vendor participation. Focusing on smart-home environments, we propose Lumos that improves interoperability by leveraging Android apps that control IoT devices. Lumos uses this information learned from IoT apps to enable "best-effort" interoperation across heterogeneous devices. Our evaluation with 15 commercial IoT devices from three major IoT platforms and in-depth user studies conducted with 24 participants demonstrate the promising efficacy of Lumos for implementing diverse interoperation scenarios. Steven Y. Ko, Sooel Son, Dongsu Han |
ICNP | 4 |
| 2020 | NEMO: enabling neural-enhanced video streaming on commodity mobile devicesabstractThe demand for mobile video streaming has experienced tremendous growth over the last decade. However, existing methods of video delivery fall short of delivering high-quality video. Recent advances in neural super-resolution have opened up the possibility of enhancing video quality by leveraging client-side computation. Unfortunately, mobile devices cannot benefit from this because it is too expensive in computation and power-hungry. Hyunho Yeo, Chan Ju Chong, Youngmok Jung, Juncheol Ye, Dongsu Han |
MobiCom | 5 |
| 2020 | Neural-Enhanced Live Streaming: Improving Live Video Ingest via Online LearningabstractLive video accounts for a significant volume of today's Internet video. Despite a large number of efforts to enhance user quality of experience (QoE) both at the ingest and distribution side of live video, the fundamental limitations are that streamer's upstream bandwidth and computational capacity limit the quality of experience of thousands of viewers. Jaehong Kim 0002, Youngmok Jung, Hyunho Yeo, Juncheol Ye, Dongsu Han |
SIGCOMM | 5 |
| 2020 | NetKernel: Making Network Stack Part of the Virtualized Infrastructure
Zhixiong Niu, Hong Xu 0001, Peng Cheng 0005, Yongqiang Xiong, Tao Wang 0088, Dongsu Han, Keith Winstein |
USENIX ATC | 7 |
| 2020 | A Secure Middlebox Framework for Enabling Visibility Over Multiple Encryption ProtocolsabstractNetwork middleboxes provide the first line of defense for enterprise networks. Many of them typically inspect packet payload to filter malicious attack patterns. However, the widespread use of end-to-end cryptographic protocols designed to promote security and privacy, either inhibits deep packet inspection in the network or forces enterprises to use solutions that are not secure. This article introduces a complete framework for building secure and practical network middleboxes, called EVE, which enables visibility over encrypted traffic. EVE securely processes encrypted traffic using a combination of hardware-based trusted execution and software security technology. For enhanced programmability and security, EVE provides a high-level programming interface based on the Rust language. The high-level APIs of EVE provide security and significantly ease the development effort by hiding the details of cryptographic operations, enclave processing, TCP reassembly, and out-of-band key sharing. Our evaluation shows EVE supports diverse use cases with multiple encryption protocols in a secure fashion while delivering high performance. Juhyeng Han, Seong-Min Kim, Daeyang Cho, Byungkwon Choi, Jaehyeong Ha, Dongsu Han |
IEEE/ACM Trans. Netw. | 6 |
| 2019 | FlowShader: a Generalized Framework for GPU-accelerated VNF Flow ProcessingabstractGPU acceleration has been widely investigated for packet processing in virtual network functions (NFs), but not for L7 flow-processing NFs. In L7 NFs, reassembled TCP messages of the same flow should be processed in order in the same processing thread, and the uneven sizes among flows pose a major challenge for full realization of GPU's parallel computation power. To exploit GPUs for L7 NF processing, this paper presents FlowShader, a GPU acceleration framework to achieve both high generality and throughput even under skewed flow size distributions. We carefully design an efficient scheduling algorithm that fully exploits available GPU and CPU capacities; in particular, we dispatch large flows which seriously break up the size balance to CPU and the rest of flows to GPU. Furthermore, FlowShader allows similar NF logic (as CPU-based NFs) to run on individual threads in a GPU, which is more generalized and easy to take on as compared to redesigning an NF for operation parallelism on GPU. We implemented a number of L7 flow processing NFs based on FlowShader. Evaluations are conducted under both synthetic and real-world traffic traces and results show that the throughput achieved by FlowShader is up to 6x that of the CPU-only baseline and 3x of the GPU-only design. Xiaodong Yi 0001, Jingpu Duan, Wei Bai 0001, Chuan Wu 0001, Yongqiang Xiong, Dongsu Han |
ICNP | 7 |
| 2019 | Congestion Control for Cross-Datacenter NetworksabstractGeographically distributed applications hosted on cloud are becoming prevalent. They run on cross-datacenter network that consists of multiple data center networks (DCNs) connected by a wide area network (WAN). Such a cross-DC network imposes significant challenges in transport design because the DCN and WAN segments have vastly distinct characteristics (e.g., butter depths, RTTs). In this paper, we find that existing DCN or WAN transports reacting to ECN or delay alone do not (and cannot be extended to) work well for such an environment. The key reason is that neither of the signals, by itself, can simultaneously capture the location and degree of congestion. This is due to the discrepancies between DCN and WAN. Motivated by this, we present the design and implementation of GEMINI that strategically integrates both ECN and delay signals for cross-DC congestion control. To achieve low latency, GEMINI bounds the inter-DC latency with delay signal and prevents the intra-DC packet loss with ECN. To maintain high throughput, GEMINI modulates the window dynamics and maintains low butter occupancy utilizing both congestion signals. GEMINI is implemented in Linux kernel and evaluated by extensive testbed experiments. Results show that GEMINI achieves up to 53%, 31% and 76% reduction of small flow average completion times compared to TCP Cubic, DCTCP and BBR; and up to 58% reduction of large flow average completion times compared to TCP Vegas. Gaoxiong Zeng, Wei Bai 0001, Kai Chen 0005, Dongsu Han, Yibo Zhu 0001 |
ICNP | 5 |
| 2019 | Cybercriminal Minds: An investigative study of cryptocurrency abuses in the Dark Web
Changhoon Yoon, Heedo Kang, Yeonkeun Kim, Yongdae Kim, Dongsu Han, Sooel Son, Seungwon Shin 0001 |
NDSS | 6 |
| 2018 | APPx: an automated app acceleration framework for low latency mobile appabstractMinimizing response time of mobile applications is critical for user experience. Existing work predominantly focuses on reducing mobile Web latency, whereas users spend more time on native mobile apps than mobile Web. Similar to Web, mobile apps contain a chain of dependencies between successive requests. However, unlike Web acceleration where object dependencies can easily be identified by parsing Web documents, App acceleration is much more difficult because the dependency is encoded in the app binary. Byungkwon Choi, Daeyang Cho, Seong-Min Kim, Dongsu Han |
CoNEXT | 5 |
| 2018 | Neural Adaptive Content-aware Internet Video Delivery
Hyunho Yeo, Youngmok Jung, Jaehong Kim 0002, Jinwoo Shin, Dongsu Han |
OSDI | 5 |
| 2018 | SGX-Tor: A Secure and Practical Tor Anonymity Network With SGX Enclaves
Seong-Min Kim, Juhyeng Han, Jaehyeong Ha, Taesoo Kim, Dongsu Han |
IEEE/ACM Trans. Netw. | 5 |
| 2017 | SGX-Box: Enabling Visibility on Encrypted Traffic using a Secure Middlebox ModuleabstractA network middlebox benefits both users and network operators by offering a wide range of security-related in-network functions, such as web firewalls and intrusion detection systems (IDS). However, the wide usage of encryption protocol restricts functionalities of network middleboxes. This forces network operators and users to make a choice between end-to-end privacy and security. This paper presents SGX-Box, a secure middlebox system that enables visibility on encrypted traffic by leveraging Intel SGX technology. The entire process of SGX-Box ensures that the sensitive information, such as decrypted payloads and session keys, is securely protected within the SGX enclave. SGX-Box provides easy-to-use abstraction and a high-level programming language, called SB lang for handling encrypted traffic in middleboxes. It greatly enhances programmability by hiding details of the cryptographic operations and the implementation details in SGX enclave processing. We implement a proof-of-concept IDS using SB lang. Our preliminary evaluation shows that SGX-Box incurs acceptable performance overhead while it dramatically reduces middlebox developer's effort. Juhyeng Han, Seong-Min Kim, Jaehyeong Ha, Dongsu Han |
APNet | 4 |
| 2017 | Combining ECN and RTT for Datacenter TransportabstractDatacenter transports should provide low average and tail flow completion times (FCT) to achieve desired application performance. While most prior datacenter transports take either ECN or RTT as congestion signal, this paper makes a case that both signals are indispensable: ECN, as a per-hop signal, is more effective to prevent packet loss; while RTT, as an end-to-end signal, controls end-to-end queueing delay better. As persistent low flow completion times imply low queueing delay and near zero packet loss, we introduce EAR, a new datacenter transport that hears and reacts to both ECN and RTT. Our preliminary results show that: 1) compared to delay-based DCTCP, EAR achieves up to 91% lower packet losses and 93% fewer timeouts; 2) compared to ECN-based DCTCP, EAR reduces RTT by up to 32% for cross-rack traffic in a 4-level fattree. As a result, EAR delivers persistent low average and tail completion times under various scenarios in large scale simulations. Gaoxiong Zeng, Wei Bai 0001, Kai Chen 0005, Dongsu Han, Yibo Zhu 0001 |
APNet | 5 |
| 2017 | Network Stack as a Service in the CloudabstractThe tenant network stack is implemented inside the virtual machines in today's public cloud. This legacy architecture presents a barrier to protocol stack innovation due to the tight coupling between the network stack and the guest OS. In particular, it causes many deployment troubles to tenants and management and efficiency problems to the cloud provider. To address these issues, we articulate a vision of providing the network stack as a service. The central idea is to decouple the network stack from the guest OS, and offer it as an independent entity implemented by the cloud provider. This re-architecting allows tenants to readily deploy any stack independent of its kernel, and the provider to offer meaningful SLAs to tenants by gaining control over the network stack. We sketch an initial design called NetKernel to accomplish this vision. Our preliminary testbed evaluation with a prototype shows the feasibility and benefits of our idea. Zhixiong Niu, Hong Xu 0001, Dongsu Han, Peng Cheng 0005, Yongqiang Xiong, Guo Chen 0001, Keith Winstein |
HotNets | 3 |
| 2017 | How will Deep Learning Change Internet Video Delivery?abstractresearch-article Share on How will Deep Learning Change Internet Video Delivery? Authors: Hyunho Yeo KAIST KAISTView Profile , Sunghyun Do KAIST KAISTView Profile , Dongsu Han KAIST KAISTView Profile Authors Info & Claims HotNets-XVI: Proceedings of the 16th ACM Workshop on Hot Topics in NetworksNovember 2017 Pages 57–64https://doi.org/10.1145/3152434.3152440Published:30 November 2017Publication History 18citation950DownloadsMetricsTotal Citations18Total Downloads950Last 12 Months83Last 6 weeks12 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Hyunho Yeo, Sunghyun Do, Dongsu Han |
HotNets | 3 |
| 2017 | Rate-aware flow scheduling for commodity data center networksabstractFlow completion times (FCTs) are critical for many cloud applications. To minimize the average FCT, recent transport designs, such as pFabric, PASE, and PIAS, approximate the Shortest Remaining Time First (SRTF) scheduling. A common, implicit assumption of these solutions is that the remaining time is only determined by the remaining flow size. However, this assumption does not hold in many real-world scenarios where applications generate data at diverse rates that are smaller than the network capacity. In this paper, we look into this issue from system perspective and find that the operating system (OS) kernel can be exploited to better estimate the remaining time of a flow. In particular, we use the rate of copying data from user space to kernel space to measure the data generation rate. We design RAX, a rate aware flow scheduling method, that calculates the remaining time of a flow more accurately, based on not only the flow size but also the data generation rate. We have implemented a RAX prototype in Linux kernel and evaluated it through testbed experiments and ns-2 simulations. Our testbed results show that RAX reduces FCT by up to 14.9%/41.8% and 7.8%/22.9% over DCTCP and PIAS for all/medium flows respectively. Ziyang Li 0003, Wei Bai 0001, Kai Chen 0005, Dongsu Han, Yiming Zhang 0003, Dongsheng Li 0001, Hong-Fang Yu |
INFOCOM | 4 |
| 2017 | SGX-Shield: Enabling Address Space Layout Randomization for SGX Programs
Jaebaek Seo, Byoungyoung Lee, Seong-Min Kim, Ming-Wei Shih, Insik Shin, Dongsu Han, Taesoo Kim |
NDSS | 6 |
| 2017 | mOS: A Reusable Networking Stack for Flow Monitoring Middleboxes
Muhammad Asim Jamshed, YoungGyoun Moon, Donghwi Kim, Dongsu Han, KyoungSoo Park |
NSDI | 4 |
| 2017 | Enhancing Security and Privacy of Tor's Ecosystem by Using Trusted Execution Environments
Seong-Min Kim, Juhyeng Han, Jaehyeong Ha, Taesoo Kim, Dongsu Han |
NSDI | 5 |
| 2017 | Credit-Scheduled Delay-Bounded Congestion Control for DatacentersabstractSmall RTTs (~tens of microseconds), bursty flow arrivals, and a large number of concurrent flows (thousands) in datacenters bring fundamental challenges to congestion control as they either force a flow to send at most one packet per RTT or induce a large queue build-up. The widespread use of shallow buffered switches also makes the problem more challenging with hosts generating many flows in bursts. In addition, as link speeds increase, algorithms that gradually probe for bandwidth take a long time to reach the fair-share. An ideal datacenter congestion control must provide 1) zero data loss, 2) fast convergence, 3) low buffer occupancy, and 4) high utilization. However, these requirements present conflicting goals. Inho Cho, Keon Jang, Dongsu Han |
SIGCOMM | 3 |
| 2017 | PIAS: Practical Information-Agnostic Flow Scheduling for Commodity Data CentersabstractMany existing data center network (DCN) flow scheduling schemes, that minimize flow completion times (FCT) assume prior knowledge of flows and custom switch functions, making them superior in performance but hard to implement in practice. By contrast, we seek to minimize FCT with no prior knowledge and existing commodity switch hardware. To this end, we present PIAS, a DCN flow scheduling mechanism that aims to minimize FCT by mimicking shortest job first (SJF) on the premise that flow size is not knowna priori. At its heart, PIAS leverages multiple priority queues available in existing commodity switches to implement a multiple level feedback queue, in which a PIAS flow is gradually demoted from higher-priority queues to lower-priority queues based on the number of bytes it has sent. As a result, short flows are likely to be finished in the first few high-priority queues and thus be prioritized over long flows in general, which enables PIAS to emulate SJF without knowing flow sizes beforehand. We have implemented a PIAS prototype and evaluated PIAS through both testbed experiments and ns-2 simulations. We show that PIAS is readily deployable with commodity switches and backward compatible with legacy TCP/IP stacks. Our evaluation results show that PIAS significantly outperforms existing information-agnostic schemes, for example, it reduces FCT by up to 50% compared to DCTCP[11]and L2DCT[32]; and it only has a 1.1% performance gap to an ideal information-aware scheme, pFabric[13], for short flows under a production DCN workload. Wei Bai 0001, Li Chen 0008, Kai Chen 0005, Dongsu Han, Chen Tian 0001, Hao Wang 0022 |
IEEE/ACM Trans. Netw. | 4 |
| 2017 | DX: Latency-Based Congestion Control for DatacentersabstractSince the advent of datacenter networking, achieving low latency within the network has been a primary goal. Many congestion control schemes have been proposed in recent years to meet the datacenters' unique performance requirement. The nature of congestion feedback largely governs the behavior of congestion control. In datacenter networks, where round trip times are in hundreds of microseconds, accurate feedback is crucial to achieve both high utilization and low queueing delay. Proposals for datacenter congestion control predominantly leverage explicit congestion notification (ECN) or even explicit in-network feedback to minimize the queuing delay. In this paper, we explore latency-based feedback as an alternative and show its advantages over ECN. Against the common belief that such implicit feedback is noisy and inaccurate, we demonstrate that latency-based implicit feedback is accurate enough to signal a single packet's queuing delay in 10 Gb/s networks. Such high accuracy enables us to design a new congestion control algorithm, DX, that performs fine-grained control to adjust the congestion window just enough to achieve very low queuing delay while attaining full utilization. Our extensive evaluation shows that: 1) the latency measurement accurately reflects the one-way queuing delay in single packet level; 2) the latency feedback can be used to perform practical and fine-grained congestion control in high-speed datacenter networks; and 3) DX outperforms DCTCP with 5.33 times smaller median queueing delay at 1 Gb/s and 1.57 times at 10 Gb/s. Chunjong Park, Keon Jang, Sue B. Moon, Dongsu Han |
IEEE/ACM Trans. Netw. | 5 |
| 2017 | Expeditus: Congestion-Aware Load Balancing in Clos Data Center NetworksabstractData center networks often use multi-rooted Clos topologies to provide a large number of equal cost paths between two hosts. Thus, load balancing traffic among the paths is important for high performance and low latency. However, it is well known that ECMP-the de facto load balancing scheme-performs poorly in data center networks. The main culprit of ECMP's problems is its congestion agnostic nature, which fundamentally limits its ability to deal with network dynamics. We propose Expeditus, a novel distributed congestion-aware load balancing protocol for general 3-tier Clos networks. The complex 3-tier Clos topologies present significant scalability challenges that make a simple per-path feedback approach infeasible. Expeditus addresses the challenges by using simple local information collection, where a switch only monitors its egress and ingress link loads. It further employs a novel two-stage path selection mechanism to aggregate relevant information across switches and make path selection decisions. Testbed evaluation on Emulab and large-scale ns-3 simulations demonstrate that, Expeditus outperforms ECMP by up to 45% in tail flow completion times (FCT) for mice flows, and by up to 38% in mean FCT for elephant flows in 3-tier Clos networks. Peng Wang 0037, Hong Xu 0001, Zhixiong Niu, Dongsu Han, Yongqiang Xiong |
IEEE/ACM Trans. Netw. | 4 |
| 2017 | Guaranteeing Deadlines for Inter-Data Center TransfersabstractInter-data center wide area networks (inter-DC WANs) carry a significant amount of data transfers that require to be completed within certain time periods, or deadlines. However, very little work has been done to guarantee such deadlines. The crux is that the current inter-DC WAN lacks an interface for users to specify their transfer deadlines and a mechanism for provider to ensure the completion while maintaining high WAN utilization. In this paper, we address the problem by introducing a deadline-based network abstraction (DNA) for inter-DC WANs. DNA allows users to explicitly specify the amount of data to be delivered and the deadline by which it has to be completed. The malleability of DNA provides flexibility in resource allocation. Based on this, we develop a system calledAmoebathat implements DNA. Our simulations and test bed experiments show thatAmoeba, by harnessing DNA’s malleability, accommodates 15% more user requests with deadlines, while achieving 60% higher WAN utilization than prior solutions. Hong Zhang 0025, Kai Chen 0005, Wei Bai 0001, Dongsu Han, Chen Tian 0001, Hao Wang 0022, Haibing Guan, Ming Zhang 0005 |
IEEE/ACM Trans. Netw. | 4 |
| 2016 | Expeditus: Congestion-aware Load Balancing in Clos Data Center NetworksabstractData center networks often use multi-rooted Clos topologies to provide a large number of equal cost paths between two hosts. Thus, load balancing traffic among the paths is important for high performance and low latency. However, it is well known that ECMP---the de facto load balancing scheme---performs poorly in data center networks. The main culprit of ECMP's problems is its congestion agnostic nature, which fundamentally limits its ability to deal with network dynamics. Peng Wang 0037, Hong Xu 0001, Zhixiong Niu, Dongsu Han, Yongqiang Xiong |
SoCC | 4 |
| 2016 | Enabling Automatic Protocol Behavior Analysis for Android ApplicationsabstractAndroid application is an important class on today's Internet. While understanding app-specific behavior is important for network operation and management, it is often difficult because it requires an in-depth application-layer protocol analysis due to the common use of HTTP(S) and standard data representations (e.g., JSON). This paper presents Extractocol, the first system to offer an automatic and comprehensive analysis of application protocol behaviors. Extractocol only uses Android application binary as input and accurately reconstructs HTTP transactions (request-response pairs) and identifies their message format and relationships using binary analysis. Our evaluation and in-depth case studies on commercial and open-source apps demonstrate that Extractocol provides high coverage and accurately characterizes network-related application behaviors. Hyunwoo Choi, Hun Namkung, Woohyun Choi, Byungkwon Choi, Hyunwook Hong, Yongdae Kim, Jonghyup Lee, Dongsu Han |
CoNEXT | 9 |
| 2016 | OpenSGX: An Open Platform for SGX Research
Prerit Jain, Soham Jayesh Desai, Ming-Wei Shih, Taesoo Kim, Seong-Min Kim, Jae-Hyuk Lee, Changho Choi, Youjung Shin, Brent ByungHoon Kang, Dongsu Han |
NDSS | 10 |
| 2016 | DFC: Accelerating String Pattern Matching for Network Applications
Byungkwon Choi, Jongwook Chae, Muhammad Asim Jamshed, KyoungSoo Park, Dongsu Han |
NSDI | 5 |
| 2016 | Application-specific Acceleration Framework for Mobile ApplicationsabstractMinimizing response times for mobile applications is critical for quality user experience that often impacts the revenue of mobile services. Generalized approaches to accelerated mobile applications (e.g., TCP acceleration, SPDY, compression) are less effective because they do not take account for application specific behaviors. In contrast, application specific approaches build application-specific proxies by leveraging the app-specific protocol behaviors to enable dynamic caching and/or prefetching. However, this is non-trivial because it requires manual analysis of application level protocols and their interactions. Therefore, only a small number of apps enjoyed the benefit. Byungkwon Choi, Dongsu Han |
SIGCOMM | 3 |
| 2015 | Scaling the Performance of Network Intrusion Detection with Many-core ProcessorsabstractIn this work, we present a highly scalable network intrusion detection system on many-core processors. To maximize the NIDS performance, we take advantage of the underlying hardware and adhere to four design principles: shared-nothing architecture, computation offloading, lightweight data structure, and flow offloading. Through the experimental results, we find that our design choices can significantly improve the NIDS performance (79 Gbps with 1514B synthetic packets). We believe that our design decisions can be easily extended to other many-core processors and programmable NICs. Jaehyun Nam, Muhammad Asim Jamshed, Byungkwon Choi, Dongsu Han, KyoungSoo Park |
ANCS | 4 |
| 2015 | Practical message-passing framework for large-scale combinatorial optimizationabstractGraphical Model (GM) has provided a popular framework for big data analytics because it often lends itself to distributed and parallel processing by utilizing graph-based ‘local’ structures. It models correlated random variables where in particular, the max-product Belief Propagation (BP) is the most popular heuristic to compute the most-likely assignment in GMs. In the past years, it has been proven that BP can solve a few classes of combinatorial optimization problems under certain conditions. Motivated by this, we explore the prospect of using BP to solve generic combinatorial optimization problems. The challenge is that, in practice, BP may converge very slowly and even if it does converge, the BP decision often violates the constraints of the original problem. This paper proposes a generic framework that enables us to apply BP-based algorithms to compute an approximate feasible solution for an arbitrary combinatorial optimization task. The main novel ingredients include (a) careful initialization of BP messages, (b) hybrid damping on BP updates, and (c) post-processing using BP beliefs. Utilizing the framework, we develop parallel algorithms for several large-scale combinatorial optimization problems including maximum weight matching, vertex cover and independent set. We demonstrate that our framework delivers high approximation ratio, speeds up the process by parallelization, and allows large-scale processing involving billions of variables. Inho Cho, Soya Park, Dongsu Han, Jinwoo Shin |
IEEE BigData | 4 |
| 2015 | Breaking and Fixing VoLTE: Exploiting Hidden Data Channels and Mis-implementationsabstractLong Term Evolution (LTE) is becoming the dominant cellular networking technology, shifting the cellular network away from its circuit-switched legacy towards a packet-switched network that resembles the Internet. To support voice calls over the LTE network, operators have introduced Voice-over-LTE (VoLTE), which dramatically changes how voice calls are handled, both from user equipment and infrastructure perspectives. We find that this dramatic shift opens up a number of new attack surfaces that have not been previously explored. To call attention to this matter, this paper presents a systematic security analysis. Dongkwan Kim 0001, Minhee Kwon, HyungSeok Han, Yeongjin Jang, Dongsu Han, Taesoo Kim, Yongdae Kim |
CCS | 6 |
| 2015 | Guaranteeing deadlines for inter-datacenter transfersabstractInter-datacenter wide area networks (inter-DC WAN) carry a significant amount of data transfers that require to be completed within certain time periods, or deadlines. However, very little work has been done to guarantee such deadlines. The crux is that the current inter-DC WAN lacks an interface for users to specify their transfer deadlines and a mechanism for provider to ensure the completion while maintaining high WAN utilization. Hong Zhang 0025, Kai Chen 0005, Wei Bai 0001, Dongsu Han, Chen Tian 0001, Hao Wang 0022, Haibing Guan, Ming Zhang 0005 |
EuroSys | 4 |
| 2015 | A First Step Towards Leveraging Commodity Trusted Execution Environments for Network ApplicationsabstractNetwork applications and protocols are increasingly adopting security and privacy features, as they are becoming one of the primary requirements. The wide-spread use of transport layer security (TLS) and the growing popularity of anonymity networks, such as Tor, exemplify this trend. Motivated by the recent movement towards commoditization of trusted execution environments (TEEs), this paper explores alternative design choices that application and protocol designers should consider. In particular, we explore the possibility of using Intel SGX to provide security and privacy in a wide range of network applications. We show that leveraging hardware protection of TEEs opens up new possibilities, often at the benefit of a much simplified application/protocol design. We demonstrate its practical implications by exploring the design space for SGX-enabled software-defined inter-domain routing, peer-to-peer anonymity networks (Tor), and middleboxes. Finally, we quantify the potential overheads of the SGX-enabled design by implementing it on top of OpenSGX, an open source SGX emulator. Seong-Min Kim, Youjung Shin, Jaehyung Ha, Taesoo Kim, Dongsu Han |
HotNets | 5 |
| 2015 | Information-Agnostic Flow Scheduling for Commodity Data Centers
Wei Bai 0001, Kai Chen 0005, Hao Wang 0022, Li Chen 0008, Dongsu Han, Chen Tian 0001 |
NSDI | 5 |
| 2015 | Haetae: Scaling the Performance of Network Intrusion Detection with Many-Core Processors
Jaehyun Nam, Muhammad Asim Jamshed, Byungkwon Choi, Dongsu Han, KyoungSoo Park |
RAID | 4 |
| 2015 | Extractocol: Autoatic Extraction of Application-level Protocol Behaviors for Android ApplicationsabstractNo abstract available. Hyunwoo Choi, Hyunwook Hong, Yongdae Kim, Jonghyup Lee, Dongsu Han |
SIGCOMM | 6 |
| 2015 | A Case for a Stateful Middlebox Networking StackabstractNo abstract available. Muhammad Asim Jamshed, Donghwi Kim, YoungGyoun Moon, Dongsu Han, KyoungSoo Park |
SIGCOMM | 4 |
| 2015 | Practical, Real-time Centralized Control for CDN-based Live Video DeliveryabstractLive video delivery is expected to reach a peak of 50 Tbps this year. This surging popularity is fundamentally changing the Internet video delivery landscape. CDNs must meet users' demands for fast join times, high bitrates, and low buffering ratios, while minimizing their own cost of delivery and responding to issues in real-time. Wide-area latency, loss, and failures, as well as varied workloads ("mega-events" to long-tail), make meeting these demands challenging. Matthew K. Mukerjee, David Naylor, Junchen Jiang, Dongsu Han, Srinivasan Seshan, Hui Zhang 0001 |
SIGCOMM | 4 |
| 2015 | Accurate Latency-based Congestion Feedback for Datacenters
Chunjong Park, Keon Jang, Sue B. Moon, Dongsu Han |
USENIX ATC | 5 |
| 2014 | PIAS: Practical Information-Agnostic Flow Scheduling for Data Center NetworksabstractMany existing data center network (DCN) flow scheduling schemes minimize flow completion times (FCT) based on prior knowledge of flows and custom switch designs, making them hard to use in practice. This paper introduces, Pias, a practical flow scheduling approach that minimizes FCT with no prior knowledge using commodity switches. At its heart, Pias leverages multiple priority queues available in commodity switches to implement a Multiple Level Feedback Queue (MLFQ), in which a PIAS flow gradually demotes from higher-priority queues to lower-priority queues based on the bytes it has sent. In this way, short flows are prioritized over long flows, which enables Pias to emulate Shortest Job First (SJF) scheduling without knowing the flow sizes beforehand. Our preliminary evaluation shows that Pias significantly outperforms all existing information-agnostic solutions. It improves average FCT for short flows by up to 50% and 40% over DCTCP [3] and L2DCT [16]. Compared to an ideal information-aware DCN transport, p-Fabric [5], it only shows 4.9% performance degradation for short flows in a production datacenter workload. Wei Bai 0001, Li Chen 0008, Kai Chen 0005, Dongsu Han, Chen Tian 0001, Weicheng Sun |
HotNets | 4 |
| 2014 | mTCP: a Highly Scalable User-level TCP Stack for Multicore Systems
Eunyoung Jeong, Shinae Woo, Muhammad Asim Jamshed, Haewon Jeong, Sunghwan Ihm, Dongsu Han, KyoungSoo Park |
NSDI | 6 |
| 2014 | MICA: A Holistic Approach to Fast In-Memory Key-Value Storage
Hyeontaek Lim, Dongsu Han, David G. Andersen, Michael Kaminsky |
NSDI | 2 |
| 2014 | Enabling near real-time central control for live video delivery in CDNsabstractUser-created live video streaming is marking a fundamental shift in the workload of live video delivery. However, live-video-specific challenges and the viral nature of user-created content makes it difficult for current CDNs to deliver 1) high-quality, 2) highly-scalable, and 3) highly-responsive service. We present the design and implementation of VDN, a new control plane for CDNs designed to optimize the delivery of live streams within the CDN. VDN satisfies these requirements by using two approaches: 1) optimizing directly for video quality (not just throughput) and 2) combining centralized control with local control, allowing VDN to adapt to traffic dynamics and network failures at fine timescales. Matthew K. Mukerjee, JungAh Hong, Junchen Jiang, David Naylor, Dongsu Han, Srinivasan Seshan, Hui Zhang 0001 |
SIGCOMM | 5 |
| 2013 | Understanding tradeoffs in incremental deployment of new network architecturesabstractDespite the plethora of incremental deployment mechanisms proposed, rapid adoption of new network-layer protocols and architectures remains difficult as reflected by the widespread lack of IPv6 traffic on the Internet. We show that all deployment mechanisms must address four key questions: How to select an egress from the source network, how to select an ingress into the destination network, how to reach that egress, and how to reach that ingress. By creating a design space that maps all existing mechanisms by how they answer these questions, we identify the lack of existing mechanisms in part of this design space and propose two novel approaches: the "4ID" and the "Smart 4ID". The 4ID mechanism utilizes new data plane technology to flexibly decide when to encapsulate packets at forwarding time. The Smart 4ID mechanism additionally adopts an SDN-style control plane to intelligently pick ingress/egress pairs based on a wider view of the local network. We implement these mechanisms along with two widely used IPv6 deployment mechanisms and conduct wide-area deployment experiments over PlanetLab. We conclude that Smart 4ID provide better overall performance and failure semantics, and that innovations in the data plane and control plane enable straightforward incremental deployment. Matthew K. Mukerjee, Dongsu Han, Srinivasan Seshan, Peter Steenkiste |
CoNEXT | 2 |
| 2013 | CAMEO: a middleware for mobile advertisement deliveryabstractAdvertisements are the de-facto currency of the Internet with many popular applications (e.g. Angry Birds) and online services (e.g., YouTube) relying on advertisement generated revenue. However, the current economic models and mechanisms for mobile advertising are fundamentally not sustainable and far from ideal. In particular, as we show, applications which use mobile advertising are capable of using significant amounts of a mobile users' critical resources without being controlled or held accountable. This paper seeks to redress this situation by enabling advertisement supported applications to become significantly more ``user-friendly''. To this end, we present the design and implementation of CAMEO, a new framework for mobile advertising that 1) employs intelligent and proactive retrieval of advertisements, using context prediction, to significantly reduce the bandwidth and energy overheads of advertising, and 2) provides a negotiation protocol and framework that empowers applications to subsidize their data traffic costs by ``bartering'' their advertisement rights for access bandwidth from mobile ISPs. Our evaluation, that uses real mobile advertising data collected from around the globe, demonstrates that CAMEO effectively reduces the resource consumption caused by mobile advertising. Azeem J. Khan, Kasthuri Jayarajah, Dongsu Han, Archan Misra, Rajesh Krishna Balan, Srinivasan Seshan |
MobiSys | 3 |
| 2013 | FCP: a flexible transport framework for accommodating diversityabstractTransport protocols must accommodate diverse application and network requirements. As a result, TCP has evolved over time with new congestion control algorithms such as support for generalized AIMD, background flows, and multipath. On the other hand, explicit congestion control algorithms have been shown to be more efficient. However, they are inherently more rigid because they rely on in-network components. Therefore, it is not clear whether they can be made flexible enough to support diverse application requirements. This paper presents a flexible framework for network resource allocation, called FCP, that accommodates diversity by exposing a simple abstraction for resource allocation. FCP incorporates novel primitives for end-point flexibility (aggregation and preloading) into a single framework and makes economics-based congestion control practical by explicitly handling load variations and by decoupling it from actual billing. We show that FCP allows evolution by accommodating diversity and ensuring coexistence, while being as efficient as existing explicit congestion control algorithms. Dongsu Han, Robert Grandl, Aditya Akella, Srinivasan Seshan |
SIGCOMM | 1 |
| 2012 | RPT: Re-architecting Loss Protection for Content-Aware Networks
Dongsu Han, Ashok Anand, Aditya Akella, Srinivasan Seshan |
NSDI | 1 |
| 2012 | XIA: Efficient Support for Evolvable Internetworking
Dongsu Han, Ashok Anand, Fahad R. Dogar, Hyeontaek Lim, Michel Machado, Arvind Mukundan, Wenfei Wu, Aditya Akella, David G. Andersen, John W. Byers, Srinivasan Seshan, Peter Steenkiste |
NSDI | 1 |
| 2012 | Supporting network evolution and incremental deployment with XIAabstracteXpressive Internet Architecture (XIA) [1] is an architecture that natively supports multiple communication types and allows networks to evolve their abstractions and functionality to accommodate new styles of communication over time. XIA embeds an elegant mechanism for handling unforeseen communication types for legacy routers. Robert Grandl, Dongsu Han, Suk-Bok Lee, Hyeontaek Lim, Michel Machado, Matthew K. Mukerjee, David Naylor |
SIGCOMM | 2 |
| 2011 | XIA: an architecture for an evolvable and trustworthy internetabstractMotivated by limitations in today's host-based IP network architecture, recent studies have proposed clean-slate network architectures centered around alternative first-class principals, such as content, services, or users. However, much like the host-centric IP design, elevating one principal type above others hinders communication between other principals and inhibits the network's capability to evolve. Our work presents the eXpressive Internet Architecture (XIA), an architecture with native support for multiple principals and the ability to evolve its functionality to accommodate new, as yet unforeseen, principals over time. XIA also provides intrinsic security: communicating entities validate that their underlying intent was satisfied correctly without relying on external databases or configuration. Ashok Anand, Fahad R. Dogar, Dongsu Han, Hyeontaek Lim, Michel Machado, Wenfei Wu, Aditya Akella, David G. Andersen, John W. Byers, Srinivasan Seshan, Peter Steenkiste |
HotNets | 3 |
| 2011 | The hare and the tortoise: taming wireless losses by exploiting wired reliabilityabstractMultiple communication channels are common in today's consumer and enterprise networks. For example, a high bandwidth but unreliable wireless network might co-exist with a reliable wired link (EWLANs and neighborhood networks). In this paper, we present a system that uses this reliable wired communication channel to boost the bandwidth of the lossy wireless link. Specifically, we propose a new, efficient partial packet recovery (PPR) technique and adaptive feedback mechanism specially designed to correct partial packets on an 802.11 wireless network using a wired backhaul. Our initial experiments demonstrate up to a 3x improvement over standalone 802.11 and upto a 30% improvement over existing PPR techniques. Anirudh Badam, Michael Kaminsky, Dongsu Han, Konstantina Papagiannaki, David G. Andersen, Srinivasan Seshan |
MobiHoc | 3 |
| 2010 | ATLAS: A scalable and high-performance scheduling algorithm for multiple memory controllersabstractModern chip multiprocessor (CMP) systems employ multiple memory controllers to control access to main memory. The scheduling algorithm employed by these memory controllers has a significant effect on system throughput, so choosing an efficient scheduling algorithm is important. The scheduling algorithm also needs to be scalable - as the number of cores increases, the number of memory controllers shared by the cores should also increase to provide sufficient bandwidth to feed the cores. Unfortunately, previous memory scheduling algorithms are inefficient with respect to system throughput and/or are designed for a single memory controller and do not scale well to multiple memory controllers, requiring significant finegrained coordination among controllers. This paper proposes ATLAS (Adaptive per-Thread Least-Attained-Service memory scheduling), a fundamentally new memory scheduling technique that improves system throughput without requiring significant coordination among memory controllers. The key idea is to periodically order threads based on the service they have attained from the memory controllers so far, and prioritize those threads that have attained the least service over others in each period. The idea of favoring threads with least-attained-service is borrowed from the queueing theory literature, where, in the context of a single-server queue it is known that least-attained-service optimally schedules jobs, assuming a Pareto (or any decreasing hazard rate) workload distribution. After verifying that our workloads have this characteristic, we show that our implementation of least-attained-service thread prioritization reduces the time the cores spend stalling and significantly improves system throughput. Furthermore, since the periods over which we accumulate the attained service are long, the controllers coordinate very infrequently to form the ordering of threads, thereby making ATLAS scalable to many controllers. We evaluate ATLAS on a wide variety of multiprogrammed SPEC 2006 workloads and systems with 4-32 cores and 1-16 memory controllers, and compare its performance to five previously proposed scheduling algorithms. Averaged over 32 workloads on a 24-core system with 4 controllers, ATLAS improves instruction throughput by 10.8%, and system throughput by 8.4%, compared to PAR-BS, the best previous CMP memory scheduling algorithm. ATLAS's performance benefit increases as the number of cores increases. Yoongu Kim, Dongsu Han, Onur Mutlu, Mor Harchol-Balter |
HPCA | 2 |
| 2009 | Access Point Localization Using Local Signal Strength Gradient
Dongsu Han, David G. Andersen, Michael Kaminsky, Konstantina Papagiannaki, Srinivasan Seshan |
PAM | 1 |
| 2008 | Mark-and-sweep: getting the "inside" scoop on neighborhood networksabstractResidential Internet connectivity is growing at a phenomenal rate. A number of recent studies have attempted to characterize this connectivity - measuring coverage and performance of last-mile broadband links - from a various vantage points on the Internet, via wireless APs, and even with user cooperation. These studies, however, sacrifice accuracy or require substantial human time. In this work, we present a novel two-pass method to characterize neighborhood networks. We demonstrate that the two pass method dramatically reduces the time spent in active measurement while retaining accuracy. A case study on two neighborhoods in Pittsburgh provide new and accurate insights into broadband connectivity, including throughput, broadband coverage (DSL vs. cable vs. fiber), NAT configurations, DHCP, DNS usage. The results further characterize 802.11 connectivity in the neighborhood. Dongsu Han, Aditiya Agarwala, David G. Andersen, Michael Kaminsky, Konstantina Papagiannaki, Srinivasan Seshan |
Internet Measurement Conference | 1 |