KyoungSoo Park

dblp:67/4936 · DBLP profile ↗
← Back
43ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 25 · 2 first-author · 7 since 2021Systems, architecture and hardware · 8 · 2 first-author · 2 since 2021Security and privacy · 4Artificial intelligence and machine learning · 3Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 PacketExpress: Fully Exploiting Large MTUs for Internet Traffic in Private Networks
abstract
Network bandwidth continues to scale rapidly, yet Internet data transmission performance remains constrained by the legacy 1500 B MTU. This small MTU translates high bandwidth into high packet rates that strain CPU processing at middleboxes and end hosts. While increasing the MTU could substantially improve performance, coordinating upgrades across arbitrary Internet paths is impractical.
Junghan Yoon, Youngmin Choi, Juyoung Park, Daehyeok Kim, Changhoon Kim, KyoungSoo Park
SIGCOMM6
2025 Towards Incremental MTU Upgrade for the Internet
abstract
This paper proposes a systematic approach to incrementally enabling large MTUs in the Internet. We demonstrate that increasing the MTU size significantly enhances the performance of both middleboxes and end hosts. To bridge MTU mismatches at network borders, we introduce PacketExpress gateway (PXGW), an MTU-translating gateway that dynamically adjusts packet sizes for cross-traffic. PXGW merges and splits TCP payloads on the fly and tunnels UDP packets, ensuring seamless adaptation. Also, we propose F-PMTUD, a new path MTU discovery algorithm that determines the path MTU within a single round-trip without relying on ICMP. Our preliminary evaluation shows that the PXGW prototype achieves 1.45 Tbps of packet forwarding throughput using only 8 CPU cores. After dynamic conversion, 94% of transmitted TCP packets are 9000 B jumbo frames, indicating that most flows were effectively converted into large segments, thereby demonstrating the system's efficiency and scalability. We also find that large-MTU packets, made available via PXGW, enhance end-host performance by up to 2.5X.
Junghan Yoon, Youngmin Choi, Juyoung Park, Daehyeok Kim, Changhoon Kim, KyoungSoo Park
HotNets6
2025 A4: Microarchitecture-Aware LLC Management for Datacenter Servers with Emerging I/O Devices
abstract
In modern server CPUs, the Last-Level Cache (LLC) serves not only as a victim cache for higher-level private caches but also as a buffer for low-latency DMA transfers between CPU cores and I/O devices through Direct Cache Access (DCA).However, prior work has shown that high-bandwidth network-I/O devices can rapidly flood the LLC with packets, often causing significant contention with co-running workloads.One step further, this work explores hidden microarchitectural properties of the Intel Xeon CPUs, uncovering two previously unrecognized LLC contentions triggered by emerging high-bandwidth I/O devices.Specifically, (C1) DMAwritten cache lines in LLC ways designated for DCA (referred to as DCA ways) are migrated to certain LLC ways (denoted as inclusive ways) when accessed by CPU cores, unexpectedly contending with non-I/O cache lines within the inclusive ways.In addition, (C2) high-bandwidth storage-I/O devices, which are increasingly common in datacenter servers, benefit little from DCA while contending with (latency-sensitive) network-I/O devices within DCA ways.To this end, we present A4, a runtime LLC management framework designed to alleviate both (C1) and (C2) among diverse co-running workloads, using a hidden knob and other hardware features implemented in those CPUs.Additionally, we demonstrate that A4 can also alleviate other previously known network-I/Odriven LLC contentions.Overall, it improves the performance of latency-sensitive, high-priority workloads by 51% without notably compromising that of low-priority workloads.
Haneul Park, Jiaqi Lou, Sangjin Lee 0003, KyoungSoo Park, Yongseok Son, Ipoom Jeong, Nam Sung Kim
ISCA5
2024 Logan: Loss-tolerant Live Video Analytics System
abstract
Cloud-based live video analytics with tight latency bound is gaining importance to support emerging applications such as UAVs and augmented reality. However, existing systems often struggle to meet stringent latency constraints under fluctuating network conditions with packet losses and late-arriving packets. We propose a loss-tolerant live video analytics system called Logan, which effectively accepts packet losses while maintaining high accuracy by utilizing the inherent resilience in DNNs. We design i) Codec-aware Inpainting, which accurately recovers the frame error from packet losses ii) Fast-Forward Recovery that prevents the remaining un-recovered error from propagating over future frames indefinitely. Our results show a 3× improvement (33.2%→99.9%) in SLO satisfaction rate compared to the reliable transmission scheme with <1% accuracy drop under a 5% packet loss rate.
Kichang Yang, Minkyung Jeong, Juheon Yi, Jingyu Lee, KyoungSoo Park, Youngki Lee 0001
MobiCom5
2024 mmTLS: Scaling the Performance of Encrypted Network Traffic Inspection
Junghan Yoon, Seunghyun Do, Duckwoo Kim, Taejoong Chung, KyoungSoo Park
USENIX ATC5
2023 Is Large MTU Beneficial to Cellular Core Networks?
abstract
The Maximum Transmission Unit (MTU) refers to the largest packet size that can be transferred on a particular layer-3 network. As the dominance of Ethernet prevails, the "de-facto" standard MTU of 1500B has become universal in the wide-area networks. Unfortunately, the current MTU size overly limits the transmission performance especially when the underlying link speed rapidly increases while the CPU advancement stagnates.
Youngmin Choi, Junghan Yoon, YoungGyoun Moon, KyoungSoo Park
APNet4
2023 ARK: GPU-driven Code Execution for Distributed Deep Learning
Changho Hwang, KyoungSoo Park, Ran Shu 0001, Xinyuan Qu, Peng Cheng 0005, Yongqiang Xiong
NSDI2
2023 Rearchitecting the TCP Stack for I/O-Offloaded Content Delivery
Deondre Martin Ng, Junzhi Gong, Youngjin Kwon, Minlan Yu, KyoungSoo Park
NSDI6
2021 Elastic Resource Sharing for Distributed Deep Learning
Changho Hwang, Jinwoo Shin, KyoungSoo Park
NSDI5
2020 A Case for SmartNIC-accelerated Private Communication
abstract
Transport Layer Security (TLS) has become a key building block for private network communication in modern Internet. While recent advancement of CPU has substantially improved the data encryption performance, TLS key exchange still remains the bottleneck for short-lived transactions. Dedicated hardware crypto accelerators promise good performance, but they often require invasive modification of the application due to its inherent architecture of asynchronous processing.
Duckwoo Kim, SeungEon Lee 0001, KyoungSoo Park
APNet3
2020 AccelTCP: Accelerating Network Applications with Stateful TCP Offloading
YoungGyoun Moon, SeungEon Lee 0001, Muhammad Asim Jamshed, KyoungSoo Park
NSDI4
2019 DPX: Data-Plane eXtensions for SDN Security Service Instantiation
Taejune Park, Yeonkeun Kim, Vinod Yegneswaran, Phillip A. Porras, Zhaoyan Xu, KyoungSoo Park, Seungwon Shin 0001
DIMVA6
2019 Hyperscan: A Fast Multi-pattern Regex Matcher for Modern CPUs
Harry Chang, KyoungSoo Park, Geoff Langdale, Heqing Zhu
NSDI4
2017 Faster Greedy MAP Inference for Determinantal Point Processes
abstract
Determinantal point processes (DPPs) are popular probabilistic models that arise in many machine learning tasks, where distributions of diverse sets are characterized by determinants of their features. In this paper, we develop fast algorithms to find the most likely configuration (MAP) of large-scale DPPs, which is NP-hard in general. Due to the submodular nature of the MAP objective, greedy algorithms have been used with empirical success. Greedy implementations require computation of log-determinants, matrix inverses or solving linear systems at each iteration. We present faster implementations of the greedy algorithms by utilizing the orthogonal benefits of two log-determinant approximation schemes: (a) first-order expansions to the matrix log-determinant function and (b) high-order expansions to the scalar log function with stochastic trace estimators. In our experiments, our algorithms are orders of magnitude faster than their competitors, while sacrificing marginal accuracy.
Insu Han, Prabhanjan Kambadur, KyoungSoo Park, Jinwoo Shin
ICML3
2017 Confident Multiple Choice Learning
abstract
Ensemble methods are arguably the most trustworthy techniques for boosting the performance of machine learning models. Popular independent ensembles (IE) relying on naive averaging/voting scheme have been of typical choice for most applications involving deep neural networks, but they do not consider advanced collaboration among ensemble models. In this paper, we propose new ensemble methods specialized for deep neural networks, called confident multiple choice learning (CMCL): it is a variant of multiple choice learning (MCL) via addressing its overconfidence issue.In particular, the proposed major components of CMCL beyond the original MCL scheme are (i) new loss, i.e., confident oracle loss, (ii) new architecture, i.e., feature sharing and (iii) new training method, i.e., stochastic labeling. We demonstrate the effect of CMCL via experiments on the image classification on CIFAR and SVHN, and the foreground-background segmentation on the iCoseg. In particular, CMCL using 5 residual networks provides 14.05\% and 6.60\% relative reductions in the top-1 error rates from the corresponding IE scheme for the classification task on CIFAR and SVHN, respectively.
Kimin Lee, Changho Hwang, KyoungSoo Park, Jinwoo Shin
ICML3
2017 APUNet: Revitalizing GPU as Packet Processing Accelerator
Younghwan Go, Muhammad Asim Jamshed, YoungGyoun Moon, Changho Hwang, KyoungSoo Park
NSDI5
2017 mOS: A Reusable Networking Stack for Flow Monitoring Middleboxes
Muhammad Asim Jamshed, YoungGyoun Moon, Donghwi Kim, Dongsu Han, KyoungSoo Park
NSDI5
2017 Cedos: A Network Architecture and Programming Abstraction for Delay-Tolerant Mobile Apps
abstract
Delay-tolerant Wi-Fi offloading is known to improve overall mobile network bandwidth at low delay and low cost. Yet, in reality, we rarely find mobile apps that fully support opportunistic Wi-Fi access. This is mainly because it is still challenging to develop delay-tolerant mobile apps due to the complexity of handling network disruptions and delays. In this paper, we present Cedos, a practical delay-tolerant mobile network access architecture in which one can easily build a mobile app. Cedos consists of three components. First, it provides a familiar socket API whose semantics conforms to TCP, while the underlying protocol, D2TP, transparently handles network disruptions and delays in mobility. Second, Cedos allows the developers to explicitly exploit delays in mobile apps. App developers can express maximum user-specified delays in content download or use the API for real-time buffer management at opportunistic Wi-Fi usage. Third, for backward compatibility to existing TCP-based servers, Cedos provides D2Prox, a protocol-translation Web proxy. D2Prox allows intermittent connections on the mobile device side, but correctly translates Web transactions with traditional TCP servers. We demonstrate the practicality of Cedos by porting mobile Firefox and VLC video streaming client to using the API. We also implement delay/disruption-tolerant podcast client and run a field study with 50 people for eight weeks. We find that up to 92.4% of the podcast traffic is offloaded to Wi-Fi, and one can watch a streaming video in a moving train while offloading 48% of the content to Wi-Fi without a single pause.
YoungGyoun Moon, Donghwi Kim, Younghwan Go, Yeongjin Kim, Yung Yi, Song Chong, KyoungSoo Park
IEEE/ACM Trans. Netw.7
2016 DFC: Accelerating String Pattern Matching for Network Applications
Byungkwon Choi, Jongwook Chae, Muhammad Asim Jamshed, KyoungSoo Park, Dongsu Han
NSDI4
2015 Scaling the Performance of Network Intrusion Detection with Many-core Processors
abstract
In this work, we present a highly scalable network intrusion detection system on many-core processors. To maximize the NIDS performance, we take advantage of the underlying hardware and adhere to four design principles: shared-nothing architecture, computation offloading, lightweight data structure, and flow offloading. Through the experimental results, we find that our design choices can significantly improve the NIDS performance (79 Gbps with 1514B synthetic packets). We believe that our design decisions can be easily extended to other many-core processors and programmable NICs.
Jaehyun Nam, Muhammad Asim Jamshed, Byungkwon Choi, Dongsu Han, KyoungSoo Park
ANCS5
2015 Practicalizing Delay-Tolerant Mobile Apps with Cedos
abstract
Delay-tolerant Wi-Fi offloading is known to improve overall mobile network bandwidth at low delay and low cost. Yet, in reality, we rarely find mobile apps that fully support opportunistic Wi-Fi access. This is mainly because it is still challenging to develop delay-tolerant mobile apps due to the complexity of handling network disruptions and delays.
YoungGyoun Moon, Donghwi Kim, Younghwan Go, Yeongjin Kim, Yung Yi, Song Chong, KyoungSoo Park
MobiSys7
2015 Haetae: Scaling the Performance of Network Intrusion Detection with Many-Core Processors
Jaehyun Nam, Muhammad Asim Jamshed, Byungkwon Choi, Dongsu Han, KyoungSoo Park
RAID5
2015 A Case for a Stateful Middlebox Networking Stack
abstract
No abstract available.
Muhammad Asim Jamshed, Donghwi Kim, YoungGyoun Moon, Dongsu Han, KyoungSoo Park
SIGCOMM5
2015 FloSIS: A Highly Scalable Network Flow Capture System for Fast Retrieval and Storage Efficiency
Jihyung Lee, Sungryoul Lee, Yung Yi, KyoungSoo Park
USENIX ATC5
2014 Effective content-based video caching with cache-friendly encoding and media-aware chunking
abstract
Caching similar videos transparently in a network is a cost-effective solution that potentially reduces redundant data transfers. Recent study shows that network redundancy elimination (NRE) on the content level could produce high bandwidth savings in ISPs. However, we find that blindly employing existing NRE techniques to video contents could lead to suboptimal redundancy suppression rates. This is because (a) randomness in the video encoding process could produce completely different binaries even when they deal with seemingly identical video clips and (b) existing NRE chunking schemes incur high overheads since they do not utilize the underlying video format.
Sangwook Bae, Giyoung Nam, KyoungSoo Park
MMSys3
2014 Gaining Control of Cellular Traffic Accounting by Spurious TCP Retransmission
Younghwan Go, Eunyoung Jeong, Jongil Won, Yongdae Kim, Denis Foo Kune, KyoungSoo Park
NDSS6
2014 mTCP: a Highly Scalable User-level TCP Stack for Multicore Systems
Eunyoung Jeong, Shinae Woo, Muhammad Asim Jamshed, Haewon Jeong, Sunghwan Ihm, Dongsu Han, KyoungSoo Park
NSDI7
2013 Meeting the real-time constraints with standard Ethernet in an in-vehicle network
abstract
Vehicular networks have traditionally focused on the real-time delivery of critical control messages for safe car operation. Unfortunately, the real-time requirements often cripple the development of flexible car applications by tying the application network stack to underlying physical networks. While popular real-time vehicular networks guarantee the timely delivery of prioritized messages, they often lack in bandwidth and flexibility, which limits the range of car network applications. In this work, we explore the idea of replacing the current vehicular network with standard switched Ethernet, the most popular LAN technology in computer networks. Ethernet is attractive in providing high bandwidth at a low cost with easy and flexible configuration. The most challenging part is to guarantee the real-time delivery of mission-critical messages. We first show that the soft message delivery latency of 10s to 100s milliseconds can be easily met in 100 Mbps switched Ethernet despite coexistence of high-bandwidth network applications. For meeting the hard delivery latency on the order of 100 microseconds for critical control messages, we propose limiting the path MTU to the destination node with priority queuing from IEEE 802.1Q. Our simulation shows that we can satisfy 100 microseconds of latency even in a rich set of vehicular applications without any modification of the application network stack.
Youngwoo Lee, KyoungSoo Park
Intelligent Vehicles Symposium2
2013 Comparison of caching strategies in modern cellular backhaul networks
abstract
Recent popularity of smartphones drives rapid growth in the demand for cellular network bandwidth. Unfortunately, due to the centralized architecture of cellular networks, increasing the physical backhaul bandwidth is challenging. While content caching in the cellular network could be beneficial, relatively few characteristics of the cellular traffic is known to come up with a highly-effetive caching strategy. In this work, we provide insight into flow and content-level characteristics of modern 3G traffic at a large cellular ISP in South Korea. We first develop a scalable deep flow inspection (DFI) system that can manage hundreds of thousands of concurrent TCP flows on a commodity multicore server. Our DFI system collects various HTTP/TCP-level statistics and produces logs for analyzing the effectiveness of conventional Web caching, prefix-based Web caching, and TCP-level redundancy elimination (RE) without a single packet drop at a 10~Gbps link. Our week-long measurements of over 370 TBs of the 3G traffic reveal that standard Web caching can reduce download bandwidth consumption up to 27.1% while simple TCP-level RE can save the bandwidth consumption up to 42.0% with a cache of 512~GB of RAM. We also find that applying TCP-level RE on the largest 9.4% flows eliminates 68.4% of the total redundancy. Most of the redundancy (52.1%~58.9%) comes from serving the same HTTP objects while the contribution by aliased URLs is up to 38.9%.
Shinae Woo, Eunyoung Jeong, Shinjo Park, Jong Min Lee 0001, Sunghwan Ihm, KyoungSoo Park
MobiSys6
2012 Kargus: a highly-scalable software-based intrusion detection system
abstract
As high-speed networks are becoming commonplace, it is increasingly challenging to prevent the attack attempts at the edge of the Internet. While many high-performance intrusion detection systems (IDSes) employ dedicated network processors or special memory to meet the demanding performance requirements, it often increases the cost and limits functional flexibility. In contrast, existing software-based IDS stacks fail to achieve a high throughput despite modern hardware innovations such as multicore CPUs, manycore GPUs, and 10 Gbps network cards that support multiple hardware queues.
Muhammad Asim Jamshed, Jihyung Lee, Insu Yun, Deokjin Kim, Sungryoul Lee, Yung Yi, KyoungSoo Park
CCS8
2012 Server-assisted Latency Management for Wide-area Distributed Systems
Wonho Kim, KyoungSoo Park, Vivek S. Pai
USENIX ATC2
2011 SSLShader: Cheap SSL Acceleration with Commodity Processors
Keon Jang, Sangjin Han, Seungyeop Han, Sue B. Moon, KyoungSoo Park
NSDI5
2010 Building a single-box 100 Gbps software router
abstract
Commodity-hardware technology has advanced in great leaps in terms of CPU, memory, and I/O bus speeds. Benefiting from the hardware innovation, recent software routers on commodity PC now report about 10 Gbps in packet routing. In this paper we map out expected hurdles and projected speed-ups to reach 100 Gbps in packet routing on a single commodity PC. With careful measurements, we identify two notable bottlenecks for our goal: CPU cycles and I/O bandwidth. For the former, we propose reducing per-packet processing overhead with software-level optimizations and buying extra computing power with GPUs. To improve the I/O bandwidth, we suggest scaling the performance of I/O hubs that limits packet routing speed to well before 50 Gbps.
Sangjin Han, Keon Jang, KyoungSoo Park, Sue B. Moon
LANMAN3
2010 PacketShader: a GPU-accelerated software router
abstract
We present PacketShader, a high-performance software router framework for general packet processing with Graphics Processing Unit (GPU) acceleration. PacketShader exploits the massively-parallel processing power of GPU to address the CPU bottleneck in current software routers. Combined with our high-performance packet I/O engine, PacketShader outperforms existing software routers by more than a factor of four, forwarding 64B IPv4 packets at 39 Gbps on a single commodity PC. We have implemented IPv4 and IPv6 forwarding, OpenFlow switching, and IPsec tunneling to demonstrate the flexibility and performance advantage of PacketShader. The evaluation results show that GPU brings significantly higher throughput over the CPU-only implementation, confirming the effectiveness of GPU for computation and memory-intensive operations in packet processing.
Sangjin Han, Keon Jang, KyoungSoo Park, Sue B. Moon
SIGCOMM3
2010 Accelerating SSL with GPUs
abstract
SSL/TLS is a standard protocol for secure Internet communication. Despite its great success, today's SSL deployment is largely limited to security-critical domains. The low adoption rate of SSL is mainly due to high computation overhead on the server side.
Keon Jang, Sangjin Han, Seungyeop Han, Sue B. Moon, KyoungSoo Park
SIGCOMM5
2010 Wide-area Network Acceleration for the Developing World
Sunghwan Ihm, KyoungSoo Park, Vivek S. Pai
USENIX ATC2
2009 HashCache: Cache Storage for the Next Billion
Anirudh Badam, KyoungSoo Park, Vivek S. Pai, Larry L. Peterson
NSDI2
2007 Supporting Practical Content-Addressable Caching with CZIP Compression
KyoungSoo Park, Sunghwan Ihm, Mic Bowman, Vivek S. Pai
USENIX ATC1
2006 Scale and Performance in the CoBlitz Large-File Distribution Service
KyoungSoo Park, Vivek S. Pai
NSDI1
2006 Connection Conditioning: Architecture-Independent Support for Simple, Robust Servers
KyoungSoo Park, Vivek S. Pai
NSDI1
2006 Securing Web Service by Automatic Robot Detection
KyoungSoo Park, Vivek S. Pai, Kang-Won Lee 0002, Seraphin B. Calo
USENIX ATC, General Track1
2004 CoDNS: Improving DNS Performance and Reliability via Cooperative Lookups
KyoungSoo Park, Vivek S. Pai, Larry L. Peterson
OSDI1
2004 Reliability and Security in the CoDeeN Content Distribution Network
Limin Wang 0010, KyoungSoo Park, Ruoming Pang, Vivek S. Pai, Larry L. Peterson
USENIX ATC, General Track2