Gyuyeong Kim

dblp:65/10137 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0003-0052-3568ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 3 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Network-Accelerated Multiget Coordination for Distributed Key-Value Stores
abstract
Modern key-value stores support multiget requests, which allow clients to retrieve the values of multiple specified keys in a single request. However, the multiget operation causes extra coordination overhead when requested keys are distributed across multiple servers. Unfortunately, existing coordination architectures suffer from high client overhead or limited scalability. To this end, we propose NetMC, a networkaccelerated multiget coordination architecture that achieves high throughput, low latency, and scalability simultaneously. Our key idea is to distribute the labor of multiget coordination between the network switch and the client. Specifically, we offload the stateless request splitting to the I/O-optimized network switch, while the stateful reply aggregation is performed on the client side. By leveraging the strengths of each system component for coordination functionality, NetMC reduces the coordination overhead significantly. We implement a NetMC prototype on a cluster of commodity servers connected via an Intel Tofino switch. Our experimental results show that NetMC outperforms existing architectures and is robust to various system conditions.
Jiyoon Bang, Gyuyeong Kim
CCGrid2
2025 Pushing the Limits of In-Network Caching for Key-Value Stores
Gyuyeong Kim
NSDI1
2023 NetClone: Fast, Scalable, and Dynamic Request Cloning for Microsecond-Scale RPCs
abstract
Spawning duplicate requests, called cloning, is a powerful technique to reduce tail latency by masking service-time variability. However, traditional client-based cloning is static and harmful to performance under high load, while a recent coordinator-based approach is slow and not scalable. Both approaches are insufficient to serve modern microsecond-scale Remote Procedure Calls (RPCs). To this end, we present NetClone, a request cloning system that performs cloning decisions dynamically within nanoseconds at scale. Rather than the client or the coordinator, NetClone performs request cloning in the network switch by leveraging the capability of programmable switch ASICs. Specifically, NetClone replicates requests based on server states and blocks redundant responses using request fingerprints in the switch data plane. To realize the idea while satisfying the strict hardware constraints, we address several technical challenges when designing a custom switch data plane. NetClone can be integrated with emerging in-network request schedulers like RackSched. We implement a NetClone prototype with an Intel Tofino switch and a cluster of commodity servers. Our experimental results show that NetClone can improve the tail latency of microsecond-scale RPCs for synthetic and real-world application workloads and is robust to various system conditions.
Gyuyeong Kim
SIGCOMM1
2023 DynaQ: Enabling Protocol-Independent Service Queue Isolation in Cloud Data Centers
abstract
Switches in cloud data centers support multiple service queues per port to provide differentiated network performance among different traffic classes. To isolate service queues, recent solutions leverage the power of Explicit Congestion Notification (ECN). However, this causes a fundamental dependency on ECN-based transport protocols, making it hard to use generic transport protocols. To this end, we design DynaQ, a protocol-independent multi-queue management solution that enables service queue isolation with generic transport protocols. The key idea of DynaQ is to adjust the packet dropping threshold of service queues dynamically. Specifically, DynaQ allows a service queue to occupy free buffer space but prevents the queue from hurting other active queues. Our solution requires only a few additional clock cycles to implement on hardware. To evaluate DynaQ comprehensively, we conduct a series of testbed experiments and large-scale simulations. Our evaluation results show that, compared to alternative schemes, DynaQ is the only solution that achieves work-conserving weighted fair sharing and low latency without protocol dependency.
Gyuyeong Kim, Wonjun Lee 0001
IEEE Trans. Cloud Comput.1
2022 Network Policy Enforcement With Commodity Multiqueue NICs for Multitenant Data Centers
abstract
Data centers are the fundamental component in the Internet of Things (IoT) system architecture. Data center servers where IoT services are co-located require hierarchical network policy enforcement to ensure fair bandwidth sharing among tenants and to prioritize latency-sensitive traffic within a tenant simultaneously. Meanwhile, emerging network interface cards (NICs) in servers make use of multiple hardware queues to drive increasing line rates. Unfortunately, multiqueue NICs make it hard to enforce hierarchical policies because the NIC packet scheduler dequeues packets in a static round-robin (RR) fashion for per-flow fairness. In this article, we enable hierarchical network policy enforcement with existing commodity multiqueue NICs. We design TONIC, a multiqueue NIC packet scheduling solution that approximates hierarchical packet scheduling by manipulating the packet dequeueing sequence of the NIC scheduler through dynamic packet enqueueing decisions. Specifically, TONIC leverages multiple hardware queues and the double-ended queue structure of qdiscs to express different tenant weights and application priorities without hardware modifications. We implement a TONIC prototype as a Linux kernel module and evaluate it on a testbed with commodity multiqueue NICs. Our experiment results show that TONIC can enforce hierarchical policies consisting of weighted fair sharing and traffic prioritization while maintaining robustness to various network conditions.
Gyuyeong Kim, Wonjun Lee 0001
IEEE Internet Things J.1
2022 In-Network Leaderless Replication for Distributed Data Stores
abstract
Leaderless replication allows any replica to handle any type of request to achieve read scalability and high availability for distributed data stores. However, this entails burdensome coordination overhead of replication protocols, degrading write throughput. In addition, the data store still requires coordination for membership changes, making it hard to resolve server failures quickly. To this end, we present NetLR, a replicated data store architecture that supports high performance, fault tolerance, and linearizability simultaneously. The key idea of NetLR is moving the entire replication functions into the network by leveraging the switch as an on-path in-network replication orchestrator. Specifically, NetLR performs consistency-aware read scheduling, high-performance write coordination, and active fault adaptation in the network switch. Our in-network replication eliminates inter-replica coordination for writes and membership changes, providing high write performance and fast failure handling. NetLR can be implemented using programmable switches at a line rate with only 5.68% of additional memory usage. We implement a prototype of NetLR on an Intel Tofino switch and conduct extensive testbed experiments. Our evaluation results show that NetLR is the only solution that achieves high throughput and low latency and is robust to server failures.
Gyuyeong Kim, Wonjun Lee 0001
Proc. VLDB Endow.1
2022 LossPass: Absorbing Microbursts by Packet Eviction for Data Center Networks
abstract
A bursty traffic pattern, called the microburst, is a key hurdle to achieve low latency for user-facing applications because it causes excessive packet losses in shallow buffered switches. Explicit Congestion Notification (ECN) can absorb microbursts by reserving buffer headroom, but the existence of headroom results in a fundamental trade-off between latency and throughput. To this end, we present LossPass, a buffer sharing mechanism that absorbs microbursts as many as possible while maintaining line-rate throughput. Specifically, LossPass evicts buffered large flow packets to make free buffer space on demand for arriving small flow packets. Our solution is inexpensive to implement on hardware. We implement a LossPass prototype and evaluate its performance through extensive testbed experiments and large-scale simulations. Our evaluation results show that LossPass reduces the FCT of small flows while maintaining line-rate throughput. For example, in testbed experiments, LossPass outperforms ECN by up to$3.20\times$in the 99th percentile FCT of small flows.
Gyuyeong Kim, Wonjun Lee 0001
IEEE Trans. Cloud Comput.1
2020 Protocol-Independent Service Queue Isolation for Multi-Queue Data Centers
abstract
To isolate service queues in switch ports, recent solutions leverage the power of Explicit Congestion Notification (ECN). However, this causes a fundamental dependency on ECN-based transport protocols, making it hard to use generic transport protocols. To this end, we design DynaQ, a protocol-independent multi-queue management solution that enables service queue isolation with generic transport protocols. DynaQ dynamically adjusts the packet dropping threshold of service queues. Our solution is inexpensive to implement on hardware. Through extensive testbed experiments and large-scale simulations, we show that compared to alternative schemes, DynaQ is the only solution that achieves work-conserving weighted fair sharing and low latency without protocol dependency.
Gyuyeong Kim, Wonjun Lee 0001
ICDCS1
2017 Video based pedestrian detection and tracking at night-time
abstract
This paper is an approach for pedestrian detection and tracking with infrared imagery. The detection phase is performed by AdaBoost algorithm based on Haar-like features. AdaBoost classifier is trained with datasets generated from infrared images. The number of negative images used for training with AdaBoost algorithm is 3000. For positive training, 1000 samples are used After detecting the pedestrian with AdaBoost classifier, we proposed the Tracking-Learning-Detection (TLD) frameworks tracking strategies. TLD frameworks are preferred in this study because of its high accuracy rate and computation speed Tracking performance comparison is made between TLD and particle filtering. Results prove that TLD performs a higher tracking rate than particle filtering.
Geun-Hoo Lee, Gyuyeong Kim, Jongkwan Song, O. Faruk Ince, Jangsik Park
HSI2