EDBT 2026 Demo / reviewers in the wild / expert
Jae-Hyun Hwang
dblp:58/3121 · also Jaehyun Hwang
· DBLP profile ↗
20ranked-venue papers
8as first author
11since 2021 · last 2026
0000-0001-8149-8397ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 5 first-author · 3 since 2021Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lockify: Understanding Linux Distributed Lock Management Overheads in Shared Storage
Taeyoung Park, Yunjae Jo, Daegyu Han, Beomseok Nam, Jae-Hyun Hwang |
FAST | 5 |
| 2026 | SMTcheck: Accurate SMT Interference Prediction to Improve Scheduling Efficiency in DatacentersabstractSimultaneous multithreading (SMT) is widely used in modern x86 processors to improve core utilization by sharing hardware resources between co-located threads. However, such resource sharing often leads to severe performance interference, making efficient workload co-scheduling difficult, especially given the complexity and diversity of modern x86 CPUs. Our analysis reveals that SMT-aware workload scheduling can significantly improve system throughput and reduce tail latency for datacenter workloads, but identifying optimal thread combinations is challenging due to the lack of visibility into platform-specific resource sharing behaviors. In this paper, we present SMTcheck, a lightweight, accurate, and platformindependent methodology for predicting SMT interference for diverse x86 processors. SMTcheck uses carefully designed code snippets (Diags) to extract hidden microarchitectural features of performance-critical shared resources. With these extracted features, SMTcheck builds per-resource microbenchmarks (Injectors) to apply pinpoint pressure to specific target resources in order to capture workload-specific contention characteristics. SMTcheck then constructs a hardware-aware contention model to predict performance interference between arbitrary workload pairs without requiring exhaustive profiling. We evaluate SMTcheck on six x86 desktop processors and five x86 server processors from Intel and AMD across different generations and show that it achieves high prediction accuracy by up to 95.5 % (94.6 % on average). We further demonstrate its effectiveness by implementing a contention-aware scheduler in the Linux kernel. Compared to the default Linux scheduler, our contention-aware scheduler significantly reduces tail latency for latency-critical workloads (e.g., database, key-value store) by up to 36.09 %, and improves the overall system throughput by up to$1.072 \times$. Finally, using real-world cluster traces from Alibaba and Google, we demonstrate that SMTcheck incurs negligible profiling overheads ($\approx 0.113 \%$), making it practical for deployment in productionscale datacenter environments. Jinhyeok Oh, Gyutae Kim, Youngsok Kim, Jae-Hyun Hwang, Joonsung Kim 0001 |
HPCA | 6 |
| 2026 | Understanding Host Network Stack Latency
Tianyu Zuo, Jae-Hyun Hwang, Ao Tang, Rachit Agarwal 0001, Qizhe Cai |
SIGCOMM | 2 |
| 2026 | Prediction-based GPU sharing for distributed trainingabstract• Formulate the inconsistent JCT problem using gSLA for the first time. • Design a new JCT increase prediction model and job scheduler for GPU sharing. • Achieve up to 47.3× better gSLA satisfaction and 50× lower gSLA excess ratio. • Improve JCT and GPU efficiency by ∼ 60% and ∼ 44% over existing methods. • Demonstrate TensorShare’s effectiveness in improving gSLA and JCT for unseen jobs. GPU sharing aims to enhance the efficiency of GPU utilization by running distributed deep learning training jobs concurrently. However, GPU sharing poses a significant challenge: the increase in job completion time (JCT) caused by interference between jobs is inconsistent, complicating job scheduling. Our experiments reveal that the degree of JCT increase varies by as much as ∼ 3.7 × . While previous studies have analyzed this JCT inconsistency problem, none of them have been able to minimize the inconsistency. We propose TensorShare, a proactive GPU sharing technique that leverages a deep learning model to predict the extent of JCT increase. This study defines a new metric, called GPU SLA, which represents the upper threshold of JCT increase. TensorShare then introduces a novel scheduler that proactively identifies which jobs meet GPU SLA while minimizing the JCT increase. Our evaluation shows that TensorShare improves GPU SLA satisfaction rates by 26.1 × –47.3 × and reduces the JCT increase by 37%–60%. Furthermore, we evaluate TensorShare with large language models that are not included in training TensorShare’s prediction model, achieving ∼ 7 × and ∼ 10.3 × improvements in GPU SLA satisfaction and JCT inconsistency, respectively. Changyong Shin, Younghun Go, Yeonho Yoo, Jae-Hyun Hwang, Gyeongsik Yang, Chuck Yoo |
Future Gener. Comput. Syst. | 5 |
| 2025 | Optimizing Receive Flow Steering for Mixed Traffic in High-Performance Cloud DatacentersabstractThe evolution of cloud datacenter infrastructure, driven by ever-increasing bandwidth demands, has shifted performance bottlenecks from network hardware to server-side protocol processing. To mitigate per-CPU processing bottlenecks, load-balancing schemes such as Receive Side Scaling (RSS) and accelerated Receive Flow Steering (aRFS) have been introduced to distribute network flows across multiple CPU cores for incoming packet processing. While improving overall core utilization and cache locality, these schemes do not consider recent advancements such as non-uniform memory access (NUMA)-based architectures and direct cache access (DCA), leading to suboptimal performance, particularly when large and small flows are mixed. In this paper, we introduce a new receive flow steering scheme, mixed Flow Steering (mFS), which disaggregates large and small flows to optimize network flow steering in NUMA-based architectures. Our approach incorporates a flow monitoring module to classify flows based on cumulative data volume, a flow core migration mechanism that aligns application processing with the appropriate NUMA node, and adaptive handling of multi-large flow contention. Our experimental evaluations demonstrate that, by leveraging DCA, mFS significantly improves total throughput for large flows by 42.12% compared to existing schemes while maintaining comparable throughput for small flows across various mixed traffic scenarios and real-world applications. Junseo Jang, Jae-Hyun Hwang |
CLOUD | 2 |
| 2024 | enCloud: Aspect-oriented trusted service migration on SGX-enabled cloud VMabstractAbstract This paper presents enCloud, a new aspect‐oriented trusted service migration with SGX‐enabled cloud VM. Addressing the challenge of reconciling end‐to‐end security with VM migration, enCloud incorporates two key aspects: (1) end‐to‐end security for enclave context migration, and (2) VM abstraction for conventional VM context migration. This paper provides a practical guideline with applicable APIs for trusted service migration. In a case study, enCloud demonstrates effective trusted DB service migration on a cloud VM, achieving end‐to‐end security with minimal trust boundaries. The framework supports pre‐copy live VM migration to minimize service downtime. This paper contributes a concise and practical solution in the form of the enCloud framework for secure service migration. See-hwan Yoo, Young-Pil Kim, Hyunchan Park, Jae-Hyun Hwang, Kitak Kim |
Softw. Pract. Exp. | 4 |
| 2023 | Development of a real-time noise estimation model for construction sitesabstractAs construction noise negatively affects the health and quality of life of stakeholders, field managers need to properly monitor and manage noise. Thus, the authors developed a model that estimates real-time noise levels at a construction site and the surroundings to enable preemptive responses to noise-related issues. To accurately estimate noise, necessary field data were collected using an unmanned aerial vehicle (UAV) and noise sensors. The noise estimation model was composed of two sub-models: the noise-customized spatial interpolation model and the noise propagation model. The noise-customized spatial interpolation model was developed to estimate the internal noise of the construction site using a few sensor noise levels. Meanwhile, the noise propagation model was developed to estimate the noise level outside the construction site using internal noise estimation results, obstacles, weather information, and noise sources information. The model was evaluated through field tests at a construction technology demonstration center, environments identical to real construction sites in South Korea. The model showed satisfactory performance, with an accuracy of 96.71% and a root mean square error (RMSE) of 2.62 for the internal construction site noise and an accuracy of 96.03% and an RMSE of 2.70 for outside the construction site. To facilitate the usage of the noise estimation results for field managers, the research team visualized the results using the Unity 3D Engine. The results will enable field managers to assess workers’ long-term noise exposure and respond to potential civil complaints, gearing up to realize environmental, social, and governance (ESG) goals in the construction industry. Gitaek Lee, Seonghyeon Moon, Jae-Hyun Hwang, Seokho Chi |
Adv. Eng. Informatics | 3 |
| 2023 | Autothrottle: Satisfying Network Performance Requirements for ContainersabstractThis article investigates how to satisfy network performance requirements that are crucial in achieving the service level objectives (SLOs) in clouds. Traditional techniques for network performance management have a limited ability to satisfy the network SLOs. Our in-depth analysis reveals that the fundamental reason comes from decoupling of the CPU scheduler and the network traffic controller as the current CPU scheduler is not aware of such network requirements but only provides a fair-share amount of CPU to all containers. Thus, the container cannot perform the amount of network processing as needed to satisfy its SLO when the CPU allocation is insufficient. In this article, we propose Autothrottle that dynamically adjusts the CPU allocation for the containers to satisfy their network SLOs. The key element of Autothrottle is a throttle algorithm that autonomously determines the amount of CPU for each container needed to satisfy the requirement. We implement Autothrottle in the Linux kernel and evaluate it with massive real-world workloads such as Apache Kafka. Our evaluation results show that Autothrottle successfully satisfies the given network SLO only with a 2% gap while the existing scheme achieves 20% less than the SLO. We further observe that Autothrottle also reduces the CPU overhead in network processing by 19%, improving the network throughput by 27% compared to the existing scheme. Kyungwoon Lee, Kwanhoon Lee, Hyunchan Park, Jae-Hyun Hwang, Chuck Yoo |
IEEE Trans. Cloud Comput. | 4 |
| 2022 | Towards μs tail latency and terabit ethernet: disaggregating the host network stackabstractDedicated, tightly integrated, and static packet processing pipelines in today's most widely deployed network stacks preclude them from fully exploiting capabilities of modern hardware. Qizhe Cai, Midhul Vuppalapati, Jae-Hyun Hwang, Christoforos E. Kozyrakis, Rachit Agarwal 0001 |
SIGCOMM | 3 |
| 2021 | Rearchitecting Linux Storage Stack for µs Latency and High Throughput
Jae-Hyun Hwang, Midhul Vuppalapati, Simon Peter 0001, Rachit Agarwal 0001 |
OSDI | 1 |
| 2021 | Understanding host network stack overheadsabstractTraditional end-host network stacks are struggling to keep up with rapidly increasing datacenter access link bandwidths due to their unsustainable CPU overheads. Motivated by this, our community is exploring a multitude of solutions for future network stacks: from Linux kernel optimizations to partial hardware offload to clean-slate userspace stacks to specialized host network hardware. The design space explored by these solutions would benefit from a detailed understanding of CPU inefficiencies in existing network stacks. Qizhe Cai, Shubham Chaudhary 0004, Midhul Vuppalapati, Jae-Hyun Hwang, Rachit Agarwal 0001 |
SIGCOMM | 4 |
| 2020 | TCP ≈ RDMA: CPU-efficient Remote Storage Access with i10
Jae-Hyun Hwang, Qizhe Cai, Ao Tang, Rachit Agarwal 0001 |
NSDI | 1 |
| 2018 | Kafe: Can OS Kernels Forward Packets Fast Enough for Software Routers?abstractIt is widely believed that software routers based on commodity operating systems cannot deliver high-speed packet processing, and a number of alternative approaches (including user-space network stacks) have been proposed. This paper revisits the inefficiency of kernel-level packet processing inside modern OS-based software routers and explores whether a redesign of kernel network stacks can improve the incompetence. We present a case contrary to the belief through a redesign: Kafe-a kernel-based advanced forwarding engine that can process packets as fast as user-space network stacks. The Kafe neither adds any new API nor depends on proprietary hardware features, but the Kafe outperforms Linux by seven times and RouteBricks by three times. The current implementation of the Kafe can forward 64-byte IPv4 packets at 28.2 Gbps using eight cores running at 2.6 GHz. Our evaluation results show that the Kafe achieves similar packet forwarding performance to Intel DPDK while consuming much less CPU and memory resources. Cheol-Ho Hong, Kyungwoon Lee, Jae-Hyun Hwang, Hyunchan Park, Chuck Yoo |
IEEE/ACM Trans. Netw. | 3 |
| 2016 | Eliminating bandwidth estimation from adaptive video streaming in wireless networks
Jae-Hyun Hwang, Chuck Yoo |
Signal Process. Image Commun. | 1 |
| 2016 | Multipath TCP: Analysis, Design, and ImplementationabstractMultipath TCP (MP-TCP) has the potential to greatly improve application performance by using multiple paths transparently. We propose a fluid model for a large class of MP-TCP algorithms and identify design criteria that guarantee the existence, uniqueness, and stability of system equilibrium. We clarify how algorithm parameters impact TCP-friendliness, responsiveness, and window oscillation and demonstrate an inevitable tradeoff among these properties. We discuss the implications of these properties on the behavior of existing algorithms and motivate our algorithm Balia (balanced linked adaptation), which generalizes existing algorithms and strikes a good balance among TCP-friendliness, responsiveness, and window oscillation. We have implemented Balia in the Linux kernel. We use our prototype to compare the new algorithm to existing MP-TCP algorithms. Qiuyu Peng, Anwar Elwalid, Jae-Hyun Hwang, Steven H. Low |
IEEE/ACM Trans. Netw. | 3 |
| 2015 | Scalable Congestion Control Protocol Based on SDN in Data Center NetworksabstractOn-line data center applications render challenging network latency demands to meet their service level requirements. These applications, however, frequently suffer from increased latency due to the packet loss and queueing delay at the network switches. These are mainly as a result of the momentary massive bursts by the Partition/Aggregation application traffic patterns, which causes incast network congestion at the network switches. In this paper, we propose a scalable congestion control protocol, called SCCP. Our scheme effectively limits the data rate of the TCP senders by leveraging the Software Defined Networking (SDN) switches, so that the total utilization does not exceed the bottleneck link capacity. Furthermore, SCCP can be easily deployed to the existing SDN data center switches by extending the OpenFlow specifications. Our Open vSwitch-based prototype experiments and ns-3 simulations show that SCCP is scalable for up to hundreds of concurrent flows traversing through the data center network switch port. Jae-Hyun Hwang, Joon Yoo, Hyun-Wook Jin |
GLOBECOM | 1 |
| 2014 | Deadline and Incast Aware TCP for cloud data center networks
Jae-Hyun Hwang, Joon Yoo, Nakjung Choi |
Comput. Networks | 1 |
| 2012 | IA-TCP: A rate based incast-avoidance algorithm for TCP in data center networksabstractIn recent years, the data center networks commonly accommodate applications such as MapReduce and web search that inherently shows the incast communication pattern; multiple workers simultaneously transmit TCP data to a single aggregator. In this environment, the TCP performance is significantly degraded in terms of goodput and query completion time, as a result of the severe packet loss at Top of Rack (ToR) switches. The TCP senders aggressively transmit packets causing throughput collapse even though the network pipe size, i.e., bandwidth-delay product, is extremely small. In this paper, we introduce a novel end-to-end congestion control algorithm called IA-TCP that avoids the TCP incast congestion problem effectively. IA-TCP employs the rate-based algorithm at the aggregator node, which controls both the window size of workers and ACK delay. Through extensive NS-2 simulations, we validate that our algorithm is scalable in terms of the number of workers achieving enhanced goodput and zero timeouts. Jae-Hyun Hwang, Joon Yoo, Nakjung Choi |
ICC | 1 |
| 2009 | TCP Feno: Enhancement for higher accuracy of loss differentiation over small buffer heterogeneous networksabstractIt is well known that TCP shows performance degradation over wired/wireless networks since TCP regards packet loss as network congestion. TCP Veno has successfully addressed this fundamental problem by proposing an end-to-end loss differentiation algorithm, which distinguishes the cause of packet loss by the number of packets in the router buffer. Unfortunately, Veno's algorithm shows very low accuracy on small buffer routers that are emerging recently as a new challenge for Internet routers. In the small buffer networks, routers can overflow easily in a short time although Veno diagnoses wireless loss, and this leads to the failure in loss differentiation. Furthermore, when congestion loss is misdiagnosed as wireless loss, Veno shows poor Reno-friendliness since it can increase sending rate by setting ssthresh to a larger value than Reno. In this paper, we propose a more accurate loss differentiation algorithm for small buffer heterogeneous networks. Our algorithm accurately distinguishes wireless and wired packet loss by newly defining congestive rate. Through extensive network simulations, we confirm that our new TCP Feno achieves not only higher accuracy, but also better Reno-friendliness while not losing performance efficiency. Jae-Hyun Hwang, See-hwan Yoo, Chuck Yoo |
LCN | 1 |
| 2008 | DR-TCP: Downloadable and reconfigurable TCP
Jae-Hyun Hwang, Jin-Hee Choi, Se-Won Kim, Chuck Yoo |
J. Syst. Softw. | 1 |