Hyun-Wook Jin

dblp:36/2516 · DBLP profile ↗
← Back
49ranked-venue papers
9as first author
8since 2021 · last 2026
0000-0002-9496-3486ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 33 · 7 first-author · 6 since 2021Computer networks · 6 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Software engineering, systems software and programming languages · 2Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Mammoth: Macro-Level MPI Offloading to Off-Path Accelerator in DPU
Jong-Bin Lee, Pu-Rum Seo, Ki-Moon Jeong, Hyun-Wook Jin
CCGrid4
2025 Spatio-Temporal Resource Control for Cloud-Native GPU Provisioning
abstract
Modern cloud platforms, such as Kubernetes, provide a service-oriented resource abstraction for explicit Quality of Service (QoS) provisioning, which guarantees resource reservations and supports resource elasticity. In regards to CPU resources, tenants can explicitly specify a guaranteed base demand and an upper bound for resources. Thus, it is desirable that GPU resource provisioning should be analogous to mature cloud-native CPU resource abstraction and can support familiar QoS classes, such as Guaranteed, Burstable, and BestEffort. Although existing research efforts have primarily focused on maximizing GPU utilization by exploiting profiling, capabilities for guaranteed reservation, precise throttling, and elastic bursting are essential to support cloud-native GPU provisioning. It is challenging to provide such features due to the GPU's asynchronous and non-preemptive characteristics. To address this issue, we introduce a spatio-temporal provisioning framework that ensures both resource guarantee and elasticity for GPUs. We define a resource model to control both spatial and temporal dimensions of GPU resources and support familiar QoS classes. Our framework features a per-container Agent for transparent resource accounting and local quota enforcement, and the central Multi-Tenant Arbitrator for global, partition-aware fair scheduling. Performance evaluation demonstrates that our framework can provide accurate GPU resource guarantee and manages dynamic mixed-QoS workloads by honoring reserved GPU resources during contention while allowing tenants to elastically burst into idle capacity up to their limit.
Hyeon-Jun Jang, Sang-Jae Kim, Weikuan Yu, Hyun-Wook Jin
SoCC4
2024 SlackAuto: Slack-Aware Vertical Autoscaling of CPU Resources for Serverless Computing
abstract
In modern serverless platforms, vertical auto scaling becomes increasingly important as a single service instance handles multiple requests concurrently. Vertical autoscaling typically relies on a prediction mechanism to proactively adjust resource allocation based on estimated future resource demands. However, inaccurate predictions result in a waste of resources or poor performance. To address this issue, we introduce slack-awareness into existing prediction/autoscaling mechanisms. Our proposed mechanism, named SlackAuto, models resource slack as a queue and controls its length to improve both CPU efficiency and application-level performance. We utilize the drift-plus-penalty algorithm of the Lyapunov optimization technique, which does not add significant run-time overheads and provides an efficient way to integrate slack-awareness with existing mechanisms. Our implementation operates in both periodic and event-driven modes, enabling an immediate response to changes in demand. Performance measurement results show that SlackAuto outper-forms existing vertical autoscaling algorithms in terms of resource efficiency and response latency.
Hyeon-Jun Jang, Hyun-Wook Jin
CLOUD2
2023 Self-adaptive end-to-end resource management for real-time monitoring in cyber-physical systems
Hyun-Chul Jo, Hyun-Wook Jin, Joongheon Kim
Comput. Networks2
2023 QoS for best-effort batch jobs in container-based cloud
abstract
Summary The resource orchestrators in the cloud provide different QoS classes. Existing studies have mainly focused on resource guarantee for high‐priority class jobs, but yet to consider delicate QoS control of low‐priority jobs in the cloud. In this paper, we propose the differentiated shares that provide weighted CPU scheduling for low‐priority batch jobs in container‐based cloud. To implement the differentiated shares, we extend resource reservation interfaces provided by Kubernetes and suggest an algorithm that maps the resource management attributes of Kubernetes into those of Linux. The proposed differentiated shares can avoid worsening interference with high‐priority containers by exploiting the hierarchical resource sharing of task groups in Linux. In addition, we suggest a node scoring policy to provide efficient inter‐node job scheduling with the consideration of the differentiated shares. The performance measurement results show that the differentiated shares can provide weighted CPU scheduling for best‐effort batch containers without interference with high‐priority containers (less than 3% with respect to application‐level performance).
Yin-Goo Yim, Hyeon-Jun Jang, Hyun-Wook Jin
Concurr. Comput. Pract. Exp.3
2023 Exploiting copy engines for intra-node MPI collective communication
abstract
Abstract As multi/many-core processors are widely deployed in high-performance computing systems, efficient intra-node communication becomes more important. Intra-node communication involves data copy operations to move messages from source to destination buffer. Researchers have tried to reduce the overhead of this copy operation, but the copy operation performed by CPU still wastes the CPU resources and even hinders overlapping between computation and communication. The copy engine is a hardware component that can move data between intra-node buffers without intervention of CPU. Thus, we can offload the copy operation performed by CPU onto the copy engine. In this paper, we aim at exploiting copy engines for MPI blocking collective communication, such as broadcast and gather operations. MPI is a messaging-based parallel programming model and provides point-to-point, collective, and one-sided communications. Research has been conducted to utilize the copy engine for MPI, but the support for collective communication has not yet been studied. We propose the asynchronism in blocking collective communication and the CE-CPU hybrid approach to utilize both copy engine and CPU for intra-node collective communication. The measurement results show that the proposed approach can reduce the overall execution time of a microbenchmark and a synthetic application that perform collective communication and computation up to 72% and 57%, respectively.
Joong-Yeon Cho, Pu-Rum Seo, Hyun-Wook Jin
J. Supercomput.3
2021 NUMA-aware I/O System Call Steering
abstract
To fully utilize the ever-increasing network and storage I/O bandwidth, there have been significant studies on interrupt steering that scatters I/O events raised by peripherals across multiple cores. However, less attention has been paid to system call steering that distributes actual processing of I/O system calls to multiple cores apart from application contexts. In this study, we suggest three different policies for I/O system call steering named Single-NUMA-Node, Per-NUMA-Node, and Cross-NUMA-Node that address the issues of parallelism between I/O operations, locality of data, and distance to I/O devices to improve utilization of network and storage I/O bandwidth on NUMA-based multi-core systems. The performance measurement results of our preliminary implementations show that the Cross-NUMA-Node policy can reduce the execution time of MapReduce applications up to 34% on a Hadoop cluster.
Chan-Gyu Lee, Hyun-Wook Jin
CLUSTER2
2021 Transparent many-core partitioning for high-performance big data I/O
abstract
Summary As the number of cores equipped in a single computing node is rapidly increasing, utilizing many cores for contemporary applications in an efficient manner is a challenging issue. We need to consider both parallelization and locality to fully exploit many cores for multifarious operations of emerging applications. In particular, big data applications perform computation and I/O intensive operations alternately. For instance, Apache Hadoop MapReduce assumes local persistent storage for each computing node. Thus, unlike traditional parallel programming models, the MapReduce framework performs not only networking but also storage I/O. In this study, we aim to improve the locality of network and storage I/O operations on many‐core systems by partitioning cores for I/O system calls and event handlers. In order to implement fine‐grained many‐core partitioning, we decouple the system call context from the user‐level process by suggesting message‐based system calls. The suggested design provides user‐level transparency and does not require any kernel‐level modifications. In addition, we propose a scheme that dynamically decides the core affinity of system calls and event handlers by considering locality, run‐time loads, and hardware architectures. The experimental results show that the proposed many‐core partitioning can improve the locality of network and storage I/O operations in an integrated manner for MapReduce applications.
Chan-Gyu Lee, Joong-Yeon Cho, Jooho Kim, Hyun-Wook Jin
Concurr. Comput. Pract. Exp.4
2020 Vertical Autoscaling of GPU Resources for Machine Learning in the Cloud
abstract
Vertical autoscaling scales the amount of resources reserved by a virtual machine in the cloud. Although there have been studies on vertical autoscaling of CPU and memory resources, these have yet to consider GPU resources. In this paper, we propose a vertical autoscaling algorithm to improve the utilization of GPU resources within budget limit by exploiting Lyapunov optimization. Our algorithm deals with the correlation between GPU and CPU resources and requires only resource utilization information to decide to scale GPU resources. The performance measurement results show that our GPU vertical autoscaler can provide optimal performance to containerized CNN-based machine learning applications, such as ResNet-50 and Yolo-v4, with respect to execution time and throughput.
Hyeon-Jun Jang, Yin-Goo Yim, Hyun-Wook Jin
IEEE BigData3
2018 Application-transparent scheduling of socket system calls on many-core systems
abstract
As the number of cores equipped in network servers is rapidly increasing, a greater number of processes or threads run concurrently. However, if these tasks invoke system calls frequently, they are not executed as concurrently as expected due to the synchronization of in-kernel data structures, cache coherence for shared data structures, and cache pollution by system calls. To address these issues, we suggest decoupling the system call context that performs network I/O from the application context and scheduling the system call contexts on a set of cores independently, which can maximize the data and instruction locality of network protocol stacks. Our design does not require any modifications of the kernel and existing applications.
Jooho Kim, Joong-Yeon Cho, Hyun-Wook Jin
ANCS3
2016 Real-Time Software Pipelining for Multidomain Motion Controllers
abstract
Motion control systems require an isochronal real-time guarantee that each control task should periodically produce outputs with no jitters. However, it is difficult to build up such a tight isochronal system with a multicore architecture and a general-purpose operating system, because the inherent resource sharing principle leads to large jitters to the control tasks. This paper proposes a software pipelining framework for an EtherCAT-based motion controller that achieves a tight isochronal guarantee with that combination. The tight guarantee is possible by multicore partitioning and reservation-aware task phasing, which reduce resource contentions between the tasks on each stage of the pipeline. Through experiments, we show that the proposed pipelining framework gives a tight isochronal guarantee with high scalability in terms of the number of motion transactions. On a real 8-axis motion control platform with two processor cores dedicated to the pipeline and a slight modification of the Linux operating system, it achieves a maximum jitter of 10 μs for four motion transactions with a common period of 1.6 ms, whereas a priority-driven method gives a maximum jitter of a few hundreds of microsecond under the same condition.
Hyeongseok Kang, Kanghee Kim, Hyun-Wook Jin
IEEE Trans. Ind. Informatics3
2015 Scalable Congestion Control Protocol Based on SDN in Data Center Networks
abstract
On-line data center applications render challenging network latency demands to meet their service level requirements. These applications, however, frequently suffer from increased latency due to the packet loss and queueing delay at the network switches. These are mainly as a result of the momentary massive bursts by the Partition/Aggregation application traffic patterns, which causes incast network congestion at the network switches. In this paper, we propose a scalable congestion control protocol, called SCCP. Our scheme effectively limits the data rate of the TCP senders by leveraging the Software Defined Networking (SDN) switches, so that the total utilization does not exceed the bottleneck link capacity. Furthermore, SCCP can be easily deployed to the existing SDN data center switches by extending the OpenFlow specifications. Our Open vSwitch-based prototype experiments and ns-3 simulations show that SCCP is scalable for up to hundreds of concurrent flows traversing through the data center network switch port.
Jae-Hyun Hwang, Joon Yoo, Hyun-Wook Jin
GLOBECOM4
2014 Towards a practical implementation of criticality mode change in RTOS
abstract
In order to address the trade-off between certification and resource efficiency, researchers are recently trying to apply a criticality mode change mechanism to mixed-criticality systems. However, the actual implementation of the criticality mode change has not been studied rigorously. In this paper, we suggest a practical design to implement the criticality mode change framework for Real-Time Operating Systems (RTOS). In particular, we aim to minimize the scheduler overheads while maximizing the resource efficiency. To the best of our knowledge, this is the first work in the literature that presents an actual implementation of the criticality mode change in RTOS.
Young-Seung Kim, Hyun-Wook Jin
ETFA2
2014 Dynamic core affinity for high-performance file upload on Hadoop Distributed File System
Joong-Yeon Cho, Hyun-Wook Jin, Min Lee, Karsten Schwan
Parallel Comput.2
2014 Resource partitioning for Integrated Modular Avionics: comparative study of implementation alternatives
abstract
Most current generation avionics systems are based on a federated architecture, where an electronic device runs a single software module or application that collaborates with other devices through a network. This architecture makes the software development process very simple, but the hardware system becomes very complicated and it is difficult to resolve issues of size, weight, and power efficiently. An integrated architecture can address the size, weight, and power issues and provide better software reusability, testability, and reliability by means of partitioning. Partitioning provides a framework that can transparently integrate several real-time applications on the same computing device, allowing the isolation of the execution environment in terms of resources and faults. Several studies on partitioning software platforms have been reported; however, to the best of our knowledge, extensive comparison and analysis of design and implementation alternatives have not been conducted owing to the extreme complexity of their implementation and measurement. In this paper, we present three design alternatives for partitioning at the user, kernel, and virtual machine monitor levels, which are compared quantitatively. In particular, we target the worldwide standard software platform for avionics systems, that is, Aeronautical Radio, Incorporated Specification 653 (ARINC 653). Overall, our study provides valuable design references and demonstrates the characteristics of design alternatives. Copyright © 2013 John Wiley & Sons, Ltd.
Hyun-Wook Jin
Softw. Pract. Exp.2
2012 A Configurable, Extensible Implementation of Inter-Partition Communication for Integrated Modular Avionics
abstract
Aerial vehicles consist of many electronic devices connected through various networks. Thus, we should be able to describe them very clearly and easily to configure network channels. It is also highly desirable to have a framework that allows adding new network devices or protocols to the existing systems while minimizing the effects on the existing software. At the same time, since there are several kinds of network protocols available, an abstraction that supports multiple protocols in a transparent manner are essential to provide the portability of avionics applications. To address these, we extend the XML-based configuration of ARINC 653 so that the description of network devices and protocols can be done very systematically. In addition, we introduce the network manager that provides a transparent abstraction over multiple networks and efficient way of adding a new network protocol without modifications of existing software. We implement our design over Ethernet, Control Area Network (CAN) and POSIX Inter-Process Communication (IPC), and show its performance in terms of communication latency and jitter.
Hyun-Wook Jin
RTCSA3
2012 Design and Implementation of a Delay-Guaranteed Motor Drive for Precision Motion Control
abstract
This paper proposes a systematic design approach for a precision-guaranteed motion control system. We develop a delay-guaranteed motor drive with our new software implementation and real-time Ethernet, which can be used as a building block to build up a multi-axis motion control system. Our drive software implementation provides a probabilistic guarantee on drive-local processing delays to motor actuation, while real-time Ethernet provides a deterministic guarantee on message communication delays from a motion control host to each drive. In the paper, we address the precision of a motion control system in two terms: host cycle time and simultaneous actuation deviation. The host cycle time is a period with which the host can periodically release motor control messages while the average drive utilization does not exceed 1, and the simultaneous actuation deviation is the difference between the earliest and the latest actuation at different drives in response to the same message. In our approach, the main objective is to minimize the periods of tasks in each drive, using our stochastic analysis, which gives us a minimum possible host cycle time. Together with an existing delay analysis of real-time Ethernet, we analyze the end-to-end delay from message release to motor actuation and in turn the simultaneous actuation deviation. Through experiments, we show that for various requirements on the deadline miss probabilities of the tasks, we can successfully reduce the host cycle time and evaluate the resulting distribution of the simultaneous actuation deviation depending on the number of drives.
Kanghee Kim, Minyoung Sung, Hyun-Wook Jin
IEEE Trans. Ind. Informatics3
2011 Temporal partitioning for mixed-criticality systems
abstract
In embedded systems, such as aerospace crafts and automobiles, it is desirable to run several real-time applications of different criticality on a single computing board by exploiting temporal partitioning. The applications, however, are usually developed by different organizations independently. Thus providing a seamless way to integrate separate applications on a control board guaranteeing real-time requirements and criticality is a very important issue. In this paper, we suggest a partition model with a mechanism that can decide each partition's period and execution time automatically preserving its criticality level. We show that the suggestion can i) prevent a task in a low-criticality partition preempting a task in a high-criticality partition (i.e., criticality inversion) and ii) provide high system throughput.
Hyun-Wook Jin
ETFA1
2011 Fieldbus virtualization for Integrated Modular Avionics
abstract
The architecture of avionics systems is moving from federated to integrated architecture called Integrated Modular Avionics (IMA). In the IMA architecture, several avionics applications can be transparently integrated on the same computing device and communicate with others without knowing whether others are running on the same node or not. However, the fieldbuses have not been studied thoroughly in the context of IMA architecture though they are quite attractive for implementing avionics data bus. In this paper, we propose a novel design of device-level virtualization to utilize Controller Area Network (CAN) efficiently on various implementations of the IMA architecture. Our proposed device-level virtualization provides virtual CAN devices and emulates the characteristics of CAN bus. We also deal with the issues to support the communication modes defined by ARINC 6531.
Jong-Seo Kim, Hyun-Wook Jin
ETFA3
2010 Isolating System Faults on Vehicular Network Gateways Using Virtualization
abstract
The traditional vehicular network gateway takes charge of communication between different internal networks and helping the electric control units in vehicle to collaborate each other. Due to the increasing requirements on innovative applications such as infotainment systems and cyber-physical systems, there are significant efforts to have an external wireless network connection on the vehicles. Accordingly, the secure architecture of the network gateway that can avoid or isolate the malicious behavior of external nodes is very critical for the next-generation vehicles. In this paper, we design a safe vehicular network gateway by exploiting full virtualization technology. Since the virtualization adds additional overheads, we try to minimize this side effect while considering the security by carefully choosing the communication mechanisms in the virtualized gateway. In our preliminary implementation, we use Virtual Box to run Linux and QNX as guest operating systems, which handles external (Wi-Fi) and internal (CAN) networks, respectively. The performance measurement results show that the virtualization-based gateway adds only 10% overhead compared with non-virtualized gateway while improving the security. We also show that the multi-core processor can leverage performance improvement.
Sung-Moon Chung, Hyun-Wook Jin
EUC2
2010 User-Level Network Protocol Stacks for Automotive Infotainment Systems
abstract
The mechanical control interfaces in automobiles are replaced rapidly by electrical interfaces enabling the x-by-wire technology. Besides the legacy automotive control parts, the infotainment systems recently come into the spotlight and become an important part of automobiles. Since the traditional automobile networks are not suitable to support these infotainment applications, new automobile network called Media Oriented System Transport (MOST) has been introduced. The MOST standard defines the network protocol stacks called Network Service. Though the industry just starts releasing new automobiles equipped with MOST-based infotainment systems, they either have yet to follow the specification of Network Service or highly dependent on the commercial implementation. In this paper, we aim to suggest a design of the Network Service protocol stacks called u-OMNiPro (user-level Open MOST Network Service Protocol Stacks), which can provide portability across different operating systems, better responsiveness, and easy interfaces to implement applications. The performance measurement results show that u-OMNiPro can run on Linux and Windows operating systems with low CPU resource requirements and high communication responsiveness.
Mu-Youl Lee, Hyun-Wook Jin
EUC2
2009 Improving TCP Goodput over Wireless Networks Using Kernel-Level Data Compression
abstract
Due to the rapid evolution of mobile processors and wireless networks, many personal mobile devices are able to support wide spectrum of Internet applications. In such systems, it is highly desirable to provide high-speed wireless communication while restraining the CPU from wasting its resources. Previous study has revealed that the bottlenecks of the communication on wireless devices are the data movements over not only the wireless link but also the bus between the memory and network controller. To overcome this performance limitation, we consider reducing the size of data traversing the bus and wireless link. In this paper, we suggest an efficient kernel-level TCP data compression scheme, which is transparent to the existing applications and can provide high-speed wireless communication. A challenging issue is that the performance gain should amortize the data compression overhead. The experimental results on realistic wireless Internet scenarios show that the modified Linux kernel can achieve better performance up to 60% and 72% than the original TCP over wireless LAN and WiMAX, respectively. Moreover, we show that the suggested scheme can save more CPU resources in spite of data compression overhead.
Moo-Yeol Lee, Hyun-Wook Jin, Ikhwan Kim, Taehyoun Kim
ICCCN2
2009 Communication Primitives for Real-Time Distributed Synchronization over Small Area Networks
abstract
Many emerging embedded systems consist of multiple embedded nodes, which are connected through an internal small area network. In such systems, the synchronization between the embedded nodes is essential to provide real-time cooperation environment. On small area networks, the severe obstacle for achieving real-time synchronization is the unpredictable overhead of the system software running on the embedded nodes. This is true especially if a traditional operating system has been installed. Despite of these drawbacks, it is quite common to run traditional operating systems on some of the embedded nodes to utilize various existing applications. In this paper, we propose a novel design of communication primitives called RTDiP-Sync for real-time distributed synchronization over small area networks. Experiment results show that the suggested design can minimize the communication overhead and provide better overhead prediction.
Hyun-Wook Jin
ISORC2
2008 Making a Case for Proactive Flow Control in Optical Circuit-Switched Networks
Mithilesh Kumar 0002, Vineeta Chaube, Pavan Balaji, Wu-chun Feng, Hyun-Wook Jin
HiPC5
2008 Designing an Efficient Kernel-Level and User-Level Hybrid Approach for MPI Intra-Node Communication on Multi-Core Systems
abstract
The emergence of multi-core processors has made MPI intra-node communication a critical component in high performance computing. In this paper, we use a three-stepmethodology to design an efficient MPI intra-node communication scheme from two popular approaches: shared memory and OS kernel-assisted direct copy. We use an Intel quad-core cluster for our study. We first run microbenchmarks to analyze the advantages and limitations of these two approaches, including the impacts of processor topology, communication buffer reuse, process skew effects, and L2 cache utilization. Based on the results and the analysis, we propose topology-aware and skew-aware thresholds to build an optimized hybrid approach. Finally, we evaluate the impact of the hybrid approach on MPI collective operations and applications using IMB, NAS, PSTSWM, and HPL benchmarks. We observe that the optimized hybrid approach can improve the performance of MPI collective operations by up to 60%, and applications by up to 17%.
Lei Chai, Ping Lai, Hyun-Wook Jin, Dhabaleswar K. Panda 0001
ICPP3
2007 Lightweight kernel-level primitives for high-performance MPI intra-node communication over multi-core systems
abstract
Modern processors have multiple cores on a chip to overcome power consumption and heat dissipation issues. As more and more compute cores become available on a single node, it is expected that node-local communication will play an increasingly greater role in overall performance of parallel applications such as MPI applications. It is therefore crucial to optimize intra-node communication paths utilized by MPI libraries. In this paper, we propose a novel design of a kernel extension, called LiMIC2, for high-performance MPI intra-node communication over multi-core systems. LiMIC2 can minimize the communication overheads by implementing lightweight primitives and provide portability across different interconnects and flexibility for performance optimization. Our performance evaluation indicates that LiMIC2 can attain 80% lower latency and more than three times improvement in bandwidth. Also the experimental results show that LiMIC2 can deliver bidirectional bandwidth greater than 11GB/s.
Hyun-Wook Jin, Sayantan Sur, Lei Chai, Dhabaleswar K. Panda 0001
CLUSTER1
2007 Impact of protocol overheads on network throughput over high-speed interconnects: measurement, analysis, and improvement
Hyun-Wook Jin, Chuck Yoo
J. Supercomput.1
2006 Design of High Performance MVAPICH2: MPI2 over InfiniBand
abstract
MPICH2 provides a layered architecture for implementing MPI-2. In this paper, we provide a new design for implementing MPI-2 over InfiniBand by extending the MPICH2 ADI3 layer. Our new design aims to achieve high performance by providing a multi-communication method framework that can utilize appropriate communication channels/devices to attain optimal performance without compromising on scalability and portability. We also present the performance comparison of the new design with our previous design based on the MPICH2 RDMA channel. We show significant performance improvements in micro-benchmarks and NAS Parallel Benchmarks.
Wei Huang 0003, Gopalakrishnan Santhanaraman, Hyun-Wook Jin, Qi Gao 0004, Dhabaleswar K. Panda 0001
CCGRID3
2006 Designing Efficient Cooperative Caching Schemes for Multi-Tier Data-Centers over RDMA-enabled Networks
abstract
Caching has been a very important technique in improving the performance and scalability of web-serving datacenters. The research community has proposed cooperation of caching servers to achieve higher performance benefits. These existing cooperative caching mechanisms often partially duplicate the cached data redundantly on multiple servers for higher performance (by optimizing the datafetch costs for multiple similar requests). With the advent of RDMA enabled interconnects these basic data-fetch cost estimates have changed significantly. Further, the effective utilization of the vast resources available across multiple tiers in today’s data-centers is of obvious interest. Hence, a systematic study of these various issues involved is of paramount importance. In this paper, we present several cooperative caching schemes that are designed to benefit in the light of the above mentioned trends. In particular, we design schemes that take advantage of the RDMA capabilities of networks and the multitude of resources available in modern multi-tier data-centers. Our designs are implemented on InfiniBand based clusters to work in conjunction with Apache based servers. Our experimental results show that our schemes achieve a throughput improvement of up to 35% as compared to the basic cooperative caching schemes and 180% better than the simple single node caching schemes. Our experimental results lead us to a new scheme which can deliver good performance in many Caching has been a very important technique in improving the performance and scalability of web-serving datacenters. The research community has proposed cooperation of caching servers to achieve higher performance benefits. These existing cooperative caching mechanisms often partially duplicate the cached data redundantly on multiple servers for higher performance (by optimizing the datafetch costs for multiple similar requests). With the advent of RDMA enabled interconnects these basic data-fetch cost estimates have changed significantly. Further, the effective utilization of the vast resources available across multiple tiers in today’s data-centers is of obvious interest. Hence, a systematic study of these various issues involved is of paramount importance. In this paper, we present several cooperative caching schemes that are designed to benefit in the light of the above mentioned trends. In particular, we design schemes that take advantage of the RDMA capabilities of networks and the multitude of resources available in modern multi-tier data-centers. Our designs are implemented on InfiniBand based clusters to work in conjunction with Apache based servers. Our experimental results show that our schemes achieve a throughput improvement of up to 35% as compared to the basic cooperative caching schemes and 180% better than the simple single node caching schemes. Our experimental results lead us to a new scheme which can deliver good performance in many scenarios.
Sundeep Narravula, Hyun-Wook Jin, Karthikeyan Vaidyanathan, Dhabaleswar K. Panda 0001
CCGRID2
2006 Exploiting RDMA operations for Providing Efficient Fine-Grained Resource Monitoring in Cluster-based Servers
abstract
Efficiently capturing the resource usage in a shared server environment has been a critical research issue in the past several years. With the amount of resources used by each application becoming more and more divergent and unpredictable, the solution to this problem is becoming increasingly important. In the past, several researchers have come up with a number of techniques which rely on coarse-grained monitoring of resources in order to avoid the overheads associated with fine-grained monitoring. In this paper, we propose a low-overhead efficient fine-grained resource monitoring scheme using the advanced Remote Direct Memory Access (RDMA) operation provided by RDMA-enabled interconnects such as InfiniBand (IBA). We evaluate the relative benefits of our approach against traditional approaches in various environments (including micro-benchmarks as well as real applications such as an auction server based on the RUBiS benchmark and the Ganglia distributed monitoring tool). Our results indicate that our approach for fine-grained monitoring can significantly improve the overall system utilization, thereby resulting in up to 25% improvement in the number of requests the cluster-system can admit
Karthikeyan Vaidyanathan, Hyun-Wook Jin, Dhabaleswar K. Panda 0001
CLUSTER2
2006 NemC: A Network Emulator for Cluster-of-Clusters
abstract
A large number of clusters are being used in all different organizations such as universities, laboratories, etc. These clusters are, however, usually independent from each other even in the same organization or building. To provide a single image of such clusters to users and utilize them in an integrated manner, cluster-of-clusters has been suggested. However, since research groups usually do not have the actual backbone networks for cluster-of-clusters, which can be reconfigured with respect to delay, packet loss, etc. as needed, it is not feasible to carry out practical research over realistic environments. Accordingly, the demand for an efficient way to emulate the backbone networks for cluster-of-clusters is overreaching. In this paper, we suggest a novel design for emulating the backbone networks of cluster-of-clusters. The emulator named NemC can support the fine-grained network delay resolution minimizing the additional overheads. The experimental results show that NemC can emulate the low delay and high bandwidth backbone networks more accurately than existing emulators such as NISTNet and NetEm. We also present a case study showing the performance of MPI applications over cluster-of-clusters environment using NemC.
Hyun-Wook Jin, Sundeep Narravula, Karthikeyan Vaidyanathan, Dhabaleswar K. Panda 0001
ICCCN1
2006 Asynchronous zero-copy communication for synchronous sockets in the sockets direct protocol (SDP) over InfiniBand
abstract
Sockets direct protocol (SDP) is an industry standard pseudo sockets-like implementation to allow existing sockets applications to directly and transparently take advantage of the advanced features of current generation networks such as InfiniBand. The SDP standard supports two kinds of sockets semantics, viz., synchronous sockets (e.g., used by Linux, BSD, Windows) and asynchronous sockets (e.g., used by Windows, upcoming support in Linux). Due to the inherent benefits of asynchronous sockets, the SDP standard allows several intelligent approaches such as source-avail and sink-avail based zero-copy for these sockets. Unfortunately, most of these approaches are not beneficial for the synchronous sockets interface. Further, due to its portability, ease of use and support on a wider set of platforms, the synchronous sockets interface is the one used by most sockets applications today. Thus, a mechanism by which the approaches proposed for asynchronous sockets can be used for synchronous sockets is highly desirable. In this paper, we propose one such mechanism, termed as AZ-SDP (asynchronous zero-copy SDP), where we memory-protect application buffers and carry out communication asynchronously while maintaining the synchronous sockets semantics. We present our detailed design in this paper and evaluate the stack with an extensive set of benchmarks. The experimental results demonstrate that our approach can provide an improvement of close to 35% for medium-message unidirectional throughput and up to a factor of 2 benefit for computation-communication overlap tests and multi-connection benchmarks
Pavan Balaji, Sitha Bhagvat, Hyun-Wook Jin, Dhabaleswar K. Panda 0001
IPDPS3
2006 Designing next generation data-centers with advanced communication protocols and systems services
abstract
Current data-centers rely on TCP/IP over fast- and gigabit-Ethernet for data communication even within the cluster environment for cost-effective designs, thus limiting their maximum capacity. Together with raw performance, such data-centers also lack in efficient support for intelligent services, such as requirements for caching documents, managing limited physical resources, load-balancing, controlling overload scenarios, and prioritization and QoS mechanisms, that are becoming a common requirement today. On the other hand, the system area network (SAN) technology is making rapid advances during the recent years. Besides high performance, these modern interconnects are providing a range of novel features and their support in hardware (e.g., RDMA, atomic operations, QoS support). In this paper, we address the capabilities of these current generation SAN technologies in addressing the limitations of existing data-centers. Specifically, we present a novel framework comprising of three layers (communication protocol support, data-center service primitives and advanced data-center services) that work together to tackle the issues associated with existing data-centers. We also present preliminary results in the various aspects of the framework, which demonstrate close to an order of magnitude performance benefits achievable by our framework as compared to existing data-centers in several cases.
Pavan Balaji, Karthikeyan Vaidyanathan, Sundeep Narravula, Hyun-Wook Jin, Dhabaleswar K. Panda 0001
IPDPS4
2006 Efficient SMP-aware MPI-level broadcast over InfiniBand's hardware multicast
abstract
Most of the high-end computing clusters found today feature multi-way SMP nodes interconnected by an ultra-low latency and high bandwidth network. InfiniBand is emerging as a high-speed network for such systems. InfiniBand provides a scalable and efficient hardware multicast primitive to efficiently implement many MPI collective operations. However, employing hardware multicast as the communication method may not perform well in all cases. This is true especially when more than one process is running per node. In this context, shared memory channel becomes the desired communication medium within the node as it delivers latencies which are of an order of magnitude lower than the inter-node message latencies. Thus, to deliver optimal collective performance, coupling hardware multicast with shared memory channel becomes necessary. In this paper we propose mechanisms to address this issue. On a 16-node 2-way SMP cluster, the Leader-based scheme proposed in this paper improves the performance of the MPI/spl I.bar/Bcast operation by a factor of as much as 2.3 and 1.8 when compared to the point-to-point and original solution employing only hardware multicast. We have also evaluated our designs on NUMA based system and obtained a performance improvement of 1.7 using our designs on 2-node 4-way system. We also propose a dynamic attach policy as an enhancement to this scheme to mitigate the impact of process skew on the performance of the collective operation.
Amith R. Mamidala, Lei Chai, Hyun-Wook Jin, Dhabaleswar K. Panda 0001
IPDPS3
2006 Shared receive queue based scalable MPI design for InfiniBand clusters
abstract
Clusters of several thousand nodes interconnected with InfiniBand, an emerging high-performance interconnect, have already appeared in the Top 500 list. The next-generation InfiniBand clusters are expected to be even larger with tens-of-thousands of nodes. A high-performance scalable MPI design is crucial for MPI applications in order to exploit the massive potential for parallelism in these very large clusters. MVAPICH is a popular implementation of MPI over InfiniBand based on its reliable connection oriented model. The requirement of this model to make communication buffers available for each connection imposes a memory scalability problem. In order to mitigate this issue, the latest InfiniBand standard includes a new feature called shared receive queue (SRQ) which allows sharing of communication buffers across multiple connections. In this paper, we propose a novel MPI design which efficiently utilizes SRQs and provides very good performance. Our analytical model reveals that our proposed designs take only 1/10/sup th/ the memory requirement as compared to the original design on a cluster sized at 16,000 nodes. Performance evaluation of our design on our 8-node cluster shows that our new design was able to provide the same performance as the existing design while requiring much lesser memory. In comparison to tuned existing designs our design showed a 20% and 5% improvement in execution time of NAS Benchmarks (Class A) LU and SP, respectively. The high performance Linpack was able to execute a much larger problem size using our new design, whereas the existing design ran out of memory.
Sayantan Sur, Lei Chai, Hyun-Wook Jin, Dhabaleswar K. Panda 0001
IPDPS3
2006 RDMA read based rendezvous protocol for MPI over InfiniBand: design alternatives and benefits
abstract
Message Passing Interface (MPI) is a popular parallel programming model for scientific applications. Most high-performance MPI implementations use Rendezvous Protocol for efficient transfer of large messages. This protocol can be designed using either RDMA Write or RDMA Read. Usually, this protocol is implemented using RDMA Write. The RDMA Write based protocol requires a two-way handshake between the sending and receiving processes. On the other hand, to achieve low latency, MPI implementations often provide a polling based progress engine. The two-way handshake requires the polling progress engine to discover multiple control messages. This in turn places a restriction on MPI applications that they should call into the MPI library to make progress. For compute or I/O intensive applications, it is not possible to do so. Thus, most communication progress is made only after the computation or I/O is over. This hampers the computation to communication overlap severely, which can have a detrimental impact on the overall application performance. In this paper, we propose several mechanisms to exploit RDMA Read and selective interrupt based asynchronous progress to provide better computation/communication overlap on InfiniBand clusters. Our evaluations reveal that it is possible to achieve nearly complete computation/communication overlap using our RDMA Read with Interrupt based Protocol. Additionally, our schemes yield around 50% better communication progress rate when computation is overlapped with communication. Further, our application evaluation with Linpack (HPL) and NAS-SP (Class C) reveals that MPI_Wait time is reduced by around 30% and 28%, respectively, for a 32 node InfiniBand cluster. We observe that the gains obtained in the MPI_Wait time increase as the system size increases. This indicates that our designs have a strong positive impact on scalability of parallel applications.
Sayantan Sur, Hyun-Wook Jin, Lei Chai, Dhabaleswar K. Panda 0001
PPoPP2
2005 Architecture for caching responses with multiple dynamic dependencies in multi-tier data-centers over InfiniBand
abstract
It has been well acknowledged in the research community that in order to design a data-center environment which is efficient and offers high performance, one of the critical issues that needs to be addressed is the effective reuse of cache content stored away from the origin server. However, for caching dynamically changing content (e.g., content involved in online banking, Internet auctions, etc.). consistency and coherency issues need to be addressed. In addition, most current real world requests have multiple dynamic dependencies, i.e., these requests might depend on multiple data objects. Further, these requests are not entirely independent; several requests might have common dependencies. While there have been previous research solutions on maintaining coherent caches for dynamic content, these solutions have several shortcomings including inability to adapt to server load or handle multiple dynamic dependencies. In this paper, we propose a load resilient architecture using one sided operations supported by several high performance interconnects such as InfiniBand, while maintaining multiple dynamic dependencies per response. Our experimental results show that our schemes to tackle the multi-dependency issue efficiently and significantly outperform the existing approaches. Further, our results demonstrate that the proposed load resilient architecture can possibly improve the performance of loaded data-centers by over an order of magnitude.
Sundeep Narravula, Pavan Balaji, Karthikeyan Vaidyanathan, Hyun-Wook Jin, Dhabaleswar K. Panda 0001
CCGRID4
2005 Supporting iWARP Compatibility and Features for Regular Network Adapters
abstract
With several recent initiatives in the protocol offloading technology present on network adapters, the user market is now distributed amongst various technology levels including regular Ethernet network adapters, TCP Offload Engines (TOEs) and the recently introduced iWARP-capable networks. While iWARP-capable networks provide all the features provided by their predecessors (TOEs and regular Ethernet network adapters) and a new richer programming interface, they lack with respect to backward compatibility. In this aspect, two important issues need to be considered. First, not all network adapters support iWARP; thus, software compatibility for regular network adapters (which have no offloaded protocol stack) with iWARP capable network adapters needs to be achieved. Second, several applications on top of regular Ethernet as well as TOE based adapters have been written with the sockets interface; rewriting such applications using the new iWARP interface is cumbersome and impractical. Thus, it is desirable to have an interface which provides a two-fold benefit: (i) it allows existing applications to run directly without any modifications and (ii) it exposes the richer feature set of iWARP to the applications to be utilized with minimal modifications. In this paper, we design and implement a software stack to handle these issues. Specifically, (i) the software stack emulates the functionality of the iWARP stack in software to provide compatibility for regular Ethernet adapters with iWARP capable networks and (ii) it provides applications with an extended sockets interface that provides the traditional sockets functionality as well as functionality extended with the rich iWARP features
Pavan Balaji, Hyun-Wook Jin, Karthikeyan Vaidyanathan, Dhabaleswar K. Panda 0001
CLUSTER2
2005 High Performance RDMA Based All-to-All Broadcast for InfiniBand Clusters
Sayantan Sur, Uday Bondhugula, Amith R. Mamidala, Hyun-Wook Jin, Dhabaleswar K. Panda 0001
HiPC4
2005 Supporting MPI-2 One Sided Communication on Multi-rail InfiniBand Clusters: Design Challenges and Performance Benefits
Abhinav Vishnu, Gopalakrishnan Santhanaraman, Wei Huang 0003, Hyun-Wook Jin, Dhabaleswar K. Panda 0001
HiPC4
2005 LiMIC: Support for High-Performance MPI Intra-node Communication on Linux Cluster
abstract
High performance intra-node communication support for MPI applications is critical for achieving best performance from clusters of SMP workstations. Present day MPI stacks cannot make use of operating system kernel support for intra-node communication. This is primarily due to the lack of an efficient, portable, stable and MPI friendly interface to access the kernel functions. In this paper we attempt to address design challenges for implementing such a high performance and portable kernel module interface. We implement a kernel module interface called LiMIC and integrate it with MVAPICH, an open source MPI over InfiniBand. Our performance evaluation reveals that the point-to-point latency can be reduced by 71% and the bandwidth improved by 405% for 64 KB message size. In addition, LiMIC can improve HPCC effective bandwidth and NAS IS class B benchmarks by 12% and 8%, respectively, on an 8-node dual SMP InfiniBand cluster.
Hyun-Wook Jin, Sayantan Sur, Lei Chai, Dhabaleswar K. Panda 0001
ICPP1
2005 On the provision of prioritization and soft qos in dynamically reconfigurable shared data-centers over infiniband
abstract
In the past few years several researchers have proposed and configured data-centers providing multiple independent services, known as shared data-centers. For example, several ISPs and other Web service providers host multiple unrelated Web-sites on their data-centers allowing potential differentiation in the service provided to each of them. Such differentiation becomes essential in several scenarios in a shared data-center environment. In this paper, we extend our previously proposed scheme on dynamic re-configurability to allow service differentiation in the shared data-center environment. In particular, we point out the issues associated with the basic dynamic configurability scheme and propose two extensions to it, namely (i) dynamic reconfiguration with prioritization and (ii) dynamic reconfiguration with prioritization and QoS. Our experimental results show that our extensions can allow the dynamic reconfigurability scheme to attain a performance improvement of up to five times for high priority Web sites irrespective of any background low priority requests. Also, these extensions are able to significantly improve the performance of low priority requests when there are minimal or no high priority requests in the system. Further, they can achieve a similar performance as a static scheme with up to 43% lesser nodes in some cases
Pavan Balaji, Sundeep Narravula, Karthikeyan Vaidyanathan, Hyun-Wook Jin, Dhabaleswar K. Panda 0001
ISPASS4
2005 Exploiting NIC architectural support for enhancing IP-based protocols on high-performance networks
Hyun-Wook Jin, Pavan Balaji, Chuck Yoo, Dhabaleswar K. Panda 0001
J. Parallel Distributed Comput.1
2004 High performance MPI-2 one-sided communication over InfiniBand
abstract
Many existing MPI-2 one-sided communication implementations are built on top of MPI send/receive operations. Although this approach can achieve good portability, it suffers front high communication overhead and dependency on remote process for communication progress. To address these problems, we propose a high performance MPI-2 one-sided communication design over the InfiniBand Architecture. In our design, MPI-2 one-sided communication operations such as MPI-Put, MPI-Get and MPI-Accumulate are directly mapped to InfiniBand Remote Direct Memory Access (RDMA) operations. Our design has been implemented based on MPICH2 over InfiniBand. We present detailed design issues for this approach and perform a set of microbenchmarks to characterize different aspects of its performance. Our performance evaluation shows that compared with the design based on MPI send/receive, our design can improve throughput up to 77%, and reduce latency and synchronization overhead up to 19% and 13%, respectively. Under certain process skew, the bad impact can be significantly reduced by new design, from 41% to nearly 0%. It also can achieve better overlap of communication and computation.
Weihang Jiang, Jiuxing Liu, Hyun-Wook Jin, Dhabaleswar K. Panda 0001, William Gropp, Rajeev Thakur
CCGRID3
2004 NIC-based offload of dynamic user-defined modules for Myrinet clusters
abstract
Many of the modern networks used to interconnect nodes in cluster-based computing systems provide network-interface cards (NICs) that offer programmable processors. Substantial research has been done with the focus of offloading processing from the host to the NIC processor. However, the research has primarily focused on the static offload of specific features to the NIC, mainly to support the optimization of common collective and synchronization-based communications. We describe the design and implementation of a framework based on MP1CH-GM to support the dynamic NIC-based offload of user-defined modules for Myrinet clusters. We evaluate our implementation on a 16-node cluster using a NIC-based version of the common broadcast operation and we find a maximum factor of improvement of 1.2 with respect to total latency as well as a maximum factor of improvement of 2.2 with respect to average CPU utilization under conditions of process skew. In addition, we see that these improvements increase with system size, indicating that our NIC-based framework offers enhanced scalability when compared to a purely host-based approach.
Adam Wagner, Hyun-Wook Jin, Dhabaleswar K. Panda 0001, Rolf Riesen
CLUSTER2
2004 Efficient and Scalable All-to-All Personalized Exchange for InfiniBand-Based Clusters
abstract
The all-to-all personalized exchange is the most dense collective communication function offered by the MPI specification. The operation involves every process sending a different message to all other participating processes. This collective operation is essential for many parallel scientific applications. With increasing system and message sizes, it becomes challenging to offer a fast, scalable and efficient implementation of this operation. InfiniBand is an emerging modern interconnect. It offers very low latency, high bandwidth and one-sided operations like RDMA write. Its advanced features like RDMA write gather allow us to design and implement all-to-all algorithms much more efficiently than in the past. Our aim in This work is to design efficient and scalable implementations of traditional personalized exchange algorithms. We present two novel approaches towards designing all-to-all algorithms for short and long messages respectively. The hypercube RDMA write gather and direct eager schemes effectively leverage the RDMA and RDMA with write gather mechanisms offered by InfiniBand. Performance evaluation of our design and implementation reveals that it is able to reduce the all-to-all communication time by upto a factor of 3.07 for 32 byte messages on a 16 node InfiniBand cluster. Our analytical models suggest that the proposed designs perform 64% better on InfiniBand clusters with 1024 nodes for 4k message size.
Sayantan Sur, Hyun-Wook Jin, Dhabaleswar K. Panda 0001
ICPP2
2003 Firmware-Level Latency Analysis on a Gigabit Network
Hyun-Wook Jin, Chuck Yoo
J. Supercomput.1
2002 Stepwise Optimizations of UDP/IP on a Gigabit Network (Research Note)
Hyun-Wook Jin, Chuck Yoo, Sung-Kyun Park
Euro-Par1
1999 Latency analysis of UDP and BPI on Myrinet
abstract
High-speed networks such as ATM, Myrinet, and Gigabit Ethernet are available today, and many researchers make efforts to enhance the performance of end-to-end communication on these high-speed networks. One of the efforts is to develop new light-weight communication primitives for high-speed network. However the latency of the new primitives has not been characterized thoroughly, partly because existing measurement methodologies do not take into account the features of high-speed networks. Therefore, there are only incomplete comparisons of the new primitives and traditional protocols, and they cannot really prove the usefulness of new primitives. In order to address this issue, this paper suggests a new measurement methodology and uses the methodology to perform a detailed latency analysis of UDP and a light-weight primitive, called BPI, on Myrinet. Our results clearly show the difference of per-byte overhead between BPI and UDP. A surprising result is that BPI is found to be slower than UDP for 4KB or larger data size.
Hyun-Wook Jin, Chuck Yoo
IPCCC1