Kun Suo

dblp:184/5707 · DBLP profile ↗
← Back
29ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0001-8562-0492ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 4 first-author · 8 since 2021Computer networks · 8 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Quantum-Enhanced Multimodal Generative AI for Early Disease Detection and Personalized Treatment in Healthcare
Gayathri Priya Kolavennu, Yong Shi 0002, Kun Suo, Tu N. Nguyen 0001
IEA/AIE (2)3
2026 Promoting Quantum-based Machine Learning through Multifaceted Activities
abstract
Quantum-based machine learning (QML) is known to integrate quantum computing (QC) and machine learning (ML) to boost performance in data analysis and decision-making. The demand for professionals in QML has increased; however, QML is not part of the curricula for most colleges and universities. In this project, we create labware in Google Colab that encompasses knowledge of QC, ML, and QML, and demonstrates QML applications for analyzing scientific and engineering data. We integrate our learning modules into courses in Computer Science, Industrial Engineering, Physics, and other disciplines, and host a faculty workshop, a student camp, and a conference tutorial session to promote the use of our learning modules. The feedback from participating faculty and students is overwhelmingly positive.
Yong Shi 0002, Dan Chia-Tien Lo, Valentina Nino, Hongmei Chi, Kun Suo, Tu N. Nguyen 0001
SIGCSE (2)5
2025 A Sparsity Predicting Approach for Large Language Models via Activation Pattern Clustering
Nobel Dhar, Bobin Deng, Md. Romyull Islam, Xinyue Zhang 0001, Kazi Fahim Ahmad Nasif, Kun Suo
Euro-Par (1)6
2025 AccelMD: A Self-Adaptive AI-Enabled Framework for Accelerating Molecular Dynamics Simulations
abstract
Molecular Dynamics (MD) simulation is a fundamental exploring approach for numerous scientific fields, including but not limited to drug discovery, biology, material science, and chemistry. Unfortunately, large-scale MD simulation generally requires long-period processing, even in resource-intensive supercomputers. Parallel computing is the primary methodology to accelerate MD simulation. However, parallel computing has a low theoretical improvable ceiling for MD simulation because the atom statuses of one timestep depend on its previous timestep. To unlock the full potential of MD simulation in drug discovery and biology, we propose an AccelMD framework for reducing processing time and computing costs. AccelMD is an AI-enabled framework that accelerates MD simulation while maintaining high predictive accuracy. In order to extend AccelMD's applicability and flexibility, we also developed an extra preprocessing stage, which allows the framework to adapt to various protein inputs automatically. Compared to conventional MD, the AccelMD framework achieves 47.59 X speedup on average (max: 63.78 X, min: 30.61 X). Regarding the prediction accuracy of protein structures, the average Mean Absolute Error (MAE) and TMScore values of AccelMD are 0.0058 and 0.99980, respectively. Our empirical experiments exihibit the robustness of AccelMD framework for the study of complex, long-timescale molecular interactions.
Kazi Fahim Ahmad Nasif, Bobin Deng, Lingtao Chen, Chloe Yixin Xie, Shaolei Teng, Liang Zhao 0024, Syed Md Shamsul Alam, Nobel Dhar, Kun Suo, Dan Chia-Tien Lo
ICTAI10
2025 Assessing and Visualizing Completeness, Co-Coverage, and Scalability in Multivariate Time-Series Data
abstract
Assessing data quality in multivariate time-series datasets is crucial for reliable analysis, particularly when dealing with missing values, inconsistent feature availability, and massive records in large-scale edge computing and IoT clusters. Existing methods often fall short of capturing intricate patterns of missingness and co-coverage, restricting the capacity to make well-informed decisions regarding the usability of the data. In order to systematically extract reliable data segments, this paper presents a comprehensive framework that combines a heuristic model with temporal coverage, period-specific missingness, and co-coverage metrics. By integrating these metrics with visualizations such as temporal coverage heatmaps and parallel coordinates plots, the framework reveals complex patterns of missingness while supporting human involvement in validating data subsets. Our approach effectively balances automation with expert judgment, enhancing the interpretability of data quality assessments. The findings show that the proposed methods satisfy the design specifications for revealing patterns, quantifying missingness impact, measuring feature availability, guiding feature selection, and facilitating scalable, multi-scale data summarization. The framework offers a solid way to improve the quality of data in multivariate time-series analysis, opening the door to more precise and trustworthy insights for assessing data gathered from edge computing infrastructures and large-scale, heterogeneous IoT deployments, where data consistency and completeness are frequently very variable.
Long Vu, Madeline Frank, Honghui Xu 0001, Sisi Chen, Tu N. Nguyen 0001, Selena He, Bobin Deng, Kun Suo
IPCCC8
2025 Characterizing and Understanding Energy Footprint and Efficiency of Small Language Model on Edges
abstract
Cloud-based large language models (LLMs) and their variants have significantly influenced real-world applications. Deploying smaller models (i.e., small language models (SLMs)) on edge devices offers additional advantages, such as reduced latency and independence from network connectivity. However, edge devices’ limited computing resources and constrained energy budgets challenge efficient deployment. This study evaluates the power efficiency of five representative SLMs — Llama 3.2, Phi-3 Mini, TinyLlama, and Gemma 2 on Raspberry Pi 5, Jetson Nano, and Jetson Orin Nano (CPU and GPU configurations). Results show that Jetson Orin Nano with GPU acceleration achieves the highest energy-to-performance ratio, significantly outperforming CPU-based setups. Llama 3.2 provides the best balance of accuracy and power efficiency, while TinyLlama is well-suited for low-power environments at the cost of reduced accuracy. In contrast, Phi-3 Mini consumes the most energy despite its high accuracy. In addition, GPU acceleration, memory bandwidth, and model architecture are key in optimizing inference energy efficiency. Our empirical analysis offers practical insights for AI, smart systems, and mobile ad-hoc platforms to leverage tradeoffs from accuracy, inference latency, and power efficiency in energy-constrained environments.
Md. Romyull Islam, Bobin Deng, Nobel Dhar, Tu N. Nguyen 0001, Selena He, Yong Shi 0002, Kun Suo
MASS7
2025 Three-stage Learning with Portable Online Hands-on Labware for Quantum-based Machine Learning Development
abstract
Quantum-based Machine Learning (QML) combines quantum computing (QC) with machine learning (ML), which can be applied in various sectors, and there is a high demand for QML professionals. However, QML is not yet in many schools' curricula. We design labware for the basic concepts of QC, ML, and QML and their applications in science and engineering fields in Google Colab, applying a three-stage learning strategy for efficient and effective student learning.
Yong Shi 0002, Dan Chia-Tien Lo, Hongmei Chi, Andrew Polisetty, Kun Suo, Tu N. Nguyen 0001
SIGCSE (2)5
2024 FTBC: Forward Temporal Bias Correction for Optimizing ANN-SNN Conversion
Xiaofeng Wu 0002, Velibor Bojkovic, Bin Gu 0001, Kun Suo
ECCV (69)4
2024 Activation Sparsity Opportunities for Compressing General Large Language Models
abstract
Deploying local AI models, such as Large Language Models (LLMs), to edge devices can substantially enhance devices’ independent capabilities, alleviate the server’s burden, and lower the response time. Owing to these tremendous potentials, many big tech companies have been actively promoting edge LLM evolution and released several lightweight Small Language Models (SLMs) to bridge this gap. However, SLMs currently only work well on limited real-world applications. We still have huge motivations to deploy more powerful (larger-scale) AI models on edge devices and enhance their smartness level. Unlike the conventional approaches for AI model compression, we investigate from activation sparsity. The activation sparsity method is orthogonal and combinable with existing techniques to maximize compression rate while maintaining great accuracy. According to statistics of open-source LLMs, their Feed-Forward Network (FFN) components typically comprise a large proportion of parameters (around $\tfrac{2}{3}$). This internal feature ensures that our FFN optimizations would have a better chance of achieving effective compression. Moreover, our findings are beneficial to general LLMs and are not restricted to ReLU-based models.This work systematically investigates the tradeoff between enforcing activation sparsity and perplexity (accuracy) on state-of-the-art LLMs. Our empirical analysis demonstrates that we can obtain around 50% of main memory and computing reductions for critical FFN components with negligible accuracy degradation. This extra 50% sparsity does not naturally exist in the current LLMs, which require tuning LLMs’ activation outputs by injecting zero-enforcing thresholds. To obtain the benefits of activation sparsity, we provide a guideline for the system architect for LLM prediction and prefetching. Moreover, we further verified the predictability of activation patterns in recent LLMs. The success prediction allows the system to prefetch the necessary weights while omitting the inactive ones and their successors (compress models from the memory’s perspective), therefore lowering cache/memory pollution and reducing LLM execution time on resource-constraint edge devices.
Nobel Dhar, Bobin Deng, Md. Romyull Islam, Kazi Fahim Ahmad Nasif, Liang Zhao 0024, Kun Suo
IPCCC6
2024 Characterizing and Understanding the Performance of Small Language Models on Edge Devices
abstract
In recent years, significant advancements in computing power, data richness, algorithmic development, and the growing demand for applications have catalyzed the rapid emergence and proliferation of large language models (LLMs) across various scenarios. Concurrently, factors such as computing resource limitations, cost considerations, real-time application requirements, task-specific customization, and privacy concerns have also driven the development and deployment of small language models (SLMs). Unlike extensively researched and widely deployed LLMs in the cloud, the performance of SLM workloads and their resource impact on edge environments remain poorly understood. More detailed studies will have to be carried out to understand the advantages, constraints, performances, and resource consumption in different settings of the edge.This paper addresses this gap by comprehensively analyzing representative SLMs on edge platforms. Initially, we provide a summary of contemporary edge hardware and popular SLMs. Subsequently, we quantitatively evaluate several widely used SLMs, including TinyLlama, Phi-3, Llama-3, etc., on popular edge platforms such as Raspberry Pi, Nvidia Jetson Orin, and Mac mini. Our findings reveal that the interaction between different hardware and SLMs can significantly impact edge AI workloads while introducing non-negligible overhead. Our experiments demonstrate that variations in performance and resource usage might constrain the workload capabilities of specific models and their feasibility on edge platforms. Therefore, users must judiciously match appropriate hardware and models based on the requirements and characteristics of the edge environment to avoid performance bottlenecks and optimize the utility of edge computing capabilities.
Md. Romyull Islam, Nobel Dhar, Bobin Deng, Tu N. Nguyen 0001, Selena He, Kun Suo
IPCCC6
2024 Fixed-point Encoding and Architecture Exploration for Residue Number Systems
abstract
Residue Number Systems (RNS) demonstrate the fascinating potential to serve integer addition/ multiplication-intensive applications. The complexity of Artificial Intelligence (AI) models has grown enormously in recent years. From a computer system’s perspective, ensuring the training of these large-scale AI models within an adequate time and energy consumption has become a big concern. Matrix multiplication is a dominant subroutine in many prevailing AI models, with an addition/multiplication-intensive attribute. However, the data type of matrix multiplication within machine learning training typically requires real numbers, which indicates that RNS benefits for integer applications cannot be directly gained by AI training. The state-of-the-art RNS real-number encodings, including floating-point and fixed-point, have defects and can be further enhanced. To transform default RNS benefits to the efficiency of large-scale AI training, we propose a low-cost and high-accuracy RNS fixed-point representation: Single RNS Logical Partition (S-RNS-Logic-P) representation with Scaling-down Postprocessing Multiplication (SD-Post-Mul) . Moreover, we extend the implementation details of the other two RNS fixed-point methods: Double RNS Concatenation and S-RNS-Logic-P representation with Scaling-down Preprocessing Multiplication . We also design the architectures of these three fixed-point multipliers. In empirical experiments, our S-RNS-Logic-P representation with SD-Post-Mul method achieves less latency and energy overhead while maintaining good accuracy. Furthermore, this method can easily extend to the Redundant Residue Number System to raise the efficiency of error-tolerant domains, such as improving the error correction efficiency of quantum computing.
Bobin Deng, Bhargava Nadendla, Kun Suo, Chloe Yixin Xie, Dan Chia-Tien Lo
ACM Trans. Archit. Code Optim.3
2024 Incendio: Priority-Based Scheduling for Alleviating Cold Start in Serverless Computing
abstract
In serverless computing, cold start results in long response latency. Existing approaches strive to alleviate the issue by reducing the number of cold starts. However, our measurement based on real-world production traces shows that the minimum number of cold starts does not equate to the minimum response latency, and solely focusing on optimizing the number of cold starts will lead to sub-optimal performance. The root cause is that functions have different priorities in terms of latency benefits by transferring a cold start to a warm start. In this paper, we proposeIncendio, a serverless computing framework exploiting priority-based scheduling to minimize the overall response latency from the perspective of cloud providers. We reveal the priority of a function is correlated to multiple factors and design a priority model based on Spearman’s rank correlation coefficient. We integrate a hybrid Prophet-LightGBM prediction model to dynamically manage runtime pools, which enables the system to prewarm containers in advance and terminate containers at the appropriate time. Furthermore, to satisfy the low-cost and high-accuracy requirements in serverless computing, we propose a Clustered Reinforcement Learning-based function scheduling strategy. The evaluations show that Incendio speeds up the native system by 1.4×, and achieves 23% and 14.8% latency reductions compared to two state-of-the-art approaches.
Xinquan Cai, Qianlong Sang, Chuang Hu, Yili Gong, Kun Suo, Xiaobo Zhou 0002, Dazhao Cheng
IEEE Trans. Computers5
2024 vKernel: Enhancing Container Isolation via Private Code and Data
abstract
Container technology is increasingly adopted in cloud environments. However, the lack of isolation in the shared kernel becomes a significant barrier to the wide adoption of containers. The challenges lie in how to simultaneously attain high performance and isolation. On the one hand, kernel-level isolation mechanisms, such asseccomp,capabilities, andapparmor, achieve good performance without much overhead, but lack the support for per-container customization. On the other hand, user-level and VM-based isolation offer superior security guarantees and allow for customization since a container is assigned a dedicated kernel, however, at the cost of high overhead. We presentvKernel, a kernel isolation framework. It maintains a minimal set of code and data that are either sensitive or are prone to interference in a virtual kernel instance (vKI). vKernel relies on inline hooks to intercept and redirect requests sent to the host kernel to a vKI, where container-specific security rules, functions, and data are implemented. Through case studies, we demonstrate that under vKernel user-defined data isolation and kernel customization can be supported with a reasonable engineering effort. An evaluation of vKernel with micro-benchmarks, cloud services, real-world applications show that vKernel achieves good security guarantees, but with much less overhead.
Hang Huang, Jia Rao, Song Wu 0001, Hao Fan 0006, Chen Yu 0003, Hai Jin 0001, Kun Suo, Lisong Pan
IEEE Trans. Computers8
2024 QoS-Aware Power Management via Scheduling and Governing Co-Optimization on Mobile Devices
abstract
Scheduling and governing are two key technologies to trade off the Quality of Service (QoS) against the power consumption on mobile devices with heterogeneous cores. However, there are still defects in the use of them, among which two of the decoupling issues are critical and need to be resolved. First, both the scheduling and governing decouple from QoS, one of the most important metrics of user experience on mobile platforms. Second, scheduling and governing also decouple from each other in mobile systems and they might weaken each other when being effective at the same time. To address the above issues, we propose Orthrus, a comprehensive QoS-aware power management approach that involves a governing approach based on deep reinforcement learning to adjust the frequency of heterogeneous cores, a scheduling algorithm based on finite state machine that assigns cores to QoS-related threads, and expert fuzzy control-based coordination mechanism between the two to manage the impact between scheduling and governing. Our proposed approach aims to minimize power consumption while guaranteeing the QoS. We implement Orthrus on Google Pixel 3 as the system service of Android and evaluate it using several widespread mobile applications. The performance evaluation demonstrates that Orthrus reduces the average power consumption by up to 35.7% compared to three state-of-the-art techniques while ensuring the QoS on mobile platforms.
Qianlong Sang, Jinqi Yan, Chuang Hu, Kun Suo, Dazhao Cheng
IEEE Trans. Mob. Comput.5
2023 Characterizing and optimizing Kernel resource isolation for containers
Kun Wang 0005, Song Wu 0001, Kun Suo, Hang Huang, Hai Jin 0001
Future Gener. Comput. Syst.3
2023 Adapt Burstable Containers to Variable CPU Resources
abstract
In the age of the cloud-native, container technology, referred as OS-level virtualization, is increasingly adopted to deploy cloud applications. Compared with virtual machines, containers are lightweight and flexible in resource management. An important quality-of-service (QoS) class in container management is burstable container, whose resource limits are higher than the actual requests allowing a container to expand whenever demands ramp up and additional resources become available. However, efficiently managing burstable containers is challenging, especially for CPU resources. On the one hand, burstable containers should maintain sufficient concurrency, in the form of threads, to utilize extendable CPU resources. On the other hand, the degree of concurrency necessary for utilizing peak CPU resources leads to suboptimal performance when a container's CPU allocation is constrained. In this paper, we recommend that the number of threads in burstable containers should always be set to the CPU limit to guarantee extensibility. However, modern operating systems (OSes) fall short of efficiently managing thread oversubscription. First, the OS CPU scheduler is inefficient for scheduling excessive threads and lacks container awareness. Second, the existing blocking synchronization supported by the OS kernel is inefficient in handling the sleep and wakeup of excessive threads. Finally, the non-blocking synchronization may waste CPUs performing busy waiting when more than one thread in the run queue. To this end, we present a user-level adaptive container scheduler and two OS mechanisms,virtual blockingandbusy-waiting detection, to avoid inefficiency in managing burstable containers without requiring program code changes. Experimental results show that our approaches can keep burstable containers efficient while allowing the applications in containers to take advantage of additional CPUs. The performance gain under high system load is up to 29.7×.
Hang Huang, Jia Rao, Song Wu 0001, Hai Jin 0001, Duoqiang Wang, Kun Suo, Lisong Pan
IEEE Trans. Computers7
2022 Energy Efficiency on Edge Computing: Challenges and Vision
abstract
The Internet of Things (IoT) has been the key to many advancements in next-generation technologies for the past few years. With a conceptual grouping of ecosystem elements such as sensors, actuators, and smart objects connected together, complex operations like environmental monitoring, intelligent transport systems, smart buildings, and smart cities are able to be performed. Edge computing technology extends the reach/scope of IoT ecosystems, offering robust and powerful computational capabilities by connecting multiple devices through the Internet. Unfortunately, this form of computation comes with a significant drawback with strict energy constraints and low power efficiency, which highly limits its potential and usage. In this paper, we present some of the challenges in the planning of energy-efficient IoT edge devices and discuss some of the recent research efforts that proposed promising solutions that address these challenges. Specifically, we first analyze the challenges and reasons for improving the energy consumption of edge platforms and IoT devices. Next, we perform case studies that outline the energy-saving techniques in smart grids, smart cities, electric vehicles (EV), smart home devices, and Virtual Reality and Augmented Reality (VR/AR). We further discuss different approaches such as computation offloading, edge devices hardware and software designs, and a number of algorithms that help reduce energy consumption. Finally, we outline possible future directions and our vision of improving energy efficiency on edge platforms.
Tyler Holmes, Charlie McLarty, Yong Shi 0002, Patrick O. Bobbie, Kun Suo
IPCCC5
2022 Keep Clear of the Edges : An Empirical Study of Artificial Intelligence Workload Performance and Resource Footprint on Edge Devices
abstract
Recently, with the advent of the Internet of everything and 5G network, the amount of data generated by various edge scenarios such as autonomous vehicles, smart industry, 4K/8K, virtual reality (VR), augmented reality (AR), etc., has greatly exploded. All these trends significantly brought real-time, hardware dependence, low power consumption, and security requirements to the facilities, and rapidly popularized edge computing. Meanwhile, artificial intelligence (AI) workloads also changed the computing paradigm from cloud services to mobile applications dramatically. Different from wide deployment and sufficient study of AI in the cloud or mobile platforms, AI workload performance and their resource impact on edges have not been well understood yet. There lacks an in-depth analysis and comparison of their advantages, limitations, performance, and resource consumptions in an edge environment. In this paper, we perform a comprehensive study of representative AI workloads on edge platforms. We first conduct a summary of modern edge hardware and popular AI workloads. Then we quantitatively evaluate three categories (i.e., classification, image-to-image, and segmentation) of the most popular and widely used AI applications in realistic edge environments based on Raspberry Pi, Nvidia TX2, etc. We find that interaction between hardware and neural network models incurs non-negligible impact and overhead on AI workloads at edges. Our experiments show that performance variation and difference in resource footprint limit availability of certain types of workloads and their algorithms for edge platforms, and users need to select appropriate workload, model, and algorithm based on requirements and characteristics of edge environments.
Kun Suo, Tu N. Nguyen 0001, Yong Shi 0002, Selena He, Chih-Cheng Hung
IPCCC1
2021 Tackling Cold Start of Serverless Applications by Efficient and Adaptive Container Runtime Reusing
abstract
During the past few years, serverless computing has changed the paradigm of application development and deployment in the cloud and edge due to its unique advantages, including easy administration, automatic scaling, built-in fault tolerance, etc. Nevertheless, serverless computing is also facing challenges such as long latency due to the cold start. In this paper, we present an in-depth performance analysis of cold start in the serverless framework and propose HotC, a container-based runtime management framework that leverages the lightweight containers to mitigate the cold start and improve the network performance of serverless applications. HotC maintains a live container runtime pool, analyzes the user input or configuration file, and provides available runtime for immediate reuse. To precisely predict the request and efficiently manage the hot containers, we design an adaptive live container control algorithm combining the exponential smoothing model and Markov chain method. Our evaluation results show that HotC introduces negligible overhead and can efficiently improve the performance of various applications with different network traffic patterns in both cloud servers and edge devices.
Kun Suo, Junggab Son, Dazhao Cheng, Wei Chen 0038, Sabur Baidya
CLUSTER1
2021 Parallelizing packet processing in container overlay networks
abstract
Container networking, which provides connectivity among containers on multiple hosts, is crucial to building and scaling container-based microservices. While overlay networks are widely adopted in production systems, they cause significant performance degradation in both throughput and latency compared to physical networks. This paper seeks to understand the bottlenecks of in-kernel networking when running container overlay networks. Through profiling and code analysis, we find that a prolonged data path, due to packet transformation in overlay networks, is the culprit of performance loss. Furthermore, existing scaling techniques in the Linux network stack are ineffective for parallelizing the prolonged data path of a single network flow.
Jiaxin Lei, Manish Munikar, Kun Suo, Hui Lu 0001, Jia Rao
EuroSys3
2020 Data Life Aware Model Updating Strategy for Stream-based Online Deep Learning
abstract
Many deep learning applications deployed in dynamic environments change over time, in which the training models are supposed to be continuously updated with streaming data in order to guarantee better descriptions on data trends. However, most of the state-of-the-art learning frameworks support well in offline training methods while omitting online model updating strategies. In this work, we propose and implement iDlaLayer, a thin middleware layer on top of existing training frameworks that streamlines the support and implementation of online deep learning applications. In pursuit of good model quality as well as fast data incorporation, we design a Data Life Aware model updating strategy (DLA), which builds training data samples according to contributions of data from different life stages, and considers the training cost consumed in model updating. We evaluate iDlaLayer's performance through both simulations and experiments based on TensorflowOnSpark with three representative online learning workloads. Our experimental results demonstrate that iDlaLayer reduces the overall elapsed time of MNIST, Criteo and PageRank by 11.3%, 28.2% and 15.2% compared to the periodic update strategy, respectively. It further achieves an average 20% decrease in training cost and brings about 5 % improvement in model quality against the traditional continuous training method.
Wei Rang, Donglin Yang, Dazhao Cheng, Kun Suo, Wei Chen 0038
CLUSTER4
2019 CNTC: A Container Aware Network Traffic Control Framework
Lin Gu 0002, Junjian Guan, Song Wu 0001, Hai Jin 0001, Jia Rao, Kun Suo, Deze Zeng
GPC6
2019 Adaptive Resource Views for Containers
abstract
As OS-level virtualization advances, containers have become a viable alternative to virtual machines in deploying applications in the cloud. Unlike virtual machines, which allow guest OSes to run atop virtual hardware, containers have direct access to physical hardware and share one OS kernel. While the absence of virtual hardware abstractions eliminates most virtualization overhead, it presents unique challenges for containerized applications to efficiently utilize the underlying hardware. The lack of hardware abstraction exposes the total amount of resources that are shared among all containers to each individual container. Parallel runtimes (e.g., OpenMP) and managed programming languages (e.g., Java) that rely on OS-exported information for resource management could suffer from suboptimal performance. In this paper, we develop a per-container view of resources to export information on the actual resource allocation to containerized applications. The central design of the resource view is a per-container sys\_namespace that calculates the effective capacity of CPU and memory in the presence of resource sharing among containers. We further create a virtual sysfs to seamlessly interface user space applications with sys\_namespace. We use two case studies to demonstrate how to leverage the continuously updated resource view to enable elasticity in the HotSpot JVM and OpenMP. Experimental results show that an accurate view of resource allocation leads to more appropriate configurations and improved performance in a variety of containerized applications.
Hang Huang, Jia Rao, Song Wu 0001, Hai Jin 0001, Kun Suo, Xiaofeng Wu 0002
HPDC5
2019 Preemptive Multi-Queue Fair Queuing
abstract
Fair queuing (FQ) algorithms have been widely adopted in computer systems to share resources among multiple users. Modern operating systems and hypervisors use variants of FQ algorithms to implement the critical OS resource management -- the thread scheduler. While the existing FQ algorithms enforce fair CPU allocation on a per-core basis, there lacks an algorithm to fairly allocate CPU on multiple cores. This common deficiency in state-of-the-art multicore schedulers causes unfair CPU allocations to parallel programs using blocking synchronization, leading to severe performance degradation. Parallel threads that frequently block due to synchronization exhibit deceptive idleness and are penalized by the thread scheduler. To this end, we propose a preemptive multi-queue fair queuing (P-MQFQ) algorithm that uses a centralized queue to fairly dispatch threads from different programs based on their received CPU bandwidth from multiple cores. We demonstrate that P-MQFQ can be approximated by augmenting the existing load balancing in the OS without requiring to implement the centralized queue or undermining scalability. We implement P-MQFQ in Linux and Xen, respectively, and show significantly improved utilization and performance for parallel programs.
Kun Suo, Xiaofeng Wu 0002, Jia Rao, Song Wu 0001, Hai Jin 0001
HPDC2
2018 Characterizing and optimizing hotspot parallel garbage collection on multicore systems
abstract
The proliferation of applications, frameworks, and services built on Java have led to an ecosystem critically dependent on the underlying runtime system, the Java virtual machine (JVM). However, many applications running on the JVM, e.g., big data analytics, suffer from long garbage collection (GC) time. The long pause time due to GC not only degrades application throughput and causes long latency, but also hurts overall system efficiency and scalability.
Kun Suo, Jia Rao, Hong Jiang 0001, Witawas Srisa-an
EuroSys1
2018 vNetTracer: Efficient and Programmable Packet Tracing in Virtualized Networks
abstract
As the scale of cloud systems continues to grow, virtualized networks that provide connectivity between services within and across data centers, are becoming increasingly important to the performance and reliability of the cloud. Despite many advantages, including fast deployment, ease of management, and programmability, virtualized networks require additional layers of abstraction and complicate monitoring and diagnosis of performance issues compared to traditional networks on physical hardware. Virtualized networks usually connect components in multiple protection domains, such as a guest OS, the hypervisor, network bridges, and separate virtualized network functions. There is no efficient means to trace packet transmission across the boundaries. Furthermore, it is challenging to reason about the performance of dynamic virtualized networks. Therefore, fine-grained, user customizable, and reconfigurable network tracing becomes a great need. To address these challenges, we built vNetTracer, an efficient and programmable packet profiler for virtualized networks. vNetTracer relies on the extended Berkeley Packet Filter (eBPF) to dynamically insert user-defined trace programs into a live virtualized network without any changes to the applications or restarts of the monitored network. Through three case studies, we demonstrate the effectiveness of vNetTracer in diagnosing various virtualized networking problems.
Kun Suo, Wei Chen 0038, Jia Rao
ICDCS1
2018 An Analysis and Empirical Study of Container Networks
abstract
Containers, a form of lightweight virtualization, provide an alternative means to partition hardware resources among users and expedite application deployment. Compared to virtual machines (VMs), containers incur less overhead and allow a much higher consolidation ratio. Container networking, a vital component in container-based virtualization, is still not well understood. Many techniques have been developed to provide connectivity between containers on a single host or across multiple machines. However, there lacks an in-depth analysis of their respective advantages, limitations, and performance in a cloud environment. In this paper, we perform a comprehensive study of representative container networks. We first conduct a qualitative comparison of their applicable scenarios, levels of security isolation, and overhead. Then we quantitatively evaluate the throughput, latency, scalability, and startup cost of various container networks in a realistic cloud environment. We find that virtualized network in containers incurs non-negligible overhead compared to physical networks. Performance degradation varies depending on the type of network protocol and packet size. Our experiments show that there is no clear winner in performance and users need to select an appropriate container network based on the requirements and characteristics of their workloads.
Kun Suo, Wei Chen 0038, Jia Rao
INFOCOM1
2017 Preserving I/O prioritization in virtualized OSes
abstract
While virtualization helps to enable multi-tenancy in data centers, it introduces new challenges to the resource management in traditional OSes. We find that one important design in an OS, prioritizing interactive and I/O-bound workloads, can become ineffective in a virtualized OS. Resource multiplexing between multiple tenants breaks the assumption of continuous CPU availability in physical systems and causes two types of priority inversions in virtualized OSes. In this paper, we present xBalloon, a lightweight approach to preserving I/O prioritization. It uses a balloon process in the virtualized OS to avoid priority inversion in both short-term and long-term scheduling. Experiments in a local Xen environment and Amazon EC2 show that xBalloon improves I/O performance in a recent Linux kernel by as much as 136% on network throughput, 95% on disk throughput, and 125x on network tail latency.
Kun Suo, Jia Rao, Luwei Cheng, Xiaobo Zhou 0002, Francis C. M. Lau 0001
SoCC1
2017 Scheduler activations for interference-resilient SMP virtual machine scheduling
abstract
The wide adoption of SMP virtual machines (VMs) and resource consolidation present challenges to efficiently executing multi-threaded programs in the cloud. An important problem is the semantic gaps between the guest OS and the hypervisor. The well-known lock-holder preemption (LHP) and lock-waiter preemption (LWP) problems are examples of such semantic gaps, in which the hypervisor is unaware of the activities in the guest OS and adversely deschedules virtual CPUs (vCPUs) that are executing in critical sections. Existing studies have focused on inferring a high-level semantic state of the guest OS to aid hypervisor-level scheduling so as to avoid the LHP and LWP problems.
Kun Suo, Luwei Cheng, Jia Rao
Middleware2