Jian Li 0021

dblp:33/5448-21 · DBLP profile ↗
← Back
50ranked-venue papers
9as first author
11since 2021 · last 2025
0000-0003-0894-0892ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 32 · 5 first-author · 5 since 2021Computer networks · 8 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSecurity and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Fault Escaping: Improving Robustness of DPU Enhanced Platform with Mutual Assisted VM Recovery
abstract
Modern cloud servers achieve significant performance improvements by exploiting data processing units (DPUs) to offload virtualization overhead.Unlike traditional monolithic hypervisors, this offloading approach splits VM state and distributes hypervisor functions across the DPU and the Host, transforming the cloud server from a single system into a sophisticated orchestration of multiple self-managed processing units.However, the failure rate of these systems has significantly increased, as errors from either the DPU or the Host can crash the entire machine.All these changes necessitate a comprehensive revamp of current VM fault tolerance and isolation mechanisms.This paper explores a novel approach to VM fault tolerance by treating the split hypervisors on the DPU SoC and the Host as redundant peers.We propose a fault-escaping scheme that enables VMs to escape from a failing SoC or Host.With hybrid synchronization, the entire VM state becomes fully accessible on either the SoC or the Host with minimum synchronization overhead (in kilobytes).Instead of recovering the failed hypervisor, this scheme migrates the verified VM state, allowing the VM to escape from the failure
Chao Zhang 0115, Tao Xu 0057, Pai Liu, Zhilang Xu, Jinhu Li, Wenhui Shu, Feifei Fan, Yibin Shen, Jianming Song, Jiesheng Wu, Jian Li 0021
ASPLOS (3)18
2025 CoINT2: A Heuristic Coordinator for Responsive Receive-Side Network I/O Virtualization in Overcommitment Cloud
Xu Huan, Jian Li 0021, Haibing Guan
INFOCOM2
2025 P4KVS: A Role-Replica Separation Offloading Method to Achieve In-Network Consistency for KV Stores Based on P4 Switches
abstract
Strong consistency, particularly linearizability, is essential for distributed DBMSs deployed in correctness-critical domains such as finance and defense. In general, an optimal linearizability DBMS system focus on two key principles: (1) matching single-node (no-consistency cost) Read/Write performance under strong consistency, and (2) practical deployability via general database compatibility. Unfortunately, existing solutions fall short on both fronts. %However, achieving strong consistency often comes with steep performance penalties. For example, etcd-a widely-used Raft-based system-achieves only ~5.9% of the throughput of LevelDB, a single-node store without consistency overhead. To achieve higher performance, software approaches adopt weaker consistency models (e.g., ZAB), rely on narrow network assumptions (e.g., NOPaxos), or expose protocol internals to clients (e.g., CURP), yet still fail to close the performance gap. Recent programmable networking hardware offers promising advances, yet current hardware solutions face practical limitations, including minimal storage and incompatibility with general-purpose databases. We propose P4KVS, the first practical Raft-based in-network consensus offloading solution leveraging programmable switches (P4) for distributed key-value stores. P4KVS offloads only the Leader role to the switch while retaining Followers on servers. Under linearizability, it achieves 74% of single-node LevelDB's throughput for write-heavy workloads, and up to 222.4% for read-heavy workloads by distributing reads across three replicas. This demonstrates that, even under strong consistency, P4KVS can match or exceed the performance of a single-node system. Compared to etcd (which also uses Raft), P4KVS delivers 37.5× higher read throughput and 3520× lower write latency. These results validate our hardware role-replica separation design in eliminating software Raft bottlenecks, while preserving compatibility via standard database interfaces (e.g., LevelDB, etcd) and scaling beyond typical switch memory constraints.
Haojuan Li, Zongpu Zhang, Chenzhen Ye, Ruohan Tang, Jian Li 0021, Haibing Guan, Qiaoling Wang, Pengpeng Zhou
Proc. ACM Manag. Data5
2024 HD-IOV: SW-HW Co-designed I/O Virtualization with Scalability and Flexibility for Hyper-Density Cloud
abstract
As the resource density of cloud servers increases, cloud providers deploy hundreds of VMs concurrently on a single server, requiring a high-performance, scalable, flexible and high-density I/O virtualization method. Hardware assisted virtualization such as device pass-through with SR-IOV can achieve near-native performance, however, at the expense of flexibility and a limited device count. Traditional software-based I/O virtualization systems tend to dedicate additional computing cores for higher performance, but suffer from critical scalability problems especially in high-density cloud.
Zongpu Zhang, Jiangtao Chen, Banghao Ying, Yahui Cao, Lingyu Liu, Jian Li 0021, Weigang Li 0002, Haibing Guan
EuroSys6
2024 vCrypto: a Unified Para-Virtualization Framework for Heterogeneous Cryptographic Resources
abstract
Transport Layer Security (TLS) connections involve costly cryptographic operations which incur significant resource consumption in the cloud. Hardware accelerators are affordable substitutes of expensive CPU cores to accommodate with the constantly increasing security requirements of datacenters. Existing accelerators virtualization mainly relies on passthrough of Single Root I/O Virtualization (SR-IOV) devices. However, deficiency of service accessibility, functionality and availability make device passthrough not an optimal solution for heterogeneous accelerators with different capabilities. To make up the gap, we propose vCrypto, a unified para-virtualization framework for heterogeneous cryptographic resources. vCrypto supports stateful crypto requests offloading and result retrieval with session lifecycle management and event driven notification. vCrypto transparently integrates virtual crypto device capabilities into the OpenSSL framework to benefit existing applications that are based on crypto library APIs without modification. Multiple physical resources can be partitioned flexibly and scheduled cooperatively to enhance the functionality, performance and robustness of virtual crypto service. Finally, vCrypto achieves an optimized performance with two layers polling and memory sharing mechanism. The comprehensive experiments show that with the same cryptographic resources used, vCrypto framework can provide 2.59x to 3.36x higher AES-CBC-HMAC-SHA1 throughput compared to passthrough SR-IOV device.
Chao Zhang 0115, Zongpu Zhang, Hubin Zhang, Weigang Li 0002, Yibin Shen, Jian Li 0021, Haibing Guan
INFOCOM10
2024 Un-IOV: Achieving Bare-Metal Level I/O Virtualization Performance for Cloud Usage With Migratability, Scalability and Transparency
abstract
I/O virtualization is utilized by cloud platforms to provide tenants with efficient, scalable, and manageable network and storage services. The de-facto industrial standard, paravirtualization, offers rich cloud functionality by introducing split front-end and back-end drivers in the guest and host operating systems, respectively. Given this fact, paravirtualization incurs host inefficiency and performance overhead. Thus, emerging hardware virtio accelerators (i.e., SRIOV-capable devices that conform to virtio specification) with device passthrough technologies mitigate the performance issue. However, adopting these devices presents the challenge of insufficient support for live migration.This paper proposes Un-IOV, a novel I/O virtualization system that simultaneously achieves bare-metal level I/O performance and migratability. The key idea is to develop a new hybrid virtualization stack with: (1) a host-bypassed direct data path for virtio accelerators, and (2) a relayed control path guaranteeing seamless live migration support. Un-IOV achieves high scalability by consuming minimum host resources. Extensive experiment results demonstrate that Un-IOV achieves superior network and storage virtualization performance than software implementations with comparable performance of direct passthrough I/O virtualization, while imposing zero guest modification (i.e., guest transparency).
Zongpu Zhang, Chenbo Xia, Cunming Liang, Jian Li 0021, Chen Yu 0003, Tiwei Bie, Roberts Martin, Dan Daly, Xiao Wang 0084, Haibing Guan
IEEE Trans. Computers4
2023 QKPT: Securing Your Private Keys in Cloud With Performance, Scalability and Transparency
abstract
Private key (e.g., RSA key) protection is a significant issue for cloud but existing keyless or keyguard solutions suffer from performance, elasticity or applicability limitations. Recently, represented by Intel KPT, a novel keyguard architecture emerges to combine trusted platform module and crypto accelerator for achieving both security and performance. However, the straight use of KPT for private key protection may not be a good fit in cloud as it incurs challenges on protection capacity, key provisioning latency and transparency. Based on KPT-like hardware, we propose QKPT, a comprehensive key management system to bring your own private keys (BYOPK) into multi-tenant clouds. QKPT introduces a carefully-designed key wrapping layer to overcome these challenges. A small symmetric wrapping key (SWK) is generated for each tenant as the master key to resolve the former two challenges, while a special private key wrapping scheme is adopted to resolve the transparency limitation. Additionally, QKPT incorporates certificate trust to enhance the security of the SWK lifecycle and provides a hardened key server solution without expensive HSM. The evaluation shows that QKPT has a low runtime overhead ($\leq$1.2% for SSL/TLS handshakes) and still greatly outperforms the software baseline (3.5x-17x) owing to the crypto offloading.
Zongpu Zhang, Hubin Zhang, Xiaokang Hu, Jian Li 0021, Weigang Li 0002, Guodong Zhu, Kapil Sood, Brian Will, Haibing Guan
IEEE Trans. Dependable Secur. Comput.5
2022 ES2: Building an Efficient and Responsive Event Path for I/O Virtualization
abstract
Hypervisor intervention in the virtual I/O event path is a main performance bottleneck for I/O virtualization because of the incurred costly VM exits. The shortcomings of prior software solutions against virtual interrupt delivery, a major source of VM exits, promoted the emergence of the hardware-based Posted-Interrupt (PI) technology. PI can provide non-exit interrupt delivery without compromising any virtualization benefit. However, it only acts on the half of the event path, i.e., the interrupt path, while guests I/O requests may also trigger a large amount of VM exits. Additionally, PI may still suffer a severe latency from the vCPU scheduling while delivering interrupts. Aiming at an optimal event path, we propose ES2 to simultaneously improve bidirectional I/O event delivery between guests and their devices. On the basis of PI, ES2 introduces hybrid I/O handling scheme for efficient I/O request delivery and intelligent interrupt redirection for enhanced I/O responsiveness. It does not require any modification to guest OS. We demonstrate that ES2 greatly reduces I/O-related VM exits with the exit handling time (EHT) below 2.5 percent for TCP streams and 0.1 percent for UDP streams, increases guest throughput by 1.9x for Memcached and 1.6x for Nginx, and keeps guest latency at a low level.
Xiaokang Hu, Jian Li 0021, Ruhui Ma, Haibing Guan
IEEE Trans. Cloud Comput.2
2021 HAVS: Hardware-accelerated Shared-memory-based VPP Network Stack
abstract
The number of requests to transfer large files is increasing rapidly in web server and remote-storage scenarios, and this increase requires a higher processing capacity from the network stack. However, to fully decouple from applications, many latest userspace network stacks, such as VPP (vector packet processing) and snap, adopt a shared-memory-based solution to communicate with upper applications. During this communication, the application or network stack needs to copy data to or from shared memory queues. In our verification experiment, these multiple copy operations incur more than 50% CPU consumption and severe performance degradation when the transferred file is larger than 32 KB. This paper adopts a hardware-accelerated solution and proposes HAVS which integrates Intel I/O Acceleration Technology into the VPP network stack to achieve high-performance memory copy offloading. An asynchronous copy architecture is introduced in HAVS to free up CPU resources. Moreover, an abstract memcpy accelerator layer is constructed in HAVS to ease the use of different types of hardware accelerators and sustain high availability with a fault-tolerance mechanism. The comprehensive evaluation shows that HAVS can provide an average 50%-60% throughput improvement over the original VPP stack when accelerating the nginx and SPDK iSCSI target application.
Shujun Zhuang, Jian Li 0021, Haibing Guan
INFOCOM3
2021 A comprehensive test framework for cryptographic accelerators in the cloud
Hubin Zhang, Xiaokang Hu, Jian Li 0021, Haibing Guan
J. Syst. Archit.3
2021 STYX: A Hierarchical Key Management System for Elastic Content Delivery Networks on Public Clouds
abstract
Hosting content delivery networks (CDNs) on clouds has the potential to improve the performance as resources and caches can be placed closer to subscribers. However, avoiding data leakage over an untrusted public cloud is critical, especially for sensitive data such as the SSL private key. The popular Keyless SSL solution allows content owners to retain on-premise custody of SSL private keys on their own key servers, but this solution likely causes performance bottlenecks and impedes the elasticity of CDNs. This paper describes a novel key management system, named STYX, for transmitting trusted data over untrusted channels and storing them on untrusted platforms. STYX accomplishes secure key provisioning for CDN scale-out and the key is securely protected with full revocation rights for CDN scale-in. STYX is implemented as a three-phase hierarchical key management scheme by leveraging Intel Software Guard Extensions (SGX) and QuickAssist Technology (QAT). Furthermore, STYX supports CDN services by integrating Nginx as the SSL termination proxy and the popular Redis/Memcached/Apache as backend caching engines. The performance evaluation shows that STYX significantly outperforms the native HTTPS servers on the CDN node due to QAT acceleration, providing up to a 5× enhancement in throughput and a 50 percent reduction in latency.
Xiaokang Hu, Jian Li 0021, Changzheng Wei, Weigang Li 0002, Haibing Guan
IEEE Trans. Dependable Secur. Comput.2
2020 RECANS: Low-Latency Network Function Chains with Hierarchical State Sharing
abstract
In this paper, we present RECANS, a low-latency NFV state management framework that supports elastic scaling of NF chains while meeting the increasingly stringent NFV performance requirements. Its design builds on the insight that an instance may be scaled to a local server or a remote server, and we thus adopt a hierarchical state-sharing approach that exploits shared memory and low-latency network (RDMA) features to realize state sharing in a server and across servers. Globally, all state data are divided into partitions spread across NF servers, and we reorganize the partitions dynamically through state migration implemented with carefully chosen RDMA primitives. Locally, each partition is maintained by a node store, which uses shared memory to share state with NFs in the same server and thus avoids remote accesses. Being aware of NF chaining, RECANS also adopts specially designed flow tables that support chain-wide batch operations. Our evaluation shows that RECANS has a submillisecond peak latency during scaling events, and outperforms traditional hash tables by 4.78 times for normal state access throughput in NF chains.
Shujun Zhuang, Jian Li 0021, Haibing Guan
HPDC3
2020 Online traffic-aware linked VM placement in cloud data centers
David S. L. Wei, Ruhui Ma, Jian Li 0021, Haibing Guan
Sci. China Inf. Sci.4
2020 Balancing Power And Performance In HPC Clouds
abstract
Abstract With energy consumption in high-performance computing clouds growing rapidly, energy saving has become an important topic. Virtualization provides opportunities to save energy by enabling one physical machine (PM) to host multiple virtual machines (VMs). Dynamic voltage and frequency scaling (DVFS) is another technology to reduce energy consumption. However, in heterogeneous cloud environments where DVFS may be applied at the chip level or the core level, it is a great challenge to combine these two technologies efficiently. On per-core DVFS servers, cloud managers should carefully determine VM placements to minimize performance interference. On full-chip DVFS servers, cloud managers further face the choice of whether to combine VMs with different characteristics to reduce performance interference or to combine VMs with similar characteristics to take better advantage of DVFS. This paper presents a novel mechanism combining a VM placement algorithm and a frequency scaling method. We formulate this VM placement problem as an integer programming (IP) to find appropriate placement configurations, and we utilize support vector machines to select suitable frequencies. We conduct detailed experiments and simulations, showing that our scheme effectively reduces energy consumption with modest impact on performance. Particularly, the total energy delay product is reduced by up to 60%.
Lixia Chen, Jian Li 0021, Ruhui Ma, Haibing Guan, Hans-Arno Jacobsen
Comput. J.2
2020 QWEB: High-Performance Event-Driven Web Architecture With QAT Acceleration
abstract
Hardware accelerators have been a promising solution to reduce the cost of cloud datacenters. This article investigates the acceleration of an important datacenter workload: the web server (or proxy) that faces high computational consumption originated from SSL/TLS processing and HTTP compression. Our study reveals that for the widely-deployed event-driven web architecture, the straight offloading of SSL/TLS or compression tasks suffers from frequent blockings in the offload I/O, leading to the underutilization of both CPU and accelerator resources. To achieve efficient acceleration, we propose QWEB, a comprehensive offload solution based on Intel QuickAssist Technology (QAT). QWEB introduces an asynchronous offload mode for SSL/TLS processing and a pipelining offload mode for HTTP compression, both allowing concurrent offload tasks from a single application process/thread. With these two novel offload modes, the blocking penalty is amortized or even eliminated, and the utilization rate of the parallel computation engines inside the QAT accelerator is greatly increased. The evaluation shows that QWEB provides up to 9x handshake performance with TLS-RSA (2048-bit) over the software baseline. Additionally, the secure data transfer throughput is enhanced by 2x for the SSL/TLS offloading only, 3.5x for the compression offloading only and 5x for the combined offloading.
Jian Li 0021, Xiaokang Hu, David Qian, Changzheng Wei, Gordon McFadden, Brian Will, Weigang Li 0002, Haibing Guan
IEEE Trans. Parallel Distributed Syst.1
2019 vDARM: Dynamic Adaptive Resource Management for Virtualized Multiprocessor Systems
abstract
Modern data center servers have been enhancing their computing capacity by increasing processor counts. Meanwhile, these servers are highly virtualized to achieve efficient resource utilization and energy savings. However, due to the shifting of server architecture to non-uniform memory access (NUMA), current hypervisor-level or OS-level resource management methods continue to be challenged in their ability to meet the performance requirement of various user applications. In this work, we first build a performance slowdown model to accurate identify the current system overheads. Based on the model, we finally design a dynamic adaptive virtual resource management method (vDARM) to eliminate the runtime NUMA overheads by re-configuring virtual-to-physical resource mappings. Experiment results show that, compared with state-of-art approaches, vDARM can bring up an average performance improvement of 42.3% on an 8-node NUMA machines. Meanwhile, vDARM only incurs extra CPU utilization no more than 4%.
Jianmin Qian, Jian Li 0021, Ruhui Ma, Haibing Guan
DATE2
2019 A Holistic Model for Performance Prediction and Optimization on NUMA-based Virtualized Systems
abstract
The non-uniform memory access (NUMA) architecture has become the dominant server architecture due to its scalable bandwidth performance. However, the NUMA architecture also introduces the complicated performance influences to the applications, because of the differentiated remote devices access latency and shared resource access contention. Secondly, quick developments of high speed networking devices make I/O resource be another important performance affecting element for I/O-intensive cloud applications on NUMA server. Thirdly, it is more critical in virtualized environment since all resources are managed uniformly and transparently to the VM, and the application behaviors in the VM are shielded from the Virtual Machine Manager (VMM). In this paper, we first give an analytic evaluation for performance influence from the various resource affinity. Motivated by the observations, we then build an accurate performance prediction model, named Resource Affinity performance Influence Estimation (RAIE). RAIE provides a novel performance prediction model with the holistic resource affinity parameters that are measured with the platform independent quantification approaches that need be executed in one-off manner. Moreover, RAIE model takes into account the actual influence of resource affinity according to the VM behaviours that can be monitored online without VM modification. Comprehensive evaluations prove that the RAIE model for a VM's performance prediction can increase the average prediction accuracy by 3.27x on a 4node NUMA server with high speed Network Interface Cards (NIC). The RAIE guided scheduling case validates that it can achieve 2.1x performance improvement for actual VMM resource management servicing a VM running the dynamic applications.
Jian Li 0021, Jianmin Qian, Haibing Guan
INFOCOM1
2019 EnclaveCache: A Secure and Scalable Key-value Cache in Multi-tenant Clouds using Intel SGX
abstract
With in-memory key-value caches such as Redis and Memcached being a key component for many systems to improve throughput and reduce latency, cloud caches have been widely adopted for small companies to deploy their own cache systems. However, data security is still a major concern, which affects the adoption of cloud caches. Tenant's data stored in a multi-tenant cloud environment faces threats from both co-located other tenants, as well as the untrusted cloud provider.
Li-Xia Chen, Jian Li 0021, Ruhui Ma, Haibing Guan, Hans-Arno Jacobsen
Middleware2
2019 QTLS: high-performance TLS asynchronous offload framework with Intel® QuickAssist technology
abstract
Hardware accelerators are a promising solution to optimize the Total Cost of Ownership (TCO) of cloud datacenters. This paper targets the costly Transport Layer Security (TLS) and investigates the TLS acceleration for the widely-deployed event-driven TLS servers or terminators. Our study reveals an important fact: the straight offloading of TLS-involved crypto operations suffers from the frequent long-lasting blockings in the offload I/O, leading to the underutilization of both CPU and accelerator resources.
Xiaokang Hu, Changzheng Wei, Jian Li 0021, Brian Will, Lu Gong, Haibing Guan
PPoPP3
2019 QZFS: QAT Accelerated Compression in File System for Application Agnostic and Cost Efficient Data Storage
Xiaokang Hu, Fuzong Wang, Weigang Li 0002, Jian Li 0021, Haibing Guan
USENIX ATC4
2019 vSimilar: A high-adaptive VM scheduler based on the CPU pool mechanism
Ruhui Ma, Jian Li 0021, Dajin Wang, Haibing Guan
J. Syst. Archit.4
2019 LG-RAM: Load-aware global resource affinity management for virtualized multicore systems
Jianmin Qian, Jian Li 0021, Ruhui Ma, Haibing Guan
J. Syst. Archit.2
2019 When I/O Interrupt Becomes System Bottleneck: Efficiency and Scalability Enhancement for SR-IOV Network Virtualization
abstract
High performance networking interface cards (NIC) have become essential networking devices in commercial cloud computing environments. Therefore, efficient and scalable I/O virtualization is one of the primary challenges on virtualized cloud computing platforms. Single Root I/O Virtualization (SR-IOV) is a network interface technology that eliminates the overhead of redundant data copies and the virtual network switches through direct I/O in order to achieve nearly natural I/O performance. However, the SR-IOV still suffers from serious problems due to the high overhead for processing excessive network interrupts as well as the unpredictable and bursty traffic load in high-speed networking connections. In this paper, the defects of SR-IOV with 10 Gigabit Ethernet networking are studied first and two major challenges are identified: excessive interrupt rate and single threaded virtual network driver. Second, two interrupt rate control optimization schemes, called coarse-grained interrupt rate (CGR) control and adaptive interrupt rate (AIR) control are proposed. The proposed control schemes can significantly reduce the overhead and enhance the SR-IOV performance compared with the traditional driver with fixed interrupt throttle rate (FIR). In addition, multi-threaded VF driver (MTVD) is proposed that allows the SR-IOV VFs to leverage multi-core resources in order to achieve high scalability. Finally, these optimizations are implemented and detailed performance evaluations are conducted. The results show that CGR and AIR can improve the throughput by 2.26× and 2.97× while saving the CPU resources by 1.23 core and 1.44 core, respectively. The MTVD can achieve 2.03× performance with additional 1.46 cores consumption for VM using the SR-IOV driver.
Jian Li 0021, Ruhui Ma, Zhengwei Qi, Haibing Guan
IEEE Trans. Cloud Comput.1
2018 Topology-aware virtual resource management for heterogeneous multicore systems
abstract
Virtualization technology consolidates multiple independent workloads on a single physical server with virtual resource management, which can result in a significant utilization improvement and energy saving. However, the management of virtual resource is becoming more and more challenging, due to the lack of accurate performance prediction model for the diverse applications' irregular resource access behaviors as well as the complicated Non-Uniform Memory Access (NUMA) server architecture. These challenges drastically affect the overall consolidation performance. This paper proposes vTRMS, a runtime Topology-aware virtual Resource Management Scheme for heterogeneous NUMA multicore systems. vTRMS can improve application performance based on the comprehensive online monitor of the application resource access behaviors as well as an accurate and platform-independent detected NUMA topology metric. Experiment results show that, compared with state-of-art approach, vTRMS can bring up an average throughput improvement of 28.3% and 36.2% on Intel and AMD NUMA machines respectively, when consolidating 32-VMs. At the same time, vTRMS only incurs a runtime overhead no more than 5%.
Jianmin Qian, Jian Li 0021, Ruhui Ma
DATE2
2018 Optimizing Virtual Resource Management for Consolidated NUMA Systems
abstract
Virtualization consolidates multiple virtual machines (VMs) on a single physical server to achieve high resource utilization and energy conservation. However, as current data center server architectures shifting to Non-Uniform Memory Access (NUMA), the complex interplay between data access affinity and shared resource access overloaded still challenge the consolidation efficiency. Existing NUMA-aware virtual resource management policies lack holistic awareness of the data access affinities as well as unable to manage the system loads, which may result in sub-optimal overall system performance. In this paper, we propose a load-aware global resource access management framework (LG-RAM) that aims to optimize VM consolidation performance on NUMA systems. Our real system based evaluations indicate that, compared with state-of-the-art approaches, LG-RAM exhibits an average throughput improvement of 36.5% on our testbed. Besides, LG-RAM only incurs an extra CPU usage of no more than 9% on average when consolidating 32 VMs.
Jianmin Qian, Jian Li 0021, Ruhui Ma, Haibing Guan
ICCD2
2018 CoINT: Proactive Coordinator for Avoiding Interruptability Holder Preemption Problem in VSMP Environment
abstract
In a Virtual Symmetric Multiprocessing (VSMP) environment, the behavior of hypervisor scheduler can significantly influence a guest's I/O responsiveness. The interrupt remapping mechanism, which can leverage multiple virtual CPUs in the VSMP guest to process I/O events, is known to be an efficient and prevalent solution to improve the I/O performance. However, in this paper we identified a novel challenge called the “Interruptability Holder Preemption” (IHP) problem in interrupt remapping mechanism. The IHP issue presents that a virtual CPU (vCPU) disabling the interruptability of guest's network device is descheduled by the hypervisor scheduler, which can easily invalidate the efficiency of the interrupt remapping mechanism. To solve this problem, we propose CoINT, a gasket coordinator residing in the hypervisor, to substantially enhance the network I/O performance by empowering the hypervisor to be proactively aware of the interruptability information of the guest's network device. ColNT completely eliminates the “Interruptability Holder Preemption” problem and largely reduces I/Ointerrupt processing delay caused by hypervisor scheduler. We implement ColNT in KVM hypervisor and evaluate its efficiency and responsiveness using both macro and microlevel benchmarks. The results show that ColNT can improve the netperf throughput up to 3x compared with native KVM, and up to 1.3x compared with traditional interrupt remapping of hypervisor-Ievel solution, in sacrifice of an negligible and reasonable overhead in the hypervisor.
Xiaokang Hu, Jian Li 0021, Haibing Guan
INFOCOM3
2017 STYX: a trusted and accelerated hierarchical SSL key management and distribution system for cloud based CDN application
abstract
Protecting the customer's SSL private key is the paramount issue to persuade the website owners to migrate their contents onto the cloud infrastructure, besides the advantages of cloud infrastructure in terms of flexibility, efficiency, scalability and elasticity. The emerging Keyless SSL solution retains on-premise custody of customers' SSL private keys on their own servers. However, it suffers from significant performance degradation and limited scalability, caused by the long distance connection to Key Server for each new coming end-user request. The performance improvements using persistent session and key caching onto cloud will degrade the key invulnerability and discourage the website owners because of the cloud's security bugs.
Changzheng Wei, Jian Li 0021, Weigang Li 0002, Haibing Guan
SoCC2
2017 ES2: Aiming at an Optimal Virtual I/O Event Path
abstract
Improving the performance of I/O virtualization is a key issue for cloud and datacenter infrastructures, especially with the rapid increase of network interconnection speeds. Previous efforts have made the performance overhead associated with the virtual I/O data path largely negligible. The remaining bottlenecks mainly lie in the event path: hypervisor interventions trigger costly virtual machine (VM) exits and lead to dramatical performance degradation. Aiming at an optimal virtual I/O event path, we propose ES2, a comprehensive scheme that simultaneously improves bidirectional I/O event delivery between guest VMs and their devices. ES2 can provide efficient I/O request delivery, non-exit interrupt delivery and enhanced I/O responsiveness. Moreover, it does not require any modification to guest operating system (OS) or compromise any virtualization benefit. We demonstrate that ES2 greatly reduces VM exit rate with the time in guest (TIG) for I/O processing above 96% for TCP streams and 99% for UDP streams, increases guest throughput by 1.8× for Memcached and 2× for Apache, and keeps guest latency at a very low level.
Xiaokang Hu, Jian Li 0021, Ruhui Ma, Haibing Guan
ICPP3
2017 MigVisor: Accurate Prediction of VM Live Migration Behavior using a Working-Set Pattern Model
abstract
Live migration of a virtual machine (VM) is a powerful technique with benefits of server maintenance, resource management, dynamic workload re-balance, etc. Modern research has effectively reduced the VM live migration (VMLM) time to dozens of milliseconds, but live migration still exhibits failures if it cannot terminate within the given time constraint. The ability to predict this type of failure can avoid wasting networking and computing resources on the VM migration, and the associated system performance degradation caused by wasting these resources. The cost of VM live migration highly depends on the application workload of the VM, which may undergo frequent changes. At the same time, the available system resources for VM migration can also change substantially and frequently. To account for these issues, we present a solution called MigVisor, which can accurately predict the behaviour of VM migration using working-set model. This can enable system managers to predict the migration cost and enhance the system management efficacy. The experimental results prove the design suitability and show that the MigVisor has a high prediction accuracy since the average relative error between the predicted value and the measured value is only 6.2%~9%.
Jinshi Zhang, Eddie Dong, Jian Li 0021, Haibing Guan
VEE3
2017 TEES: An Efficient Search Scheme over Encrypted Data on Mobile Cloud
abstract
Cloud storage provides a convenient, massive, and scalable storage at low cost, but data privacy is a major concern that prevents users from storing files on the cloud trustingly. One way of enhancing privacy from data owner point of view is to encrypt the files before outsourcing them onto the cloud and decrypt the files after downloading them. However, data encryption is a heavy overhead for the mobile devices, and data retrieval process incurs a complicated communication between the data user and cloud. Normally with limited bandwidth capacity and limited battery life, these issues introduce heavy overhead to computing and communication as well as a higher power consumption for mobile device users, which makes the encrypted search over mobile cloud very challenging. In this paper, we propose traffic and energy saving encrypted search (TEES), a bandwidth and energy efficient encrypted search architecture over mobile cloud. The proposed architecture offloads the computation from mobile devices to the cloud, and we further optimize the communication between the mobile clients and the cloud. It is demonstrated that the data privacy does not degrade when the performance enhancement methods are applied. Our experiments show that TEES reduces the computation time by 23 to 46 percent and save the energy consumption by 35 to 55 percent per file retrieval, meanwhile the network traffics during the file retrievals are also significantly reduced.
Jian Li 0021, Ruhui Ma, Haibing Guan
IEEE Trans. Cloud Comput.1
2017 Accurate CPU Proportional Share and Predictable I/O Responsiveness for Virtual Machine Monitor: A Case Study in Xen
abstract
In cloud computing, the performance of applications is heavily dependent on resource services provided by the virtualized environment. However, in some virtualized environment, such as Xen, the accuracy of CPU proportional share and the responsiveness of I/O processing are heavily dependent on the proportion of the allocated CPU resource. In this paper, we study how inaccurate share ratio of CPU proportional share and proportion dependent responsiveness of I/O affect the performance of Xen, and discover that they lead to unstable performance and is thus not able to conform service-level agreements (SLA). We conclude that the scheduling scheme and the coarse grained time-slice are the major negative impacts on this issue. Therefore, we propose a novel scheduling scheme, named Predictable Resource Guarantee Scheduler (PRGS), that achieves accurate CPU proportional share and predictable I/O responsiveness. We implement a PRGS prototype on Xen virtualization platform and carry out a thorough evaluation via experimentation. The experimental results show that PRGS achieves accurate CPU proportional share and predictable I/O responsiveness. Also, with only slight overhead, PRGS controls PING packet delay to a specified fixed time threshold (e.g. 30 ms in our experiments).
Jian Li 0021, Ruhui Ma, Haibing Guan, David S. L. Wei
IEEE Trans. Cloud Comput.1
2017 ForenVisor: A Tool for Acquiring and Preserving Reliable Data in Cloud Live Forensics
abstract
Live forensics is an important technique in cloud security but is facing the challenge of reliability. Most of the live forensic tools in cloud computing run either in the target Operating System (OS), or as an extra hypervisor. The tools in the target OS are not reliable, since they might be deceived by the compromised OS. Furthermore, traditional general purpose hypervisors are vulnerable due to their huge code size. However, some modules of a general purpose hypervisor, such as device drivers, are indeed unnecessary for forensics. In this paper, we propose a special purpose hypervisor, called ForenVisor, which is dedicated to reliable live forensics. The reliability is improved in three ways: reducing Trusted Computing Base (TCB) size by leveraging a lightweight architecture, collecting evidence directly from the hardware, and protecting the evidence and other sensitive files with Filesafe module. We have implemented a proof-of-concept prototype on the Windows platform, which can acquire the process data, raw memory, and I/O data, such as keystrokes and network traffic. Furthermore, we evaluate ForenVisor in terms of code size, functionality, and performance. The experiment results show that ForenVisor has a relatively small TCB size of about 13 KLOC, and only causes less than 10 percent performance reduction to the target system. In particular, our experiments verify that ForenVisor can guarantee that the protected files remain untampered, even when the guest OS is compromised by viruses, such as `ILOVEYOU' and Worm.WhBoy. Also, our system can be loaded as a hypervisor without needing to pause the target OS. This allows it to not only avoid destructing but also to gather the live evidence of the target OS. We also posted the source code of ForenVisor on Github.
Zhengwei Qi, Chengcheng Xiang, Ruhui Ma, Jian Li 0021, Haibing Guan, David S. L. Wei
IEEE Trans. Cloud Comput.4
2016 Optimizations for High Performance Network Virtualization
Fanfu Zhou, Ruhui Ma, Jian Li 0021, Li-Xia Chen, Weidong Qiu, Haibing Guan
J. Comput. Sci. Technol.3
2015 vINT: Hardware-Assisted Virtual Interrupt Remapping for SMP VM with Scheduling Awareness
abstract
Symmetric Multi-Processing (SMP) virtual machine (VM), or virtual SMP for short, enables a single virtual machine to span multiple processors, thereby supporting the virtual machine to run resource-intensive applications. In addition to offering higher computing capacity, virtual SMP also offers the opportunity to alleviate the problem of unpredictable I/O responsiveness. To this end, we propose vINT (scheduling status based virtual INterrupt remapping adapTer), a scheme that leverages hardware-assisted interrupt mapping. vINT obtains high efficiency and flexibility by adding a lightweight module in virtual machine monitor (VMM) with no need of changing VMM scheduler and is transparent to guest OS. We implement the prototype in XEN 4.3.0 and conduct evaluations with both micro-benchmarks and macro-benchmarks. The experimental results show that vINT can increase the networking throughput by 5x and can reduce the required execute time of disk I/O by 17.5%, while introducing only a light overhead.
Jian Li 0021, Ruhui Ma, Haibing Guan, David S. L. Wei
CloudCom1
2015 DaSS: Dynamic Time Slice Scheduler for Virtual Machine Monitor
Ruhui Ma, Jian Li 0021, Haibing Guan
ICA3PP (1)2
2014 Workload-Aware Credit Scheduler for Improving Network I/O Performance in Virtualization Environment
abstract
Single-root I/O virtualization (SR-IOV) has become the de facto standard of network virtualization in cloud infrastructure. Owing to the high interrupt frequency and heavy cost per interrupt in high-speed network virtualization, the performance of network virtualization is closely correlated to the computing resource allocation policy in Virtual Machine Manager (VMM). Therefore, more sophisticated methods are needed to process irregularity and the high frequency of network interrupts in high-speed network virtualization environment. However, the I/O-intensive and CPU-intensive applications in virtual machines are treated in the same manner since application attributes are transparent to the scheduler in hypervisor, and this unawareness of workload makes virtual systems unable to take full advantage of high performance networks. In this paper, we discuss the SR-IOV networking solution and show by experiment that the current credit scheduler in Xen does not utilize high performance networks efficiently. Hence we propose a novel workload-aware scheduling model with two optimizations to eliminate the bottleneck caused by scheduler. In this model, guest domains are divided into I/O-intensive domains and CPU-intensive domains according to their monitored behaviour. I/O-intensive domains can obtain extra credits that CPU-intensive domains are willing to share. In addition, the total number of credits available is adjusted to accelerate the I/O responsiveness. Our experimental evaluations show that the new scheduling models improve bandwidth and reduce response time, by keeping the fairness between I/O-intensive and CPU-intensive domains. This enables virtualization infrastructure to provide cloud computing services more efficiently and predictably.
Haibing Guan, Ruhui Ma, Jian Li 0021
IEEE Trans. Cloud Comput.3
2013 COLO: COarse-grained LOck-stepping virtual machines for non-stop service
abstract
Virtual machine (VM) replication provides a software solution of for business continuity and disaster recovery through application-agnostic hardware fault tolerance by replicating the state of primary VM (PVM) to secondary VM (SVM) on a different physical node. Unfortunately, current VM replication approaches suffer from excessive overhead, which severely limit their applicability and suitability. In this paper, we leverage the practical effect of networked server-client system that PVM and SVM are considered as in the same state only if they can generate the same response from the clients' point of view, and this is exploited to optimize performance. To this end, we propose a generic and highly efficient non-stop service solution, named as "COLO" (COarse-grained LOck-stepping virtual machine) utilizing on-demand VM replication. COLO monitors the output responses of the PVM and SVM, and rules the SVM as a valid replica of the PVM according to the output similarity between PVM and SVM. If the responses do not match, the commit of network response is withheld until PVM's state has been synchronized to SVM. Hence, we ensure that the system is always capable of failover by SVM. Although non-determinism may mean a different internal state of SVM from that of the PVM, it is equally valid and remains consistent from external observations. Unlike earlier instruction level lock-stepping deterministic execution approaches, COLO can easily support Multi-Processors (MP) involving workloads with the satisfying performance. Results show that COLO significantly outperforms existing approaches, particularly on server-client workloads such as online databases and web server applications.
Yaozu Dong, Yunhong Jiang, Ian Pratt 0001, Shiqing Ma, Jian Li 0021, Haibing Guan
SoCC6
2013 Cache isolation for virtualization of mixed general-purpose and real-time systems
Ruhui Ma, Alei Liang, Haibing Guan, Jian Li 0021
J. Syst. Archit.5
2013 SR-IOV Based Network Interrupt-Free Virtualization with Event Based Polling
abstract
Along with the developments of networking and virtualization technologies, high speed network connections have become one of the key components in cloud computing and data-centers. Single-Root I/O Virtualization (SR-IOV) enhances the network throughput to the extent of becoming close to the line rate and achieving high scalability in the 10Gbps and higher network environments. However, the overhead of SR-IOV interrupt virtualization remains significant due to some additional trap-and-emulation overhead on the virtual interrupt controller. The higher the virtualization network connection is, the higher the interrupt frequency becomes through high bandwidth network. To mitigate this problem, we propose a smart Event-Based Polling model (sEBP), which leverages existing system events to trigger a regular packet polling such that network interrupts are eliminated from the critical I/O paths in the virtual environment. Due to the many varieties of system events, sEBP can deal with the network workload in a configurable and flexible manner. Based on a hierarchical virtualized environment, it can also be implemented either at the guest OS kernel level or at the Virtual Machine Manager (VMM) level. Since polling is much lighter than interrupt processing, sEBP significantly reduces the network processing overhead. The experimental results prove the efficiency of sEBP, which can achieve up to a 59% performance improvement and a 23% improved scalability ratio.
Haibing Guan, Yaozu Dong, Jian Li 0021
IEEE J. Sel. Areas Commun.4
2013 Performance Enhancement for Network I/O Virtualization with Efficient Interrupt Coalescing and Virtual Receive-Side Scaling
abstract
Virtualization is a key technology in cloud computing; it can accommodate numerous guest VMs to provide transparent services, such as live migration, high availability, and rapid checkpointing. Cloud computing using virtualization allows workloads to be deployed and scaled quickly through the rapid provisioning of virtual machines on physical machines. However, I/O virtualization, particularly for networking, suffers from significant performance degradation in the presence of high-speed networking connections. In this paper, we first analyze performance challenges in network I/O virtualization and identify two problems-conventional network I/O virtualization suffers from excessive virtual interrupts to guest VMs, and the back-end driver does not efficiently use the computing resources of underlying multicore processors. To address these challenges, we propose optimization methods for enhancing the networking performance: 1) Efficient interrupt coalescing for network I/O virtualization and 2) virtual receive-side scaling to effectively leverage multicore processors. These methods are implemented and evaluated with extensive performance tests on a Xen virtualization platform. Our experimental results confirm that the proposed optimizations can significantly improve network I/O virtualization performance and effectively solve the performance challenges.
Haibing Guan, Yaozu Dong, Ruhui Ma, Dongxiao Xu, Jian Li 0021
IEEE Trans. Parallel Distributed Syst.6
2012 Adjustable Credit Scheduling for High Performance Network Virtualization
abstract
Virtualization technology is now widely adopted in cloud computing to support heterogeneous and dynamic workload. The scheduler in a virtual machine monitor (VMM) plays an important role in allocating resources. However, the type of applications in virtual machines (VM) is unknown to the scheduler, and I/O-intensive and CPU-intensive applications are treated the same. This makes virtual systems unable to take full advantage of high performance networks such as 10-Gigabit Ethernet. In this paper, we review the SR-IOV networking solution and show by experiment that the current credit scheduler in Xen does not utilize high performance networks efficiently. For this reason, we propose a novel scheduling model with two optimizations to eliminate the bottleneck caused by scheduler. In this model, guest domains are divided into I/O-intensive domains and CPU-intensive domains according to their monitored behaviour. I/O-intensive domains can obtain extra credits that CPU-intensive domains are willing to share. Besides, the total available credits is adjusted agilely to accelerate the I/O responsiveness. Our experimental evaluation with benchmarks shows that the new scheduling model improves bandwidth even when the system's load is very high.
Zhibo Chang, Jian Li 0021, Ruhui Ma, Haibing Guan
CLUSTER2
2012 Adaptive and Scalable Optimizations for High Performance SR-IOV
abstract
High performance networking interfaces, such as 10-Gigabit Ethernet (10GE), are now widely deployed in commercial Cloud computing environments. Virtualization is a standard technique for these environments, one of whose key challenges is to achieve highly efficient and scalable I/O virtualization. Single Root I/O Virtualization (SR-IOV) eliminates the overhead of redundant data copies and the virtual network switch through direct I/O, but needs more work on performance and scalability. In this paper, we first study the defects of SR-IOV with 10GE networking and find two major challenges. Due to multiplexing of traffic from different virtual machines, SR-IOV may generate redundant interrupts unexpectedly and thus result in high CPU overhead. SR-IOV also suffers from single-threaded NAPI which prevents it from fully utilizing multi-core machines. Then we propose two optimizations for enhancing the SR-IOV performance. The first uses adaptive interrupt rate control (AIRC) to reduce CPU overhead caused by excessive interrupts. The second is a multi-threaded network driver (MTND) which allows SR-IOV to make full use of multi-core resources. We implement these optimizations and carry out a detailed performance evaluation. The results show that AIRC can reduce CPU overhead by up to 143% and MTND can improve SR-IOV performance by up to 38%.
Ruhui Ma, Jian Li 0021, Zhibo Chang, Haibing Guan
CLUSTER3
2011 A novel traffic shaping algorithm with delay jitter constraints for real-time multimedia networks
abstract
Data traversing packet networks experience varying delays, resulting in noticeable delay jitters. This significantly degrades overall system performance in a real-time multimedia network. This paper proposes TSJC, a novel traffic shaping algorithm with delay jitter constraints for real-time multimedia networks, on the basis of traditional shaping algorithms and their traffic characteristics. TSJC computes the queuing delay and queuing delay variation on line by monitoring the token bucket states, such as the queue length and token arrival rate. TSJC adaptively configures system parameters based on the queuing delay and time jitter, to provide universally low jitter outputs. Simulations have shown that TSJC can both smooth traffic fluctuation and decrease the delay jitter.
Hairui Zhou, Jian Li 0021, Guangyu Hu, Yeqiong Song
ETFA2
2011 Control Theoretic Analysis of eXplicit Control Protocol with Short-Lived Traffic
abstract
The eXplicit Control Protocol (XCP) is a promising congestion control protocol that outperforms TCP in terms of efficiency, fairness, persistent queue length, and packet loss rate. XCP quantitatively informs senders how to adjust their sending rate, and assumes that all senders will respond. However, short-lived flows are often on the order of a few segments, they may be unresponsive to XCP feedback, and have negative impact on the system control loop. In this paper, a control theoretic analysis of the XCP properties is conducted in the presence of short-lived traffic. First, the original XCP model is modified to account for these unresponsive flows. Then, with the M/G/∞ model of short-lived traffic, their impact on the XCP control loop is analyzed, and theoretical results show that short-lived bursty traffic have the effect of reducing the available bandwidth and increasing the variances of long-lived XCP flows. Finally, theoretical results are verified using packet level simulations.
Hairui Zhou, Chengchen Hu, Jian Li 0021
GLOBECOM3
2011 End-to-End Delay Analysis in Wireless Network Coding: A Network Calculus-Based Approach
abstract
Network coding provides a powerful mechanism for improving performance of wireless networks. In this paper, we present an analytical approach for end-to-end delay analysis in wireless networks that employ inter-session coding. Prior work on performance analysis in wireless network coding mainly focuses on the throughput of the overall network. Our approach aims to analyze the end-to-end delay performance of each flow in the network. The theoretical basis of our approach is network calculus. In order to apply network calculus to the analysis of wireless network coding, we address three specific problems: identifying traffic flows, characterizing broadcast links, and measuring coding opportunities. We make three main contributions. First, we obtain theoretical formulations for computing the delay bounds of bursty flows in wireless networks employing network coding. Second, based on the formulations, we figure out the factors that affect the end-to-end delays, and find an interesting phenomenon that, as traffic grows, the overall delay can potentially decrease. Third, in order to exploit the benefit of our findings, we introduce a new scheduling scheme that can improve the performance of current practical wireless network coding.
Huanzhong Li, Xue (Steve) Liu, Wenbo He 0003, Jian Li 0021, Wenhua Dou
ICDCS4
2011 Improving the stability of eXplicit Control Protocol under heterogeneous delays
abstract
The eXplicit Control Protocol (XCP) is a promising congestion control protocol that outperforms TCP in terms of efficiency, fairness, persistent queue length, and packet loss rate. However, XCP will behave in a noticeably unstable manner if the maximum round trip time of a flow is much larger than average round trip time of all flows, which is a typical nonlinear instability. In this paper, eXCP is proposed to stabilize the system and enable XCP to deal with heterogeneous feedback delays. With the aid of the exponentially weighted moving average filter, eXCP directly reduces the volatility of the control interval and effectively improves the stability of the aggregate input traffic at a bottleneck link. Simulations also have shown that the variance of per-flow throughput and round trip time decrease dramatically.
Hairui Zhou, Chengchen Hu, Jian Li 0021
LCN3
2010 Online adaptive utilization control for real-time embedded multiprocessor systems
Jianguo Yao 0002, Xue (Steve) Liu, Zonghua Gu 0001, Jian Li 0021
J. Syst. Archit.5
2006 DLB: A Novel Real-time QoS Control Mechanism for Multimedia Transmission
abstract
This paper presents a new QoS guarantee scheme called R-(m,k)-firm (Relaxed-(m,k)-firm) which provides the guarantee on transmission delay of at least m out of any k consecutive packets (m/spl les/k). It has several advantages: (1) during network congestion, packets are dropped according to the (m,k) model rather than uncontrollably as the case of TD and RED, avoiding thus undesirable long consecutive packet drops; (2) it allows to admit more real-time flows than the traditional over-provisioning approach. A new mechanism, called DLB (double leaks bucket) is also proposed for dropping a proportion of packets of a flow or of aggregated-flows in case of network congestion while still guaranteeing the R-(m,k)-firm constraint. The sufficient condition for this guarantee is given for configuring the DLB parameters. It is easy to implement DLB in the actual IntServ and Diffserv architectures (by simply replacing the actual leaky, bucket by DLB) for providing respectively per flow and per class (m,k) guarantee, or event per flow R-(m,k)-firm guarantee in Diffserv.
Jian Li 0021, Yeqiong Song
AINA (1)1
2006 Relaxed (m, k)-firm Constraint to Improve Real-time Streams Admission Rate under Non Pre-emptive Fixed Priority Scheduling
abstract
Comparing with hard real-time approach, (m,k)-firm constraint and its related scheduling policies are considered as an efficient way to increase the admission rate of real-time streams to a network thanks to the possibility to drop until k-m out of any k consecutive processing requirements, reducing thus the workload. Although it is interesting for probabilistic (m,k)-firm guarantee, we show however in this paper that for deterministic (m,k)-firm guarantee, reducing the workload by a factor of m/k does not contribute to reducing the resource requirement in general. So the relaxed (m,k)-firm constraint is proposed. The sufficient schedulability condition under non preemptive fixed priority scheduling is derived and its practical interest in terms of the resource requirement reduction is demonstrated.
Jian Li 0021, Yeqiong Song
ETFA1
2006 Providing Real-Time Applications With Graceful Degradation of QoS and Fault Tolerance According to(m, k)-Firm Model
abstract
The$(m, k)$-firm model has recently drawn a lot of attention. It provides a flexible real-time system with graceful degradation of the quality of service (QoS), thus achieving the fault tolerance in case of system overload. In this paper, we focus on the distance-based priority (DBP) algorithm as it presents the interesting feature of dynamically assigning the priorities according to the system's current state (QoS-aware scheduling). However, DBP cannot readily be used for systems requiring a deterministic$(m, k)$-firm guarantee since the schedulability analysis was not done in the original proposition. In this paper, a sufficient schedulability condition is given to deterministically guarantee a set of periodic or sporadic activities (jobs) sharing a common non-preemptive server. This condition is applied to two case studies showing its practical usefulness for both bandwidth dimensioning of the communication system providing graceful degradation of QoS and the task scheduling in an in-vehicle embedded system allowing fault tolerance.
Jian Li 0021, Yeqiong Song, Françoise Simonot-Lion
IEEE Trans. Ind. Informatics1