Xiaoning Ding

dblp:36/679 · DBLP profile ↗
← Back
62ranked-venue papers
13as first author
18since 2021 · last 2025
0000-0002-9947-0437ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 37 · 9 first-author · 13 since 2021Computer networks · 8 · 3 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2025 Optimizing Task Scheduling in Cloud VMs with Accurate vCPU Abstraction
abstract
The paper shows that task scheduling in Cloud VMs hasn't evolved quickly to handle the dynamic vCPU resources. The existing vCPU abstraction cannot accurately depict the vCPU dynamics in capacity, activity, and topology, and these mismatches can mislead the scheduler, causing performance degradation and system anomalies. The paper proposes a novel solution, vSched, which probes accurate vCPU abstraction through a set of lightweight microbenchmarks (vProbers) without modifying the hypervisor, and leverages the probed information to optimize task scheduling in cloud VMs with three new techniques: biased vCPU selection, intra-VM harvesting, and relaxed work conservation. Our evaluation of vSched's implementation in x86 Linux Kernel demonstrates that it can effectively improve both system throughput and workload latency across various VM types in the dynamic multi-cloud environment.
Edward Guo, Weiwei Jia 0001, Xiaoning Ding, Jianchen Shan
EuroSys3
2025 Byte vSwitch: A High-Performance Virtual Switch for Cloud Networking
abstract
Virtual switch is a fundamental component of cloud computing as it provides core networking functionalities for VMs and containers. Open vSwitch (OVS) is widely adopted in cloud environments due to its open-source nature, programmability, and rich set of features. At ByteDance, we initially adopted OVS in our public cloud, but as our cloud business grew, its generic design along with its complex code-base quickly became obstacles to improvements. Hence, we developed Byte vSwitch (BVS), a high-performance virtual switch that was specifically designed to address the performance, scalability, serviceability, and operational efficiency needs of our cloud services. More specifically, BVS adopts a simple architecture with an optimized hash table to maximize forwarding performance. In addition, we introduced several optimizations to improve BVS scalability, operability, and serviceability in cloud environments. Our evaluations show that BVS achieves up to 3.3× higher PPS and 25% lower latency compared to OVS. BVS has been deployed at scale across all regions of the ByteDance public cloud for over four years, and this paper presents our experience in designing, deploying, and operating BVS in production.
Xin Wang 0261, Deguo Li, Lidong Jiang, Shubo Wen, Daxiang Kang, Engin Arslan, Peng He 0003, Xinyu Qian, Jianwen Pi, Xiaoning Ding, Hao Luo 0013
EuroSys12
2025 ByteDance Jakiro: Enabling RDMA and TCP over Virtual Private Cloud
abstract
A Virtual Private Cloud (VPC) that enables both RDMA and TCP provides advantages for both tenants and cloud providers. It serves the flexible demands of RDMA and TCP of tenant applications while delivering a cost-effective solution compared to the construction of two distinct overlay networks. In this study, we introduce Jakiro, an innovative framework of vNIC design that supports both RDMA and TCP within ByteDance Cloud. Jakiro holds the capability to support fundamental VPC features such as QoS, security groups, etc., for both RDMA and TCP streams while maintaining compatibility with applications and intra-host RDMA optimization techniques. We benchmark Jakiro's performance using basic test cases and real-world high-performance computing applications and distributed machine learning training. The results indicate that the RDMA performance of Jakiro is close to that of the physical RDMA. Concurrently, Jakiro guarantees a weighted fair QoS between RDMA and TCP. Jakiro has been deployed in ByteDance Cloud for one year, we share our critical design and deployment decisions, as well as experiences and lessons from production.
Yirui Liu 0001, Lidong Jiang, Deguo Li, Daxiang Kang, Zhaoyang Wei, Yuqi Chai, Xiaoning Ding, Jianwen Pi, Hao Luo 0013
SIGCOMM9
2024 Exploring Performance and Cost Optimization with ASIC-Based CXL Memory
abstract
As memory-intensive applications continue to drive the need for advanced architectural solutions, Compute Express Link (CXL) has risen as a promising interconnect technology that enables seamless high-speed, low-latency communication between host processors and various peripheral devices. In this study, we explore the application performance of ASIC CXL memory in various data-center scenarios. We then further explore multiple potential impacts (e.g., throughput, latency, and cost reduction) of employing CXL memory via carefully designed policies and strategies. Our empirical results show the high potential of CXL memory, reveal multiple intriguing observations of CXL memory and contribute to the wide adoption of CXL memory in real-world deployment environments. Based on our benchmarks, we also develop an Abstract Cost Model that can estimate the cost benefit from using CXL memory.
Yupeng Tang, Henry Hu, Tongping Liu, Jiaxin Shan, Ruoyun Huang, Cheng Zhao 0001, Cheng Chen 0008, Xiaoning Ding, Jianjun Chen 0001
EuroSys15
2024 Effective Huge Page Strategies for TLB Miss Reduction in Nested Virtualization
abstract
Huge page strategies, such as Linux Transparent Huge Page (THP), have become a prevalent solution to mitigate the performance bottleneck caused by increasingly high memory address translation overhead. However, in cloud environments, virtualization presents a two-fold challenge, exacerbating address translation overhead and undermining the effectiveness of huge page strategies. To effectively reduce address translation overhead, huge page strategies in the host and guest virtual machines (VMs) must work in concert for “proper huge page alignment”, i.e., huge pages in guest VMs being backed by host huge pages. This requires a cross-layer coordinating mechanism, which has been designed targeting non-nested virtualization settings. The paper introduces XGEMINI as an efficient solution targeting nested virtualization settings, where addressing these issues is particularly challenging, given the additional obstacles in creating synergy between host and guest VMs, due to an extra layer of page mappings by guest hypervisors. XGEMINI addresses these challenges by improving the shadow paging mechanism. Evaluation based on the KVM/Linux prototype implementation and diverse real-world applications shows XGEMINI greatly reduces TLB misses and enhances application performance in nested virtualization.
Weiwei Jia 0001, Jiyuan Zhang 0003, Jianchen Shan, Xiaoning Ding
IEEE Trans. Computers4
2023 HugeGPT: Storing Guest Page Tables on Host Huge Pages to Accelerate Address Translation
abstract
Expensive page table walks triggered by frequent TLB misses have incurred major performance bottlenecks for data-intensive workloads that are dominated by memory accesses with weak locality. Since it is hard to reduce TLB misses for such workloads, reducing page table walk overhead (i.e., the overhead of each TLB miss) is an increasingly important direction for improving application performance. The direction is more compelling for workloads running in virtual machines (VMs). In virtualized environments, each TLB miss triggers a two-dimensional page table walk, which has a significantly higher overhead than that on native systems. This paper presents HugeGPT, a software approach to reducing two-dimensional page table walk overhead in virtualized environments. HugeGPT ensures that page tables used in guest systems are physically held in the huge pages formed in the host. This brings two-fold benefits: 1) the number of steps walking down the host page table is reduced; 2) the misses of page walk caches incurred by accessing the leaf nodes on host page tables can be eliminated. Extensive evaluation based on the prototype implementation and diverse real-world applications shows that HugeGPT can efficiently reduce address translation overhead and improve application performance in virtualized clouds.
Weiwei Jia 0001, Jiyuan Zhang 0003, Jianchen Shan, Yiming Du, Xiaoning Ding, Tianyin Xu
PACT5
2023 Making Dynamic Page Coalescing Effective on Virtualized Clouds
abstract
Using huge pages has become a mainstream method to reduce address translation overhead for big memory workloads in modern computer systems. To create huge pages, system software usually uses page coalescing methods to dynamically combine contiguous base pages. Though page coalescing methods help effectively reduce address translation overhead on native systems, as the paper shows, their effectiveness is substantially undermined on virtualized platforms.
Weiwei Jia 0001, Jiyuan Zhang 0003, Jianchen Shan, Xiaoning Ding
EuroSys4
2023 Reestablishing Page Placement Mechanisms for Nested Virtualization
abstract
Page placement mechanisms have long been used to reduce cache conflict misses. They become more important in clouds where the emerging way-based cache partitioning is used for better workload isolation but at a cost of increased cache conflicts. However, page placement mechanisms become ineffective in virtualized environments, such as clouds, because the real locations of memory pages (i.e., their host physical addresses) are hidden from guest OSs. The paper proposesxPlaceas a solution to reestablish page placement mechanisms under the nested virtualization configuration. To keep high portability and low overhead,xPlacefollows an approach that creates a synergy between the host and guest VMs, such that the page placement mechanism inside each guest VM becomes effective even if its page placement decisions are made based on the guest physical addresses of memory pages. The paper addresses the technical issues for implementing this approach in the nested virtualization setting, particularly how to create the synergy with the obstacle created by guest hypervisors sitting between the host and guest VMs. Evaluation based on the prototype implementation and diverse real world applications shows thatxPlacecan greatly reduce cache conflicts and improve application performance in the nested environment.
Xiaowei Shang, Weiwei Jia 0001, Jianchen Shan, Xiaoning Ding, Cristian Borcea
IEEE Trans. Cloud Comput.4
2022 Achieving low latency in public edges by hiding workloads mutual interference
abstract
On multi-tenant platforms, such as public clouds and edges, workloads interfere with each other through shared resources. The performance degradation caused by such interference is a notoriously challenging problem. Though many solutions have been proposed for clouds, they can hardly help the application in edges, where workloads are mostly latency-critical, highly dynamic, and more sensitive to interference. Aggressive resource over-provisioning looks to be the only practical solution, albeit it causes significant resource waste.
Weiwei Jia 0001, Jiyuan Zhang 0003, Jianchen Shan, Jing Li 0025, Xiaoning Ding
SoCC5
2022 DFPS: A Distributed Mobile System for Free Parking Assignment
abstract
Cruising for vacant curbside parking spaces causes waste of time, frustration, waste of fuel, and pollution. This problem has been addressed by centralized solutions that perform parking assignments and communicate them to drivers’ smart phones. These solutions suffer, however, from two intrinsic problems: scalability, as the server has to perform intensive computation and communication with the drivers; and privacy, as the drivers have to disclose their destinations to the server. This article proposes DFPS, a distributed mobile system for free parking assignment. DFPS solves the scalability problem by using the drivers’ smart phones to cooperatively compute the parking assignments, and a centralized dispatcher to receive and distribute parking requests to the network of smart phones. The phones of the parked drivers in DFPS are structured in a K-D tree to serve parking requests in a distributed fashion. DFPS removes the computation from the dispatcher and substantially reduces its communication load. DFPS solves the privacy problem through an entropy-based cloaking technique that runs on drivers’ smart phones and conceals drivers’ destinations from the dispatcher. The evaluation demonstrates that DFPS is scalable and obtains better travel time than a centralized system, while protecting the privacy of drivers’ destinations.
Abeer Hakeem, Reza Curtmola, Xiaoning Ding, Cristian Borcea
IEEE Trans. Mob. Comput.3
2021 A Novel Middleware for Efficiently Implementing Complex Cloud-Native SLOs
abstract
Service Level Objectives (SLOs) guide the elasticity of cloud applications, e.g., by deciding when and how much the resources provisioned to an application should be changed. Evaluating SLOs requires metrics, which can be directly measured on the application or system, or, more elaborately, be composed from multiple low-level metrics. The implementation of such metrics and SLOs, the triggering of elasticity strategies, and allowing configurability by the user deploying an application, requires a flexible middleware. In this paper, we present a middleware that provides an orchestrator-independent SLO controller for periodically evaluating SLOs and triggering elasticity strategies, while decoupling SLOs from the elasticity strategies to increase flexibility, and provider-independent services for obtaining low-level metrics and composing them into higher-level metrics. We evaluate our middleware by implementing a motivating use case, featuring a cost efficiency SLO for an application deployed on Kubernetes.
Thomas W. Pusztai, Andrea Morichetta 0002, Víctor Casamayor-Pujol, Schahram Dustdar, Stefan Nastic, Xiaoning Ding, Deepak Vij
CLOUD6
2021 CoPlace: Effectively Mitigating Cache Conflicts in Modern Clouds
abstract
Substantial renovations in hardware cache have been focused on reducing cache interference between workloads recently. However, cache conflicts within each workload are surprisingly overlooked. The paper identifies that cache conflicts cannot be effectively reduced in virtualized clouds. Enhancements for cache partitioning, such as Intel cache allocation technology, make cache conflicts even more serious for cloud workloads. The paper proposes CoPlace as a low overhead and highly portable solution for virtualized clouds. CoPlace enhances the page placement mechanisms implemented in the host OS, such that it can collaborate with the guest OS to reduce cache conflicts. With CoPlace, the guest OS makes page placement decisions; and the host OS helps enforce the decisions. Evaluation based on the prototype implementation in Linux and KVM and diverse real world applications shows that CoPlace can significantly reduce cache conflicts and improve application performance.
Xiaowei Shang, Weiwei Jia 0001, Jianchen Shan, Xiaoning Ding
PACT4
2021 Paratick: Reducing Timer Overhead in Virtual Machines
abstract
To this day, efficient timer management is a major challenge in virtualized environments. Contemporary timekeeping techniques in guest kernels frequently interact with timer hardware, which requires continual and costly hypervisor interference.
Stijn Schildermans, Kris Aerts, Jianchen Shan, Xiaoning Ding
ICPP4
2021 SLO Script: A Novel Language for Implementing Complex Cloud-Native Elasticity-Driven SLOs
abstract
Service Level Objectives (SLOs) allow defining expected performance of cloud services, such that cloud service providers know what they guarantee and service consumers know what to expect. Most approaches focus on low-level SLOs, closely related to resources, e.g., average CPU or memory usage, and are usually bound to specific elasticity controllers. We present SLO Script, a language and accompanying framework, motivated by real-world, industrial needs to allow service providers to define complex, high-level SLOs in an orchestrator-independent manner. The main features of SLO Script include: i) novel abstractions (StronglyTypedSLO) with type safety features, ensuring compatibility between SLOs and elasticity strategies, ii) abstractions that enable decoupling of SLOs from elasticity strategies, iii) a strongly typed metrics API, and iv) an orchestrator-independent object model that enables language extensibility. We present a case study about a real-world, cloud-native application and evaluate our language while implementing a realistic Cost Efficiency SLO.
Thomas W. Pusztai, Andrea Morichetta 0002, Víctor Casamayor-Pujol, Schahram Dustdar, Stefan Nastic, Xiaoning Ding, Deepak Vij
ICWS6
2021 Diagnosing the Interference on CPU-GPU Synchronization Caused by CPU Sharing in Multi-Tenant GPU Clouds
abstract
The GPU-accelerated cloud, enabled by maturing GPU virtualization techniques, has become the most attractive platform for high-performance computing and machine learning workloads. However, it is notoriously challenging to build the multi-tenant GPU cloud where resources, like CPUs and GPUs, can be shared. One well-known and heavily studied reason is that workloads suffer from poor performance isolation and low GPU utilization when GPUs are shared. But little attention has been paid to another fundamental yet under studied problem: how sharing CPUs among GPU instances could affect the workload performance?Targeting this problem, the paper conducts experiments to measure the performance slowdown and vGPU utilization decrease under interference from CPU sharing. The results show that GPU workloads suffer from poor and unpredictable performance and heavy vGPU under-utilization because of CPU sharing. We find that such interference is the result of the complex interplay between the characteristics of CPU-GPU interactions and the special behavior of shared vCPUs: vCPU discontinuity. To diagnose how vCPU discontinuity causes the interference, the paper leverages NVIDIA Nsight Systems for fine-grained profiling and has the following findings: 1) vCPU discontinuity causes inefficient CPU-GPU synchronizations; 2) vCPU discontinuity delays task offloading to the vGPU; 3) Polling-based CPU-GPU synchronization suffers from interference more than blocking-based CPU-GPU synchronization; 4) GPU workloads with frequent task offloads and synchronizations are more vulnerable. Based on the findings, the paper proposes a novel polling-then-blocking CPU-GPU synchronization primitive. Evaluation shows that it can improve the performance by 4.2x.
Youssef Elmougy, Weiwei Jia 0001, Xiaoning Ding, Jianchen Shan
IPCCC3
2021 In-situ workflow auto-tuning through combining component models
abstract
In-situ parallel workflows couple multiple component applications via streaming data transfer to avoid data exchange via shared file systems. Such workflows are challenging to configure for optimal performance due to the huge space of possible configurations. Here, we propose an in-situ workflow auto-tuning method, ALIC, which integrates machine learning techniques with knowledge of in-situ workflow structures to enable automated workflow configuration with a limited number of performance measurements. Experiments with real applications show that ALIC identify better configurations than existing methods given a computer time budget.
Tong Shu, Yanfei Guo, Justin M. Wozniak, Xiaoning Ding, Ian T. Foster, Tahsin M. Kurç
PPoPP4
2021 Bootstrapping in-situ workflow auto-tuning via combining performance models of component applications
abstract
In an in-situ workflow, multiple components such as simulation and analysis applications are coupled with streaming data transfers. The multiplicity of possible configurations necessitates an auto-tuner for workflow optimization. Existing auto-tuning approaches are computationally expensive because many configurations must be sampled by running the whole workflow repeatedly in order to train the auto-tuner surrogate model or otherwise explore the configuration space. To reduce these costs, we instead combine the performance models of component applications by exploiting the analytical workflow structure, selectively generating test configurations to measure and guide the training of a machine learning workflow surrogate model. Because the training can focus on well-performing configurations, the resulting surrogate model can achieve high prediction accuracy for good configurations despite training with fewer total configurations. Experiments with real applications demonstrate that our approach can identify significantly better configurations than other approaches for a fixed computer time budget.
Tong Shu, Yanfei Guo, Justin M. Wozniak, Xiaoning Ding, Ian T. Foster, Tahsin M. Kurç
SC4
2021 Virtualization Overhead of Multithreading in X86 State-of-the-Art & Remaining Challenges
abstract
Despite great advancements in hardware-assisted virtualization of the x86 architecture, certain workloads still suffer significant overhead. This article dissects said overhead in the context of multi-threading. We describe the state-of-the-art, pinpoint challenges, and suggest improvements, aiming to provide a valuable reference to developers and users of virtualization systems alike. We study the virtualization overhead of the PARSEC and SPLASH2X multithreaded benchmarks in a variety of scenarios using a state-of-the-art system. Through controlled experiments, source code analysis and literature review, we quantify the virtualization overhead multithreading still induces and link it to its root causes, after which we suggest possible mitigation strategies. Multithreading still induces high virtualization overhead, mainly caused by synchronization, spinning at user level and NUMA management. The overhead is diverse in nature and embodiment as it is a function of many system and workload properties. System-level solutions are feasible, but often imply difficult trade-offs. Systematic workload optimization is a promising alternative.
Stijn Schildermans, Jianchen Shan, Kris Aerts, Jason Jackrel, Xiaoning Ding
IEEE Trans. Parallel Distributed Syst.5
2020 vSMT-IO: Improving I/O Performance and Efficiency on SMT Processors in Virtualized Clouds
Weiwei Jia 0001, Jianchen Shan, Tsz On Li, Xiaowei Shang, Heming Cui, Xiaoning Ding
USENIX ATC6
2020 Design and Implementation of an Overlay File System for Cloud-Assisted Mobile Apps
abstract
With cloud assistance, mobile apps can offload their resource-demanding computation tasks to the cloud. This leads to a scenario where computation tasks in the same program run concurrently on both the mobile device and the cloud. An important challenge is to ensure that the tasks are able to access and share the files on both the mobile and the cloud in a manner that is efficient, consistent, and transparent to locations. Existing distributed file systems and network file systems do not satisfy these requirements. Current systems for offloading tasks either do not support file access for offloaded tasks or do not offload tasks with file access. The paper addresses this issue by designing and implementing an application-level file system called Overlay File System (OFS). To improve efficiency, OFS maintains and buffers local copies of data sets on both the cloud and the mobile device. OFS ensures consistency and guarantees that all the reads get the latest data. It combines write-invalidate and write-update policies to effectively reduce the network traffic incurred by invalidating/updating stale data copies and to reduce the execution delay when the latest data cannot be accessed locally. To guarantee location transparency, OFS creates a unified view of the data that is location independent and is accessible as local storage. We overcome the challenges caused by the special features of mobile systems on an application-level file system, like the lack of root privilege and state loss when application is killed due to the shortage of resource and implement an easy to deploy prototype of OFS. The paper tests the OFS prototype on Android OS with a real mobile app and real mobile user traces. Extensive experiments show that OFS can effectively support consistent file accesses from computation tasks, no matter whether they are on a mobile device or offloaded to the cloud. In addition, OFS reduce both file access latency and network traffic incurred by file accesses.
Nafize R. Paiker, Jianchen Shan, Cristian Borcea, Narain H. Gehani, Reza Curtmola, Xiaoning Ding
IEEE Trans. Cloud Comput.6
2019 Multi-destination vehicular route planning with parking and traffic constraints
abstract
This paper aims to provide an efficient solution for people in a city who drive their cars to visit several destinations, where they need to park for a while, but do not care about the visiting order. This instance of the multi-destination route planning problem is novel in terms of its constraints: the real-time traffic conditions and the real-time free parking conditions in the city. The paper proposes a novel Multi-Destination Vehicle Route Planning (MDVRP) system to optimize the travel time for all drivers. MDVRP's design has two components: a mobile app running on the drivers' smart phones that submits real-time route requests and guides the drivers toward destinations, and a server in the cloud that optimizes the routes by finding the most efficient order to visit the destinations. MDVRP uses TDTSP-FPA, an algorithm that finds the fastest route to the next destination and also assigns free curbside parking spaces that minimize the total travel time for drivers. We evaluate MDVRP using a driver trip dataset that contains real vehicular mobility traces of over two million drivers from the city of Cologne, Germany. By learning the spatio-temporal distribution of real driver destinations from this dataset, we build a novel experimental platform that simulates real, multi-destination driver trips. Extensive simulations executed over this platform demonstrate that TDTSP-FPA delivers the best performance when compared to three baseline algorithms.
Abeer Hakeem, Narain H. Gehani, Xiaoning Ding, Reza Curtmola, Cristian Borcea
MobiQuitous3
2019 MEFS: Mobile Edge File System for Edge-Assisted Mobile Apps
abstract
Computation offloading is employed by mobile apps running over resource-constrained devices to leverage the cloud in overcoming their resource limits. The advent of the Multi-access Edge Computing (MEC) paradigm further extends the potential opportunities of mobile-cloud offloading, allowing new service provisioning scenarios, such as mobile gaming and multimedia, where responsiveness of mobile devices at the network edge significantly benefits from low latency interactions. However, state-of-the-art offloading platforms for MEC architectures have not addressed the technical challenge of supporting specific file systems for this MEC-enabled class of applications, with components running at three hosting environments, i.e., mobile, edge, and cloud. This paper proposes the Mobile Edge File System (MEFS), an application-level distributed file system designed to be highly resilient and able to efficiently maintain consistency among the mobile, edge, and cloud entities. MEFS supports application handoff through live migration as end devices move between edges. The cloud transparently helps with recovery from faulty edge nodes or in the case of unavailability of edges in the user's proximity. We implemented a MEFS prototype in Android along with MEFS-based MEC-enabled mobile apps. The experimental results show how MEFS can achieve low latency and low overhead.
Domenico Scotece, Nafize R. Paiker, Luca Foschini 0001, Paolo Bellavista, Xiaoning Ding, Cristian Borcea
WOWMOM5
2019 The Moitree middleware for distributed mobile-cloud computing
Hillol Debnath, Mohammad A. Khan, Nafize R. Paiker, Xiaoning Ding, Narain H. Gehani, Reza Curtmola, Cristian Borcea
J. Syst. Softw.4
2019 Multivariate modeling and two-level scheduling of analytic queries
Amit Kumar Nath, Xiaoning Ding, Huansong Fu, Muhib Khan, Weikuan Yu
Parallel Comput.3
2018 Context-Aware File Discovery System for Distributed Mobile-Cloud Apps
abstract
Recent research has proposed middleware to enable efficient distributed apps over mobile-cloud platforms. This paper presents a Context-Aware File Discovery Service (CAFDS) that allows distributed mobile-cloud applications to find and access files of interest shared by collaborating users. CAFDS enables programmers to search for files defined by context and content features, such as location, creation time, or the presence of certain object types within an image file. CAFDS provides low-latency through a cloud-based metadata server, which uses a decision tree to locate the nearest files that satisfy the context and content features requested by applications. We implemented CAFDS in Android and Linux. Experimental results show CAFDS achieves substantially lower latency than peer-to-peer solutions that cannot leverage context information.
Nafize R. Paiker, Xiaoning Ding, Reza Curtmola, Cristian Borcea
CloudCom2
2018 Sentio: Distributed Sensor Virtualization for Mobile Apps
abstract
This paper presents Sentio, a distributed middle-ware designed to provide mobile apps with seamless connectivity to remote sensors when the sensing code and the sensors are not physically on the same device, e.g., when the sensing code is offloaded to the cloud. Sentio presents the apps with virtual sensors that are mapped to remote physical sensors. Virtual sensors can be composed into higher-level sensors, which fuse sensing data from multiple physical sensors. Furthermore, they are mapped to the best available physical sensors when the app starts and re-mapped transparently to other physical sensors at runtime in response to context changes. Sentio was designed to work without modifications to the operating system and to provide low-latency access to remote sensors, which is beneficial to apps with real time-requirements such as mobile games. We have built a prototype of Sentio on Android. We have also developed four apps based on Sentio to understand the programming effort and evaluate the performance. The development of the apps shows that complex sensing tasks can be implemented quickly, benefiting from Sentio's high-level API. The experimental results show that Sentio achieves good real-time performance.
Hillol Debnath, Narain H. Gehani, Xiaoning Ding, Reza Curtmola, Cristian Borcea
PerCom3
2018 Effectively Mitigating I/O Inactivity in vCPU Scheduling
Weiwei Jia 0001, Cheng Wang 0021, Xusheng Chen, Jianchen Shan, Xiaowei Shang, Heming Cui, Xiaoning Ding, Luwei Cheng, Francis C. M. Lau 0001, Yuangang Wang
USENIX ATC7
2017 Rethinking Multicore Application Scalability on Big Virtual Machines
abstract
Virtual machine (VM) sizes keep increasing in the cloud. However, little attention has been paid to analyze and understand the scalability of multicore applications on big VMs with multiple virtual CPUs (VCPUs), assuming that application scalability on VMs can be analyzed in the same ways as that on physical machines (PMs). The paper demonstrates that, since hardware CPU resource is dynamically allocated to VCPUs, the executions of multicore applications on VMs show different scalability from those on PMs. The paper systematically studies how the virtualization of CPU resource changes execution scalability, identifies key application features and system factors that affect execution scalability on VMs, and investigates possible directions to improve scalability. The paper presents a few important findings. First, the execution scalability of applications on VMs is determined by different factors than those on PMs. Second, virtualization and resource sharing can improve scalability by nature. Thus, applications may show better scalability on VMs than on PMs. Linear scalability can be achieved even when there is substantial sequential computation. Third, there is still much space to further improve execution scalability by enhancing system designs. Better scalability can be achieved by increasing allocation period length and/or matching resource allocation and workload distribution.
Jianchen Shan, Weiwei Jia 0001, Xiaoning Ding
ICPADS3
2017 APPLES: Efficiently Handling Spin-lock Synchronization on Virtualized Platforms
abstract
Spin-locks are widely used in software for efficient synchronization. However, they cause serious performance degradation on virtualized platforms, such as the Lock Holder Preemption (LHP) problem and the Lock Waiter Preemption (LWP) problem, due to excessive spinning by virtual CPUs (VCPUs). The excessive spinning occurs when a VCPU waits to acquire a spin-lock. To address the performance degradation, hardware facilities, such as Intel PLE and AMD PF, are provided on processors to preempt VCPUs when they spin excessively. Although these facilities have been predominantly used on mainstream virtualization systems, using them in a manner that achieves the highest performance is still a challenging issue. There are two core problems in using these hardware facilities to reduce excessive spinning. One is to determine the best time to preempt a spinning VCPU (i.e., the selection of spinning thresholds). The other is which VCPU should be scheduled to run after the spinning VCPU is descheduled. Due to the semantic gap between different software layers, the virtual machine monitor (VMM) does not have information about the computation characteristics on VCPUs, which is needed to address the above problems. This makes the problems inherently challenging. We propose a framework named AdPtive Pause-Loop Exiting and Scheduling (APPLES) to address these problems. APPLES monitors the overhead caused by excessive spinning and preempting spinning VCPUs, and periodically adjusts spinning thresholds to reduce the overhead. APPLES also evaluates and schedules “ready” VCPUs in a VM by their potential to reduce the spinning incurred by the spin-lock synchronization. The evaluation is based on the causality and the time of VCPU preemptions. The implementation of APPLES incurs only minimal changes to existing systems (about 100 lines of code in KVM). Experiments show that APPLES can improve performance by 3$\sim$49 percent (14 percent on average) for the workloads with frequent spin-lock operations.
Jianchen Shan, Xiaoning Ding, Narain H. Gehani
IEEE Trans. Parallel Distributed Syst.2
2016 An Overlay File System for cloud-assisted mobile applications
abstract
With cloud assistance, a mobile application can offload its resource-demanding computation tasks to the cloud (public cloud, cloudlet, or personal cloud, etc). This leads to a scenario where computation tasks in the same application run concurrently on both the mobile device and the cloud. These tasks need to save, read, and write files on both the mobile device and the cloud. An important challenge is to ensure that the tasks are able to access and share the files in a manner that is efficient, consistent, and transparent to locations. The paper addresses this issue by designing an application-level file system called Overlay File System (OFS). To improve efficiency, OFS maintains and buffers local copies of data sets on both the cloud and the mobile device. OFS ensures consistency and guarantees that all the reads get the latest data. It combines write-invalidate and write-update policies to effectively reduce the network traffic incurred by invalidating/updating stale data copies and to reduce the application delay when the latest data cannot be accessed locally. To guarantee location transparency, OFS creates an unified view of the data that is location independent and is accessible as local storage. Our experiments show that OFS can effectively support task offloading and efficient execution of offloaded tasks by significantly decreasing both file access latency and network traffic incurred by file accesses.
Jianchen Shan, Nafize R. Paiker, Xiaoning Ding, Narain H. Gehani, Reza Curtmola, Cristian Borcea
MSST3
2016 BCC: Reducing False Aborts in Optimistic Concurrency Control with Low Cost for In-Memory Databases
abstract
The Optimistic Concurrency Control (OCC) method has been commonly used for in-memory databases to ensure transaction serializability --- a transaction will be aborted if its read set has been changed during execution. This simple criterion to abort transactions causes a large proportion of false positives, leading to excessive transaction aborts. Transactions aborted false-positively (i.e. false aborts) waste system resources and can significantly degrade system throughput (as much as 3.68x based on our experiments) when data contention is intensive. Modern in-memory databases run on systems with increasingly parallel hardware and handle workloads with growing concurrency. They must efficiently deal with data contention in the presence of greater concurrency by minimizing false aborts. This paper presents a new concurrency control method named Balanced Concurrency Control (BCC) which aborts transactions more carefully than OCC does. BCC detects data dependency patterns which can more reliably indicate unserializable transactions than the criterion used in OCC. The paper studies the design options and implementation techniques that can effectively detect data contention by identifying dependency patterns with low overhead. To test the performance of BCC, we have implemented it in Silo and compared its performance against that of the vanilla Silo system with OCC and two-phase locking (2PL). Our extensive experiments with TPC-W-like, TPC-C-like and YCSB workloads demonstrate that when data contention is intensive, BCC can increase transaction throughput by more than 3x versus OCC and more than 2x versus 2PL; meanwhile, BCC has comparable performance with OCC for workloads with low data contention.
Yuan Yuan 0014, Kaibo Wang, Rubao Lee, Xiaoning Ding, Spyros Blanas, Xiaodong Zhang 0001
Proc. VLDB Endow.4
2016 A General Approach to Scalable Buffer Pool Management
abstract
In high-end data processing systems, such as databases, the execution concurrency level rises continuously since the introduction of multicore processors. This happens both on premises and in the cloud. For these systems, a buffer pool management of high scalability plays an important role on overall system performance. The scalability of buffer pool management is largely determined by its data replacement algorithm, which is a major component in the buffer pool management. It can seriously degrade the scalability if not designed and implemented properly. The root cause is its use of lock-protected data structures that incurs high contention with concurrent accesses. A common practice is to modify the replacement algorithm to reduce the contention on the lock(s), such as approximating the LRU replacement with the CLOCK algorithm or partitioning the data structures and using distributed locks. Unfortunately, the modification usually compromises the algorithm's hit ratio, a major performance goal. It may also involve significant effort on overhauling the original algorithm design and implementation. This paper provides a general solution to improve the scalability of a buffer pool management using any replacement algorithms for the data processing systems on physical on-premises machines and virtual machines in the cloud. Instead of making a difficult trade-off between the high hit ratio of a replacement algorithm and the low lock contention of its approximation, we design a system framework, called BP-Wrapper, that eliminates almost all lock contention without requiring any changes to an existing algorithm. In BP-Wrapper, we use a dynamic batching technique and a prefetching technique to reduce lock contention and to retain high hit ratio. The implementation of BP-Wrapper in PostgreSQL adds only about 300 lines of C code. It can increase the throughput by up to two folds compared with the replacement algorithms with lock contention when running TPC-C-like and TPC-W-like workloads.
Xiaoning Ding, Jianchen Shan, Song Jiang 0001
IEEE Trans. Parallel Distributed Syst.1
2015 Diagnosing Virtualization Overhead for Multi-threaded Computation on Multicore Platforms
abstract
Hardware-assisted virtualization, as an effective approach to low virtualization overhead, has been dominantly used. However, existing hardware assistance mainly focuses on single-thread performance. Much less attention has been paid to facilitate the efficient interaction between threads, which is critical to the execution of multi-threaded computation on virtualized multicore platforms. This paper aims to answer two questions: 1) what is the performance impact of virtualization on multi-threaded computation, and 2) what are the factors impeding multi-threaded computation from gaining full speed on virtualized platforms. Targeting the first question, the paper measures the virtualization overhead for computation-intensive applications that are designed for multicore processors. We show that some multicore applications still suffer significant performance losses in virtual machines. Even with hardware assistance for reducing virtualization overhead fully enabled, the execution time may be increased by more than 150% when the system is not over-committed, and the system throughput can be reduced by 6x when the system is over-committed. To answer the second question, with experiments, the paper diagnoses the main causes for the performance losses. Focusing the interaction between threads and between VCPUs, the paper identifies and examines a few performance factors, including the intervention of the virtual machine monitor (VMM) to schedule/switch virtual CPUs (VCPUs) and to handle interrupts required by inter-core communication, excessive spinning in user space, and cache-unaware data sharing.
Xiaoning Ding, Jianchen Shan
CloudCom1
2015 APLE: Addressing Lock Holder Preemption Problem with High Efficiency
abstract
On virtualized platforms, Lock Holder Preemption (LHP) is known as a serious problem, which makes virtual CPUs (VCPUs) spin excessively while waiting for locks and seriously degrades performance. To address this problem, hardware facilities, such as Intel PLE and AMD PF, are provided on processors to preempt spinning VCPUs. Though these facilities have been predominantly used on mainstreamvirtualization systems, using them in a manner that achieves the highest performance is still a challenging issue. The core issue in dealing with the LHP problem is to determine the best time to preempt spinning VCPUs (i.e., spinning thresholds). Due to the semantic gap between different software layers, the virtual machine monitor (VMM) does not have the information about whether a VCPU is spinning normally (i.e., waiting for a lock to be released quickly) or is spinning excessively (i.e., waiting for a lock which is currently held by a preempted VCPU and cannot be released quickly). Thus, it cannot determine adequate thresholds for preempting spinning VCPUs to achieve high performance. Preempting spinning VCPUs late wastes system resources. Preempting them prematurely incurs costly context switches between VCPUs and delays lock acquisition. The paper addresses the issue of preempting spinning VCPUs with an end-to-end approach named Adaptive PLE (APLE). APLE monitors the execution efficiency of each VM by collecting the overhead incurred by wasteful spinning and wasteful VCPU switches. Then, it periodically adjusts the spinning threshold to reduce the overhead and increase the execution efficiency of the VM. The implementation of APLE incurs only minimal changes to existing systems (about 80 lines of code in KVM). The experiments with multicore workloads show that APLE can improve throughput by up to 68%.
Jianchen Shan, Xiaoning Ding, Narain H. Gehani
CloudCom2
2015 Hetero-DB: Next Generation High-Performance Database Systems by Best Utilizing Heterogeneous Computing and Storage Resources
Kai Zhang 0006, Feng Chen 0005, Xiaoning Ding, Yin Huai, Rubao Lee, Kaibo Wang, Yuan Yuan 0014, Xiaodong Zhang 0001
J. Comput. Sci. Technol.3
2014 GDM: device memory management for gpgpu computing
abstract
GPGPUs are evolving from dedicated accelerators towards mainstream commodity computing resources. During the transition, the lack of system management of device memory space on GPGPUs has become a major hurdle. In existing GPGPU systems, device memory space is still managed explicitly by individual applications, which not only increases the burden of programmers but can also cause application crashes, hangs, or low performance.
Kaibo Wang, Xiaoning Ding, Rubao Lee, Shinpei Kato, Xiaodong Zhang 0001
SIGMETRICS2
2014 Gleaner: Mitigating the Blocked-Waiter Wakeup Problem for Virtualized Multicore Applications
Xiaoning Ding, Phillip B. Gibbons, Michael A. Kozuch, Jianchen Shan
USENIX ATC1
2014 Concurrent Analytical Query Processing with GPUs
abstract
In current databases, GPUs are used as dedicated accelerators to process each individual query. Sharing GPUs among concurrent queries is not supported, causing serious resource underutilization. Based on the profiling of an open-source GPU query engine running commonly used single-query data warehousing workloads, we observe that the utilization of main GPU resources is only up to 25%. The underutilization leads to low system throughput. To address the problem, this paper proposes concurrent query execution as an effective solution. To efficiently share GPUs among concurrent queries for high throughput, the major challenge is to provide software support to control and resolve resource contention incurred by the sharing. Our solution relies on GPU query scheduling and device memory swapping policies to address this challenge. We have implemented a prototype system and evaluated it intensively. The experiment results confirm the effectiveness and performance advantage of our approach. By executing multiple GPU queries concurrently, system throughput can be improved by up to 55% compared with dedicated processing.
Kaibo Wang, Kai Zhang 0006, Yuan Yuan 0014, Rubao Lee, Xiaoning Ding, Xiaodong Zhang 0001
Proc. VLDB Endow.6
2013 A Prefetching Scheme Exploiting both Data Layout and Access History on Disk
abstract
Prefetching is an important technique for improving effective hard disk performance. A prefetcher seeks to accurately predict which data will be requested and load it ahead of the arrival of the corresponding requests. Current disk prefetch policies in major operating systems track access patterns at the level of file abstraction. While this is useful for exploiting application-level access patterns, for two reasons file-level prefetching cannot realize the full performance improvements achievable by prefetching. First, certain prefetch opportunities can only be detected by knowing the data layout on disk, such as the contiguous layout of file metadata or data from multiple files. Second, nonsequential access of disk data (requiring disk head movement) is much slower than sequential access, and the performance penalty for mis-prefetching a randomly located block, relative to that of a sequential block, is correspondingly greater.
Song Jiang 0001, Xiaoning Ding, Yuehai Xu, Kei Davis
ACM Trans. Storage2
2012 BWS: balanced work stealing for time-sharing multicores
abstract
Running multithreaded programs in multicore systems has become a common practice for many application domains. Work stealing is a widely-adopted and effective approach for managing and scheduling the concurrent tasks of such programs. Existing work-stealing schedulers, however, are not effective when multiple applications time-share a single multicore---their management of steal-attempting threads often causes unbalanced system effects that hurt both workload throughput and fairness.
Xiaoning Ding, Kaibo Wang, Phillip B. Gibbons, Xiaodong Zhang 0001
EuroSys1
2011 SRM-buffer: an OS buffer management technique to prevent last level cache from thrashing in multicores
abstract
Buffer caches in operating systems keep active file blocks in memory to reduce disk accesses. Related studies have been focused on how to minimize buffer misses and the caused performance degradation. However, the side effects and performance implications of accessing the data in buffer caches (i.e. buffer cache hits) have not been paid attention. In this paper, we show that accessing buffer caches can cause serious performance degradation on multicores, particularly with shared last level caches (LLCs). There are two reasons for this problem. First, data in files normally have weaker localities than data objects in virtual memory spaces. Second, due to the shared structure of LLCs on multicore processors, an application accessing the data in a buffer cache may flush the to-be-reused data of its co-running applications from the shared LLC and significantly slow down these applications.
Xiaoning Ding, Kaibo Wang, Xiaodong Zhang 0001
EuroSys1
2011 ULCC: a user-level facility for optimizing shared cache performance on multicores
abstract
Scientific applications face serious performance challenges on multicore processors, one of which is caused by access contention in last level shared caches from multiple running threads. The contention increases the number of long latency memory accesses, and consequently increases application execution times. Optimizing shared cache performance is critical to reduce significantly execution times of multi-threaded programs on multicores. However, there are two unique problems to be solved before implementing cache optimization techniques on multicores at the user level. First, available cache space for each running thread in a last level cache is difficult to predict due to access contention in the shared space, which makes cache conscious algorithms for single cores ineffective on multicores. Second, at the user level, programmers are not able to allocate cache space at will to running threads in the shared cache, thus data sets with strong locality may not be allocated with sufficient cache space, and cache pollution can easily happen. To address these two critical issues, we have designed ULCC (User Level Cache Control), a software runtime library that enables programmers to explicitly manage and optimize last level cache usage by allocating proper cache space for different data sets of different threads. We have implemented ULCC at the user level based on a page-coloring technique for last level cache usage management. By means of multiple case studies on an Intel multicore processor, we show that with ULCC, scientific applications can achieve significant performance improvements by fully exploiting the benefit of cache optimization algorithms and by partitioning the cache space accordingly to protect frequently reused data sets and to avoid cache pollution. Our experiments with various applications show that ULCC can significantly improve application performance by nearly 40%.
Xiaoning Ding, Kaibo Wang, Xiaodong Zhang 0001
PPoPP1
2010 Splitter: a proxy-based approach for post-migration testing of web applications
abstract
The benefits of virtualized IT environments, such as compute clouds, have drawn interested enterprises to migrate their applications onto new platforms to gain the advantages of reduced hardware and energy costs, increased flexibility and deployment speed, and reduced management complexity. However, the process of migrating a complex application takes a considerable amount of effort, particularly when performing post-migration testing to verify that the application still functions correctly in the target environment. The traditional approach of test case generation and execution can take weeks and synthetic test cases may not adequately reflect actual application usage.
Xiaoning Ding, Hai Huang 0002, Yaoping Ruan, Anees Shaikh, Brian Peterson, Xiaodong Zhang 0001
EuroSys1
2009 Soft-OLP: Improving Hardware Cache Performance through Software-Controlled Object-Level Partitioning
abstract
Performance degradation of memory-intensive programs caused by the LRU policy's inability to handle weak-locality data accesses in the last level cache is increasingly serious for two reasons. First, the last-level cache remains in the CPU's critical path, where only simple management mechanisms, such as LRU, can be used, precluding some sophisticated hardware mechanisms to address the problem. Second, the commonly used shared cache structure of multi-core processors has made this critical path even more performance-sensitive due to intensive inter-thread contention for shared cache resources. Researchers have recently made efforts to address the problem with the LRU policy by partitioning the cache using hardware or OS facilities guided by run-time locality information. Such approaches often rely on special hardware support or lack enough accuracy. In contrast, for a large class of programs, the locality information can be accurately predicted if access patterns are recognized through small training runs at the data object level. To achieve this goal, we present a system-software framework referred to as Soft-OLP (Software-based Object-Level cache Partitioning). We first collect per-object reuse distance histograms and inter-object interference histograms via memory-trace sampling. With several low-cost training runs, we are able to determine the locality patterns of data objects. For the actual runs, we categorize data objects into different locality types and partition the cache space among data objects with a heuristic algorithm, in order to reduce cache misses through segregation of contending objects. The object-level cache partitioning framework has been implemented with a modified Linux kernel, and tested on a commodity multi-core processor. Experimental results show that in comparison with a standard L2 cache managed by LRU, Soft-OLP significantly reduces the execution time by reducing L2 cache misses across inputs for a set of single- and multi-threaded programs from the SPEC CPU2000 benchmark suite, NAS benchmarks and a computational kernel set.
Qingda Lu, Jiang Lin, Xiaoning Ding, Zhao Zhang 0010, Xiaodong Zhang 0001, P. Sadayappan
PACT3
2009 BP-Wrapper: A System Framework Making Any Replacement Algorithms (Almost) Lock Contention Free
abstract
In a high-end database system, the execution concurrency level rises continuously in a multiprocessor environment due to the increase in number of concurrent transactions and the introduction of multi-core processors. A new challenge for buffer management to address is to retain its scalability in responding to the highly concurrent data processing demands and environment. The page replacement algorithm, a major component in the buffer management, can seriously degrade the system's performance if the algorithm is not implemented in a scalable way. A lock-protected data structure is used in most replacement algorithms, where high contention is caused by concurrent accesses. A common practice is to modify a replacement algorithm to reduce the contention, such as to approximate the LRU replacement with the clock algorithm. Unfortunately, this type of modification usually hurts hit ratios of original algorithms. This problem may not exist or can be tolerated in an environment of low concurrency, thus has not been given enough attention for a long time. In this paper, instead of making a trade-off between the high hit ratio of a replacement algorithm and the low lock contention of its approximation, we propose a system framework, called BP-Wrapper, that (almost) eliminates lock contention for any replacement algorithm without requiring any changes to the algorithm. In BP-Wrapper, we use batching and prefetching techniques to reduce lock contention and to retain high hit ratio. The implementation of BP-Wrapper in PostgreSQL version 8.2 adds only about 300 lines of C code. It can increase the throughput up to two folds compared with the replacement algorithms with lock contention when running TPC-C-like and TPC-W-like workloads.
Xiaoning Ding, Song Jiang 0001, Xiaodong Zhang 0001
ICDE1
2009 Enabling software management for multicore caches with a lightweight hardware support
abstract
The management of shared caches in multicore processors is a critical and challenging task. Many hardware and OS-based methods have been proposed. However, they may be hardly adopted in practice due to their non-trivial overheads, high complexities, and/or limited abilities to handle increasingly complicated scenarios of cache contention caused by many-cores.
Jiang Lin, Qingda Lu, Xiaoning Ding, Zhao Zhang 0010, Xiaodong Zhang 0001, P. Sadayappan
SC3
2009 MCC-DB: Minimizing Cache Conflicts in Multi-core Processors for Databases
abstract
In a typical commercial multi-core processor, the last level cache (LLC) is shared by two or more cores. Existing studies have shown that the shared LLC is beneficial to concurrent query processes with commonly shared data sets. However, the shared LLC can also be a performance bottleneck to concurrent queries, each of which has private data structures, such as a hash table for the widely used hash join operator, causing serious cache conflicts. We show that cache conflicts on multi-core processors can significantly degrade overall database performance. In this paper, we propose a hybrid system method called MCC-DB for accelerating executions of warehouse-style queries, which relies on the DBMS knowledge of data access patterns to minimize LLC conflicts in multi-core systems through an enhanced OS facility of cache partitioning. MCC-DB consists of three components: (1) a cacheaware query optimizer carefully selects query plans in order to balance the numbers of cache-sensitive and cache-insensitive plans; (2) a query execution scheduler makes decisions to co-run queries with an objective of minimizing LLC conflicts; and (3) an enhanced OS kernel facility partitions the shared LLC according to each query's cache capacity need and locality strength. We have implemented MCC-DB by patching the three components in PostgreSQL and Linux kernel. Our intensive measurements on an Intel multi-core system with warehouse-style queries show that MCC-DB can reduce query execution times by up to 33%.
Rubao Lee, Xiaoning Ding, Feng Chen 0005, Qingda Lu, Xiaodong Zhang 0001
Proc. VLDB Endow.2
2008 A Concurrency Control Mechanism for Composite Service Supporting User-Defined Relaxed Atomicity
abstract
To ensure the relaxed atomicity of transactional composite service (TCS), a lot of relaxed atomicity models have been proposed, but they don't handle the concurrency control problem specific to their relaxed atomicity model, which can leads to unexpected results when multiple TCS run concurrently. In this paper, we propose a novel concurrency control mechanism to support the user-defined relaxed atomicity model for service composition. Based on the analysis of the two typical kinds of inconsistency caused by cascade compensation and cycle dependency, concurrency control mechanism to confine the forming and changing of dependency is presented, which can eliminate the inconsistencies while maintaining the relaxed atomicity of individual TCS.
Zongtao Zhao, Jun Wei 0001, Xiaoning Ding
COMPSAC4
2008 Gaining insights into multicore cache partitioning: Bridging the gap between simulation and real systems
abstract
Cache partitioning and sharing is critical to the effective utilization of multicore processors. However, almost all existing studies have been evaluated by simulation that often has several limitations, such as excessive simulation time, absence of OS activities and proneness to simulation inaccuracy. To address these issues, we have taken an efficient software approach to supporting both static and dynamic cache partitioning in OS through memory address mapping. We have comprehensively evaluated several representative cache partitioning schemes with different optimization objectives, including performance, fairness, and quality of service (QoS). Our software approach makes it possible to run the SPEC CPU2006 benchmark suite to completion. Besides confirming important conclusions from previous work, we are able to gain several insights from whole-program executions, which are infeasible from simulation. For example, giving up some cache space in one program to help another one may improve the performance of both programs for certain workloads due to reduced contention for memory bandwidth. Our evaluation of previously proposed fairness metrics is also significantly different from a simulation-based study. The contributions of this study are threefold. (1) To the best of our knowledge, this is a highly comprehensive execution- and measurement-based study on multicore cache partitioning. This paper not only confirms important conclusions from simulation-based studies, but also provides new insights into dynamic behaviors and interaction effects. (2) Our approach provides a unique and efficient option for evaluating multicore cache partitioning. The implemented software layer can be used as a tool in multicore performance evaluation and hardware design. (3) The proposed schemes can be further refined for OS kernels to improve performance.
Jiang Lin, Qingda Lu, Xiaoning Ding, Zhao Zhang 0010, Xiaodong Zhang 0001, P. Sadayappan
HPCA3
2008 Automatic Software Fault Diagnosis by Exploiting Application Signatures
Xiaoning Ding, Hai Huang 0002, Yaoping Ruan, Anees Shaikh, Xiaodong Zhang 0001
LISA1
2007 DiskSeen: Exploiting Disk Layout and Access History to Enhance I/O Prefetch
Xiaoning Ding, Song Jiang 0001, Feng Chen 0005, Kei Davis, Xiaodong Zhang 0001
USENIX ATC1
2007 A performance study of BitTorrent-like peer-to-peer systems
abstract
This paper presents a performance study of BitTorrent-like P2P systems by modeling, based on extensive measurements and trace analysis. Existing studies on BitTorrent systems are single-torrent based and usually assume the process of request arrivals to a torrent is Poisson-like. However, in reality, most BitTorrent peers participate in multiple torrents and file popularity changes over time. Our study of representative BitTorrent traffic provides insights into the evolution of single-torrent systems and several new findings regarding the limitations of BitTorrent systems: (1) Due to the exponentially decreasing peer arrival rate in a torrent, the service availability of the corresponding file becomes poor quickly, and eventually it is hard to locate and download this file. (2) Client performance in the BitTorrent-like system is unstable, and fluctuates significantly with the changes of the number of online peers. (3) Existing systems could provide unfair services to peers, where a peer with a higher downloading speed tends to download more and upload less. Motivated by the analysis and modeling results, we have further proposed a graph based model to study interactions among multiple torrents. Our model quantitatively demonstrates that inter-torrent collaboration is much more effective than stimulating seeds to serve longer for addressing the service unavailability in BitTorrent systems. An architecture for inter-torrent collaboration under an exchange based instant incentive mechanism is also discussed and evaluated by simulations.
Lei Guo 0004, Songqing Chen, Enhua Tan, Xiaoning Ding, Xiaodong Zhang 0001
IEEE J. Sel. Areas Commun.5
2007 Cooperative Relay Service in a Wireless LAN
abstract
As a family of wireless local area network (WLAN) protocols between physical layer and higher layer protocols, IEEE 802.11 has to accommodate the features and requirements of both ends. However, current practice has addressed the problems of these two layers separately and is far from satisfactory. On one end, due to varying channel conditions, WLANs have to provide multiple physical channel rates to support various signal qualities. A low channel rate station not only suffers low throughput, but also significantly degrades the throughput of other stations. On the other end, the power saving mechanism of 802.11 is ineffective in TCP-based communications, in which the wireless network interface (WNI) has to stay awake to quickly acknowledge senders, and hence, the energy is wasted on channel listening during idle awake time. In this paper, considering the needs of both ends, we utilize the idle communication power of the WNI to provide a Cooperative Relay Service (CRS) for WLANs with multiple channel rates. We characterize energy efficiency as energy per bit, instead of energy per second. In CRS, a high channel rate station relays data frames as a proxy between its neighboring stations with low channel rates and the Access Point, improving their throughput and energy efficiency. Different from traditional relaying approaches, CRS compensates a proxy for the energy consumed in data forwarding. The proxy obtains additional channel access time from its clients, leading to the increase of its own throughput without compromising its energy efficiency. Extensive experiments are conducted through a prototype implementation and ns-2 simulations to evaluate our proposed CRS. The experimental results show that CRS achieves significant performance improvements for both low and high channel rate stations
Lei Guo 0004, Xiaoning Ding, Haining Wang 0001, Qun Li 0001, Songqing Chen, Xiaodong Zhang 0001
IEEE J. Sel. Areas Commun.2
2007 A buffer cache management scheme exploiting both temporal and spatial localities
abstract
On-disk sequentiality of requested blocks, or their spatial locality, is critical to real disk performance where the throughput of access to sequentially-placed disk blocks can be an order of magnitude higher than that of access to randomly-placed blocks. Unfortunately, spatial locality of cached blocks is largely ignored, and only temporal locality is considered in current system buffer cache managements. Thus, disk performance for workloads without dominant sequential accesses can be seriously degraded. To address this problem, we propose a scheme called DULO ( DU al LO cality) which exploits both temporal and spatial localities in the buffer cache management. Leveraging the filtering effect of the buffer cache, DULO can influence the I/O request stream by making the requests passed to the disk more sequential, thus significantly increasing the effectiveness of I/O scheduling and prefetching for disk performance improvements. We have implemented a prototype of DULO in Linux 2.6.11. The implementation shows that DULO can significantly increases disk I/O throughput for real-world applications such as a Web server, TPC benchmark, file system benchmark, and scientific programs. It reduces their execution times by as much as 53%.
Xiaoning Ding, Song Jiang 0001, Feng Chen 0005
ACM Trans. Storage1
2006 A Locality-Aware Cooperative Cache Management Protocol to Improve Network File System Performance
abstract
In a distributed environment the utilization of file buffer caches in different clients may vary greatly. Cooperative caching is used to increase cache utilization by coordinating the usage of distributed caches. Existing cooperative caching protocols mainly address organizational issues, paying little attention to exploiting locality of file access patterns. We propose a locality-aware cooperative caching protocol, called LAC, that is based on analysis and manipulation of data block reuse distance to effectively predict cache utilization and the probability of data reuse. Using a dynamically controlled synchronization technique, we make local information consistently comparable among clients. The system is highly scalable in the sense that global coordination is achieved without centralized control.
Song Jiang 0001, Fabrizio Petrini, Xiaoning Ding, Xiaodong Zhang 0001
ICDCS3
2006 User-Defined Atomicity Constraint: A More Flexible Transaction Model for Reliable Service Composition
Xiaoning Ding, Jun Wei 0001, Tao Huang 0001
ICFEM1
2006 Exploiting Idle Communication Power to Improve Wireless Network Performance and Energy Efficiency
abstract
Abstract — As a family of wireless local area network (WLAN) protocols between physical layer and higher-layer protocols, IEEE 802.11 has to accommodate the features and requirements of both ends. However, current practice has addressed the problems separately and is far from being satisfactory. On the one end, due to varying channel conditions, WLANs have to provide multiple data channel rates to support various bit error rates. A low channel rate station not only suffers low throughput itself, but also significantly degrades the throughput of other stations. On the other end, TCP is not energy efficient running on 802.11. This is because a wireless network interface (WNI) has to stay awake to generate timely acknowledgments during a TCP session, and hence, the energy consumed during idle awake time is wasted for channel listening. In this paper, considering the needs of both ends, we utilize the idle communication power of the WNI to improve the throughput and energy efficiency of stations in WLANs supporting multiple channel rates. We characterize the energy efficiency as energy per bit, instead of energy per second. Based on modeling and analysis, we propose a data forwarding mechanism and an energy-aware channel allocation mechanism. In such a system, a high channel rate station relays data frames between its neighboring stations with low channel rates and Access Point, improving their throughput and energy efficiency. Different from traditional relaying approaches, our scheme compensates for the energy consumption for data forwarding. The forwarding station gets additional channel access time from its beneficiaries, leading to the increase of its own throughput without compromising its energy efficiency. We implement a prototype of our proposed system and evaluate it through extensive experiments. Our results show significant performance improvements for both low and high channel rate stations. I.
Lei Guo 0004, Xiaoning Ding, Haining Wang 0001, Qun Li 0001, Songqing Chen, Xiaodong Zhang 0001
INFOCOM2
2006 MESA: reducing cache conflicts by integrating static and run-time methods
abstract
The paper proposes MESA (Multicoloring with Embedded Skewed Associativity), a novel cache indexing scheme that integrates dynamic page coloring with static skewed associativity to reduce conflicts in L2/L3 caches with a small degree of associativity. MESA associates multiple cache pages (colors) with each virtual memory page and uses two-level skewed associativity, first to map a page to a different color in each bank of the cache, and then to disperse the lines of a page across the banks and within the colors of the page. MESA is a multi-grained cache indexing scheme that combines the best of two worlds, page coloring and skewed associativity. We also propose a novel cache management scheme based on page remapping, which uses cache miss imbalance between colors in each bank as the metric to track conflicts and trigger remapping. We evaluate MESA using 24 benchmarks from multiple application domains and with various degrees of sensitivity to conflict misses, on both an in-order issue processor (using complete system simulation) and an out-of-order issue processor (using SimpleScalar). MESA outperforms skewed associativity, prime modulo hashing, and dynamic page coloring schemes proposed earlier. Compared to a 4-way associative cache, MESA can provide as much as 76% improvement in IPC.
Xiaoning Ding, Dimitrios S. Nikolopoulos, Song Jiang 0001, Xiaodong Zhang 0001
ISPASS1
2006 An application-semantics-based relaxed transaction model for internetware
Tao Huang 0001, Xiaoning Ding, Jun Wei 0001
Sci. China Ser. F Inf. Sci.2
2005 DULO: An Effective Buffer Cache Management Scheme to Exploit Both Temporal and Spatial Localities
Song Jiang 0001, Xiaoning Ding, Feng Chen 0005, Enhua Tan, Xiaodong Zhang 0001
FAST2
2005 Multigrain parallel Delaunay Mesh generation: challenges and opportunities for multithreaded architectures
abstract
Given the importance of parallel mesh generation in large-scale scientific applications and the proliferation of multilevel SMT-based architectures, it is imperative to obtain insight on the interaction between meshing algorithms and these systems. We focus on Parallel Constrained Delaunay Mesh (PCDM) generation. We exploit coarse-grain parallelism at the subdomain level and fine-grain at the element level. This multigrain data parallel approach targets clusters built from low-end, commercially available SMTs. Our experimental evaluation shows that current SMTs are not capable of executing fine-grain parallelism in PCDM. However, experiments on a simulated SMT indicate that with modest hardware support it is possible to exploit fine-grain parallelism opportunities. The exploitation of fine-grain parallelism results to higher performance than a pure MPI implementation and closes the gap between the performance of PCDM and the state-of-the-art sequential mesher on a single physical processor. Our findings extend to other adaptive and irregular multigrain, parallel algorithms.
Christos D. Antonopoulos, Xiaoning Ding, Andrey N. Chernikov, Filip Blagojevic, Dimitrios S. Nikolopoulos, Nikos Chrisochoides
ICS2
2005 Measurements, Analysis, and Modeling of BitTorrent-like Systems
Lei Guo 0004, Songqing Chen, Enhua Tan, Xiaoning Ding, Xiaodong Zhang 0001
Internet Measurement Conference5