EDBT 2026 Demo / reviewers in the wild / expert
Song Wu 0001
dblp:23/1092-1
· DBLP profile ↗
148ranked-venue papers
26as first author
46since 2021 · last 2026
0000-0001-8690-127XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 100 · 19 first-author · 31 since 2021Software engineering, systems software and programming languages · 12 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 2 since 2021Computer networks · 6 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5Security and privacy · 4 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AEP: Achieving Hierarchical Fault Tolerance in DSM Through Atomic Execution ProtectionabstractRecent advances in CXL 3.0 have revitalized Distributed Shared Memory (DSM), enabling zero-copy sharing across machine nodes at near-NUMA latency and making shared-memory communication practical for distributed applications. However, DSM memory allocators remain vulnerable to partial failures: client crashes during critical, non-atomic memory operations can corrupt metadata and stall the entire system. Existing solutions either block all clients during recovery or impose heavy runtime overhead due to distributed fault-tolerance protocols. We propose Atomic Execution Protection (AEP), a kernel-assisted fault tolerance mechanism that guarantees atomic completion of user-space critical operations by deferring termination until protected regions finish. AEP tolerates intra-node partial failures without costly distributed coordination. To extend beyond a single node, we further design AEP-DSM, a hierarchical DSM allocator that combines AEP-protected node-level allocators with cross-node coordination. Our Linux AEP prototype integrated with a cross-node DSM consistency protocol (CXL-SHM) achieves near-native local performance and improves throughput by up to 13.3x over state-of-the-art DSM allocators. Zixuan Wang 0030, Hang Huang, Jia Rao, Hui Lu 0001, Hao Fan 0006, Song Wu 0001, Hai Jin 0001 |
EuroSys | 8 |
| 2026 | Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric ApproachabstractServerless computing offers a compelling paradigm for deploying machine learning inference workflows composed of heterogeneous CPU and GPU functions. However, existing data-passing solutions in serverless systems primarily rely on host memory for data exchange (host-centric), leading to substantial data movement and salient I/O overhead. Moreover, modern GPU communication libraries (e.g., NCCL, NVSHMEM, UCX) are ill-suited to serverless environments, suffering from redundant data copies, underutilized transfer bandwidth, and inefficient temporary GPU storage. Hao Wu 0032, Yaochen Liu, Minchen Yu, Qizhen Weng 0001, Junxiao Deng, Hao Fan 0006, Song Wu 0001, Wei Wang 0030, Hai Jin 0001 |
EuroSys | 8 |
| 2026 | SRAG: A Lightweight and Specialized Retrieval-augmented Generation System at the EdgeabstractRetrieval-augmented generation (RAG) has shown strong potential for deploying large language models at the edge, yet existing designs largely rely on generic and monolithic knowledge bases that are poorly matched to the heterogeneous queries and resource-constrained edge computing environments. Through extensive empirical analysis, we find that domain-specialized knowledge bases, when deployed on individual edge servers, deliver substantially higher retrieval accuracy and generation quality than generic knowledge bases under identical resource budgets. Based on this, we propose SRAG, a distributed RAG system that enforces knowledge specialization at the edge. Each edge server maintains a domain-aware specialized knowledge base by retaining domain-aligned knowledge and decoupling out-of-domain content. SRAG uses a buffer-based knowledge migration mechanism to redistribute out-of-domain content to better-matched edge servers, enabling efficient global knowledge utilization without central coordination. To handle domain-mismatched queries, SRAG employs lightweight cross-node routing guided by compact metadata summaries, avoiding full knowledge replication. Together, these mechanisms form an end-to-end workflow for decentralized edge RAG. Experiments show that SRAG improves retrieval relevance, generation quality, and storage efficiency, while reducing end-to-end latency. Ruikun Luo, Zihan Xing, Lin Gu 0002, Song Wu 0001, Hai Jin 0001, Xiaoyu Xia 0001 |
SIGIR | 4 |
| 2026 | KPT-Fork: Enhancing User Data Isolation via Kernel Page Table ForkabstractModern operating systems (OS) utilize a shared kernel address design, enabling the OS kernel to access physical memory addresses conveniently and efficiently. However, this design introduces vulnerabilities that attackers can exploit to arbitrarily read/write content from any kernel-space address, known as data-oriented attacks. Existing MMU-based isolation approaches face inherent challenges in scalability, compatibility, and context switch overhead. Inspired by the reference kernel page table, we propose to use user-private reference kernel page tables (UP-KPTs) as the foundation for data isolation in a shared kernel address space. A UP-KPT is a customized kernel page table tailored for a specific container, encompassing address mappings of private user data that are not visible to other containers. To protect sensitive system-wide data, an administrator can unmap such data from UP-KPTs and make it inaccessible to potentially malicious containers. UP-KPT provides strong data isolation while maintaining full compatibility with legacy applications and existing kernel components. We develop KPT-fork, a kernel isolation framework that enables the creation and management of UP-KPTs. KPT-fork entails two important designs: 1) a UP-KPT structure that enables data isolation while preserving the simplicity, convenience, and efficiency of a shared kernel page table; 2) a private memory allocator that efficiently transforms static and direct address mappings of private data in shared kernel addresses into dynamic user-private mappings. We demonstrate through three case studies that KPT-fork can be extended to protect various types of data. Our evaluation shows that KPT fork can defend against a wide range of data-oriented attacks while incurring negligible overhead. Zixuan Wang 0030, Lisong Pan, Jia Rao, Hao Fan 0006, Song Wu 0001 |
IEEE Trans. Cloud Comput. | 7 |
| 2026 | Disguiser: A Privacy-Preserving Scheme for Efficient Edge User AllocationabstractMulti-access edge computing (MEC) has garnered increasing attention from users due to low web service latency. Edge servers deployed near base stations have limited resources and can only serve users within their coverage areas. Therefore, efficiently allocating users to the appropriate edge servers is crucial for significantly enhancing system performance. Traditional edge user allocation (EUA) strategies often rely on precise user location, leading to significant privacy leakage risks and weakening users' trust. To tackle this challenge, this paper presents Disguiser, a novel privacy-preserving scheme designed to achieve efficient edge user allocation while safeguarding user location privacy. Disguiser employs a Laplace noise-based location obfuscation mechanism to ensure users' privacy. To resolve the user allocation problem after location obfuscation, which involves balancing real-time performance and accuracy, we design a two-stage user allocation algorithm, TEUA, consisting of the LR-EUA initial allocation algorithm and the MR-EUA reallocation algorithm. First, a novel EUA algorithm, LR-EUA, is integrated into Disguiser, innovatively combining linear relaxation with a greedy approach. Additionally, Disguiser further minimizes resource wastage in MEC systems caused by location obfuscation by employing the MR-EUA algorithm at base stations. Experimental results show that Disguiser effectively balances privacy protection with system performance, substantially outperforming existing state-of-the-art methods. Ruikun Luo, Qiang He 0001, Feifei Chen 0001, Song Wu 0001, Hai Jin 0001, Jing Yang 0051, Yuan Gao 0031, Iqbal Gondal, Xiaoyu Xia 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2026 | Puffer: A Serverless Platform Based on Vertical Memory ScalingabstractThis paper quantitatively analyses the potential of vertical scaling MicroVMs in serverless computing. Our analysis shows that under real-world serverless workloads, vertical scaling can significantly improve execution performance and resource utilization. However, we also find that the memory scaling of MicroVMs is the bottleneck that hinders vertical scaling from reaching the performance ceiling. We propose Faascale, a novel mechanism that efficiently scales the memory of MicroVMs for serverless applications. Faascale employs a series of techniques to tackle this bottleneck: 1) it sizes up/down the memory for a MicroVM by blocks that bind with a function instance instead of general pages; and 2) it pre-populates physical memory for function instances to reduce the delays introduced by the lazy-population. Compared with existing memory scaling mechanisms, Faascale improves the memory scaling efficiency by 2 to 3 orders of magnitude. Based on Faascale, we realize a serverless platform, named Puffer. Experiments conducted on eight serverless benchmark functions demonstrate that compared with horizontal scaling strategies, Puffer reduces time for cold-starting MicroVMs by 89.01%, improves memory utilization by 17.66%, and decreases functions execution time by 23.93% on average. Hao Fan 0006, Kun Wang 0059, Haibo Mi, Song Wu 0001, Chen Yu 0003 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2026 | Arcus: Fast and Reliable Function State I/O for Serverless Computing With Log-Cache Co-Design
Hanxiang Huang, Junhui Peng, Song Wu 0001, Hai Jin 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2026 | Popularity Uncertainty-Aware Online Edge Data MigrationabstractAs an emerging distributed computing paradigm, edge computing (EC) brings computing and storage capabilities to the network edge to support low-latency services. Service providers can cache popular data on edge servers to greatly reduce service latency and improve service quality. However, data popularity in EC environments is often dynamic and stochastic. Given the limited storage resources of edge servers, when data popularity changes, data may need to be migrated from edge servers with lower data popularity to those with higher data popularity. Data migration between edge servers can effectively reduce the high transmission costs associated with transferring data from the cloud to the edge. Most existing data migration approaches are designed for static EC environments, assuming that user access patterns and data demands remain constant over time. Even in dynamic EC environments, many approaches oversimplify data popularity by treating it as static, failing to model its fluctuations over time. In realistic EC scenarios, unexpected situations, such as sudden shifts in user preferences for specific data at certain times, can cause significant fluctuations in data popularity, leading to increased data retrieval latency and inefficient storage management. To address this uncertainty, this paper proposes a popularity uncertainty-aware online data migration approach that combines distributed robust optimization and Lyapunov optimization to minimize the data retrieval latency and migration cost for service providers. Specifically, it handles data popularity uncertainty in a data-driven manner through robust optimization. Then, it employs Lyapunov optimization to decompose the continuous optimization problem into multiple single-slot online optimization problems. Extensive experiments on a widely used real-world dataset confirm the effectiveness of this approach and its significant advantages over state-of-the-art approaches. Ruikun Luo, Zixiao Feng, Qiang He 0001, Feifei Chen 0001, Song Wu 0001, Hai Jin 0001, Xiaoyu Xia 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2026 | EdgeDup: Popularity-Aware Communication-Efficient Decentralized Edge Data DeduplicationabstractData deduplication, originally designed for cloud storage systems, is increasingly popular in edge storage systems due to the costly and limited resources and prevalent data redundancy in edge computing environments. The geographical distribution of edge servers poses a challenge in aggregating all data storage information for global decision-making. Existing edge data deduplication (EDD) methods rely on centralized cloud control, which faces issues of timeliness and system scalability. Additionally, these methods overlook data popularity, leading to significantly increased data retrieval latency. A promising approach to this challenge is to implement distributed EDD without cloud control, performing regional deduplication with the edge server requiring deduplication as the center. However, our investigation reveals that existing distributed EDD approaches either fail to account for the impact of collaborative caching on data availability or generate excessive information exchange between edge servers, leading to high communication overhead. To tackle this challenge, this paper presents EdgeDup, which attempts to implement effective EDD in a distributed manner. Additionally, to ensure data availability, EdgeDup aims to maintain low data retrieval latency. EdgeDup achieves its goals by: 1) identifying data redundancies across different edge servers in the system; 2) deduplicating data based on their popularity; and 3) reducing communication overheads using a novel data dependency index. Extensive experimental results show that EdgeDup significantly enhances performance, i.e., reducing data retrieval latency by an average of 47.78% compared to state-of-the-art EDD approaches while maintaining a comparable deduplication ratio. Ruikun Luo, Qiang He 0001, Feifei Chen 0001, Song Wu 0001, Hai Jin 0001, Yun Yang 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2026 | HyDuo: A Low-Latency and Cost-Efficient Storage System for Stateful Serverless Applications
Hao Fan 0006, Hanxiang Huang, Song Wu 0001 |
IEEE Trans. Serv. Comput. | 7 |
| 2026 | Latency Uncertainty-Aware User Allocation in Mobile Edge ComputingabstractMobile edge computing (MEC) is an emerging distributed paradigm where edge servers are deployed near base stations or access points to support low-latency services. The edge user allocation (EUA) problem, which aims to minimize system cost while meeting constraints like data transmission latency, has become a critical challenge for service providers. Existing studies typically focus on static MEC scenarios, assuming predictable latency between users and edge servers. However, network congestion introduces latency uncertainty, which, if not addressed, increases the risk of allocation failures or excessive service latency. This paper addresses the latency uncertainty-aware edge user allocation (uEUA) problem. We model uEUA as an integer programming problem and apply chance-constrained programming to convert uncertain latency into probabilistic constraints, which we then transform into deterministic conditions using Chebyshev's inequality. We propose two methods to solve the problem: BD-uEUA, an exact approach based on Benders decomposition, and LR-uEUA, an approximate method based on linear relaxation. Extensive experiments on a real-world dataset show that BD-uEUA reduces system costs by 19.70% compared to the state-of-the-art method, while LR-uEUA achieves a 3.05% reduction with just 0.09% of the system overhead. Ruikun Luo, Qiang He 0001, Feifei Chen 0001, Song Wu 0001, Hai Jin 0001, Yun Yang 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2025 | WAF: An Efficient WebAssembly-Based Execution Environment for User-Defined FunctionsabstractUser-Defined Functions (UDFs) have long served as the standard method for extending the capabilities of data management systems. With the advent of WebAssembly (WASM), UDFs' dependencies, such as language runtimes and libraries, can be compiled into a WASM module, which is then instantiated to execute the UDF. This approach offers several key advantages: 1) it allows developers to write UDFs in their preferred programming language, rather than being limited to those natively supported by the database engine; 2) it isolates UDFs' dependencies within the WASM module, mitigating the risk of errors caused by conflicting dependencies on the same host; and 3) it promotes cross-platform compatibility, enabling seamless execution of UDFs across different engines, operating systems, and architectures. However, our analysis reveals that executing a WASM-based UDF incurs overhead due to data transfer between the database engine and the WASM runtime. This process involves data copying and data layout adjustments, which can significantly impact performance. To address these challenges, we present WAF, a WASM-based UDF execution environment. WAF leverages shared memory to eliminate data copying and shifts data layout adjustments from the execution phase to the compilation phase. Experimental results show that WAF reduces the execution overhead of WASM-based UDFs by 3.1x and achieves an 18.1x speedup compared to the container-based approach, eliminating nearly all data transfer delays. Hao Fan 0006, Junhui Peng, Song Wu 0001, Chen Yu 0003, Hai Jin 0001, Wei Yang 0013 |
ICDE | 5 |
| 2025 | Sim-LLM: Optimizing LLM Inference at the Edge through Inter-Task KV ReuseabstractKV cache technology, by storing key-value pairs, helps reduce the computational overhead incurred by *large language models* (LLMs). It facilitates their deployment on resource-constrained edge computing nodes like edge servers. However, as the complexity and size of tasks increase, KV cache usage leads to substantial GPU memory consumption. Existing research has focused on mitigating KV cache memory usage through sequence length reduction, task-specific compression, and dynamic eviction policies. However, these methods are computationally expensive for resource-constrained edge computing nodes. To tackle this challenge, this paper presents Sim-LLM, a novel inference optimization mechanism that leverages task similarity to reduce KV cache memory consumption for LLMs. By caching KVs from processed tasks and reusing them for subsequent similar tasks during inference, Sim-LLM significantly reduces memory consumption while boosting system throughput and increasing maximum batch size, all with minimal accuracy degradation. Evaluated on both A40 and A100 GPUs, Sim-LLM achieves a system throughput improvement of up to 39.40\% and a memory reduction of up to 34.65%, compared to state-of-the-art approaches. Our source code is available at https://github.com/CGCL-codes/SimLLM. Ruikun Luo, Changwei Gu, Qiang He 0001, Feifei Chen 0001, Song Wu 0001, Hai Jin 0001, Yun Yang 0001 |
NeurIPS | 5 |
| 2025 | EDDE: Container Deployment Framework Beyond the CloudabstractContainers, renowned for their lightweight nature and flexibility, have seen growing adoption for deploying edge services such as web applications. However, existing cloud-oriented container deployment frameworks fail to address the unique challenges of edge environments, including geographical distribution, device heterogeneity, and resource constraints. This oversight leads to suboptimal performance for latency-sensitive edge services like HPC/AI-powered autonomous driving and edge gaming, which demand rapid startup and immediate responsiveness. Hao Fan 0006, Shadi Ibrahim, Lin Gu 0002, Song Wu 0001 |
SC | 5 |
| 2025 | System log isolation for containersabstractAbstract Container-based virtualization is increasingly popular in cloud computing due to its efficiency and flexibility. Isolation is a fundamental property of containers and weak isolation could cause significant performance degradation and security vulnerability. However, existing works have almost not discussed the isolation problems of system log which is critical for monitoring and maintenance of containerized applications. In this paper, we present a detailed isolation analysis of system log in current container environment. First, we find several system log isolation problems which can cause significant impacts on system usability, security, and efficiency. For example, system log accidentally exposes information of host and co-resident containers to one container, causing information leakage. Second, we reveal that the root cause of these isolation problems is that containers share the global log configuration, the same log storage, and the global log view. To address these problems, we design and implement a system named private logs (POGs). POGs provides each container with its own log configuration and stores logs individually for each container, avoiding log configuration and storage sharing, respectively. In addition, POGs enables private log view to help distinguish which container the logs belong to. The experimental results show that POGs can effectively enhance system log isolation for containers with negligible performance overhead. Kun Wang 0005, Song Wu 0001, Yanxiang Cui, Hao Fan 0006, Hai Jin 0001 |
Frontiers Comput. Sci. | 2 |
| 2025 | CBuild: Cluster-Oriented Collaborative Image Building for ContainersabstractStarting a container needs to build a container image layer-by-layer if the required image is not available. However, the image building involves downloading a large amount of data, which significantly delays the development and deployment of containerized services. To reduce data downloads and accelerate image building, current methods typically focus on improving data sharing through reconstructing images. Unfortunately, these approaches show limited performance improvement in clusters as they only improve data sharing on a single node. In this paper, we find that there are significant duplicated remote file downloads between nodes in a cluster. Accordingly, we propose cBuild, a distributed file cache to minimize costly image data downloads in cluster environments. Specifically, to enable inter-node image data sharing, cBuild designs a non-intrusive interception mechanism based on network namespace, instead of directly detecting building instructions that dirty images. Based on the distribution characteristics of duplicated files in layers, cBuild places image files among nodes in a balanced manner to prevent transfer bottlenecks caused by hotspot nodes and employs a layer-aware searching strategy to quickly locate the desired files. We implement cBuild on the basis of Docker. Experiments show that cBuild improves building speed by up to 15.3 × and reduces the data downloading by 80%. Hao Fan 0006, Song Wu 0001, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Computers | 4 |
| 2025 | Cost-Effective Edge Data Caching With Failure Tolerance and Popularity AwarenessabstractIn the mobile edge computing environment, caching data in edge storage systems can significantly reduce data retrieval latency for users while saving the costs incurred by cloud-edge data transmissions for app vendors. Existingedge data caching(EDC) methods prioritize popular data and aim to minimize users’ data retrieval latency and system storage costs jointly. However, these EDC methods often rely on the assumption that data popularity always follows certain distributions. As a result, they cannot properly adapt to the fluctuations in data popularity due to user mobility or unexpected increases in user demands. Meanwhile, unlike cloud data centers, complex and fragile edge servers are more likely to experience physical failures or network outages, presenting new challenges for EDC strategies. Specifically, when an edge server fails or experiences an outage, cached data may become temporarily unavailable, leading to increased latency as requests are redirected to alternative servers or the cloud. In this paper, to enableuncertainty-aware edge data caching(uEDC), we first model the problem as a robust optimization problem and propose an optimal algorithm named uEDC-B to find the optimal uEDC solution. To address the high computational complexity of uEDC-B, we introduce an approximate algorithm named uEDC-L based on linear decision rules. Theoretical analysis and extensive experiments on a real-world dataset demonstrate that the proposed methods outperform two state-of-the-art approaches in handling the uncertainties in data popularity and edge server failure with a significant performance improvement of 59.27% in data retrieval latency and 55.07% in data caching cost. Ruikun Luo, Zujia Zhang, Qiang He 0001, Mengxi Xu, Feifei Chen 0001, Xiaohai Dai, Song Wu 0001, Hai Jin 0001 |
IEEE Trans. Mob. Comput. | 7 |
| 2025 | Ripple: Enabling Decentralized Data Deduplication at the EdgeabstractWith its advantages in ensuring low data retrieval latency and reducing backhaul network traffic, edge computing is becoming a backbone solution for many latency-sensitive applications. An increasingly large number of data is being generated at the edge, stretching the limited capacity of edge storage systems. Improving resource utilization for edge storage systems has become a significant challenge in recent years. Existing solutions attempt to achieve this goal through data placement optimization, data partitioning, data sharing, etc. These approaches overlook the data redundancy in edge storage systems, which produces substantial storage resource wastage. This motivates the need for an approach for data deduplication at the edge. However, existing data deduplication methods rely on centralized control, which is not always feasible in practical edge computing environments. This article presents Ripple, the first approach that enables edge servers to deduplicate their data in a decentralized manner. At its core, it builds a data index for each edge server, enabling them to deduplicate data without central control. With Ripple, edge servers can 1) identify data duplicates; 2) remove redundant data without violating data retrieval latency constraints; and 3) ensure data availability after deduplication. The results of trace-driven experiments conducted in a testbed system demonstrate the usefulness of Ripple in practice. Compared with the state-of-the-art approach, Ripple improves the deduplication ratio by up to 16.79% and reduces data retrieval latency by an average of 60.42%. Ruikun Luo, Qiang He 0001, Feifei Chen 0001, Song Wu 0001, Hai Jin 0001, Yun Yang 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | Edge Data Deduplication Under Uncertainties: A Robust Optimization ApproachabstractThe emergence ofmobile edge computing(MEC) in distributed systems has sparked increased attention toward edge data management. A conflict arises from the disparity between limited edge resources and the continuously expanding data requests for data storage, making the reduction of data storage costs a critical objective. Despite the extensive studies of edge data deduplication as a data reduction technique, existing deduplication methods encounter numerous challenges in MEC environments. These challenges stem from disparities between edge servers and cloud data center edge servers, as well as uncertainties such as user mobility, leading to insufficient robustness in deduplication decision-making. Consequently, this paper presents a robust optimization-based approach for the edge data deduplication problem. By accounting for uncertainties including the number of data requirements and edge server failures, we propose two distinct solving algorithms: uEDDE-C, a two-stage algorithm based on column-and-constraint generation, and uEDDE-A, an approximation algorithm to address the high computation overhead of uEDDE-C. Our method facilitates efficient data deduplication in volatile edge network environments and maintains robustness across various uncertain scenarios. We validate the performance and robustness of uEDDE-C and uEDDE-A through theoretical analysis and experimental evaluations. The extensive experimental results demonstrate that our approach significantly reduces data storage cost and data retrieval latency while ensuring reliability in real-world MEC environments. Ruikun Luo, Qiang He 0001, Mengxi Xu, Feifei Chen 0001, Song Wu 0001, Jing Yang 0051, Yuan Gao 0031, Hai Jin 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2025 | Popularity-Aware Data Placement in Erasure Coding-Based Edge Storage SystemsabstractEdge computing allows app vendors to store popular data on edge servers, enabling users to retrieve data with low latency. However, edge servers may become unavailable at runtime due to expected exceptions. Data requests are routed to cloud servers, resulting in increased data retrieval latency. To address this issue,erasure coding(EC) has been employed to improve data availability, aiming to ensure full data access for all the users in anedge storage system(ESS). However, in real-world scenarios, data popularity differs and varies. Existing approaches for edge data placement place coded blocks across the entire system without considering data popularity. As a result, they often suffer from high data retrieval latency. In addition, they are designed to process data items individually. Data placed earlier will limit the placement options for subsequent files because edge servers with the most neighbors in the system can be easily exhausted. Some files cannot be placed properly to accommodate user demands. This increases users' data retrieval latency further. This paper tries to study the placement of multiple files in an edge storage system, considering their popularity. We first model theedge data placement(EDP) problem as a mixed-integer programming problem and prove its$\mathcal {NP}$-hardness. Then, we present an optimal algorithm named EDP-O, decoupling the EDP problem into three convex optimization subproblems for solving with an iterative algorithm. In addition, we propose an approximation algorithm named EDP-A that quickly solves the EDP problem in large-scale scenarios with a guaranteed approximation ratio of$\ln N$. The results of experiments conducted on a real-world dataset show that EDP-O and EDP-A reduce the average data retrieval latency against four representative approaches by an average of 18.4% and 15.6% in small-scale scenarios. EDP-A reduces the average data retrieval latency against four representative approaches by an average of 54.7% and reduces the data discard rate by an average of 34.9% in large-scale scenarios. Ruikun Luo, Jiadong Zhao, Qiang He 0001, Feifei Chen 0001, Song Wu 0001, Hai Jin 0001, Yun Yang 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2025 | KubeSPT: Stateful Pod Teleportation for Service Resilience With Live MigrationabstractContainer orchestration systems, such as Kubernetes, streamline containerized application deployment. As more and more applications are being deployed in Kubernetes, there is an increasing need for rescheduling - relocating a running pod to different nodes - due to system upgrades, node failures, and load-balancing optimizations. Live migration, which transfers services from source nodes to target nodes with minimal downtime, is the ideal support for rescheduling. However, implementing live migration for pods that run stateful services is challenging, because Kubernetes manages pods as stateless. First, the current pod's network namespace initialization process causes a mismatch in the network state between the migrated pod and internal containers. Second, migrating the memory state results in extended downtime. Third, Kubernetes operations on pods do not consider preserving the state of the pods. Therefore, we propose KubeSPT to achieve live migration of stateful pods in rescheduling scenarios. Firstly, we synchronize the network state of pods and internal containers by controlling packet flow and implement fast service redirection. Secondly, we introduce a Hot Data and Lazy-Restore method for memory restoration to reduce migration downtime. Finally, we decouple pod migration operations from other Kubernetes operations to ensure compatibility with live migration. Experimental results show that KubeSPT reduces downtime by 86%-93% compared to current rescheduling methods. Hansheng Zhang, Song Wu 0001, Hao Fan 0006, Weibin Xue, Chen Yu 0003, Shadi Ibrahim, Hai Jin 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2025 | ChestBox: Enabling Fast State Sharing for Stateful Serverless Computing With State FunctionsabstractThis paper presents ChestBox, a novel approach that utilizesstate functionsto facilitate low-latency state sharing for stateful serverless computing. When anapplication functionneeds to share a state, the state function creates a memory space with Linux's shared memory object to store the state. Other application functions can then read the state directly from the shared memory. ChestBox enables fast state sharing that avoids excessive memory overhead without compromising on-demand resource allocation compared to existing solutions. This effectively reduces the energy consumption of serverless computing and promotes sustainable computing. The implementation of ChestBox on Apache OpenWhisk unearths two major implementation challenges, which we address with respective optimization techniques, i.e., state function channel and state swapping. The evaluation of ChestBox with four real-world applications shows that compared with the state-of-the-art approach, it can reduce state-sharing latency by up to 99.71%, while reducing execution costs by 24.59% and storage costs by 99.76%. Song Wu 0001, Lin Gu 0002, Qiang He 0001, Hai Jin 0001 |
IEEE Trans. Sustain. Comput. | 2 |
| 2024 | Faascale: Scaling MicroVM Vertically for Serverless Computing with Memory ElasticityabstractThis paper quantitatively analyses the potential of vertical scaling MicroVMs in serverless computing. Our analysis shows that under real-world serverless workloads, vertical scaling can significantly improve execution performance and resource utilization. However, we also find that the memory scaling of MicroVMs is the bottleneck that hinders vertical scaling from reaching the performance ceiling. We propose Faascale, a novel mechanism that efficiently scales the memory of MicroVMs for serverless applications. Faascale employs a series of techniques to tackle this bottleneck: 1) it sizes up/down the memory for a MicroVM by blocks that bind with a function instance instead of general pages; and 2) it pre-populates physical memory for function instances to reduce the delays introduced by the lazy-population. Compared with existing memory scaling mechanisms, Faascale improves the memory scaling efficiency by 2 to 3 orders of magnitude. We implement Faascale on Amazon Firecracker to evaluate its gains for the serverless platform. The results of experiments conducted on eight serverless benchmark functions demonstrate that compared with horizontal scaling strategies based the state-of-the-art snapshots technique, Faascale reduces time for cold-starting MicroVMs by 89.01% and functions execution time by 23.93% on average. Qiang He 0001, Hao Fan 0006, Song Wu 0001 |
SoCC | 4 |
| 2024 | StreamBox: A Lightweight GPU SandBox for Serverless Inference Workflow
Hao Wu 0010, Junxiao Deng, Shadi Ibrahim, Song Wu 0001, Hao Fan 0006, Ziyue Cheng, Hai Jin 0001 |
USENIX ATC | 5 |
| 2024 | Precise control of page cache for containers
Kun Wang 0005, Song Wu 0001, Shengbang Li, Hao Fan 0006, Chen Yu 0003, Hai Jin 0001 |
Frontiers Comput. Sci. | 2 |
| 2024 | QoS-pro: A QoS-enhanced Transaction Processing Framework for Shared SSDsabstractSolid State Drives (SSDs) are widely used in data-intensive scenarios due to their high performance and decreasing cost. However, in shared environments, concurrent workloads can interfere with each other, leading to a violation of Quality of Service (QoS). While QoS mechanisms like fairness guarantees and latency constraints have been integrated into SSDs, existing transaction processing frameworks offer limited QoS guarantees and can significantly degrade overall performance in a shared environment. The reason is that the internal components of an SSD, originally designed to exploit parallelism, struggle to coordinate effectively when QoS mechanisms are applied to them. This article proposes a novel QoS -enhanced transaction pro cessing framework, called QoS-pro, which enhances QoS guarantees for concurrent workloads while maintaining high parallelism for SSDs. QoS-pro achieves this by redesigning transaction processing procedures to fully exploit the parallelism of shared SSDs and enhancing QoS-oriented transaction translation and scheduling with parallelism features in mind. In terms of fairness guarantees, QoS-pro outperforms state-of-the-art methods by achieving 96% fairness improvement and 64% maximum latency reduction. QoS-pro also shows almost no loss in throughput when compared with parallelism-oriented methods. Additionally, QoS-pro triggers the fewest Garbage Collection (GC) operations and minimally affects concurrently running workloads during GC operations. Hao Fan 0006, Yiliang Ye, Shadi Ibrahim, Xingru Li, Weibin Xue, Song Wu 0001, Chen Yu 0003, Xuanhua Shi, Hai Jin 0001 |
ACM Trans. Archit. Code Optim. | 7 |
| 2024 | vKernel: Enhancing Container Isolation via Private Code and DataabstractContainer technology is increasingly adopted in cloud environments. However, the lack of isolation in the shared kernel becomes a significant barrier to the wide adoption of containers. The challenges lie in how to simultaneously attain high performance and isolation. On the one hand, kernel-level isolation mechanisms, such asseccomp,capabilities, andapparmor, achieve good performance without much overhead, but lack the support for per-container customization. On the other hand, user-level and VM-based isolation offer superior security guarantees and allow for customization since a container is assigned a dedicated kernel, however, at the cost of high overhead. We presentvKernel, a kernel isolation framework. It maintains a minimal set of code and data that are either sensitive or are prone to interference in a virtual kernel instance (vKI). vKernel relies on inline hooks to intercept and redirect requests sent to the host kernel to a vKI, where container-specific security rules, functions, and data are implemented. Through case studies, we demonstrate that under vKernel user-defined data isolation and kernel customization can be supported with a reasonable engineering effort. An evaluation of vKernel with micro-benchmarks, cloud services, real-world applications show that vKernel achieves good security guarantees, but with much less overhead. Hang Huang, Jia Rao, Song Wu 0001, Hao Fan 0006, Chen Yu 0003, Hai Jin 0001, Kun Suo, Lisong Pan |
IEEE Trans. Computers | 4 |
| 2024 | Multi-Grained Trace Collection, Analysis, and Management of Diverse Container ImagesabstractContainer technology is getting popular in cloud environments due to its lightweight feature and convenient deployment. Container Registry plays a critical role in container-based clouds, as many container startups involve downloading layer-structured container images from Container Registry. However, Container Registry is struggling to efficiently manage images (i.e., transfer and store) with the emergence of diverse services and new image formats. The reason is that Container Registry manages images uniformly at layer granularity. On the one hand, such uniform layer-level management probably cannot fit the various requirements of different kinds of containerized services well. On the other hand, new image formats organizing data in blocks or files cannot benefit from such uniform layer-level image management. In this paper, we perform the first analysis of image traces at multiple granularities (i.e., image-, layer-, and file-level) for various services and provide an in-depth comparison of different image formats. The traces were collected from a production-level Container Registry, amounting to 24 million requests and involving more than 184 TB of transferred data. We provide a number of valuable insights, including request patterns of services, file-level access patterns, and bottlenecks associated with different image formats. Based on these insights, we propose two optimizations to improve image transfer. Both the traces and toolkit for trace collection will be open-sourced. Qi Zhang 0009, Hao Fan 0006, Song Wu 0001, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Computers | 4 |
| 2023 | Duo: Improving Data Sharing of Stateful Serverless Applications by Efficiently Caching Multi-Read DataabstractA growing number of applications are moving to serverless architectures for high elasticity and fine-grained billing. For stateful applications, however, the use of serverless architectures is likely to lead to significant performance degradation, as frequent data sharing between different execution stages involves time-consuming remote storage access. Current platforms leverage memory cache to speed up remote access. However, conventional caching strategies show limited performance improvement. We experimentally find that the reason is that current strategies overlook the stage-dependent access patterns of stateful serverless applications, i.e., data that are read multiple times across stages (denoted as multi-read data) are wrongly evicted by data that are read only once (denoted as read-once data), causing a high cache miss ratio.Accordingly, we propose a new caching strategy, Duo, whose design principle is to cache multi-read data as long as possible. Specifically, Duo contains a large cache list and a small cache list, which act as Leader list and Wingman list, respectively. Leader list ignores the data that is read for the first time to prevent itself from being polluted by massive read-once data at each stage. Wingman list inspects the data that are ignored or evicted by Leader list, and pre-fetches the data that will probably be read again based on the observation that multi-read data usually appear periodically in groups. Compared to the state-of-the-art works, Duo improves hit ratio by 1.1×-2.1× and reduces the data sharing overhead by 25%-62%. Hao Fan 0006, Chaoyi Cheng, Song Wu 0001, Hai Jin 0001 |
IPDPS | 4 |
| 2023 | QoS-Aware and Cost-Efficient Dynamic Resource Allocation for Serverless ML WorkflowsabstractMachine Learning (ML) workflows are increasingly deployed on serverless computing platforms to benefit from their elasticity and fine-grain pricing. Proper resource allocation is crucial to achieve fast and cost-efficient execution of serverless ML workflows (specially for hyperparameter tuning and model training). Unfortunately, existing resource allocation methods are static, treat functions equally, and rely on offline prediction, which limit their efficiency. In this paper, we introduce CE-scaling – a Cost-Efficient autoscaling framework for serverless ML work-flows. During the hyperparameter tuning, CE-scaling partitions resources across stages according to their exact usage to minimize resource waste. Moreover, it incorporates an online prediction method to dynamically adjust resources during model training. We implement and evaluate CE-scaling on AWS Lambda using various ML models. Evaluation results show that compared to state-of-the-art static resource allocation methods, CE-scaling can reduce the job completion time and the monetary cost by up to 63% and 41% for hyperparameter tuning, respectively; and by up to 58% and 38% for model training. Hao Wu 0010, Junxiao Deng, Hao Fan 0006, Shadi Ibrahim, Song Wu 0001, Hai Jin 0001 |
IPDPS | 5 |
| 2023 | PVM: Efficient Shadow Paging for Deploying Secure Containers in Cloud-native EnvironmentabstractIn cloud-native environments, containers are often deployed within lightweight virtual machines (VMs) to ensure strong security isolation and privacy protection. With the growing demand for customized cloud services, third-party vendors are turning to infrastructure-as-a-service (IaaS) cloud providers to build their own cloud-native platforms, necessitating the need to run a VM or a guest that hosts containers inside another VM instance leased from an IaaS cloud. State-of-the-art nested virtualization in the x86 architecture relies heavily on the host hypervisor to expose hardware virtualization support to the guest hypervisor, not only complicating cloud management but also raising concerns about an increased attack surface at the host hypervisor. Hang Huang, Jiangshan Lai, Jia Rao, Hui Lu 0001, Wenlong Hou, Zhengyu He, Weidong Han 0003, Tao Ma 0006, Song Wu 0001 |
SOSP | 15 |
| 2023 | Characterizing and optimizing Kernel resource isolation for containers
Kun Wang 0005, Song Wu 0001, Kun Suo, Hang Huang, Hai Jin 0001 |
Future Gener. Comput. Syst. | 2 |
| 2023 | Adapt Burstable Containers to Variable CPU ResourcesabstractIn the age of the cloud-native, container technology, referred as OS-level virtualization, is increasingly adopted to deploy cloud applications. Compared with virtual machines, containers are lightweight and flexible in resource management. An important quality-of-service (QoS) class in container management is burstable container, whose resource limits are higher than the actual requests allowing a container to expand whenever demands ramp up and additional resources become available. However, efficiently managing burstable containers is challenging, especially for CPU resources. On the one hand, burstable containers should maintain sufficient concurrency, in the form of threads, to utilize extendable CPU resources. On the other hand, the degree of concurrency necessary for utilizing peak CPU resources leads to suboptimal performance when a container's CPU allocation is constrained. In this paper, we recommend that the number of threads in burstable containers should always be set to the CPU limit to guarantee extensibility. However, modern operating systems (OSes) fall short of efficiently managing thread oversubscription. First, the OS CPU scheduler is inefficient for scheduling excessive threads and lacks container awareness. Second, the existing blocking synchronization supported by the OS kernel is inefficient in handling the sleep and wakeup of excessive threads. Finally, the non-blocking synchronization may waste CPUs performing busy waiting when more than one thread in the run queue. To this end, we present a user-level adaptive container scheduler and two OS mechanisms,virtual blockingandbusy-waiting detection, to avoid inefficiency in managing burstable containers without requiring program code changes. Experimental results show that our approaches can keep burstable containers efficient while allowing the applications in containers to take advantage of additional CPUs. The performance gain under high system load is up to 29.7×. Hang Huang, Jia Rao, Song Wu 0001, Hai Jin 0001, Duoqiang Wang, Kun Suo, Lisong Pan |
IEEE Trans. Computers | 4 |
| 2023 | Object Fingerprint Cache for Heterogeneous Memory SystemabstractHeterogeneous memory systems promise to provide both large storage capacity and high performance. Prior DRAM cache systems have metadata scalability issue and suffer from low cache hit rate, low DRAM space utilization, and significant data migration overhead. We observe that instances of an object type exhibit stable and predictable memory access patterns in its constituent cachelines and these patterns are referred to as object fingerprints. We propose the hardware-assisted cache that manages DRAM at the object type level and fetches data block at the granularity of a cacheline, by exploiting object fingerprint. To address its design challenges, we first present a software-hardware co-design to convey the software information to hardware. Second, we design multiple granularity sector caches that can be dynamically adjusted to adapt to changing behaviors and improve DRAM cache utilization. To address the challenges of large metadata storage overhead, we propose to bound possible sizes for each sector cache. Experimental results show our designs improve DRAM cache hit rate by 21.6%, boost IPC by 19.8%, and reduce data migration traffic by 51.6% on average, compared with state-of-art DRAM caches. More importantly, our online object fingerprint learning method is 2.3% inferior to the offline one in terms of IPC. Song Wu 0001, Jianhui Yue, Hai Jin 0001, Jiangqiu Shen |
IEEE Trans. Computers | 2 |
| 2023 | Enabling Balanced Data Deduplication in Mobile Edge ComputingabstractIn themobile edge computing(MEC) environment, edge servers with storage and computing resources are deployed at base stations within users’ geographic proximity to extend the capabilities of cloud computing to the network edge.Edge storage system(ESS), is comprised by connected edge servers in a specific area, which ensures low-latency services for users. However, high data storage overheads incurred by edge servers’ limited storage capacities is a key challenge in ensuring the performance of applications deployed on an ESS. Data deduplication, as a classic data reduction technology, has been widely applied in cloud storage systems. It also offers a promising solution to reducing data redundancy in ESSs. However, the unique characteristics of MEC, such as edge servers’ geographic distribution and coverage, render cloud data deduplication mechanisms obsolete. In addition, data distribution must be balanced over edge storage systems to accommodate future data demands, which cannot be undermined by data deduplication. Thus,balanced edge data deduplication(BEDD) must consider deduplication ratio, data storage benefits, and resource balance systematically under the latency constraint. In this article, we model the novel BEDD problem formally and prove its$\mathcal {NP}$-hardness. Then, we propose an optimal approach for solving the BEDD problem exactly in small-scale scenarios and a sub-optimal approach to solve large-scale BEDD problems with a theoretical performance guarantee. Extensive and comprehensive experiments conducted on a real-world dataset demonstrate the significant performance improvements of our approaches against four representative approaches. Ruikun Luo, Hai Jin 0001, Qiang He 0001, Song Wu 0001, Xiaoyu Xia 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | Cost-Effective Data Placement in Edge Storage Systems With Erasure CodeabstractEdge computing, as a new computing paradigm, brings cloud computing’s computing and storage capacities to network edge for providing low latency services for users. The networked edge servers in a specific area constituteedge storage systems(ESSs), where popular data can be stored to serve the users in the area. The novel ESSs raise many new opportunities as well as unprecedented challenges. Most existing studies of ESSs focus on the storage of data replicas in the system to ensure low data retrieval latency for users. However, replica-based edge storage strategies can easily incur high storage costs. It is not cost-effective to store massive replicas of large-size data, especially those that do not require real-time access at the edge, e.g., system upgrade files, popular app installation files, videos in online games. It may not even be possible due to the constrained storage resources on edge servers. In this article, we make the first attempt to investigate the use of erasure codes in cost-effective data storage at the edge. The focus is to find the optimal strategy for placing coded data blocks on the edge servers in an ESS, aiming to minimize the storage cost while serving all the users in the system. We first model this novelErasure Coding based Edge Data Placement(EC-EDP) problem as an integer linear programming problem and prove its$\mathcal {NP}$-hardness. Then, we propose an optimal approach named EC-EDP-O based on integer programming. Another approximation algorithm named EC-EDP-V is proposed to address the high computation complexity of large-scale EC-EDP scenarios efficiently. The extensive experimental results demonstrate that EC-EDP-O and EC-EDP-V can save an average of 68.58% (and up to 81.16% in large-scale scenarios) storage cost compared with replica-based storage approaches. Hai Jin 0001, Ruikun Luo, Qiang He 0001, Song Wu 0001, Zilai Zeng, Xiaoyu Xia 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2022 | Shared Incentive System for Clinical Pathway ExperienceabstractThe phenomenon of unbalanced regional medical resources has led to large differences in the implementation experience of clinical pathways in hospitals with different medical levels. However, due to the fear of privacy leakage and the lack of sharing motivation, there is a lack of effective collaborative communication channels for clinical pathways among medical institutions. In response to these problems, we propose a blockchain-based sharing incentive scheme for clinical pathway experience, which uses the clinical pathway implementation effect evaluation system to evaluate the quality of clinical pathway experience data. Specifically, we designe a two-phase cross-domain shared transaction model for the transaction of clinical pathway experience data. Moreover, we introduce the shared transaction alliance committee to verify and review the transaction, and solve the problems in the transaction in the arbitration phase. Finally, the functional test results show that the system meets the experience sharing incentive requirements, and the performance test results show that the TPS of sharing clinical pathway experience read and writed can reach around 600 and 1000, and the transaction latency of each phase is within 3 s. Weiqi Dai, Wenhao Zhao, Xia Xie 0001, Song Wu 0001, Hai Jin 0001 |
TrustCom | 4 |
| 2022 | Container-aware I/O stack: bridging the gap between container storage drivers and solid state devicesabstractSolid State Devices (SSDs) have been widely adopted in containerized cloud platforms as they provide parallel and high-speed data accesses for critical data-intensive applications. Unfortunately, the I/O stack of the physical host overlooks the layered and independent nature of containers, thus I/O operations require expensive file redirect (between the storage driver, Overlay2/EXT4, and the virtual file system, VFS) and are scheduled sequentially. Moreover, containers suffer from significant I/O contention as resources at the native file system are shared between them. This paper presents a Container-aware I/O stack (CAST). CAST is made up of Layer-aware VFS (LaVFS) and Container-aware Native File System (CaFS). LaVFS locates files based on layer information and enables simultaneous Copy-on-Write (CoW) operations and thus avoids the overhead of searching and modifying files. CaFS, on the other hand, provides contention-free access by designing fine-grain resource allocation at the native file system. Experimental results using a NVMe SSD with micro-benchmarks and real-world applications show that CAST achieves 216%-219% (38%-98%, respectively) improvement over the original I/O stack. Song Wu 0001, Hao Fan 0006, Shadi Ibrahim, Hai Jin 0001 |
VEE | 1 |
| 2022 | Container lifecycle-aware scheduling for serverless computingabstractAbstract Elastic scaling in response to changes on demand is a main benefit of serverless computing. When bursty workloads arrive, a serverless platform launches many new containers and initializes function environments (known as cold starts), which incurs significant startup latency. To reduce cold starts, platforms usually pause a container after it serves a request, and reuse this container for subsequent requests. However, this reuse strategy cannot efficiently reduce cold starts because the schedulers are agnostic of container lifecycle. For example, it may ignore soon available containers or evict soon needed containers. We propose a container lifecycle‐aware scheduling strategy for serverless computing, CAS. The key idea is to control distribution of requests and determine creation or eviction of containers according to different lifecycle phases of containers. We implement a prototype of CAS on OpenWhisk. Our evaluation shows that CAS reduces 81% cold starts and therefore brings a 63% reduction at 95th percentile latency compared with native scheduling strategy in OpenWhisk when there is worker contention between workloads, and does not add significant performance overhead. Song Wu 0001, Zhiheng Tao, Hao Fan 0006, Hai Jin 0001, Chen Yu 0003, Chun Cao |
Softw. Pract. Exp. | 1 |
| 2022 | Cost-Effective Edge Server Network Design in Mobile Edge Computing EnvironmentabstractMobile edge computing(MEC) deploys edge servers at the base station in the proximity of users to provide cloud computing-like computing and storage functionalities, which can achieve applications’ low latency requirement at the network edge. Theedge server network(ESN), constituted by edge servers in an area and the links between them, can host app vendors’ services for serving nearby users. Many existing studies have demonstrated that a high ESN density allows for high service performance because edge servers can communicate and share resources with each other effectively over the ESN. However, in the real-world MEC environment, constructing a high-density ESN may incur high construction costs. The trade-off between construction cost and network density plays a vital role in the design of an ESN. Unfortunately, existing studies of MEC have commonly and simply assumed the densities of the ESNs in their experiments. In this paper, we make the first attempt to study the design of cost-effective ESNs with the aim to trade off between the network construction cost and the network density. We model this novelEdge Server Network Design(ESND) problem as a constrained optimization problem and prove its$\mathcal {NP}$-hardness. ESND-O as an optimal approach is proposed based on integer programming to solve small-scale ESND problems. Another approximation approach named ESND-A is designed to solve large-scale ESND problems efficiently. We conduct extensive experiments to test the performance of ESND-O and ESND-A on a real-world dataset, and the experimental results demonstrate their effectiveness and efficiency against four representative approaches. Ruikun Luo, Hai Jin 0001, Qiang He 0001, Song Wu 0001, Xiaoyu Xia 0001 |
IEEE Trans. Sustain. Comput. | 4 |
| 2021 | Towards Exploiting CPU Elasticity via Efficient Thread OversubscriptionabstractElasticity is an essential feature of cloud computing, which allows users to dynamically add or remove resources in response to workload changes. However, building applications that truly exploit elasticity is non-trivial. Traditional applications need to be modified to efficiently utilize variable resources. This paper explores thread oversubscription, i.e., provisioning more threads than the available cores, to exploit CPU elasticity in the cloud. While maintaining sufficient concurrency allows applications to utilize additional CPUs when more are made available, it is widely believed that thread oversubscription introduces prohibitive overheads due to excessive context switches, loss of locality, and contention on shared resources. Hang Huang, Jia Rao, Song Wu 0001, Hai Jin 0001, Hong Jiang 0001, Hao Che, Xiaofeng Wu 0002 |
HPDC | 3 |
| 2021 | Gear: Enable Efficient Container Storage and Deployment with a New Image FormatabstractContainers have been widely used in various cloud platforms as they enable agile and elastic application deployment through their process-based virtualization and layered image system. However, different layers of a container image may contain substantial duplicate and unnecessary data, which slows down its deployment due to long image downloading time and increased burden on the image registry. To accelerate the deployment and reduce the size of the registry, we propose a new image format, named Gear image, that consists of two parts: a Gear index describing the structure of the image's file system and a set of files that are required when running an application. The Gear index is represented as a single-layer image compatible with the existing deployment framework. Containers can be launched by pulling a Gear index and on demand retrieving files pointed to by the index. Furthermore, the Gear image enables a file-level sharing mechanism, which helps remove duplicate data in the registry and avoid repeated downloading of identical files by a client. We implement a prototype of the container framework, named Gear, supporting the new image format. Evaluation shows that Gear saves 54 % storage capacity in the registry, speeds up container startup by up to${5\times}$, and reduces 84 % bandwidth demands. Hao Fan 0006, Shengwei Bian, Song Wu 0001, Song Jiang 0001, Shadi Ibrahim, Hai Jin 0001 |
ICDCS | 3 |
| 2021 | Graph-Based Data Deduplication in Mobile Edge Computing Environment
Ruikun Luo, Hai Jin 0001, Qiang He 0001, Song Wu 0001, Zilai Zeng, Xiaoyu Xia 0001 |
ICSOC | 4 |
| 2021 | Hardware-supported remote persistence for distributed persistent memoryabstractThe advent of Persistent Memory (PM) necessitates an evolution of Remote Direct Memory Access (RDMA) technologies for supporting remote data persistence. Previous software-based solutions require remote CPU intervention and postpone the visibility of remote persistence. In this paper, we design several hardware-supported RDMA primitives to flush data from the volatile cache of RDMA Network Interface Cards (RNICs) to the PM. We also propose durable RPCs based on the proposed RDMA Flush primitives to support remote data persistence and fast failure recovery. We emulate the performance of RDMA Flush primitives through other RDMA primitives, and compare our proposals with several state-of-the-art RPCs in a real testbed equipped with PM and InfiniBand networks. Experimental results show that our proposals can improve the throughput of RPCs by up to 90%, and reduce the 99th percentile latency by up to 49%. The experimental studies also provide instructive guidelines for designing RDMA-based distributed PM systems. Zhuohui Duan, Haodi Lu, Haikun Liu, Xiaofei Liao, Hai Jin 0001, Yu Zhang 0027, Song Wu 0001 |
SC | 7 |
| 2021 | Accelerating Parallel Applications in Cloud Platforms via Adaptive Time-Slice ControlabstractCloud platforms can provide flexible and cost-effective environments for parallel applications. However, the resource over-commitment issues, i.e., cloud providers often provide much more executable virtual CPUs than available physical CPUs, still impede the synchronization operations of parallel applications, causing severe performance degradation. Existing methods optimize parallel applications by promoting the priorities of involved VMs. They cannot fully explore the performance of parallel applications, because they ignore the time-slice requirements of different phases of parallel applications. Furthermore, non-parallel applications experience unsatisfied performance because of low scheduling priorities. Given empirical analysis on time-slices of virtual machines (VMs), we find that shortening time-slices can mitigate synchronization overhead which incurs during communication phases, while over-short time-slices cause frequent cache misses in computation phases. Accordingly, we propose an Adaptive Time-slice Control (ATC) mechanism. ATC first detects the phases of parallel applications based on lock latency or cache misses. Then, ATC shortens time-slices during communication phases and prolongs time-slices during computation phases for parallel applications, and sets a uniform time-slice for non-parallel applications. We evaluate ATC using seven well-known benchmarks with 25+ applications. Experiments show that ATC obtains 1.5-75× performance gain for running parallel applications than state-of-the-art solutions, with nearly unaffected impact on non-parallel applications. Hao Fan 0006, Song Wu 0001, Zhenjiang Xie, Sheng Di, Jiang Xiao 0001, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Computers | 2 |
| 2021 | Precise Power Capping for Latency-Sensitive Applications in DatacenterabstractPower capping is widely used in cloud datacenters to mitigate power over-provisioning problem, thus improve datacenter capacity and cut off their operation cost. However, inappropriate or aggressive power capping may lead to performance degradation of applications (especially latency-sensitive ones), and there are few effective methods that can accurately evaluate and control such negative impact caused by aggressive power capping. In this paper, we proposeFine-Grained Differential Method(FGD) to quantitatively analyze how inappropriate power capping degrades the performance of latency-sensitive applications. By using FGD, we can minimize the provisioned power for each server by setting a precise power budget according to application’sService Level Agreement(SLA). And we further proposePrecise Power Capping(PPCapping) which is designed to increase the datacenter capacity with a fixed power supply by means of FGD. Our research also provides an insight of precise tradeoff between applications’ SLAs and datacenter capacity. We verify FGD and PPCapping by using real world traces from Tencent’s datecenter with 25,328 servers. The experimental results show that FGD can accurately analyze the impact of power capping on the performance of latency-sensitive applications, and PPCapping can effectively increase datacenter capacity compared with the typical power provisioning strategy. Song Wu 0001, Xinhou Wang, Hai Jin 0001, Fangming Liu, Haibao Chen, Chuxiong Yan |
IEEE Trans. Sustain. Comput. | 1 |
| 2020 | Sledge: Towards Efficient Live Migration of Docker ContainersabstractModern large-scale cloud platforms require live migration technique on Docker containers with stateful workload to support load balancing, host maintenance, and Quality of Service (QoS) improvement. Efficient and scalable Docker live migration is expected to guarantee the component-integrity (image, runtime, and management context) with negligible downtime. In this paper, we present a highly efficient live migration system called Sledge, which ensures the component-integrity by integrating both images and management context during runtime migration. The key insight is that the layered image can be leveraged to reduce the migration overhead, and appropriately selective migration of management context will effectively improve QoS with negligible downtime. To achieve good scalability, a lightweight container registry mechanism for end-to-end image migration is designed to avoid the redundant layers transmission. In addition, a dynamic context loading scheme is proposed to precisely load the management context into the running daemon, which can significantly reduce downtime. Experiments show that, compared with the state-of-the-art, Sledge reduces 57% of total migration time, 55% of image migration time, and 70% downtime. Song Wu 0001, Jiang Xiao 0001, Hai Jin 0001, Yingxi Zhang, Guoqiang Shi, Tingyu Lin 0001, Jia Rao, Jizhong Jiang |
CLOUD | 2 |
| 2020 | BED: A Block-Level Deduplication-Based Container Deployment Framework
Shiqiang Zhang, Song Wu 0001, Hao Fan 0006, Deqing Zou, Hai Jin 0001 |
GPC | 2 |
| 2020 | Layup: Layer-adaptive and Multi-type Intermediate-oriented Memory Optimization for GPU-based CNNsabstractAlthough GPUs have emerged as the mainstream for the acceleration of convolutional neural network (CNN) training processes, they usually have limited physical memory, meaning that it is hard to train large-scale CNN models. Many methods for memory optimization have been proposed to decrease the memory consumption of CNNs and to mitigate the increasing scale of these networks; however, this optimization comes at the cost of an obvious drop in time performance. We propose a new memory optimization strategy named Layup that realizes both better memory efficiency and better time performance. First, a fast layer-type-specific method for memory optimization is presented, based on the new finding that a single memory optimization often shows dramatic differences in time performance for different types of layers. Second, a new memory reuse method is presented in which greater attention is paid to multi-type intermediate data such as convolutional workspaces and cuDNN handle data. Experiments show that Layup can significantly increase the scale of extra-deep network models on a single GPU with lower performance loss. It even can train ResNet with 2,504 layers using 12GB memory, outperforming the state-of-the-art work of SuperNeurons with 1,920 layers (batch size = 16). Wenbin Jiang 0001, Bo Liu 0057, Haikun Liu, Bing Bing Zhou, Song Wu 0001, Hai Jin 0001 |
ACM Trans. Archit. Code Optim. | 7 |
| 2020 | Doris: An Adaptive Soft Real-Time Scheduler in Virtualized EnvironmentsabstractWith the development of cloud computing and virtualization technologies, more and more soft real-time applications, such as Voice over Internet Protocol (VoIP) server and cloud gaming, are running in virtualized data centers. Though previous studies optimize CPU schedulers of hypervisors to support these applications in virtualized environments, there are some important challenges in designing an efficient CPU scheduler which is suitable for real-world clouds. On one hand, hypervisors do not know whether an application in a virtual machine (VM) has real-time requirements, so manually setting the scheduling parameters is a common case for CPU schedulers, which probably increases users' burden, lacks flexibility, and causes misconfigurations. On the other hand, it has been reported that most of existing CPU schedulers designed for soft real-time applications have an obvious propensity to such applications which prevents them from being applied in practical multi-tenant cloud environments. In this paper, we design and implement an adaptive soft real-time scheduler based on Xen, named Doris, to address these challenges. It identifies the VMs running soft real-time applications (RT-VMs) and infers their scheduling parameters according to the communication behaviors of VMs adaptively. Then, it promotes the priorities of VCPUs of the RT-VMs temporarily according to I/O events and the inferred scheduling parameters of RT-VMs to support soft real-time applications adaptively while minimizing the impacts on non-real-time applications. Finally, considering the importance of privileged entities (such as Domain0 in Xen) in I/O processing, Doris sets their types and scheduling parameters dynamically, which enables the adaptive scheduling of them to guarantee the performance of soft real-time applications. Our evaluation shows Doris can support soft real-time applications adaptively and efficiently, and only introduces very slight overhead. Song Wu 0001, Like Zhou, Hai Jin 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2019 | CNTC: A Container Aware Network Traffic Control Framework
Lin Gu 0002, Junjian Guan, Song Wu 0001, Hai Jin 0001, Jia Rao, Kun Suo, Deze Zeng |
GPC | 3 |
| 2019 | Adaptive Resource Views for ContainersabstractAs OS-level virtualization advances, containers have become a viable alternative to virtual machines in deploying applications in the cloud. Unlike virtual machines, which allow guest OSes to run atop virtual hardware, containers have direct access to physical hardware and share one OS kernel. While the absence of virtual hardware abstractions eliminates most virtualization overhead, it presents unique challenges for containerized applications to efficiently utilize the underlying hardware. The lack of hardware abstraction exposes the total amount of resources that are shared among all containers to each individual container. Parallel runtimes (e.g., OpenMP) and managed programming languages (e.g., Java) that rely on OS-exported information for resource management could suffer from suboptimal performance. In this paper, we develop a per-container view of resources to export information on the actual resource allocation to containerized applications. The central design of the resource view is a per-container sys\_namespace that calculates the effective capacity of CPU and memory in the presence of resource sharing among containers. We further create a virtual sysfs to seamlessly interface user space applications with sys\_namespace. We use two case studies to demonstrate how to leverage the continuously updated resource view to enable elasticity in the HotSpot JVM and OpenMP. Experimental results show that an accurate view of resource allocation leads to more appropriate configurations and improved performance in a variety of containerized applications. Hang Huang, Jia Rao, Song Wu 0001, Hai Jin 0001, Kun Suo, Xiaofeng Wu 0002 |
HPDC | 3 |
| 2019 | Preemptive Multi-Queue Fair QueuingabstractFair queuing (FQ) algorithms have been widely adopted in computer systems to share resources among multiple users. Modern operating systems and hypervisors use variants of FQ algorithms to implement the critical OS resource management -- the thread scheduler. While the existing FQ algorithms enforce fair CPU allocation on a per-core basis, there lacks an algorithm to fairly allocate CPU on multiple cores. This common deficiency in state-of-the-art multicore schedulers causes unfair CPU allocations to parallel programs using blocking synchronization, leading to severe performance degradation. Parallel threads that frequently block due to synchronization exhibit deceptive idleness and are penalized by the thread scheduler. To this end, we propose a preemptive multi-queue fair queuing (P-MQFQ) algorithm that uses a centralized queue to fairly dispatch threads from different programs based on their received CPU bandwidth from multiple cores. We demonstrate that P-MQFQ can be approximated by augmenting the existing load balancing in the OS without requiring to implement the centralized queue or undermining scalability. We implement P-MQFQ in Linux and Xen, respectively, and show significantly improved utilization and performance for parallel programs. Kun Suo, Xiaofeng Wu 0002, Jia Rao, Song Wu 0001, Hai Jin 0001 |
HPDC | 5 |
| 2019 | When FPGA-Accelerator Meets Stream Data Processing in the EdgeabstractToday, stream data applications represent the killer applications for Edge computing: placing computation close to the data source facilitates real-time analysis. Previous efforts have focused on introducing light-weight distributed stream processing (DSP) systems and dividing the computation between Edge servers and the clouds. Unfortunately, given the limited computation power of Edge servers, current efforts may fail in practice to achieve the desired latency of stream data applications. In this vision paper, we argue that by introducing FPGAs in Edge servers and integrating them into DSP systems, we might be able to realize stream data processing in Edge infrastructures. We demonstrate that through the design, implementation, and evaluation of F-Storm, an FPGA-accelerated and general-purpose distributed stream processing system on Edge servers. F-Storm integrates PCIe-based FPGAs into Edge-based stream processing systems and provides accelerators as a service for stream data applications. We evaluate F-Storm using different representative stream data applications. Our experiments show that compared to Storm, F-Storm reduces the latency by 36% and 75% for matrix multiplication and grep application. It also obtains 1.4x and 2.1x improvement for these two applications, respectively. We expect this work to accelerate progress in this domain. Song Wu 0001, Shadi Ibrahim, Hai Jin 0001, Jiang Xiao 0001, Haikun Liu |
ICDCS | 1 |
| 2019 | NCQ-Aware I/O Scheduling for Conventional Solid State DrivesabstractWhile current fairness-driven I/O schedulers are successful in allocating equal time/resource share to concurrent workloads, they ignore the I/O request queueing or reordering in storage device layer, such as Native Command Queueing (NCQ). As a result, requests of different workloads cannot have an equal chance to enter NCQ (NCQ conflict) and fairness is violated. We address this issue by providing the first systematic empirical analysis on how NCQ affects I/O fairness and SSD utilization and accordingly proposing a NCQ-aware I/O scheduling scheme, NASS. The basic idea of NASS is to elaborately control the request dispatch of workloads to relieve NCQ conflict and improve NCQ utilization. NASS builds on two core components: an evaluation model to quantify important features of the workload, and a dispatch control algorithm to set the appropriate request dispatch of running workloads. We integrate NASS into four state-of-the-art I/O schedulers and evaluate its effectiveness using widely used benchmarks and real world applications. Results show that with NASS, I/O schedulers can achieve 11-23% better fairness and at the same time improve device utilization by 9-29%. Hao Fan 0006, Song Wu 0001, Shadi Ibrahim, Hai Jin 0001, Jiang Xiao 0001, Haibing Guan |
IPDPS | 2 |
| 2019 | FastBuild: Accelerating Docker Image Building for Efficient Development and Deployment of ContainerabstractDocker containers have been increasingly adopted on various computing platforms to provide a lightweight virtualized execution environment. Compared to virtual machines, this technology can often reduce the launch time from a few minutes to less than 10 seconds, assuming the Docker image has been locally available. However, Docker images are highly customizable, and are mostly built at runtime from a remote base image by running instructions in a script (the Dockerfile). During the instruction execution, a large number of input files may have to be retrieved via the Internet. The image building may be an iterative process as one may need to repeatedly modify the Dockerfile until a desired image composition is received. In the process, every input file required by an instruction has to be remotely retrieved, even if it has been recently downloaded. This can make the process of building of an image and launching of a container unexpectedly slow. To address the issue, we propose a technique, named FastBuild, that maintains a local file cache to minimize the expensive file downloading. By non-intrusively intercepting remote file requests, and supplying files locally, FastBuild enables file caching in a manner transparent to image building. To further accelerate the image building, FastBuild overlaps operations of instructions' execution and writing intermediate image layers to the disk. We have implemented FastBuild. And experiments with images and Dockerfiles obtained from Docker Hub show that the system can improve building speed by up to 10 times, and reduce downloaded data by 72%. Song Wu 0001, Song Jiang 0001, Hai Jin 0001 |
MSST | 2 |
| 2019 | N-Docker: A NVM-HDD Hybrid Docker Storage Framework to Improve Docker Performance
Lin Gu 0002, Qizhi Tang, Song Wu 0001, Hai Jin 0001, Yingxi Zhang, Guoqiang Shi, Tingyu Lin 0001, Jia Rao |
NPC | 3 |
| 2019 | VAIL: A Victim-Aware Cache Policy to improve NVM Lifetime for hybrid memory system
Song Wu 0001, Youchuang Jia, Hai Jin 0001, Xiaofei Liao, Pingpeng Yuan |
Parallel Comput. | 2 |
| 2019 | Dual-Page Checkpointing: An Architectural Approach to Efficient Data Persistence for In-Memory ApplicationsabstractData persistence is necessary for many in-memory applications. However, the disk-based data persistence largely slows down in-memory applications. Emerging non-volatile memory (NVM) offers an opportunity to achieve in-memory data persistence at the DRAM-level performance. Nevertheless, NVM typically requires a software library to operate NVM data, which brings significant overhead. This article demonstrates that a hardware-based high-frequency checkpointing mechanism can be used to achieve efficient in-memory data persistence on NVM. To maintain checkpoint consistency, traditional logging and copy-on-write techniques incur excessive NVM writes that impair both performance and endurance of NVM; recent work attempts to solve the issue but requires a large amount of metadata in the memory controller. Hence, we design a new dual-page checkpointing system, which achieves low metadata cost and eliminates most excessive NVM writes at the same time. It breaks the traditional trade-off between metadata space cost and extra data writes. Our solution outperforms the state-of-the-art NVM software libraries by 13.6× in throughput, and leads to 34% less NVM wear-out and 1.28× higher throughput than state-of-the-art hardware checkpointing solutions, according to our evaluation with OLTP, graph computing, and machine-learning workloads. Song Wu 0001, Hai Jin 0001, Jinglei Ren |
ACM Trans. Archit. Code Optim. | 1 |
| 2019 | Towards Low-Latency Batched Stream Processing by Pre-SchedulingabstractMany stream processing frameworks have been developed to meet the requirements of real-time processing. Among them, batched stream processing frameworks are widely advocated with the consideration of their fault-tolerance, high throughput and unified runtime with batch processing. In batched stream processing frameworks, straggler, happened due to the uneven task execution time, has been regarded as a major hurdle of latency-sensitive applications. Existing straggler mitigation techniques, operating in either reactive or proactive manner, are all post-scheduling methods, and therefore inevitably result in high resource overhead or long job completion time. We notice that batched stream processing jobs are usually recurring with predictable characteristics. By exploring such a heuristic, we present a pre-scheduling straggler mitigation framework called Lever. Lever first identifies potential stragglers and evaluates nodes’ capacity by analyzing execution information of historical jobs. Then, Lever carefully pre-schedules job input data to each node before task scheduling so as to mitigate potential stragglers. We implement Lever and contribute it as an extension of Apache Spark Streaming. Our experimental results show that Lever can reduce job completion time by 30.72 to 42.19 percent over Spark Streaming, a widely adopted batched stream processing system and outperforms traditional techniques significantly. Hai Jin 0001, Song Wu 0001, Yin Yao, Zhiyi Liu, Lin Gu 0002, Yongluan Zhou |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | Container-Based Customization Approach for Mobile Environments on Clouds
Jiahuan Hu, Song Wu 0001, Hai Jin 0001, Hanhua Chen |
GPC | 2 |
| 2018 | Network-Aware Grouping in Distributed Stream Processing Systems
Song Wu 0001, Hai Jin 0001 |
ICA3PP (1) | 2 |
| 2018 | TurboStream: Towards Low-Latency Data Stream ProcessingabstractData Stream Processing (DSP) applications are often modelled as a directed acyclic graph: operators with data streams among them. Inter-operator communications can have a significant impact on the latency of DSP applications, accounting for 86% of the total latency. Despite their impact, there has been relatively little work on optimizing inter-operator communications, focusing on reducing inter-node traffic but not considering inter-process communication (IPC) inside a node, which often generates high latency due to the multiple memory-copy operations. This paper describes the design and implementation of TurboStream, a new DSP system designed specifically to address the high latency caused by inter-operator communications. To achieve this goal, we introduce (1) an improved IPC framework with OSRBuffer, a DSP-oriented buffer, to reduce memory-copy operations and waiting time of each single message when transmitting messages between the operators inside one node, and (2) a coarse-grained scheduler that consolidates operator instances and assigns them to nodes to diminish the inter-node IPC traffic. Using a prototype implementation, we show that our improved IPC framework reduces the end-to-end latency of intra-node IPC by 45.64% to 99.30%. Moreover, TurboStream reduces the latency of DSP by 83.23% compared to JStorm. Song Wu 0001, Shadi Ibrahim, Hai Jin 0001, Lin Gu 0002, Zhiyi Liu |
ICDCS | 1 |
| 2018 | Dual-Paradigm Stream ProcessingabstractExisting stream processing frameworks operate either under data stream paradigm processing data record by record to favor low latency, or under operation stream paradigm processing data in micro-batches to desire high throughput. For complex and mutable data processing requirements, this dilemma brings the selection and deployment of stream processing frameworks into an embarrassing situation. Moreover, current data stream or operation stream paradigms cannot handle data burst efficiently, which probably results in noticeable performance degradation. This paper introduces a dual-paradigm stream processing, called DO (Data and Operation) that can adapt to stream data volatility. It enables data to be processed in micro-batches (i.e., operation stream) when data burst occurs to achieve high throughput, while data is processed record by record (i.e., data stream) in the remaining time to sustain low latency. DO embraces a method to detect data bursts, identify the main operations affected by the data burst and switch paradigms accordingly. Our insight behind DO's design is that the trade-off between latency and throughput of stream processing frameworks can be dynamically achieved according to data communication among operations in a fine-grained manner (i.e., operation level) instead of framework level. We implement a prototype stream processing framework that adopts DO. Our experimental results show that our framework with DO can achieve 5x speedup over operation stream under low data stream sizes, and outperforms data stream on throughput by 2.1x to 3.2x under data burst. Song Wu 0001, Zhiyi Liu, Shadi Ibrahim, Lin Gu 0002, Hai Jin 0001 |
ICPP | 1 |
| 2018 | Disk Failure Prediction in Data Centers via Online LearningabstractDisk failure has become a major concern with the rapid expansion of storage systems in data centers. Based on SMART (Self-Monitoring, Analysis and Reporting Technology) attributes, many researchers derive disk failure prediction models using machine learning techniques. Despite the significant developments, the majority of works rely on offline training and thereby hinder their adaption to the continuous update of forthcoming data, suffering from the 'model aging' problem. We are therefore motivated to uncover the root cause -- the dynamic SMART distribution for 'model aging', aiming to resolve the performance degradation as to pave a comprehensive study in practice. Jiang Xiao 0001, Song Wu 0001, Yusheng Yi, Hai Jin 0001, Kan Hu |
ICPP | 3 |
| 2018 | Dynamic vertical memory scalability for OpenJDK cloud applicationsabstractThe cloud is an increasingly popular platform to deploy applications as it lets cloud users to provide resources to their applications as needed. Furthermore, cloud providers are now starting to offer a "pay-as-you-use" model in which users are only charged for the resources that are really used instead of paying for a statically sized instance. This new model allows cloud users to save money, and cloud providers to better utilize their hardware. Rodrigo Bruno, Paulo Ferreira 0001, Ruslan Synytsky, Tetiana Fydorenchyk, Jia Rao, Hang Huang, Song Wu 0001 |
ISMM | 7 |
| 2018 | Android Unikernel: Gearing mobile code offloading towards edge computing
Song Wu 0001, Hai Jin 0001, Duoqiang Wang |
Future Gener. Comput. Syst. | 1 |
| 2018 | Dynamic Resource Scheduling in Mobile Edge Cloud with Cloud Radio Access NetworkabstractNowadays, by integrating the cloud radio access network (C-RAN) with the mobile edge cloud computing (MEC) technology, mobile service provider (MSP) can efficiently handle the increasing mobile traffic and enhance the capabilities of mobile devices. But the power consumption has become skyrocketing for MSP and it gravely affects the profit of MSP. Previous work often studied the power consumption in C-RAN and MEC separately while less work had considered the integration of C-RAN with MEC. In this paper, we present an unifying framework for the power-performance tradeoff of MSP by jointly scheduling network resources in C-RAN and computation resources in MEC to maximize the profit of MSP. To achieve this objective, we formulate the resource scheduling issue as a stochastic problem and design a new optimization framework by using an extended Lyapunov technique. Specially, because the standard Lyapunov technique critically assumes that job requests have fixed lengths and can be finished within each decision making interval, it is not suitable for the dynamic situation where the mobile job requests have variable lengths. To solve this problem, we extend the standard Lyapunov technique and design the VariedLen algorithm to make online decisions in consecutive time for job requests with variable lengths. Our proposed algorithm can reach time average profit that is close to the optimum with a diminishing gap (1/V) for the MSP while still maintaining strong system stability and low congestion. With extensive simulations based on a real world trace, we demonstrate the efficacy and optimality of our proposed algorithm. Xinhou Wang, Kezhi Wang, Song Wu 0001, Sheng Di, Hai Jin 0001, Kun Yang 0001, Shumao Ou |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2017 | Lever: towards low-latency batched stream processing by pre-schedulingabstractWith the vast involvement of streaming big data in many applications (e.g., stock market data, sensor data, social network data, etc.), quickly mining and analyzing such data is becoming more and more important. To provide fault tolerance and efficient stream processing at scale, recent stream processing frameworks have proposed to adapt batch processing systems, such as MapReduce and Spark, to handle streaming data by putting the streams into micro-batches and treating the workloads as a continuous series of small jobs [1]. Song Wu 0001, Hai Jin 0001, Yin Yao, Zhiyi Liu, Lin Gu 0002, Yongluan Zhou |
SoCC | 2 |
| 2017 | SCVD: A New Semantics-Based Approach for Cloned Vulnerable Code Detection
Deqing Zou, Hanchao Qi, Zhen Li 0027, Song Wu 0001, Hai Jin 0001, Guozhong Sun, Sujuan Wang, Yuyi Zhong |
DIMVA | 4 |
| 2017 | Maximizing the Profit of Cloud Broker with Priority Aware PricingabstractA practical problem facing Infrastructure-as-a-Service (IaaS) cloud users is how to minimize their costs by choosing different pricing options based on their own demands. Recently, cloud brokerage service is introduced to tackle this problem. But due to the perishability of cloud resources, there still exists a large amount of idle resource waste during the reservation period of reserved instances. This idle resource waste problem is challenging cloud broker when buying reserved instances to accommodate users' job requests. To solve this challenge, we find that cloud users always have low priority jobs (e.g., non latency-sensitive jobs) which can be delayed to utilize these idle resources. With considering the priority of jobs, two problems need to be solved. First, how can cloud broker leverage jobs' priorities to reserve resources for profit maximization? Second, how to fairly price users' job requests with different priorities when previous studies either adopt pricing schemes from IaaS clouds or just ignore the pricing issue. To solve these problems, we first design a fair and priority aware pricing scheme, PriorityPricing, for the broker which charges users with different prices based on priorities. Then we propose three dynamic algorithms for the broker to make resource reservations with the objective of maximizing its profit. Experiments show that the broker's profit can be increased up to 2.5× than that without considering priority for offline algorithm, and 3.7× for online algorithm. Xinhou Wang, Song Wu 0001, Kezhi Wang, Sheng Di, Hai Jin 0001, Kun Yang 0001, Shumao Ou |
ICPADS | 2 |
| 2017 | MURS: Mitigating Memory Pressure in Service-Oriented Data Processing SystemabstractAlthough a data processing system often works as a batch processing system, many enterprises deploy such a system as a service, which we call the service-oriented data processing system. It has been shown that in-memory data processing systems suffer from serious memory pressure. The situation becomes even worse for the service-oriented data processing systems due to various reasons. For example, in a service-oriented system, multiple submitted tasks are launched at the same time and executed in the same context in the resources, compared with the batch processing mode where the tasks are processed one by one. Therefore, the memory pressure will affect all submitted tasks, including the tasks that only incur the light memory pressure when they are run alone. In this paper, we find that the reason why memory pressure arises is because the running tasks produce massive long-living data objects in the limited memory space. Our studies further reveal that the long-living data objects are generated by the API functions that are invoked by the in-memory processing frameworks. Based on these findings, we propose a method to classify the API functions based on the memory usage rate. Further, we design a scheduler called MURS to mitigate the memory pressure. We implement MURS in Spark and conduct the experiments to evaluate the performance of MURS. The results show that when comparing to Spark, MURS can 1) decrease the execution time of the submitted jobs by up to 65.8%, 2) mitigate the memory pressure in the server by decreasing the garbage collection time by up to 81%, and 3) reduce the data spilling, and hence disk I/O, by approximately 90%. Xuanhua Shi, Ligang He, Hai Jin 0001, Zhixiang Ke, Song Wu 0001 |
ICWS | 6 |
| 2017 | Container-Based Cloud Platform for Mobile Computation OffloadingabstractWith the explosive growth of smartphones and cloud computing, mobile cloud, which leverages cloud resource to boost the performance of mobile applications, becomes attrac- tive. Many efforts have been made to improve the performance and reduce energy consumption of mobile devices by offloading computational codes to the cloud. However, the offloading cost caused by the cloud platform has been ignored for many years. In this paper, we propose Rattrap, a lightweight cloud platform which improves the offloading performance from cloud side. To achieve such goals, we analyze the characteristics of typical of- floading workloads and design our platform solution accordingly. Rattrap develops a new runtime environment, Cloud Android Container, for mobile computation offloading, replacing heavy- weight virtual machines (VMs). Our design exploits the idea of running operating systems with differential kernel features inside containers with driver extensions, which partially breaks the limitation of OS-level virtualization. With proposed resource sharing and code cache mechanism, Rattrap fundamentally improves the offloading performance. Our evaluation shows that Rattrap not only reduces the startup time of runtime environments and shows an average speedup of 16x, but also saves a large amount of system resources such as 75% memory footprint and at least 79% disk capacity. Moreover, Rattrap improves offloading response by as high as 63% over the cloud platform based on VM, and thus saving the battery life. Song Wu 0001, Jia Rao, Hai Jin 0001, Xiaohai Dai |
IPDPS | 1 |
| 2017 | Evolution of Cloud Operating System: From Technology to Ecosystem
Zuoning Chen, Kang Chen 0001, Jinlei Jiang, Lufei Zhang, Song Wu 0001, Zhengwei Qi, Chunming Hu, Yongwei Wu 0001, Yuzhong Sun, Aobing Sun, Zilu Kang |
J. Comput. Sci. Technol. | 5 |
| 2017 | ACStor: Optimizing Access Performance of Virtual Disk Images in CloudsabstractIn virtualized data centers, virtual disk images (VDIs) serve as the containers in virtual environment, so their access performance is critical for the overall system performance. Some distributed VDI chunk storage systems have been proposed in order to alleviate the I/O bottleneck for VM management. As the system scales up to a large number of running VMs, however, the overall network traffic would become unbalanced with hot spots on some VMs inevitably, leading to I/O performance degradation when accessing the VMs. In this paper, we propose an adaptive and collaborative VDI storage system (ACStor) to resolve the above performance issue. In comparison with the existing research, our solution is able to dynamically balance the traffic workloads in accessing VDI chunks, based on the run-time network state. Specifically, compute nodes with lightly loaded traffic will be adaptively assigned more chunk access requests from remote VMs and vice versa, which can effectively eliminate the above problem and thus improves the I/O performance of VMs. We implement a prototype based on our ACStor design, and evaluate it by various benchmarks on a real cluster with 32 nodes and a simulated platform with 256 nodes. Experiments show that under different network traffic patterns of data centers, our solution achieves up to 2-8× performance gain on VM booting time and VM's I/O throughput, in comparison with the other state-of-the-art approaches. Song Wu 0001, Sheng Di, Haibao Chen, Hai Jin 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | A Performance Study of Containers in Cloud Environment
Bowen Ruan, Hang Huang, Song Wu 0001, Hai Jin 0001 |
APSCC | 3 |
| 2016 | vProbe: Scheduling Virtual Machines on NUMA SystemsabstractWith the development of multi-core platforms andcloud computing, Non-Uniform Memory Access (NUMA) architecturehas been dominant in cloud data centers in recentyears. However, NUMA architecture is not well supported invirtualized environments. Because of the semantic gap introducedby the virtualization layer, hypervisors know little aboutthe characteristics of applications running in virtual machines (VMs). More importantly, in order to guarantee hypervisors' applicability, load balance strategies of virtual CPU (VCPU) schedulers do not consider the memory access characteristicsof applications running in VMs, which probably introducessignificant shared resource contention and unnecessary remotememory accesses. In this paper, we propose a NUMA-aware VCPU schedulerbased on Xen, named vProbe, to improve the performanceof memory-intensive applications while maintaining the transparencyof the virtualization layer in NUMA-based servers. It collects performance monitoring units (PMU) data for eachVCPU and analyzes their memory access characteristics. Then, according to the memory access characteristics of each VCPU, it periodically reassigns all memory-intensive VCPUs to eachNUMA node evenly while preferentially allocating them to theirlocal nodes, which aims to alleviate shared resource contentionand reduce unnecessary remote memory accesses. Moreover, when a physical CPU (PCPU) becomes idle, it preferentiallysteals a VCPU from the run queues of PCPUs in the local nodeto this PCPU, which helps to maintain balanced last-level cache (LLC) contention and reduce extra remote memory accesses. Our evaluation shows that vProbe can significantly improve theperformance of memory-intensive applications (e.g., up to 45.2% performance improvement compared with the Credit scheduler) while introducing negligible overheads. Song Wu 0001, Huahua Sun, Like Zhou, Qingtian Gan, Hai Jin 0001 |
CLUSTER | 1 |
| 2016 | Dynamic Acceleration of Parallel Applications in Cloud Platforms by Adaptive Time-Slice ControlabstractTightly-coupled parallel applications in cloud systems may suffer from significant performance degradation because of the resource over-commitment issue. In this paper, we propose a dynamic approach based on the adaptive control over time-slice for virtual clusters, in order to mitigate the performance degradation for parallel applications in cloud and avoid the negative impact effectively on other non-parallel applications meanwhile. The key idea is to reduce the synchronization overhead inside and across virtual machines (VMs) in cloud systems, by dynamically adjusting the time-slices of VMs in terms of the spinlock latency at runtime. Such a design is motivated by our experimental finding that VM's time slice is a key factor determining the synchronization overhead as well as the parallel execution performance. We perform the evaluation on a real cluster environment deployed with XEN, using five well-known benchmarks with 10+ applications. Experiments show that our approach obtains 1.5-10× performance gain for running parallel applications, than other state-of-the-art solutions (including Credit Scheduling of Xen and the well-known methods like Co-Scheduling and Balance Scheduling), with nearly unaffected impact on the performance of non-parallel applications. Song Wu 0001, Zhenjiang Xie, Haibao Chen, Sheng Di, Hai Jin 0001 |
IPDPS | 1 |
| 2016 | HybridScaler: Handling Bursting Workload for Multi-tier Web Applications in CloudabstractCloud elasticity allows users to dynamically allocate resources for their applications to adapt with the fluctuant demand. But allocating right amount of resources at right time to handle the bursting workload is still challenging. Most practical auto scaling approaches allocate resources in horizontal or vertical manner while the horizontal scaling usually causes considerable overhead and extra cost for short-term bursting workload. Vertical scaling is lightweight and timely but lacks scalability and capacity guarantee in public cloud. In order to find an accurate and cost-effective auto scaling method, we propose a hybrid auto scaling solution called HybridScaler which combines long-term predictive horizontal scaling and timely reactive vertical scaling properly. Specially, our method is based on a resource-pressure model which can provide suitable amount of resources matching the changing workload. We implement our prototype in an OpenStack private cloud. We evaluate the effectiveness and efficiency of HybridScaler by comparing with mainstream auto scaling methods. HybridScaler decreases 16-39% average response time and 34-50% SLO violation rate than both static threshold-based scheme and prediction-based scheme. Meanwhile, it uses less instance-hours than static threshold-based scaling method and keeps CPU utilization almost between 60% and 70% which can avoid significant resource waste and SLO violation in the other methods. Song Wu 0001, Binji Li, Xinhou Wang, Hai Jin 0001 |
ISPDC | 1 |
| 2016 | Dynamic resource scheduling in cloud radio access network with mobile cloud computingabstractNowadays, by integrating the cloud radio access network (C-RAN) with the mobile cloud computing (MCC) technology, mobile service provider (MSP) can efficiently handle the increasing mobile traffic and enhance the capabilities of mobile users' devices to provide better quality of service (QoS). But the power consumption has become skyrocketing for MSP as it gravely affects the profit of MSP. Previous work often studied the power consumption in C-RAN and MCC separately while less work had considered the integration of C-RAN with MCC. In this paper, we present a unifying framework for optimizing the power-performance tradeoff of MSP by jointly scheduling network resources in C-RAN and computation resources in MCC to minimize the power consumption of MSP while still guaranteeing the QoS for mobile users. Our objective is to maximize the profit of MSP. To achieve this objective, we first formulate the resource scheduling issue as a stochastic problem and then propose a Resource onlIne sCHeduling (RICH) algorithm using Lyapunov optimization technique to approach a time average profit that is close to the optimum with a diminishing gap (1/V) for MSP while still maintaining strong system stability and low congestion to guarantee the QoS for mobile users. With extensive simulations, we demonstrate that the profit of RICH algorithm is 3.3× (18.4×) higher than that of active (random) algorithm. Xinhou Wang, Kezhi Wang, Song Wu 0001, Sheng Di, Kun Yang 0001, Hai Jin 0001 |
IWQoS | 3 |
| 2016 | iShare: Balancing I/O performance isolation and disk I/O efficiency in virtualized environmentsabstractSummary Performance isolation has long been a challenging problem for disk resource allocation in virtualized environments. While there have been many researches working on I/O performance isolation and disk utilization, none of them addresses the I/O performance isolation and disk utilization as a whole. To this end, we investigate the impact of current disk I/O performance isolation schemes on disk I/O utilization. Interestingly, our studies report that current isolation schemes bring unnecessary disk idle and reduce the overall disk I/O performance because of ignoring the disk states and characteristics of requests. Accordingly, we propose an adaptive proportional‐share I/O scheduling framework, namediShare, in virtualized environments.iSharenot only ensures I/O performance isolation through proportionally allocating time slices according to the weights of virtual machines but also preserves high disk efficiency by detecting disk states and adaptively adjusting the time slice size based on characteristics of requests. We implement a prototype ofiShareon the Xen platform. The experimental results show thatiShareensures I/O performance isolation while improving disk I/O efficiency, compared withBlkio(i.e., the default I/O performance isolation method in Xen),iShareincreases disk I/O bandwidth by 58% and slightly improves the I/O performance isolation for the sequential write applications. Copyright © 2015 John Wiley & Sons, Ltd. Song Wu 0001, Songqiao Tao, Hao Fan 0006, Hai Jin 0001, Shadi Ibrahim |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | A survey of cloud resource management for complex engineering applications
Haibao Chen, Song Wu 0001, Hai Jin 0001, Jidong Zhai, Yingwei Luo, Xiaolin Wang 0001 |
Frontiers Comput. Sci. | 2 |
| 2016 | Time Donating Barrier for efficient task scheduling in competitive multicore systems
Song Wu 0001, Yaqiong Peng, Hai Jin 0001 |
Future Gener. Comput. Syst. | 1 |
| 2016 | FITDOC: fast virtual machines checkpointing with delta memory compression
Yunjie Du, Xuanhua Shi, Hai Jin 0001, Song Wu 0001, Laurence T. Yang |
J. Supercomput. | 4 |
| 2016 | A Performance Debugging Framework for Unnecessary Lock Contentions with Record/Replay TechniquesabstractLocks have been widely used as an effective synchronization mechanism among processes and threads. However, we observe that, a large number of false inter-thread dependencies (i.e., unnecessary lock contentions) exist during the program execution on multicore processors, incurring significant performance overhead. This paper presents a performance debugging framework, PERFPLAY, to facilitate the identification of unnecessary lock contentions and to guide programmers to improve the program performance by eliminating the unnecessary lock contentions. Since the performance debugging of unnecessary lock contentions is input-sensitive, we first identify the representative inputs for performance debugging. Next, PERFPLAY quantifies the performance impact of unnecessary lock contention code regions for each candidate input. Taking into account conflicting attribute of performance impact and input coverage in the real world, we finally make the tradeoff between performance impact and input coverage to recommend the optimal unnecessary lock contention code regions. Our final results on five real-world programs and PARSEC benchmarks demonstrate the significant performance overhead of unnecessary lock contentions, and the effectiveness of PERFPLAY in troubleshooting the target unnecessary lock contention code regions with the consideration of both performance impact and input coverage. Xiaofei Liao, Long Zheng 0003, Bingsheng He, Song Wu 0001, Hai Jin 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | Robinhood: Towards Efficient Work-Stealing in Virtualized EnvironmentsabstractWork-stealing, as a common user-level task scheduler for managing and scheduling tasks of multithreaded applications, suffers from inefficiency in virtualized environments, because the steal attempts of thief threads may waste CPU cycles that could be otherwise used by busy threads. This paper contributes a novel scheduling framework named Robinhood. The basic idea of Robinhood is to use the time slices of thieves to accelerate busy threads with no available tasks (referred to as poor workers) at both the guest Operating System (OS) level and Virtual Machine Monitor (VMM) level. In this way, Robinhood can reduce the cost of steal attempts and accelerate the threads doing useful work, so as to put the CPU cycles to better use. We implement Robinhood based on BWS, Linux and Xen. Our evaluation with various benchmarks demonstrates that Robinhood paves a way to efficiently run work-stealing applications in virtualized environments. Compared to Cilk++ and BWS, Robinhood can reduce up to 90 and 72 percent execution time of work-stealing applications, respectively. Yaqiong Peng, Song Wu 0001, Hai Jin 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Poris: A Scheduler for Parallel Soft Real-Time Applications in Virtualized EnvironmentsabstractWith the prevalence of cloud computing and virtualization, more and more cloud services including parallel soft real-time applications (PSRT applications) are running in virtualized data centers. However, current hypervisors do not provide adequate support for them because of soft real-time constraints and synchronization problems, which result in frequent deadline misses and serious performance degradation. CPU schedulers in underlying hypervisors are central to these issues. In this paper, we identify and analyze CPU scheduling problems in hypervisors. Then, we design and implement a parallel soft real-time scheduler according to the analysis, namedPoris, based on Xen. It addresses both soft real-time constraints and synchronization problems simultaneously. In our proposed method,priority promotionanddynamic time slicemechanisms are introduced to determine when to schedulevirtual CPUs(VCPUs) according to the characteristics of soft real-time applications. Besides, considering that PSRT applications may run in avirtual machine(VM) or multiple VMs, we presentparallel scheduling,group schedulingandcommunication-driven group schedulingto accelerate synchronizations of these applications and make sure that tasks are finished before their deadlines under different scenarios. Our evaluation showsPoriscan significantly improve the performance of PSRT applications no matter how they run in a VM or multiple VMs. For example, compared to the Credit scheduler,Porisdecreases the response time of web search benchmark by up to 91.6 percent. Song Wu 0001, Like Zhou, Huahua Sun, Hai Jin 0001, Xuanhua Shi |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | A Software-Defined Cloud Resource Management Framework
Aaqif Afzaal Abbasi, Hai Jin 0001, Song Wu 0001 |
APSCC | 3 |
| 2015 | GPU-based multifrontal optimizing method in sparse Cholesky factorizationabstractIn many scientific computing applications, sparse Cholesky factorization is used to solve large sparse linear equations in distributed environment. GPU computing is a new way to solve the problem. However, sparse Cholesky factorization on GPU is hardly to achieve excellent performance due to the structure irregularity of matrix and the low GPU resource utilization. A hybrid CPU-GPU implementation of sparse Cholesky factorization is proposed based on multifrontal method. A large sparse coefficient matrix is decomposed into a series of small dense matrices (frontal matrices) in the method, and then multiple GEMM (General Matrix-matrix Multiplication) operations are computed. GEMMs are the main operations in sparse Cholesky factorization, but they are hardly to perform better in parallel on GPU. In order to improve the performance, the scheme of multiple task queues is adopted when performing multiple GEMMs parallelized with multifrontal method; all GEMM tasks are scheduled dynamically on GPU and CPU based on computation scales for load balance and computing-time reduction. Experimental results show that the approach can outperform the implementations of BLAS and cuBLAS, achieving up to 3.15× and 1.98× speedup, respectively. Wei Wang 0088, Hai Jin 0001, Song Wu 0001, Yong Chen 0004, Han Jiang 0010 |
ASAP | 4 |
| 2015 | Towards Efficient Work-Stealing in Virtualized EnvironmentsabstractWork-stealing, as a common user-level task scheduler for managing and scheduling tasks among worker threads, has been widely adopted in multithreaded applications. With work-stealing, worker threads attempt to steal tasks from other threads' queue when they run out of their own tasks. Though work-stealing based applications can achieve good performance due to the dynamic load balancing, these steal attempting operations may frequently fail especially when available tasks are scarce, thus wasting CPU resources of busy worker threads and consequently making work-stealing less efficient in competitive environments, such as traditional multi programmed and virtualized environments. Although there are some optimizations for reducing the cost of steal-attempting threads by having such threads yield their computing resources in traditional multi programmed environments, it is more challenging to enhance the efficiency of work-stealing in virtualized environments due to the semantic gap between the virtual machine monitor (VMM) and virtual machines (VMs). In this paper, we first analyze the challenges of enhancing the efficiency of work-stealing in virtualized environments, and then propose Robin hood, a scheduling framework that reduces the cost of virtual CPUs (vCPUs) running steal-attempting threads and the scheduling delay of vCPUs running busy threads. Different from traditional scheduling methods, if the steal attempting failure occurs, Robin hood can supply the CPU time of vCPUs running steal-attempting threads to their sibling vCPUs running busy threads, which can not only improve the CPU resource utilization but also guarantee better fairness among different VMs sharing the same physical node. We implement Robin hood based on BWS, Linux and Xen. Our evaluation with various benchmarks demonstrates that Robin hood paves a way to efficiently run work-stealing based applications in virtualized platform. It can reduce up to 64% and 30% execution time of work-stealing benchmarks compared to Cilk++ and BWS respectively, and outperform credit scheduler and co-scheduling for average system throughput by 1.91× and 1.3× respectively, while guaranteeing the performance fairness among applications in virtualized environments. Yaqiong Peng, Song Wu 0001, Hai Jin 0001 |
CCGRID | 2 |
| 2015 | On performance debugging of unnecessary lock contentions on multicore processors: a replay-based approachabstractLocks have been widely used as an effective synchronization mechanism among processes and threads. However, we observe that a large number of false inter-thread dependencies (i.e., unnecessary lock contentions) exist during the program execution on multicore processors, thereby incurring significant performance overhead. This paper presents a performance debugging framework, PerfPlay, to facilitate a comprehensive and in-depth understanding of the performance impact of unnecessary lock contentions. The core technique of our debugging framework is trace replay. Specifically, PerfPlay records the program execution trace, on the basis of which the unnecessary lock contentions can be identified through trace analysis. We then propose a novel technique of trace transformation to transform these identified unnecessary lock contentions in the original trace into the correct pattern as a new trace free of unnecessary lock contentions. Through replaying both traces, PerfPlay can quantify the performance impact of unnecessary lock contentions. To demonstrate the effectiveness of our debugging framework, we study five real-world programs and PARSEC benchmarks. Our experimental results demonstrate the significant performance overhead of unnecessary lock contentions, and the effectiveness of PerfPlay in identifying the performance critical unnecessary lock contentions in real applications. Long Zheng 0003, Xiaofei Liao, Bingsheng He, Song Wu 0001, Hai Jin 0001 |
CGO | 4 |
| 2015 | Evaluating Latency-Sensitive Applications: Performance Degradation in Datacenters with Restricted Power BudgetabstractFor data centers with limited power supply, restricting the servers' power budget (i.e., The maximal power provided to servers) is an efficient approach to increase the server density (the server quantity per rack), which can effectively improve the cost-effectiveness of the data centers. However, this approach may also affect the performance of applications in servers. Hence, the prerequisite of adopting the approach in data centers is to precisely evaluate the application performance degradation caused by restricting the servers' power budget. Unfortunately, existing evaluation methods are inaccurate because they are either improper or coarse-grained, especially for the latency-sensitive applications widely deployed in data centers. In this paper, we analyze the reasons why state-of-the-art methods are not appropriate for evaluating the performance degradation of latency-sensitive applications in case of power restriction, and we propose a new evaluation method which can provide a fine-grained way to precisely describe and evaluate such degradation. We verify our proposed method by a real-world application and the traces from Ten cent's date enter with 25328 servers. The experimental results show that our method is much more accurate compared with the state of the art, and we can significantly increase datacenter efficiency by saving servers' power budget while maintaining the applications' performance degradation within controllable and acceptable range. Song Wu 0001, Chuxiong Yan, Haibao Chen, Hai Jin 0001, Deqing Zou |
ICPP | 1 |
| 2015 | Optimization strategies for inter-thread synchronization overhead on NUMA machineabstractOverhead caused by data consistence issue in inter-thread synchronization probably degrades the performance of parallel applications. Non-Uniform Memory Access (NUMA), as the mainstream architecture in today's multicore processor, further exacerbates this issue due to the significant overhead incurred by Remote Memory Reference (RMR). Therefore, to reduce synchronization overhead, it is important to solve the data consistence issue. In this paper, we classify the overhead into two kinds: (1) overhead incurred by algorithms themselves, and (2) overhead incurred by critical sections. To reduce two kinds of overhead on NUMA machine, we present two optimization strategies called search and backtrace (SAB) and reorder critical section and non-critical section (RCAN), respectively. In SAB, a server thread tries to search a thread coming from master NUMA node, and designates it as the new server thread. In this way, most of the time, shared data resides in the cache of master NUMA node, resulting in lower overhead caused by data consistence issue in critical section. In RCAN, each thread consecutively posts synchronization requests, followed by consecutively executing non-critical section. In this way, server threads could serve enough requests, resulting in better data locality. We design an algorithm named R-Synch based on SAB, while designing an algorithm named H-STA based on RCAN. Our evaluation with representative synchronization algorithms demonstrates the effectiveness of R-Synch and H-STA. Song Wu 0001, Yaqiong Peng, Hai Jin 0001, Wenbin Jiang 0001 |
IPCCC | 1 |
| 2015 | Understanding and identifying latent data races cross-thread interleaving
Long Zheng 0003, Xiaofei Liao, Song Wu 0001, Xuepeng Fan, Hai Jin 0001 |
Frontiers Comput. Sci. | 3 |
| 2015 | Rethink the storage of virtual machine images in clouds
Hai Jin 0001, Song Wu 0001 |
Future Gener. Comput. Syst. | 3 |
| 2015 | Towards Optimized Fine-Grained Pricing of IaaS Cloud PlatformabstractAlthough many pricing schemes in IaaS platform are already proposed with pay-as-you-go and subscription/spot market policy to guarantee service level agreement, it is still inevitable to suffer from wasteful payment because of coarse-grained pricing scheme. In this paper, we investigate an optimized fine-grained and fair pricing scheme. Two tough issues are addressed: (1) the profits of resource providers and customers often contradict mutually; (2) VM-maintenance overhead like startup cost is often too huge to be neglected. Not only can we derive an optimal price in the acceptable price range that satisfies both customers and providers simultaneously, but we also find a best-fit billing cycle to maximize social welfare (i.e., the sum of the cost reductions for all customers and the revenue gained by the provider). We carefully evaluate the proposed optimized fine-grained pricing scheme with two large-scale real-world production traces (one from Grid Workload Archive and the other from Google data center). We compare the new scheme to classic coarse-grained hourly pricing scheme in experiments and find that customers and providers can both benefit from our new approach. The maximum social welfare can be increased up to 72.98 and 48.15 percent with respect to DAS-2 trace and Google trace respectively. Hai Jin 0001, Xinhou Wang, Song Wu 0001, Sheng Di, Xuanhua Shi |
IEEE Trans. Cloud Comput. | 3 |
| 2015 | Spatial Locality Aware Disk Scheduling in Virtualized EnvironmentabstractExploiting spatial locality, a key technique for improving disk I/O utilization and performance, faces additional challenges in the virtualized cloud because of the transparency feature of virtualization. This paper contributes a novel disk I/O scheduling framework, named Pregather, to improve disk I/O efficiency through exposure and exploitation of the special spatial locality in the virtualized environment, thereby improving the performance of disk-intensive applications without harming the transparency feature of virtualization. The key idea behind Pregatheris to implement an intelligent model to predict the access regularity of spatial locality for each VM. Moreover, Pregather embraces an adaptive time slice allocation scheme to further reduce the resource contention and ensure fairness among VMs. We implement the Pregather disk scheduling framework and perform extensive experiments that involve multiple simultaneous applications of both synthetic benchmarks and MapReduce applications on Xen-based platforms. Our experiments demonstrate the accuracy of our prediction model and indicate that Pregather results in the high disk spatial locality and a significant improvement in disk throughput and application performance. Shadi Ibrahim, Song Wu 0001, Hai Jin 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | Mammoth: Gearing Hadoop Towards Memory-Intensive MapReduce ApplicationsabstractThe MapReduce platform has been widely used for large-scale data processing and analysis recently. It works well if the hardware of a cluster is well configured. However, our survey has indicated that common hardware configurations in small- and medium-size enterprises may not be suitable for such tasks. This situation is more challenging for memory-constrained systems, in which the memory is a bottleneck resource compared with the CPU power and thus does not meet the needs of large-scale data processing. The traditional high performance computing (HPC) system is an example of the memory-constrained system according to our survey. In this paper, we have developed Mammoth, a new MapReduce system, which aims to improve MapReduce performance using global memory management. In Mammoth, we design a novel rule-based heuristic to prioritize memory allocation and revocation among execution units (mapper, shuffler, reducer, etc.), to maximize the holistic benefits of the Map/Reduce job when scheduling each memory unit. We have also developed a multi-threaded execution engine, which is based on Hadoop but runs in a single JVM on a node. In the execution engine, we have implemented the algorithm of memory scheduling to realize global memory management, based on which we further developed the techniques such as sequential disk accessing, multi-cache and shuffling from memory, and solved the problem of full garbage collection in the JVM. We have conducted extensive experiments to compare Mammoth against the native Hadoop platform. The results show that the Mammoth system can reduce the job execution time by more than 40 percent in typical cases, without requiring any modifications of the Hadoop programs. When a system is short of memory, Mammoth can improve the performance by up to 5.19 times, as observed for I/O intensive applications, such as PageRank. We also compared Mammoth with Spark. Although Spark can achieve better performance than Mammoth for interactive and iterative applications when the memory is sufficient, our experimental results show that for batch processing applications, Mammoth can adapt better to various memory environments and outperform Spark when the memory is insufficient, and can obtain similar performance as Spark when the memory is sufficient. Given the growing importance of supporting large-scale data processing and analysis and the proven success of the MapReduce platform, the Mammoth system can have a promising potential and impact. Xuanhua Shi, Ligang He, Lu Lu 0006, Hai Jin 0001, Yong Chen 0001, Song Wu 0001 |
IEEE Trans. Parallel Distributed Syst. | 8 |
| 2015 | Synchronization-Aware Scheduling for Virtual Clusters in CloudabstractDue to high flexibility and cost-effectiveness, cloud computing is increasingly being explored as an alternative to local clusters by academic and commercial users. Recent research already confirmed the feasibility of running tightly-coupled parallel applications with virtual clusters. However, such types of applications suffer from significant performance degradation, especially as the over-commitment is common in cloud. That is, the number of executable Virtual CPUs (VCPUs) is often larger than that of available Physical CPUs (PCPUs) in the system. The performance degradation is mainly due to the fact that the current virtual machine monitors (VMMs) are unaware of the synchronization requirements of the VMs which are running parallel applications. In this paper, There are two key contributions. (1) We propose an autonomous synchronization-aware VM scheduling (SVS) algorithm, which can effectively mitigate the performance degradation of tightly-coupled parallel applications running atop them in over-committed situation. (2) We integrate the SVS algorithm into Xen VMM scheduler, and rigorously implement a prototype. We evaluate our design on a real cluster environment with NPB benchmark and real-world trace. Experiments show that our solution attains better performance for tightly-coupled parallel applications than the state-of-the-art approaches like Xen's Credit scheduler, balance scheduling, and hybrid scheduling. Song Wu 0001, Haibao Chen, Sheng Di, Bing Bing Zhou, Zhenjiang Xie, Hai Jin 0001, Xuanhua Shi |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Iteration Based Collective I/O Strategy for Parallel I/O SystemsabstractMPI collective I/O is a widely used I/O method that helps data-intensive scientific applications gain better I/O performance. However, it has been observed that existing collective I/O strategies do not perform well due to the access contention problem. Existing collective I/O optimization strategies mainly focus on the I/O phase efficiency and ignore the shuffle cost that may limit the potential of their performance improvement. We observe that as the size of I/O becomes larger, one I/O operation from the upper application would be separated into several iterations to complete. So, I/O requests in each file domain do not necessarily issue to the parallel file system simultaneously unless they are carried out within the same iteration step. Based on that observation, this paper proposes a new collective I/O strategy that reorganizes I/O requests within each file domain instead of coordinating requests across file domains, such that we can eliminate access contentions without introducing extra shuffle cost between aggregators and computing processes. Using benchmark workloads IOR, we evaluate our new strategy and compare with the conventional one. The proposed strategy achieves up to 47%-63% I/O bandwidth improvement compared to the existing ROMIO collective I/O strategy. Xuanhua Shi, Hai Jin 0001, Song Wu 0001, Yong Chen 0001 |
CCGRID | 4 |
| 2014 | GIRAFFE: A scalable distributed coordination service for large-scale systemsabstractThe scale of cloud services keeps increasing over time, significantly introducing huge challenges in system manageability and reliability. Designing coordination services in cloud is the right track to solve the above problems. However, existing coordination services (e.g., Chubby and ZooKeeper) only perform well in read-intensive scenario and small ensemble scales. To this end, we propose Giraffe, a scalable distributed coordination service. There are three important contributions in our design. (1) Giraffe organizes coordination servers using interior-node-disjoint trees for better scalability. (2) Giraffe employs a novel Paxos protocol for strong consistency and fault-tolerance. (3) Giraffe supports hierarchical data organization and in-memory storage for high throughput and low latency. We evaluate Giraffe on a high performance computing test-bed. The experimental results show that Giraffe gains much better write performance than ZooKeeper when server ensemble is large. Giraffe is nearly 300% faster than ZooKeeper on update operations when ensemble size is 50 servers. Experiments also show that Giraffe reacts and recovers more quickly than ZooKeeper against node failures. Xuanhua Shi, Haohong Lin, Hai Jin 0001, Bing Bing Zhou, Zuoning Yin, Sheng Di, Song Wu 0001 |
CLUSTER | 7 |
| 2014 | Communication-driven scheduling for virtual clusters in cloudabstractDue to high flexibility and cost-effectiveness, cloud computing is increasingly being explored as an alternative to local clusters by academic and commercial users. Recent research already confirmed the feasibility of running tightly-coupled parallel applications with virtual clusters. However, such types of applications suffer from significant performance degradation, especially as the overcommitment is common in cloud. That is, the number of executable Virtual CPUs (VCPUs) is often larger than that of available Physical CPUs (PCPUs) in the system. The performance degradation mainly results from that the current Virtual Machine Monitors (VMMs) cannot co-schedule (or coordinate at the same time) the VCPUs that host parallel application threads/processes with synchronization requirements. Haibao Chen, Song Wu 0001, Sheng Di, Bing Bing Zhou, Zhenjiang Xie, Hai Jin 0001, Xuanhua Shi |
HPDC | 2 |
| 2014 | A Real-Time Scheduling Framework Based on Multi-core Dynamic Partitioning in Virtualized Environment
Song Wu 0001, Like Zhou, Danqing Fu, Hai Jin 0001, Xuanhua Shi |
NPC | 1 |
| 2014 | Page Classifier and Placer: A Scheme of Managing Hybrid Caches
Xuanhua Shi, Hai Jin 0001, Xiaofei Liao, Song Wu 0001, Xiaoming Li 0010 |
NPC | 5 |
| 2014 | MECOM: Live migration of virtual machines by adaptively compressing memory pages
Hai Jin 0001, Song Wu 0001, Xuanhua Shi, Hanhua Chen |
Future Gener. Comput. Syst. | 3 |
| 2014 | Morpho: A decoupled MapReduce framework for elastic cloud computing
Lu Lu 0006, Xuanhua Shi, Hai Jin 0001, Qiuyue Wang, Daxing Yuan, Song Wu 0001 |
Future Gener. Comput. Syst. | 6 |
| 2014 | Peacock: a customizable MapReduce for multicore platform
Song Wu 0001, Yaqiong Peng, Hai Jin 0001 |
J. Supercomput. | 1 |
| 2013 | Cost-Aware Client-Side File Caching for Data-Intensive ApplicationsabstractParallel and distributed file systems are widely used to provide high throughput in high-performance computing and Cloud computing systems. To increase the parallelism, I/O requests are partitioned into multiple sub-requests (or `flows') and distributed across different data nodes. The performance of file systems is extremely poor if data nodes have highly unbalanced response time. Client-side caching offers a promising direction for addressing this issue. However, current work has primarily used client-side memory as a read cache and employed a write-through policy which requires synchronous update for every write and significantly under-utilizes the client-side cache when the applications are write-intensive. Realizing that the cost of an I/O request depends on the struggler sub-requests, we propose a cost-aware client-side file caching (CCFC) strategy, that is designed to cache the sub-requests with high I/O cost on the client end. This caching policy enables a new trade-off across write performance, consistency guarantee and cache size dimensions. Using benchmark workloads MADbench2, we evaluate our new cache policy alongside conventional write-through. We find that the proposed CCFC strategy can achieve up to 110% throughput improvement compared to the conventional write-through policies with the same cache size on an 85-node cluster. Yaning Huang, Hai Jin 0001, Xuanhua Shi, Song Wu 0001, Yong Chen 0001 |
CloudCom (2) | 4 |
| 2013 | vMerger: Server Consolidation in Virtualized EnvironmentabstractIn virtualized clusters, applications are deployed in virtual machines (VMs). Resource can be configured dynamically with the varying workloads. When workloads become low, applications can be consolidated on smaller number of nodes. Much more tasks can be processed in clusters. Virtualized clusters have better scalability than traditional clusters. In this paper, we present a server consolidation manager vMerger to improve the scalability of clusters. Linear programming is used for finding new mapping of VMs to smaller number of nodes. An optimized topological-sorting-based migration order generation method is designed to change VM distribution from current mapping to new mapping. The experimental results show that, vMerger effectively improves the scalability of clusters using server consolidation. Hai Jin 0001, Song Wu 0001 |
DASC | 3 |
| 2013 | Supporting parallel soft real-time applications in virtualized environment
Like Zhou, Song Wu 0001, Huahua Sun, Hai Jin 0001, Xuanhua Shi |
HPDC | 2 |
| 2013 | Exploiting Spatial Locality to Improve Disk Efficiency in Virtualized EnvironmentsabstractVirtualization has become a prominent tool in data centers and is extensively leveraged in cloud environments: it enables multiple virtual machines (VMs) - with multiple operating systems and applications - to run within a physical server. However, virtualization introduces the challenging issue of preserving the high disk utilization (i.e., reducing the seek delay and rotation overhead) when allocating disk resources to VMs. Exploiting spatial locality, a key technique for improving disk utilization and performance, faces additional challenges in the virtualized cloud because of the transparency feature of virtualization (hyper visors do not have the information about the access patterns of applications running within each VM). To this end, this paper contributes a novel disk I/O scheduling framework, named Pregather, to improve disk I/O efficiency through exposure and exploitation of the special spatial locality in the virtualized environment (regional and sub-regional spatial locality corresponds to the virtual disk space and applications' access patterns, respectively), thereby improving the performance of disk-intensive applications without harming the transparency feature of virtualization (without a priori knowledge of the applications' access patterns). The key idea behind Pregather is to implement an intelligent model to predict the access regularity of sub-regional spatial locality for each VM. We implement the Pregather disk scheduling framework and perform extensive experiments that involve multiple simultaneous applications of both synthetic benchmarks and a MapReduce application on Xen-based platforms. Our experiments demonstrate the accuracy of our prediction model and indicate that Pregather results in the high disk spatial locality and a significant improvement in disk throughput and application performance. Shadi Ibrahim, Hai Jin 0001, Song Wu 0001, Songqiao Tao |
MASCOTS | 4 |
| 2013 | Virtual Machine Scheduling for Parallel Soft Real-Time ApplicationsabstractWith the prevalence of multicore processors in computer systems, many soft real-time applications, such as media-based ones, use parallel programming models to utilize hardware resources better and possibly shorten response time. Meanwhile, virtualization technology is widely used in cloud data centers. More and more cloud services including such parallel soft real-time applications are running in virtualized environment. However, current hyper visors do not provide adequate support for them because of soft real-time constraints and synchronization problems, which result in frequent deadline misses and serious performance degradation. CPU schedulers in underlying hyper visors are central to these issues. In this paper, we identify and analyze CPU scheduling problems in hyper visors, and propose a novel scheduling algorithm considering both soft real-time constraints and synchronization problems. In our proposed method, real-time priority is introduced to accelerate event processing of parallel soft real-time applications, and dynamic time slice is used to schedule virtual CPUs. Besides, all runnable virtual CPUs of virtual machines running parallel soft real-time applications are scheduled simultaneously to address synchronization problems. We implement a parallel soft real-time scheduler, named Poris, based on Xen. Our evaluation shows Poris can significantly improve the performance of parallel soft real-time applications. For example, compared to the Credit scheduler, Poris improves the performance of media player by up to a factor of 1.35, and shortens the execution time of PARSEC benchmark by up to 44.12%. Like Zhou, Song Wu 0001, Huahua Sun, Hai Jin 0001, Xuanhua Shi |
MASCOTS | 2 |
| 2013 | Flubber: Two-level disk scheduling in virtualized environment
Hai Jin 0001, Shadi Ibrahim, Wenzhi Cao, Song Wu 0001, Gabriel Antoniu |
Future Gener. Comput. Syst. | 5 |
| 2013 | Dependable Grid Workflow Scheduling Based on Resource Availability
Yongcai Tao, Hai Jin 0001, Song Wu 0001, Xuanhua Shi, Lei Shi 0001 |
J. Grid Comput. | 3 |
| 2013 | Handling partitioning skew in MapReduce using LEEN
Shadi Ibrahim, Hai Jin 0001, Lu Lu 0006, Bingsheng He, Gabriel Antoniu, Song Wu 0001 |
Peer-to-Peer Netw. Appl. | 6 |
| 2013 | Petri net based Grid workflow verification and optimization
Haijun Cao, Hai Jin 0001, Song Wu 0001, Shadi Ibrahim |
J. Supercomput. | 3 |
| 2013 | A VMM-based intrusion prevention system in cloud computing environment
Hai Jin 0001, Guofu Xiang, Deqing Zou, Song Wu 0001, Feng Zhao 0003, Weide Zheng |
J. Supercomput. | 4 |
| 2013 | Adapting grid computing environments dependable with virtual machines: design, implementation, and evaluations
Xuanhua Shi, Hai Jin 0001, Song Wu 0001 |
J. Supercomput. | 3 |
| 2012 | Maestro: Replica-Aware Map Scheduling for MapReduceabstractMapReduce has emerged as a leading programming model for data-intensive computing. Many recent research efforts have focused on improving the performance of the distributed frameworks supporting this model. Many optimizations are network-oriented and most of them mainly address the data shuffling stage of MapReduce. Our studies with Hadoop demonstrate that, apart from the shuffling phase, another source of excessive network traffic is the high number of map task executions which process remote data. That leads to an excessive number of useless speculative executions of map tasks and to an unbalanced execution of map tasks across different machines. All these factors produce a noticeable performance degradation. We propose a novel scheduling algorithm for map tasks, named Maestro, to improve the overall performance of the MapReduce computation. Maestro schedules the map tasks in two waves: first, it fills the empty slots of each data node based on the number of hosted map tasks and on the replication scheme for their input data, second, runtime scheduling takes into account the probability of scheduling a map task on a given machine depending on the replicas of the task's input data. These two waves lead to a higher locality in the execution of map tasks and to a more balanced intermediate data distribution for the shuffling phase. In our experiments on a 100-node cluster, Maestro achieves around 95% local map executions, reduces speculative map tasks by 80% and results in an improvement of up to 34% in the execution time. Shadi Ibrahim, Hai Jin 0001, Lu Lu 0006, Bingsheng He, Gabriel Antoniu, Song Wu 0001 |
CCGRID | 6 |
| 2012 | Efficient Disk I/O Scheduling with QoS Guarantee for Xen-based Hosting PlatformsabstractIn this paper, we address the problem of allocating disk resources to guarantee specified latency and throughput targets of VMs while keeping efficient disk I/O. Accordingly, we present two-level scheduling framework, namely Flubber, in Xen-based hosting platform that decouples latency and throughput allocation. The high-level throughput control regulates the pending requests from the VMs, in order to meet the throughput requirements of different VMs and ensure isolation. Meanwhile, the low-level latency control, by the virtue of the batch and delay EDF mechanism, reorders all pending requests from VMs based on the their deadlines, and batches them to the disk device considering the locality of accesses across VMs. We have implemented Flubber with intensive evaluations on Xen-based host. The results show that Flubber can simultaneously meet the different service requirements of VMs while improving the efficiency of the physical disk. In contrast to CFQ, besides that Flubber achieves the desired QoS of each VM, Flubber speeds up the sequential and random read by 17% and 25% due to the efficient physical disk utilization. Hai Jin 0001, Shadi Ibrahim, Wenzhi Cao, Song Wu 0001 |
CCGRID | 5 |
| 2012 | Effectively deploying services on virtualization infrastructure
Hai Jin 0001, Song Wu 0001, Xuanhua Shi, Jinyan Yuan |
Frontiers Comput. Sci. | 3 |
| 2011 | Adaptive Disk I/O Scheduling for MapReduce in Virtualized EnvironmentabstractVirtual machine (VM) interference has long been a challenging problem for performance predictability and system throughput for large-scale virtualized environments in the cloud. Such interferences are contributed by intertwined factors including the application's type, the number of con current VMs, and the VM scheduling algorithms used within the host. Since MapReduce has become an important data processing platform in the cloud, we investigate the impact of disk schedulers in Hadoop. Interestingly, our experimental results report a noticeable variation of the Hadoop performance between different applications when applying different disk pairs' schedulers in both the hypervisor and the virtual machines. Furthermore, a typical Hadoop application consists of different interleaving stages, each requiring different I/O workloads and patterns. As a result, the disk pairs' schedulers are not only sub-optimal for different MapReduce applications, but also sub-optimal for different sub-phases of the whole job. Accordingly, this paper presents a novel approach for adaptively tuning the disk pairs' schedulers in both the hypervisor and the virtual machines during the execution of a single MapReduce job. Our results show that MapReduce performance can be significantly improved; specifically, adaptive tuning of disk pairs' schedulers achieves a 25% performance improvement on a sort benchmark with Hadoop. Shadi Ibrahim, Hai Jin 0001, Lu Lu 0006, Bingsheng He, Song Wu 0001 |
ICPP | 5 |
| 2011 | Optimizing the live migration of virtual machine by CPU scheduling
Hai Jin 0001, Song Wu 0001, Xuanhua Shi, Xiaoxin Wu 0001 |
J. Netw. Comput. Appl. | 3 |
| 2010 | LEEN: Locality/Fairness-Aware Key Partitioning for MapReduce in the CloudabstractThis paper investigates the problem of Partitioning Skew in MapReduce-based system. Our studies with Hadoop, a widely used MapReduce implementation, demonstrate that the presence of partitioning skew causes a huge amount of data transfer during the shuffle phase and leads to significant unfairness on the reduce input among different data nodes. As a result, the applications experience performance degradation due to the long data transfer during the shuffle phase along with the computation skew, particularly in reduce phase. We develop a novel algorithm named LEEN for locality-aware and fairness-aware key partitioning in MapReduce. LEEN embraces an asynchronous map and reduce scheme. All buffered intermediate keys are partitioned according to their frequencies and the fairness of the expected data distribution after the shuffle phase. We have integrated LEEN into Hadoop-0.18.0. Our experiments demonstrate that LEEN can efficiently achieve higher locality and reduce the amount of shuffled data. More importantly, LEEN guarantees fair distribution of the reduce inputs. As a result, LEEN achieves a performance improvement of up to 40% on different workloads. Shadi Ibrahim, Hai Jin 0001, Lu Lu 0006, Song Wu 0001, Bingsheng He |
CloudCom | 4 |
| 2010 | MR-scope: a real-time tracing tool for MapReduceabstractMapReduce programming model is emerging as an efficient tool for data-intensive applications. Hadoop, an open-source implementation of MapReduce, has been widely adopted and experienced by both academia and enterprise. Recently, lots of efforts have been done on improving the performance of MapReduce system and on analyzing the MapReduce process based on the log files generated during the Hadoop execution. Visualizing log files seems to be a very useful tool to understand the behavior of the Hadoop process. In this paper, we present MR-Scope, a real-time MapReduce tracing tool. MR-Scope provides a real-time insight of the MapReduce process, including the ongoing progress of every task hosted in Task Tracker. In addition, it displays the health of the Hadoop cluster data nodes, the distribution of the file system's blocks and their replicas and the content of the different block splits of the file system. We implement MR-Scope in native Hadoop 0.1. Experimental results demonstrat that MR-Scope's overhead is less than 4% when running wordcount benchmark. Dachuan Huang, Xuanhua Shi, Shadi Ibrahim, Lu Lu 0006, Hongzhang Liu, Song Wu 0001, Hai Jin 0001 |
HPDC | 6 |
| 2010 | VirtCFT: A Transparent VM-Level Fault-Tolerant System for Virtual ClustersabstractA virtual cluster consists of a multitude of virtual machines and software components that are doomed to fail eventually. In many environments, such failures can result in unanticipated, potentially devastating failure behavior and in service unavailability. The ability of failover is essential to the virtual cluster's availability, reliability, and manageability. Most of the existing methods have several common disadvantages: requiring modifications to the target processes or their OSes, which is usually error prone and sometimes impractical; only targeting at taking checkpoints of processes, not whole entire OS images, which limits the areas to be applied. In this paper we present VirtCFT, an innovative and practical system of fault tolerance for virtual cluster. VirtCFT is a system-level, coordinated distributed checkpointing fault tolerant system. It coordinates the distributed VMs to periodically reach the globally consistent state and take the checkpoint of the whole virtual cluster including states of CPU, memory, disk of each VM as well as the network communications. When faults occur, VirtCFT will automatically recover the entire virtual cluster to the correct state within a few seconds and keep it running. Superior to all the existing fault tolerance mechanisms, VirtCFT provides a simpler and totally transparent fault tolerant platform that allows existing, unmodified software and operating system (version unawareness) to be protected from the failure of the physical machine on which it runs. We have implemented this system based on the Xen virtualization platform. Our experiments with real-world benchmarks demonstrate the effectiveness and correctness of VirtCFT. Minjia Zhang, Hai Jin 0001, Xuanhua Shi, Song Wu 0001 |
ICPADS | 4 |
| 2010 | Virtual Machine Management Based on Agent ServiceabstractWith the popularity of virtualization, the problem that how to manage hundreds even thousands of virtual machines running on multiple physical computing nodes becomes important. Current virtual machine management systems only can obtain basic information of virtual machines and execute simple operations on them, such as start, reboot and shutdown. In this paper, we design a virtual machine management approach based on agent service. Agent service can provide detail running status information inside virtual machines. It also has been a bridge for host machines and virtual machines to interact with each other. Agent service is designed to automatically start when virtual machine boots up. By agent service we can get real-time information about virtual machines. We evaluate monitoring overhead and the performance of batch operations when using agent. The experimental results show that agent mechanism outperforms methods using Libvirt or VMware tools. Song Wu 0001, Hai Jin 0001, Xuanhua Shi, Yankun Zhao, Jianyin Zhang |
PDCAT | 1 |
| 2010 | Scalable DHT- and ontology-based information service for large-scale grids
Yongcai Tao, Hai Jin 0001, Song Wu 0001, Xuanhua Shi |
Future Gener. Comput. Syst. | 3 |
| 2010 | Dependency-aware maintenance for highly available service-oriented grid
Hai Jin 0001, Yaqin Luo, Song Wu 0001 |
J. Syst. Softw. | 5 |
| 2010 | ServiceFlow: QoS-based hybrid service-oriented grid workflow system
Haijun Cao, Hai Jin 0001, Xiaoxin Wu 0001, Song Wu 0001 |
J. Supercomput. | 4 |
| 2010 | DAGMap: efficient and dependable scheduling of DAG workflow job in Grid
Haijun Cao, Hai Jin 0001, Xiaoxin Wu 0001, Song Wu 0001, Xuanhua Shi |
J. Supercomput. | 4 |
| 2009 | Evaluating MapReduce on Virtual Machines: The Hadoop Case
Shadi Ibrahim, Hai Jin 0001, Lu Lu 0006, Song Wu 0001, Xuanhua Shi |
CloudCom | 5 |
| 2009 | Live virtual machine migration with adaptive, memory compressionabstractLive migration of virtual machines has been a powerful tool to facilitate system maintenance, load balancing, fault tolerance, and power-saving, especially in clusters or data centers. Although pre-copy is a predominantly used approach in the state of the art, it is difficult to provide quick migration with low network overhead, due to a great amount of transferred data during migration, leading to large performance degradation of virtual machine services. This paper presents the design and implementation of a novel memory-compression-based VM migration approach (MECOM) that first uses memory compression to provide fast, stable virtual machine migration, while guaranteeing the virtual machine services to be slightly affected. Based on memory page characteristics, we design an adaptive zero-aware compression algorithm for balancing the performance and the cost of virtual machine migration. Pages are quickly compressed in batches on the source and exactly recovered on the target. Experiment demonstrates that compared with Xen, our system can significantly reduce 27.1% of downtime, 32% of total migration time and 68.8% of total transferred data on average. Hai Jin 0001, Song Wu 0001, Xuanhua Shi |
CLUSTER | 3 |
| 2009 | CLOUDLET: towards mapreduce implementation on virtual machinesabstractThe existing MapReduce framework in virtualized environment suffers from poor performance, due to the heavy overhead of I/O virtualization, and management difficulty for storage and computation. To address the problems, we propose Cloudlet, a novel MapReduce framework on virtual machines. The aim of Cloudlet design is to overcome the overhead of VM while benefiting of the other features of VM (i.e. management and reliability issues). Shadi Ibrahim, Hai Jin 0001, Bin Cheng 0001, Haijun Cao, Song Wu 0001 |
HPDC | 5 |
| 2008 | WAGA: A Flexible Web-Based Framework for Grid ApplicationsabstractThe research about the interaction between grid environments and users is becoming popular. Many scientists propose the integration for web 2.0 technologies and grid computing. In this paper, we propose a flexible web-based framework for grid applications - WAGA, which tries to bridge the gap between grid middleware and grid applications. WAGA provides a WYSIWYG (what you see is what you get) way for the programming for grid applications by adopting participation, interaction and sharing features of web 2.0 technology. With WAGA, a grid user is able to use the grid resources and to develop grid applications without understanding the underlying complexity of grids. WAGA is composed with a web GUI (called WAGA-designer) and some web-based APIs. WAGA-designer is used by grid users to develop application-based web portal, and the webAPIs are used by the WAGA-designer. The use case study shows that WAGA is flexible for grid users, and the performance evaluation shows that WAGA works with high efficiency. Xuanhua Shi, Hai Jin 0001, Song Wu 0001 |
APSCC | 4 |
| 2008 | Effectively Deploying Virtual Machines on ClusterabstractVirtualization technology has provided an opportunity to the efficient usage of computing resources. However, the management of VMs on cluster is still in the preliminary stage. How to construct user’s task environments fastly and efficiently remains a significant challenge. This paper presents a Multiple-VM Deployment System (MVDS)for creating and configuring users’ task environments on-demand. The system provides a template management model and all the VMs are created based on the templates including operating systems and applications. To improve the deployment performance, we explore some strategies about incremental mechanism and deployment tactics. We evaluate both the deployment time and I/O performance with proposed incremental mechanism. The experimental results show that the incremental mechanism outperforms the clone tactic. Song Wu 0001, Jinyan Yuan, Xuanhua Shi, Hai Jin 0001 |
APSCC | 1 |
| 2008 | A Trusted Group Signature Architecture in Virtual Computing Environment
Deqing Zou, Yunfa Li 0001, Song Wu 0001, Weizhong Qiang |
ATC | 3 |
| 2008 | PGWFT: A Petri Net Based Grid Workflow Verification and Optimization Toolkit
Haijun Cao, Hai Jin 0001, Song Wu 0001, Yongcai Tao |
GPC | 3 |
| 2008 | ChinaV: Building Virtualized Computing SystemabstractVirtualization technology has attracted much attention in recent years. This paper describes the vision and mission of ChinaV, which is the national fundamental research program for virtualization technology in China. Furthermore, related topics about single host virtualization, multiple VM management schemes and desktop virtualization will be introduced. We first describe a remote memory virtualization scheme and a VCPU management scheme for efficient use of physical resource. Then we describe a novel live VM migration approach based on deterministic replay with execution trace. Multiple VM management schemes are also introduced for multi-VM virtualization. In desktop virtualization field, we present the LVD, a system that combines the virtualization technology and inexpensive personal computers to realize a lightweight virtual desktop system. All of those schemes and systems are good practices of virtualization solution and they have become a strong foundation of our future work. Hai Jin 0001, Xiaofei Liao, Song Wu 0001, Zhiyuan Shao, Yingwei Luo |
HPCC | 3 |
| 2008 | A Data Storage Mechanism for P2P VoD Based on Multi-channel Overlay
Xiaofei Liao, Song Wu 0001, Hai Jin 0001 |
NPC | 3 |
| 2007 | Data Interoperation Between ChinaGrid and SRB
Muzhou Xiong, Hai Jin 0001, Song Wu 0001 |
ICA3PP | 3 |
| 2007 | Dependency-aware Maintenance for Dynamic Grid ServicesabstractAny mistaken maintenance for the complicated and distributed grid can bring unpredictable disaster. Here we focus on the system availability issues caused by service dependencies during the maintenance in grid. A novel mechanism, called Cobweb Guardian, is proposed in this paper. It provides multiple granularities (service-, container-, and node-level) maintenance for service components in grid. By using the Cobweb Guardian, grid administrators can execute the maintaining task safely in runtime with high availability. The evaluation results show that our proposed dependency-aware maintenance can make the grid management more automatic and available. Hai Jin 0001, Song Wu 0001, Yaqin Luo |
ICPP | 3 |
| 2007 | UCIPE: Ubiquitous Context-Based Image Processing Engine for Medical Image Grid
Aobing Sun, Hai Jin 0001, Ruhan He, Qin Zhang 0004, Song Wu 0001 |
UIC | 7 |
| 2006 | ServiceFlow: QoS Based Service Composition in CGSPabstractOpen, standard-based, loosely coupled Web services and WSRF services are dynamically discoverable and composable entities, which encapsulate the diverse implementations of physical and software resources in the interface that standardizes the business function. Therefore how to integrate these services to generate new value-added services is gaining attentions. In this paper, we present the architecture of ServiceFlow, a system for service composition in our grid platform, and discuss some related issues Haijun Cao, Hai Jin 0001, Song Wu 0001 |
EDOC | 3 |
| 2005 | 2-Layered Metadata Service Model in Grid Environment
Muzhou Xiong, Hai Jin 0001, Song Wu 0001 |
ICA3PP | 3 |
| 2003 | Symmetrical Declustering: A Load Balancing and Fault Tolerant Strategy for Clustered Video Servers
Song Wu 0001, Hai Jin 0001, Guang Tan |
ICCSA (1) | 1 |
| 2003 | Fault-Tolerant Grid Architecture and Practice
Hai Jin 0001, Deqing Zou, Hanhua Chen, Jianhua Sun 0002, Song Wu 0001 |
J. Comput. Sci. Technol. | 5 |
| 2002 | Study of Load Balancing Issues Based on Intra-Movie Skewness for Parallel Video ServersabstractIn order to avoid the load imbalance problem caused by video popularity, parallel video servers divide video objects into small segments and put them across multiple server nodes. However, due to users ’ various viewing time, access numbers of movie segments are quite different. Some segments are more popular than others. This instance is called intra-movie skewness, which may leads to load imbalance among server nodes in parallel video servers. In this paper we analyze and model the intra-movie skewness. Then, a novel data placement strategy, SPS (Symmetrical Pair Scheme), is proposed. It is proved that SPS prevents the impact of intra-movie skewness and has better load balancing performance than traditional round robin data placement, especially in a large-scale parallel video server. 1. Song Wu 0001, Hai Jin 0001 |
CCGRID | 1 |