VLDB 2026 Research / reviewers in the wild / expert
Yang Wang 0006
dblp:w/YangWang6
· DBLP profile ↗
120ranked-venue papers
25as first author
53since 2021 · last 2026
0000-0001-9438-6060ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 69 · 16 first-author · 27 since 2021Computer networks · 13 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Security and privacy · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Theory of computation · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Experiential Fairness: Bridging the Gap Between User Experience and Resource-Centric Fairness in Online LLM ServicesabstractConventional fairness in multi-tenant Large Language Model (LLM) inference services is typically defined by system-centric metrics such as equitable resource allocation. We argue that this is unilateral and it creates a gap between measured system performance and actual user-perceived quality. We challenge this notion by introducing and formalizing Experiential Fairness, a user-centric paradigm that shifts the objective from equality of opportunity (resource access) to equity of outcome (user experience). With this motivation we propose ExFairS, a lightweight scheduling framework that perceives each user's satisfaction as a composite measure of Service Level Objective (SLO) compliance and resource consumption, and dynamically re-orders the serving queue guided by a credit-based priority mechanism. Extensive experiments on an 8-GPU NVIDIA V100 node show that ExFairS reduces the SLO violation rate by up to 100% and improves system throughput by 14-21.9%, outperforming state-of-the-art schedulers and delivering a demonstrably higher degree of Experiential Fairness. Jiahua Huang, Wentai Wu, Yongheng Liu, Guozhi Liu, Yang Wang 0006, Weiwei Lin 0001 |
AAAI | 5 |
| 2026 | HoneyGPT: Breaking the trilemma in honeypots with large language models
Jianzhou You, Haining Wang 0001, Tianwei Yuan, Shichao Lv, Yang Wang 0006, Limin Sun 0001 |
Comput. Networks | 6 |
| 2026 | TriHID: Towards verifiable domain adaptation-based IoT intrusion detection in heterogeneous environment
Jiashu Wu, Yang Wang 0006 |
Expert Syst. Appl. | 2 |
| 2026 | Accelerated co-movement patterns mining: A heterogeneous framework based on GPU clusters
Chaowei Wu, Sasa Duan, Yang Wang 0006 |
Future Gener. Comput. Syst. | 4 |
| 2025 | Cross-Sectional Characteristic-driven Deep Reinforcement Learning
Huanghao Chen, Jerome Yen, Yang Wang 0006, Yain-Whar Si |
IEEE Big Data | 3 |
| 2025 | ModelMap: Kernel-Level Zero-Copy On-Demand Model Loading for Edge AIabstractWith the growing deployment of AI at the edge, running inference on resource-constrained devices increasingly requires efficient access to large models. However, modern models contain billions of parameters, far exceeding local memory, while TCP/IP-based edge networks make bulk transfers inefficient. Existing approaches, such as full downloads or network file systems, incur high latency, memory overhead, and lack fine-grained control. We present ModelMap, a zero-copy remote memory-mapped storage system that enables transparent, on-demand access to large AI models over TCP/IP. Instead of preloading entire models, ModelMap intercepts page faults in the Linux kernel and fetches only the needed parameter pages from a remote server, allowing applications to treat remote models as local. By operating entirely in kernel space, ModelMap reduces latency compared to user-space methods and avoids redundant copies. It further leverages kernel caching, prefetching, and sequential access optimizations to mitigate network bottlenecks. Crucially, ModelMap requires no changes to inference code, offering an efficient, scalable foundation for remote model access in edge environments. Fei Yi, Yang Wang 0006 |
CloudCom | 2 |
| 2025 | FlashFox: a secret-sharing approach to securing data deletion for Flash-based SSDabstractAbstract The ‘out-of-place’ update nature of Solid State Drives (SSDs) introduces a security risk. Scrubbing, a secure deletion method, mitigates this issue but negatively affects SSD endurance and requires redundant data for recovery due to page errors. Previously, the RAID-5 based scheme was used to manage redundant data. While effective for reading, it causes significant write latency due to the channel blocking issue. In response, we propose FlashFox, which integrates secret sharing and Reed-Solomon coding into SSDs, enabling the application of scrubbing within an encrypted storage environment. This innovation ensures secure deletion by cleaning up sensitive data, thereby reducing wear on SSD endurance. Furthermore, we have developed a RAID-4-based scheme for implementing FlashFox on SSDs. This scheme, by assigning specific channels to manage redundant data, successfully avoids the channel blocking issue prevalent in RAID-5. Experimental results show FlashFox reduces endurance wear by at least 15% compared to traditional scrubbing methods and writing response delay by at least 8$\times $ compared to RAID-5 based scheme. Wen Cheng 0003, Shengxia Tu, Yi Liu 0090, Lingfang Zeng, Yang Wang 0006, André Brinkmann |
Comput. J. | 5 |
| 2025 | DuoSQL: towards elastic data warehousing via separated data management and processingabstractAbstract Moving data warehouses (DWs) to the cloud is what today’s companies consider a trend towards cost-effective data management. To fully achieve the goal, the cloud DW system is supposed to adjust its resource provisioning to adapt to changing workload requirements. However, traditional data warehousing architecture lacks the flexibility for on-demand resource control, which severely restricts cost optimization and quality of service for both cloud providers and users. To build cloud DWs, new architectures are needed. This paper explores an architecture that decouples data management and processing to enable on-demand resource control. This optimized design enhances system elasticity and adaptability. However, this separation design is not without cost, as cooperation overhead can be high if not well optimized. For proof of concept, we build a prototype system, DuoSQL, using PostgreSQL for data management and Spark for data processing. To optimize cooperation, we conduct joint parameter tuning to improve overall system performance. We validate the system with the TPC-H benchmark. Results show the decoupling approach is flexible and offers significant performance potential. Weikang zhang, Tongxin Bai, Furong Zheng, Wenming Jin, Yang Wang 0006 |
Comput. J. | 6 |
| 2025 | ClusterHopper: Cross-region order dispatching optimization for ride-hailing drivers
Sasa Duan, Joseph Yen, Yang Wang 0006 |
Expert Syst. Appl. | 6 |
| 2025 | Efficient security interface for high-performance Ceph storage systemsabstractCeph portrays a resilient clustered storage solution with supporting object, block, and file storage capabilities with no single point of failure. Despite these qualifications, data confidentiality defines a concern in the system, as authentication and access control are the only data protection security services in Ceph. CephArmor was proposed as a third-party security interface to protect data confidentiality by adding an extra protection layer to data at rest. Despite the added layer, the initial design of the API needed to be more efficient in addressing security and performance simultaneously. In this study, we propose a new architectural design to address the associated issues with the preliminary prototype. Comprehensive performance and security analysis verify the improvement of the proposed method compared to the initial approach. The benchmark result has indicated a 37% improvement on average in IOPS, elapsed time, and bandwidth for the write benchmark compared to the initial model. Fatemeh Khoda Parast, Seyed Alireza Damghani, Brett Kelly, Yang Wang 0006, Kenneth B. Kent |
Future Gener. Comput. Syst. | 4 |
| 2025 | 9Ring: A 3D-Stacked Memory-Based Accelerator for Flexible and Efficient Deep CNN ApplicationsabstractThe massive computational and memory requirements of deep convolutional neural networks (DCNNs) have led to the development of neural network (NN) accelerators. However, as DCNN models grow in size, the demands on NN accelerators in terms of performance, memory bandwidth, and power efficiency continue to increase. We, therefore, present 9Ring , a flexible and efficient DCNN accelerator that takes full advantage of 3D-stacked memory, focusing on its hardware architecture, software scheduling, and optimization strategy. In particular, we first show that the mismatch between DCNN accelerators and DCNN models can lead to increased energy consumption and performance bottlenecks. We then present three flexible dataflow scheduling strategies to mitigate this mismatch. Afterward, we introduce an energy efficiency analysis tool that can automatically search for the optimal scheduling scheme with respect to different DCNN models for energy efficiency. Finally, we conduct an empirical study showing that 9Ring can reduce energy consumption by 31.4% and 43.9% on average, and improve performance by 12% and 10% on average, compared with Tetris and the NN accelerators on conventional low-power DRAM memory systems, respectively. Wen Cheng 0003, Qianya Cheng, Yi Liu 0090, Lingfang Zeng, André Brinkmann, Yang Wang 0006 |
ACM Trans. Archit. Code Optim. | 6 |
| 2025 | SCC: Synchronization Congestion Control for Multi-Tenant Learning Over Geo-Distributed CloudsabstractDistributed machine learning over geo-distributed clouds enables joint training of data located in different regions, alleviating the burden of transferring large volumes of training datasets, which greatly saves bandwidth. However, the limited capacity of WAN links slows down the inter-cloud communications, which significantly decelerates the synchronization of distributed machine learning over geo-distributed clouds. Besides, the multi-tenancy in clouds results in multiple training tasks running simultaneously, whose synchronizations consistently compete for the limited WAN bandwidth with each other, which further aggravates the training performance of each task. While existing works optimize synchronizations through techniques like gradient compression, multi-resource interleaving and so on, none of them targets at the synchronization congestion especially due to multi-tenant learning, which results in inferior training performance.To solve these problems, we propose a simple but effective scheme, SCC, for fast and efficient multi-tenant learning via synchronization congestion control. SCC monitors the cross-cloud network conditions and evaluates the synchronization congestion level based on the round-trip transmission time for each synchronization. Then SCC alleviates synchronization congestion via controlling the synchronization frequency according to the synchronization congestion level in a probabilistic way. Extensive experiments are conducted within our testbeds consisted of 16 NVIDIA V100 GPUs to evaluate the performance of SCC, and comparison results show that SCC can reduce the average training completion time and makespan by up to 28.6% and 43.2% over SAP-SGD [1]. Targeted experiments are conducted to demonstrate the effectiveness and robustness of SCC. Chengxi Gao, Fuliang Li, Kejiang Ye, Yang Wang 0006, Pengfei Wang 0013, Xingwei Wang 0001, Cheng-Zhong Xu 0001 |
IEEE Trans. Computers | 4 |
| 2025 | Load Balancing Scheduling for Batch-Ordered Job-Store: Online vs. OfflineabstractEfficient resource utilization is crucial in real-world applications, especially for balancing loads across machines handling specific job types. This paper introduces a novel batch-ordered job-store scheduling model, where jobs in a batch are scheduled sequentially, with their operations allocated in a round-robin fashion across two scenarios. We establish that this problem is NP-hard and analyze it in both online and offline settings. In the online case, we first examine the exclusive scenario, where operations within the same job must be scheduled on different machines, and show that a load greedy (LG) algorithm achieves a tight competitive ratio of$2 - \frac{1}{m}$, withmrepresenting the number of machines. Next, we consider the circular scenario, which requires maintaining the circular order of operations across ordered machines. In this context, we analyze potential anomalies in load distribution during local optimality achieved by the ordered load greedy (OLG) algorithm and provide bounds on the occurrence of these anomalies and the maximum load in each local scheduling round. In the offline case, we abstract each OLG scheduling process as a generalized circular sequence alignment (CSA) problem and develop a dynamic programming-based matching (DPM) algorithm to solve it. To further enhance load balancing, we develop a dynamic programming-based optimization (DPO) algorithm to schedule multiple jobs simultaneously in both scenarios. Experimental results confirm the efficiency of DPM for the CSA problem, and we validate the load balancing effectiveness of both online and offline algorithms using real traffic datasets. These theoretical findings and algorithmic implementations lay a solid groundwork for future practical advancements. Mengbing Zhou, Yang Wang 0006, Bocong Zhao, Cheng-Zhong Xu 0001 |
IEEE Trans. Computers | 2 |
| 2025 | Understanding Serverless Inference in Mobile-Edge Networks: A Benchmark ApproachabstractAlthough the emerging serverless paradigm has the potential to become a dominant way of deploying cloud-service tasks across millions of mobile and IoT devices, the overhead characteristics of executing these tasks on such a volume of mobile devices remain largely unclear. To address this issue, this paper conducts a deep analysis based on the OpenFaaS platform—a popular open-source serverless platform for mobile edge environments—to investigate the overhead of performing deep learning inference tasks on mobile devices. To thoroughly evaluate the inference overhead, we develop a performance benchmark, namedESBench, whereby a set of comprehensive experiments are conducted with respect to a bunch of simulated mobile devices associated with an edge cluster. Our investigation reveals that the performance of deep learning inference tasks is significantly influenced by the model size and resource contention in mobile devices, leading to up to$3\times$degradation in performance. Moreover, we observe that the network environment can negatively impact the performance of mobile inference, increasing the CPU overhead under poor network conditions. Based on our findings, we further propose some recommendations for designing efficient serverless platforms and resource management strategies as well as for deploying serverless computing in the mobile edge environment. Yanying Lin, Shuaipeng Wu, Kenneth B. Kent, Kejiang Ye, Yang Wang 0006 |
IEEE Trans. Cloud Comput. | 8 |
| 2025 | Towards Hybrid Architectures for Big Data Analytics: Insights From Spark-MPI IntegrationabstractHigh-Performance Data Analytics (HPDA) combines high-performance computing (HPC) with data analytics to uncover patterns and insights in dual-intensive applications that are both data-intensive and compute-intensive. Traditional big data frameworks and HPC technologies often struggle to address these demands independently, prompting researchers to explore their integration. Spark, known for its efficient in-memory computing with RDDs, and MPI, a foundational standard in HPC, are prominent candidates for such integration. This survey explores the integration of Spark and MPI for HPDA, highlighting their potential for unified data processing and computation. We first classify application workloads and review the characteristics and limitations of traditional frameworks. Then, we analyze the challenges and requirements of integrated architectures, focusing on the specific implementations of typical middleware-level architectures. Through comparative analysis, we highlight their advantages and limitations. Finally, we present application examples, outline key challenges and future research directions, and briefly discuss progress in integration approaches for other technology combinations. Mengbing Zhou, Qiuyan Li, Mingyuan Cai, Cheng-Zhong Xu 0001, Yang Wang 0006 |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | TollHelper: A Safe and Efficient Traffic Control Approach on Toll Plaza via Constrained Load BalancingabstractTraffic congestion at toll plazas is a critical issue in urban infrastructure, which is often exacerbated by surges in vehicle volume during peak hours. The congestion typically arises from imbalances in traffic demand and toll booth efficiency, often resulting in safety hazards and delays. Existing solutions, while addressing efficiency or safety aspects, often lack a comprehensive approach for efficient traffic management at toll plazas. To address this challenge, in this paper, we propose TollHelper, a framework designed to optimize vehicle scheduling and load balancing at toll plazas as well as improve safety. Our approach treats concurrently arriving vehicles as a single scheduling batch, guiding different vehicles from the same batch to different toll booths to enhance safety and reduce congestion. We address this scheduling constraint in both general and heterogeneous toll booth scenarios, introducing effective load balancing algorithms to minimize toll booth service loads and optimize user driving experiences. Based on empirical studies, we demonstrate that our methods achieve improvements in standard deviation compared to the baselines, ranging from 51.6% to 97.0% improvement in terms of load balancing effects. Mengbing Zhou, Bocong Zhao, Minxian Xu, Yang Wang 0006 |
ISPA | 5 |
| 2024 | Market Sentiment Analysis Based on Image Processing With Put-Call Volatility Gap SurfaceabstractAnalyzing the market sentiment and forecasting movements in asset prices is extremely important and has been attempted by researchers and market practitioners. Asset volatilities, regardless of historical ones or implied from the option prices, are crucial barometers of the market. And this research proposes a new approach that combines image processing and machine learning to capture the relationships between sentiment-related features and asset movement. The proposed research is based on the tick-level SPY options transactions, and the dataset contains around 1.5 million trading records. Specially, we obtained the gap between the call surface and the put surface as the second-level implied volatility surface (IVS). After adopting the traditional convolutional neural network (CNN) to compare the predictive effects among implied volatility (IV) call surface, IV put surface, and IV gap surface, the results indicate that the IV gap provides the most significant predictability. Besides, our project creatively proposes the interframe difference approach to optimized CNN and a recurrent CNN (RCNN) to fully utilize spatial and temporal features of the IV surface data. According to experiment results, the directional accuracy of prediction ranges from 67.20% to 72.05% for asset movement forecasting at the millisecond level. Guoxiang Guo, Yang Wang 0006, Jerome Yen |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | MetroBUX: A Topology-Based Visual Analytics for Bus Operational Uncertainty EXplorationabstractIn the public transportation system, punctuality benefits both bus operation and passengers’ travel experience. However, uncertainty exists due to complex traffic conditions and heterogeneous driving behaviors. To analyze bus operational uncertainty, transport planners and bus operators need a tool that supports multi-granular modeling, spatio-temporal representation, and interactive exploration. To meet the requirement, we present MetroBUX, a visual analytics system for$B$us operational$U$ncertainty e$X$ploration. MetroBUX aligns daily bus trips and models stop-level uncertainty of bus arrival time. It has a consolidated interface with three main views: Map View for presenting the spatial distribution of uncertainty, Temporal View for tracking the evolution of uncertainty, and Trip View for inspecting uncertainty propagation. Specifically, MetroBUX enables integrated spatio-temporal analysis by connecting topological uncertainty distribution at different periods in a nested tracking graph. Furthermore, it supports interactive and hierarchical exploration, including region-, route-, trip-, and stop-level analysis. Case studies on real-world bus operational data and domain experts’ feedback demonstrate the efficiency of MetroBUX. Shishi Xiao, Lingdan Shao, Bo Du 0004, Yang Wang 0006, Qiaomu Shen, Wei Zeng 0004 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Open Set Dandelion Network for IoT Intrusion DetectionabstractAs Internet of Things devices become widely used in the real-world, it is crucial to protect them from malicious intrusions. However, the data scarcity of IoT limits the applicability of traditional intrusion detection methods, which are highly data-dependent. To address this, in this article, we propose the Open-Set Dandelion Network (OSDN) based on unsupervised heterogeneous domain adaptation in an open-set manner. The OSDN model performs intrusion knowledge transfer from the knowledge-rich source network intrusion domain to facilitate more accurate intrusion detection for the data-scarce target IoT intrusion domain. Under the open-set setting, it can also detect newly-emerged target domain intrusions that are not observed in the source domain. To achieve this, the OSDN model forms the source domain into a dandelion-like feature space in which each intrusion category is compactly grouped and different intrusion categories are separated, i.e., simultaneously emphasising inter-category separability and intra-category compactness. The dandelion-based target membership mechanism then forms the target dandelion. Then, the dandelion angular separation mechanism achieves better inter-category separability, and the dandelion embedding alignment mechanism further aligns both dandelions in a finer manner. To promote intra-category compactness, the discriminating sampled dandelion mechanism is used. Assisted by the intrusion classifier trained using both known and generated unknown intrusion knowledge, a semantic dandelion correction mechanism emphasises easily-confused categories and guides better inter-category separability. Holistically, these mechanisms form the OSDN model that effectively performs intrusion knowledge transfer to benefit IoT intrusion detection. Comprehensive experiments on several intrusion datasets verify the effectiveness of the OSDN model, outperforming three state-of-the-art baseline methods by 16.9%. The contribution of each OSDN constituting component, the stability and the efficiency of the OSDN model are also verified. Jiashu Wu, Kenneth B. Kent, Jerome Yen, Cheng-Zhong Xu 0001, Yang Wang 0006 |
ACM Trans. Internet Techn. | 6 |
| 2023 | Neighborhood-Oriented Decentralized Learning Communication in Multi-Agent System
Jiashu Wu, André Brinkmann, Yang Wang 0006 |
ICANN (3) | 4 |
| 2023 | The Potential of RISC-V Platform in Financial Computing on Option Pricing and Energy EfficiencyabstractThe fifth version of the Reduced Instruction Set Computer (RISC-V) is a popular instruction set architecture (ISA) featured for low energy consumption. Currently, a growing number of industrial applications are based on RISC-V platforms, especially Internet of Things (IoT) devices. Those applications pursue low power and just sufficient computing capacity. However, low power consumption shall not be directly regarded as weak in computation. Recent advancement in RISC- V shows the potential for building a computing platform capable of handling tasks requiring considerable computing power. Traditional financial computing platforms are generally based on Complex Instruction Set Computers (CISC), like x86 platforms. As green computing is sweeping, it is meaningful to handle financial computing tasks with less energy consumption. To explore the potential of RISC- V in financial computing, we set up a typical financial computing task - American option implied volatility calculation, and examine the performance and power consumption of x86 and RISC-V platforms. The result shows that the RISC-V CPU is sufficient for some financial computing scenarios considering actual requirements. A heterogeneous computing system composed of x86 and RISC-V platforms could significantly improve energy efficiency. Guoxiang Guo, Minhao Zhu, Yang Wang 0006, Jerome Yen |
SMC | 4 |
| 2023 | Online data caching in edge computingabstractSummary Data caching is an effective method to reduce traffic and improve the quality of service in network. Traditionally, users' requests are offloaded to the cloud for centralized computing. However, due to security and privacy, these tasks are executed in the nearest server, so that the data and service needed by the task are also essential. After the task is completed, in case the next arriving request needs the same data, resulting in transmission cost, the data need to be stored for a period of time, because we know nothing about the coming request information under an online request stream. In this article, we study data caching problem by extending single data item to multiple data items among servers. About the homogeneous model and the submodular model with constraint, we propose a data caching strategy minimizing the total transfer and caching costs of the system. Moreover, we also solve the semiheterogeneous model by the anticipatory caching (AC) algorithm in Reference 21. Meanwhile we find it is more efficient for our three models in this article to improve the performance. Xinxin Han, Guichen Gao, Yang Wang 0006, Hing-Fung Ting, Ilsun You, Yong Zhang 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | How does solid-state drives cluster perform for distributed file systems: An empirical studyabstractSummary As the capacity of Solid‐State Drives (SSDs) is constantly being optimised and boosted with gradually reduced cost, the SSD cluster is now widely deployed as part of the hybrid storage system in various scenarios such as cloud computing and big data processing. However, despite its rapid developments, the performance of the SSD cluster remains largely under‐investigated, leaving its sub‐optimal applications in reality. To address this issue, in this paper we conduct extensive empirical studies for a comprehensive understanding of the SSD cluster in diverse settings. To this end, we configure a real SSD cluster and gather the generated trace data based on some often‐used benchmarks, then adopt analytical methods to analyse the performance of the SSD cluster with different configurations. In particular, regression models are built to provide better performance predictability under broader configurations, and the correlations between influential factors and performance metrics with respect to different numbers of nodes are investigated, which reveal the high scalability of the SSD cluster. Additionally, the cluster's network bandwidth is inspected to explain the performance bottleneck. Finally, the knowledge gained is summarised to benefit the SSD cluster deployment in practice. Jiashu Wu, Yang Wang 0006, Hekang Wang, Taorui Lin |
Concurr. Comput. Pract. Exp. | 2 |
| 2023 | Heterogeneous Domain Adaptation for IoT Intrusion Detection: A Geometric Graph Alignment ApproachabstractData scarcity hinders the usability of data-dependent algorithms when tackling IoT intrusion detection (IID). To address this, we utilize the data-rich network intrusion detection (NID) domain to facilitate more accurate intrusion detection for IID domains. In this article, a geometric graph alignment (GGA) approach is leveraged to mask the geometric heterogeneities between domains for better intrusion knowledge transfer. Specifically, each intrusion domain is formulated as a graph where vertices and edges represent intrusion categories and category-wise inter-relationships, respectively. The overall shape is preserved via a confused discriminator incapable to identify adjacency matrices between different intrusion domain graphs. A rotation avoidance mechanism and a center point matching mechanism are used to avoid graph misalignment due to rotation and symmetry, respectively. Besides, category-wise semantic knowledge is transferred to act as vertex-level alignment. To exploit the target data, a pseudo-label (PL) election mechanism that jointly considers network prediction, geometric property, and neighborhood information is used to produce fine-grained PL assignment. Upon aligning the intrusion graphs geometrically from different granularities, the transferred intrusion knowledge can boost IID performance. Comprehensive experiments on several intrusion data sets demonstrate state-of-the-art performance of the GGA approach and validate the usefulness of GGA-constituting components. Jiashu Wu, Yang Wang 0006, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Internet Things J. | 3 |
| 2023 | Adaptive Bi-Recommendation and Self-Improving Network for Heterogeneous Domain Adaptation-Assisted IoT Intrusion DetectionabstractAs Internet of Things (IoT) devices become prevalent, using intrusion detection to protect IoT from malicious intrusions is of vital importance. However, the data scarcity of IoT hinders the effectiveness of traditional intrusion detection methods. To tackle this issue, in this article, we propose the adaptive bi-recommendation and self-improving network (ABRSI) based on unsupervised heterogeneous domain adaptation (HDA). The ABRSI transfers enrich intrusion knowledge from a data-rich network intrusion source domain to facilitate effective intrusion detection for data-scarce IoT target domains. The ABRSI achieves fine-grained intrusion knowledge transfer via adaptive bi-recommendation matching. Matching the bi-recommendation interests of two recommender systems (RSs) and the alignment of intrusion categories in the shared feature space form a mutual-benefit loop. Besides, the ABRSI uses a self-improving mechanism, autonomously improving the intrusion knowledge transfer from four ways. A hard pseudo label (PL) voting mechanism jointly considers RS decision and label relationship information to promote more accurate hard PL assignment. To promote diversity and target data participation during intrusion knowledge transfer, target instances failing to be assigned with a hard PL will be assigned with a probabilistic soft PL, forming a hybrid pseudo-labeling strategy. Meanwhile, the ABRSI also makes soft pseudo-labels globally diverse and individually certain. Finally, an error knowledge learning mechanism is utilized to adversarially exploit factors that causes detection ambiguity and learns through both current and previous error knowledge, preventing error knowledge forgetfulness. Holistically, these mechanisms form the ABRSI model that boosts IoT intrusion detection accuracy via HDA-assisted intrusion knowledge transfer. Comprehensive experiments on several intrusion data sets demonstrate the state-of-the-art performance of the ABRSI method, outperforming its counterparts by 9.2%, and also verify the effectiveness of ABRSI constituting components and ABRSI’s overall efficiency. Jiashu Wu, Yang Wang 0006, Cheng-Zhong Xu 0001, Kenneth B. Kent |
IEEE Internet Things J. | 2 |
| 2023 | Joint Semantic Transfer Network for IoT Intrusion DetectionabstractIn this article, we propose a joint semantic transfer network (JSTN) toward effective intrusion detection (ID) for large-scale scarcely labeled Internet of Things (IoT) domain. As a multisource heterogeneous domain adaptation (MS-HDA) method, the JSTN integrates a knowledge-rich network intrusion (NI) domain and another small-scale IoT intrusion (II) domain as source domains and preserves intrinsic semantic properties to assist target II domain ID. The JSTN jointly transfers the following three semantics to learn a domain-invariant and discriminative feature representation. The scenario semantic endows source NI and II domains with characteristics from each other to ease the knowledge transfer process via a confused domain discriminator and categorical distribution knowledge preservation. It also reduces the source–target discrepancy to make the shared feature space domain invariant. Meanwhile, the weighted implicit semantic transfer boosts discriminability via a fine-grained knowledge preservation, which transfers the source categorical distribution to the target domain. The source–target divergence guides the importance weighting during knowledge preservation to reflect the degree of knowledge learning. Additionally, the hierarchical explicit semantic alignment performs centroid-level and representative-level alignment with the help of a geometric similarity-aware pseudo-label refiner, which exploits the value of the unlabeled target II domain and explicitly aligns feature representations from a global and local perspective in a concentrated manner. Comprehensive experiments on various tasks verify the superiority of the JSTN against state-of-the-art comparing methods, on average a 10.3% of accuracy boost is achieved. The statistical soundness of each constituting component and the computational efficiency is also verified. Jiashu Wu, Yang Wang 0006, Binhui Xie, Shuang Li 0008, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Internet Things J. | 2 |
| 2023 | A comprehensive survey of cryptography key management systems
Subhabrata Rana, Fatemeh Khoda Parast, Brett Kelly, Yang Wang 0006, Kenneth B. Kent |
J. Inf. Secur. Appl. | 4 |
| 2023 | PackCache: An Online Cost-Driven Data Caching Algorithm in the CloudabstractIn this paper, we study a data caching problem in the cloud environment, where multiple frequently co-utilised data items could be packed as a single item being transferred to serve a sequence of data requests dynamically with reduced cost. To this end, we propose an online algorithm with respect to a homogeneous cost model, calledPackCache, that can leverage the FP-Tree technique to mine those frequently co-utilised data items for packing whereby the incoming requests could be cost-effectively served online by exploiting the concept of anticipatory caching. We show the algorithm is$2/\alpha$competitive, reaching the lower bound of the competitive ratio for any deterministic online algorithm on the studied caching problem, and also time and space efficient to serve the requests. Finally, we evaluate the performance of the algorithm via experimental studies to show its actual cost-effectiveness and scalability. Jiashu Wu, Yang Wang 0006, Yong Zhang 0001, Cheng-Zhong Xu 0001 |
IEEE Trans. Computers | 3 |
| 2023 | Tianji: Securing a Practical Asynchronous Multi-User ORAMabstractOblivious Random Access Machines (ORAMs) allow cloud users to access remote data without leaking access patterns. Current ORAM solutions achieve this goal at expense of either increasing bandwidth consumption by a factor of$O(\log N)$, where$N$is the number of data blocks, or relying on homomorphic encryption for bandwidth amplification reduction to$O(1)$. Furthermore, most ORAMs are only effective for a single user, while the solutions for multi-user scenarios often induce security or performance problems. This article introducesTianji— an asynchronous multi-user Shamir-based ORAM system — which supports asynchronous network access scenarios for multiple users with improved security and performance.Tianjiis implemented on top ofS$^{3}$3ORAM$^+$+—an extension of the state-of-the-art Shamir-based S$^{3}$ORAM with a new non-eviction data write-back scheme to achieve$O(1)$consumption in both bandwidth amplification and storage capacity. Our experimental results show that the proposedTianjiwithS$^{3}$3ORAM$^+$+can significantly outperform the state-of-the-art multi-userTaoStorein terms of access latency and client scalability. Additionally, its average response time is relatively stable when client loads increase. Wen Cheng 0003, Dazou Sang, Lingfang Zeng, Yang Wang 0006, André Brinkmann |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2023 | Cost-Driven Data Caching in Edge-Based Content Delivery NetworksabstractIn this paper, we studied a data caching problem in edge-based CDNs to facilitate the content delivery to serve a sequence of requests, off-line and online, with minimum costs as a goal based on a semi-homo cost model. To this end, we first designed an O(mn \log(mn)) time and space optimal proactive off-line algorithm,called pro-caching, by reducing the problem to a simple shortest path problem in a directed weighted network graph, and then extended the idea of anticipatory caching to develop an 2-competitive reactive online algorithm, called re-caching, for this problem and showed its tightness by proving that no deterministic online algorithm can do better than 2-o(1) in its worst case. Finally, to combine the advantages of both algorithms, we also presented a hybrid algorithm, called hy-caching, to fully utilize the power and benefits of edge-based CDNs while reducing their service costs. Our results improve the previous results not only in the cost model being used but also in the time complexity, competitive ratio, and the quality of the solutions. We provably achieve these results with our deep insights into the problem and the careful analysis, together with an empirical evaluation. Yang Wang 0006, Xinxin Han, Pengfei Wang 0013, Yong Zhang 0001, Cheng-Zhong Xu 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Cost-Efficient Sharing Algorithms for DNN Model Serving in Mobile Edge NetworksabstractWith the fast growth of mobile edge computing (MEC), the deep neural network (DNN) has gained more opportunities in application to various mobile services. Given the tremendous number of learning parameters and large model size, the DNN model is often trained in cloud center and then dispatched to end devices for inference via edge network. Therefore, maximizing the cost-efficiency of learned model dispatch in the edge network would be a critical problem for the model serving in various application contexts. To reach this goal, in this article we focus mainly on reducing the total model dispatch cost in the edge network while maintaining the efficiency of the model inference. We first study this problem in its off-line form as a baseline where a sequence of$n$requests can be pre-defined in advance and exploit dynamic programming techniques to obtain a fast optimal algorithm in time complexity of$O(m^{2}n)$under a semi-homogeneous cost model in a$m$-sized network. Then, we design and implement a 2.5-competitive algorithm for its online case with a provable lower bound of 2 for any deterministic online algorithm. We verify our results through careful algorithmic analysis and validate their actual performance via a trace-based study based on a public open international mobile network dataset. Jiashu Wu, Yang Wang 0006, Jerome Yen, Yong Zhang 0001, Cheng-Zhong Xu 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Serverless Computing: State-of-the-Art, Challenges and OpportunitiesabstractServerless computing is growing in popularity by virtue of its lightweight and simplicity of management. It achieves these merits by reducing the granularity of the computing unit to the function level. Specifically, serverless allows users to focus squarely on the function itself while leaving other cumbersome management and scheduling issues to the platform provider, who is responsible for striking a balance between high-performance scheduling and low resource cost. In this article, we conduct a comprehensive survey of serverless computing with a particular focus on its infrastructure characteristics. Whereby some existing challenges are identified, and the associated cutting-edge solutions are analyzed. With these results, we further investigate some typical open-source frameworks and study how they address the identified challenges. Given the great advantages of serverless computing, it is expected that its deployment would dominate future cloud platforms. As such, we also envision some promising research opportunities that need to be further explored in the future. We hope that our work in this article can inspire those researchers and practitioners who are engaged in related fields to appreciate serverless computing, thereby setting foot in this promising area and making great contributions to its development. Yongkang Li 0003, Yanying Lin, Yang Wang 0006, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | Ordis: A Dynamic Order-Dispatch Algorithm for Ridehailing and Ridesharing in a Large Region
Juanjuan Zhao 0001, Yang Wang 0006, Cheng-Zhong Xu 0001 |
ICA3PP | 3 |
| 2022 | DPLFS: A Dual-Mode PCM-based Log-Structured File SystemabstractDual-mode phase change memory (PCM) allows each PCM cell to operate concurrently in two modes: multi-level cell (MLC) and single-level cell (SLC), each with its own distinct features in transfer speeds and power consumption. This paper describes the design and implementation of a log-structured file system, called DPLFS, which is proposed to provide a storage system with high throughput and low latencies. DPLFS achieves this merit by exploiting the dual-mode PCM. In particular, file data in DPLFS is stored in MLC, while its metadata is maintained in SLC, and will be written back to MLC when the space in SLC is not enough. DPLFS adapts log-structured file system techniques to exploit fast random accesses in SLC and large capacity in MLC and maintains a file directory in SLC to speed up the data retrieval speed. Our experimental results show a 21% in write-intensive workloads and a 11% in read-intensive workloads, compared to the state-of-the-art file systems. Wen Cheng 0003, Telong Zheng, Lingfang Zeng, Yang Wang 0006, André Brinkmann |
ICCD | 4 |
| 2022 | Image Processing Based Implied Volatility Surface Analysis for Asset movement ForecastingabstractNowadays, people are showing growing attention to the market movements. With more demand for market sentiment analysis and risk management, advanced investment tools are needed to assist the high frequency trading activities. Machine learning as a fast-growing tool provides people a new perspective to handle complex problems. Although financial data contains various information and is usually regarded as hard to concentrate into one unified dimension, our research aims to fuse the image processing method with the high frequency implied-volatility-based market sentiment analysis. In this way, our research implemented the real-time processing of the market data and proposes an innovative idea, applying the machine learning method to regress the market price using the two-dimensional discrete financial data, which is traditionally viewed as images. The proposed method shows satisfying performance in testing with tick-level S&P500 option dataset containing around 1.5 million trading record. To go further with the improvement of the economic image classification and represent the momentum factors of the implied volatility surface images, we also introduce the speed and acceleration of sequence images. Overall, we have reached 61.23% accuracy for implied volatility image classification, and 63.22% & 65.52% accuracy for financial image considering velocity and acceleration. Guoxiang Guo, Yang Wang 0006, Jerome Yen |
INDIN | 3 |
| 2022 | MetaWBC: POSIX-Compliant Metadata Write-Back Caching for Distributed File SystemsabstractIn parallel and distributed file systems, caching can improve data performance and metadata operations. Currently, most distributed file systems adopt a write-back data cache for performance and a write-through metadata cache for simplifying consistency. However, with modern file systems scales and workloads, write-through metadata caching can impact overall file system performance, e.g., through lock contention and heavy RPC loads required for namespace synchronization and transaction serialization. This paper proposes a novel metadata write-back caching (MetaWBC) mechanism to improve the performance of metadata operations in distributed environments. To achieve extreme metadata performance, we developed a fast, lightweight, and POSIXcompatible memory file system as a metadata cache. Further, we designed a file caching state machine and included other performance optimizations. We coupled MetaWbc with Lustre and evaluated that MetaWbc can outperform the native parallel file system by up to 8x for metadata-intensive benchmarks, and up to 7x for realistic workloads in throughput. Yingjin Qian, Wen Cheng 0003, Lingfang Zeng, Marc-Andre Vef, Oleg Drokin, Andreas Dilger, Shuichi Ihara, Wusheng Zhang, Yang Wang 0006, André Brinkmann |
SC | 9 |
| 2022 | PECCO: A profit and cost-oriented computation offloading scheme in edge-cloud environment with improved Moth-flame optimizationabstractSummary With the fast growing quantity of data generated by smart devices and the exponential surge of processing demand in the Internet of Things (IoT) era, the resource‐rich cloud centers have been utilized to tackle these challenges. To relieve the burden on cloud centers, edge‐cloud computation offloading becomes a promising solution since shortening the proximity between the data source and the computation by offloading computation tasks from the cloud to edge devices can improve performance and quality of service. Several optimization models of edge‐cloud computation offloading have been proposed that take computation costs and heterogeneous communication costs into account. However, several important factors are not jointly considered, such as heterogeneities of tasks, load balancing among nodes and the profit yielded by computation tasks, which lead to the profit and cost‐oriented computation offloading optimization modelPECCOproposed in this article. Considering that the model is hard in nature and the optimization objective is not differentiable, we propose an improved Moth‐flame optimizerPECCO‐MFIwhich addresses some deficiencies of the original Moth‐flame optimizer and integrate it under the edge‐cloud environment. Comprehensive experiments are conducted to verify the superior performance of the proposed method when optimizing the proposed task offloading model under the edge‐cloud environment. Jiashu Wu, Yang Wang 0006, Shigen Shen, Cheng-Zhong Xu 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | Toward fast theta-join: A prefiltering and amalgamated partitioning approachabstractAbstract As one of the most useful online processing techniques, the theta‐join operation has been utilized by many applications to fully excavate the relationships between data streams in various scenarios. As such, constant research efforts have been put to optimize its performance in the distributed environment, which is typically characterized by reducing the number of Cartesian products as much as possible. In this article, we design and implement a novel fast theta‐join algorithm, calledPrefap, by developing two distinct techniques—prefilteringandamalgamated partitioning—based on the state‐of‐the‐art FastThetaJoin algorithm to optimize the efficiency of the theta‐join operation. Firstly, we develop a prefiltering strategy before data streams are partitioned to reduce the amount of data to be involved and benefit a more fine‐grained partitioning. Secondly, to avoid the data streams being partitioned in a coarse‐grained isolated manner and improve the quality of the partition‐level filtering, we introduce an amalgamated partitioning mechanism that can amalgamate the partitioning boundaries of two data streams to assist a fine‐grained partitioning. With the integration of these two techniques into the existing FastThetaJoin algorithm, we design and implement a new framework to achieve a decreased number of Cartesian products and a higher theta‐join efficiency. By comparing with existing algorithms, FastThetaJoin in particular, we evaluate the performance ofPrefapon both synthetic and real data streams from two‐way to multiway theta‐join to demonstrate its superiority. Jiashu Wu, Yang Wang 0006, Xiaopeng Fan 0002, Kejiang Ye, Cheng-Zhong Xu 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2022 | The strong substructure and feature attention mechanism for image semantic segmentationabstractAbstract Semantic segmentation is a hot topic in computer vision and various deep learning networks are designed to achieve higher accuracy on that by fully exploring the capability of neural networks. This paper aims to address the issue and proposes the substructures with novelty for popular networks. Meanwhile, we present a cross‐channel structure, which simultaneously reduces parameter while the kernel size becomes larger. After that, to overcome the weakness of insufficient dataset which refers to satellite image data, we propose a feature attention mechanism with generative adversarial network to enhance the images' features. We show the recognition result on the satellite image dataset with a large picture. This paper evaluates substructures on the PASCAL VOC2012 dataset and improves the mIOU from 74.68% to 88.15%. Yuhang Zhang 0010, Hongshuai Ren, Wensi Yang, Yang Wang 0006, Kejiang Ye, Cheng-Zhong Xu 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2022 | Lifespan-based garbage collection to improve SSD's reliability and performance
Wen Cheng 0003, Mi Luo, Lingfang Zeng, Yang Wang 0006, André Brinkmann |
J. Parallel Distributed Comput. | 4 |
| 2022 | Towards scalable and efficient Deep-RL in edge computing: A game-based partition approach
Jiashu Wu, Yang Wang 0006, Cheng-Zhong Xu 0001 |
J. Parallel Distributed Comput. | 3 |
| 2022 | Deadlock Avoidance Algorithms for Recursion-Tree Modeled Requests in Parallel ExecutionsabstractWe present an extension of the bankers algorithm to resolve deadlock for programs whose resource-request graph can be modeled as a recursion tree for parallel execution. Our algorithm implements the bankers logic, with the key difference being that some properties of the tree are fully exploited to improve the resource utilization and safety check in deadlock avoidance. For an n-node tree modeled program making requests to m types of resources, our recursion-tree based algorithm can obtain a time complexity of O(mn loglogn) on average in safety check while reducing the conservativeness in resource utilization. We reap these benefits by proposing a concept of the resource critical tree and leverage it to localize the maximum claim associated with each node in the tree. To tackle the case when the tree model is not statically known, we relax the definition of a local maximum claim by sacrificing some resource utilization. With this trade-off, the algorithm can resolve the deadlock and achieve more efficient safety checks within time of O(m loglogn). Our empirical studies on a two-dimensional integration problem on sparse grids show that the proposed algorithms can reduce resource utilization conservativeness and improve avoidance performance by minimizing the number of safety checks. Yang Wang 0006, Kenneth B. Kent, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Computers | 1 |
| 2022 | Multi-Tier Workload Consolidations in the Cloud: Profiling, Modeling and OptimizationabstractReducing tail latency becomes increasingly important to improve the user-perceived service experience. User-facing latency-sensitive cloud applications typically contain multiple interactive tiers (e.g., Web, App, Database) running in different virtual machines (VMs) with complex interaction patterns. However, such interactions between VMs in different tiers are often neglected in previous VM consolidation methods, resulting in poor application performance. In this article, we study the consolidation of multi-tier interactive workloads from a new perspective of user-perceived tail latency. We propose a novel profiling-based consolidation methodology to satisfy tail latency requirements while reducing the number of used physical machines. To achieve such a goal, we first perform large-scale profiling experiments under various consolidation settings in a KVM virtualized private cluster to establish the empirical performance values. We consider two key factors that affect the tail latency of multi-tier workloads:interferencewith co-located VMs andinteractionbetween tiers. We model the consolidation of multi-tier workloads as an optimization problem with different objectives and constraints, and derive the consolidation schedule. We implement and evaluate the proposed models, as well as comparing with other methods (i.e.,withoutprofiling orwithoutconsidering interaction influence). Extensive experimental results show that the proposed method is able to reduce up to5Xtail latency, compared with the methodwithoutprofiling and up to1.3Xtail latency, compared with the methodwithoutconsidering the interaction influence between different tiers. Kejiang Ye, Haiying Shen, Yang Wang 0006, Cheng-Zhong Xu 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2022 | The State of the Art of Metadata Managements in Large-Scale Distributed File Systems - Scalability, Performance and AvailabilityabstractFile system metadata is the data in charge of maintaining namespace, permission semantics and location of file data blocks. Operations on the metadata can account for up to 80% of total file system operations. As such, the performance of metadata services significantly impacts the overall performance of file systems. A large-scale distributed file system (DFS) is a storage system that is composed of multiple storage devices spreading across different sites to accommodate data files, and in most cases, to provide users with location independent access interfaces. Large-scale DFSs have been widely deployed as a substrate to a plethora of computing systems, and thus their metadata management efficiency is crucial to a massive number of applications, especially with the advent of the Big Data age, which poses tremendous pressure on underlying storage systems. This paper reports the state-of-the-art research on metadata services in large-scale distributed file systems, which is conducted from three indicative perspectives that are always used to characterize DFSs: high-scalability, high-performance, and high-availability, with special focus on their respective major challenges as well as their developed mainstream technologies. Additionally, the paper also identifies and analyzes several existing problems in the research, which could be used as a reference for related studies. Yang Wang 0006, Kenneth B. Kent, Lingfang Zeng, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | An Online Algorithm for Data Caching Problem in Edge Computing
Xinxin Han, Guichen Gao, Yang Wang 0006, Yong Zhang 0001 |
AAIM | 3 |
| 2021 | xBCBench: A Benchmarking Tool for Analyzing the Performance of Blockchain Systems
Kejiang Ye, Yang Wang 0006, Cheng-Zhong Xu 0001 |
BlockSys | 3 |
| 2021 | RPTCN: Resource Prediction for High-dynamic Workloads in Clouds based on Deep LearningabstractResource management is challenging in clouds due to the dynamics and sharing characteristics. The crucial problem is how to allocate resources accurately and satisfy demands of workloads timely. The traditional solution is to use historical data to predict future resource usage. Although these resource prediction methods can predict the periodicity, they can not accurately predict mutation points due to the high dynamics and uncertainty of resource usage. To tackle this issue, in this paper we propose a resource usage prediction method - RPTCN, which is based on a deep learning method - temporal convolutional networks (TCNs) in cloud systems. We add a fully connected layer and attention mechanism to TCNs to improve the prediction accuracy. In order to explore the relationship between the usage of different resources in the temporal dimension, we use correlation analysis to screen performance indicators as multidimensional feature input for prediction. Finally, we evaluate the performance of this method on Alibaba trace v2018. Evaluations show that RPTCN improves the overall MAE and MSE by 6.50%~89.03% and 0.41%~68.82% respectively compared to baselines in dynamic and long-term prediction of resource usage. Moreover, the convergence and generalization of RPTCN are also better than the baselines. Wenyan Chen 0001, Chengzhi Lu, Kejiang Ye, Yang Wang 0006, Cheng-Zhong Xu 0001 |
CLUSTER | 4 |
| 2021 | A Strategy-based Optimization Algorithm to Design Codes for DNA Data Storage System
Abdur Rasool, Qiang Qu 0001, Qingshan Jiang, Yang Wang 0006 |
ICA3PP (2) | 4 |
| 2021 | Cost-Driven Data Caching in the Cloud: An Algorithmic ApproachabstractData caching in the cloud is an efficient way to improve the QoS of diverse data applications. However, this benefit is not freely available, given monetary cost to manage the caches in the cloud. In this paper, we study the data caching problem in the cloud that is driven by the monetary cost reduction, instead of the hit rate under limited capacity as in traditional cases. In particular, given a stream of requestsRto a shared data item, we present a shortest-path based optimal algorithm that can minimize the total transfer and caching costs within O(mn) time for off-line case, here m represents the number of nodes in the network, while n is the length of the request stream. The cost model in this computation is semi-homo, which indicates that all pairs of nodes have the same transfer cost, but each cache server node has its own caching cost rate. Our off-line algorithm improves the previous results not only in reducing the time complexity from O(m2n) to O(mn), but also in relaxing the cost model to be semi-homogeneous, rendering the algorithm more practical in reality. Furthermore, we also study this problem in its online form, and by extending the anticipatory caching idea, we propose a 2-competitive online algorithm based on the same cost model and show its tightness by giving a lower bound of the competitive ratio as 2 - o(1) for any deterministic online algorithm. We provably achieve these results with our deep insights into the problem and careful analysis of the solution algorithms, together with a trace-based study to evaluate their performance in reality. Yang Wang 0006, Yong Zhang 0001, Xinxin Han, Pengfei Wang 0013, Cheng-Zhong Xu 0001, Joseph Horton, Joseph C. Culberson |
INFOCOM | 1 |
| 2021 | Hercules: Intelligent Coupling of Dual-Mode Flash Memory and Hard Disk DriveabstractAbstract The write performance of multi-level cell (MLC) is several times slower than single-level cell (SLC); however, the cost per bit of MLC is much lower than SLC. Dual-mode flash (the medium can be partially switched to SLC mode by programming only 1 bit in some cells) can combine SLC and MLC to provide trading density opportunity for performance. In this paper, we present Hercules—a hybrid storage system that couples dual-mode flash memory and hard drive disk (HDD)—based on the content locality principle for high storage performance. The data are divided into two types: the reference data for read operation and the delta data for write operation. The reference data are stored in SLC and the delta data in MLC or HDD in sequential orders. Hercules organizes the metadata for the mapping of the physical locations of the reference blocks and the delta data of the original blocks, intelligently identifies hot/cold data and performs the data migration between MLC and disk for performance improvements. To validate our findings, we implemented Hercules and made evaluation to show that Hercules can effectively improve the data access speed and reduce the response time, compared with the Flashcache storage structure, and in particular, with Hercules, we can achieve 10% performance improvement over the system in absence of hot delta data caching. Wen Cheng 0003, Yuqi Zou, Lingfang Zeng, Yang Wang 0006 |
Comput. J. | 4 |
| 2021 | AucSwap: A Vickrey auction modeled decentralized cross-blockchain asset transfer protocol
Huaming Wu, Tianhui Meng, Yang Wang 0006, Cheng-Zhong Xu 0001 |
J. Syst. Archit. | 5 |
| 2021 | AIOC2: A deep Q-learning approach to autonomic I/O congestion control in Lustre
Wen Cheng 0003, Shijun Deng, Lingfang Zeng, Yang Wang 0006, André Brinkmann |
Parallel Comput. | 4 |
| 2021 | Sova: A Software-Defined Autonomic Framework for Virtual Network AllocationsabstractWith the rise of network virtualization, the workloads deployed on data center are dramatically changed to support diverse service-oriented applications, which are in general characterized by the time-bounded service response that in turn puts great burden on the data-center networks. Although there have been numerous techniques proposed to optimize the virtual network allocation in data center, the research on coordinating them in a flexible and effective way to autonomically adapt to the workloads for service time reduction is few and far between. To address these issues, in this article we propose Sova, an autonomic framework that can combine the virtual dynamic SR-IOV (DSR-IOV) and the virtual machine live migration (VLM) for virtual network allocations in data centers. DSR-IOV is a SR-IOV-based virtual network allocation technology, but its operation scope is very limited to a single physical machine, which could lead to the local hotspot issue in the course of computation and communication, likely increasing the service response time. In contrast, VLM is an often-used virtualization technique to optimize global network traffic via VM migration. Sova exploits the software-defined approach to combine these two technologies with reducing the service response time as a goal. To realize the autonomic coordination, the architecture of Sova is designed based on the MAPE-K loop in autonomic computing. With this design, Sova can adaptively optimize the network allocation between different services by coordinating DSR-IOV and VLM in autonomic way, depending on the resource usages of physical servers and the network characteristics of VMs. To this end, Sova needs to monitor the network traffic as well as the workload characteristics in the cluster, whereby the network properties are derived on the fly to direct the coordination between these two technologies. Our experiments show that Sova can exploit the advantages of both techniques to match and even beat the better performance of each individual technology by adapting to the VM workload changes. Zhiyong Ye, Yang Wang 0006, Shuibing He, Cheng-Zhong Xu 0001, Xian-He Sun |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | FastThetaJoin: An Optimization on Multi-way Data Stream θ-join with Range Constraints
Ziyue Hu, Xiaopeng Fan 0002, Yang Wang 0006, Cheng-Zhong Xu 0001 |
ICA3PP (1) | 3 |
| 2020 | Optimizing Multi-way Theta Join for Data Skew in Sub-second Stream ComputingabstractIn sub-second stream computing, the answer to a complex query usually depends on operations of aggregation or join on streams, especially multi-way theta join. Some attribute keys are not distributed uniformly, which is called the data intrinsic skew problem, such as taxi car plate in GPS trajectories and transaction records, or stock code in stock quotes and investment portfolios etc. In this paper, we define the concept of key redundancy for single stream as the degree of data intrinsic skew, and joint key redundancy for multi-way streams. We present an execution model for multi-way stream theta joins with a fine-grained cost model to evaluate its performance. We propose a solution named Group Join (GroJoin) to make use of key redundancy during transmission and execution in a cluster. GroJoin is adaptive to data intrinsic skew in the way that it depends on the grouping condition we find out, i.e., the selectivity of theta join results should be smaller than 25%. Experiments are carried out by our MS-Generator to produce multi-way streams, and the simulation results show that GroJoin can decrease at most 45% transmission overheads with different key redundancies and value-key proportionality coefficients, and reduce at most 70% query delay with different key distributions. We further implement GroJoin in Multi-Way Stream Theta Join by Spark Streaming. The experimental results demonstrate that there are about 40%~50% join latency reduced after our optimization with a very small computation cost. Xiaopeng Fan 0002, Xinchun Liu, Yang Wang 0006, Youjun Wang, Jing Li 0047 |
ICPADS | 3 |
| 2020 | AOAM: Automatic Optimization of Adjacency Matrix for Graph Convolutional NetworkabstractGraph Convolutional Network (GCN) is adopted to tackle the problem of convolution operation in non-Euclidean space. Previous works on GCN have made some progress, however, one of their limitations is that the design of Adjacency Matrix (AM) as GCN input requires domain knowledge and such process is cumbersome, tedious and error-prone. In addition, entries of a fixed Adjacency Matrix are generally designed as binary values (i.e., ones and zeros) which can not reflect the real relationship between nodes. Meanwhile, many applications require a weighted and dynamic Adjacency Matrix instead of an unweighted and fixed AM, and there are few works focusing on designing a more flexible Adjacency Matrix. To that end, we propose an end-to-end algorithm to improve the GCN performance by focusing on the Adjacency Matrix. We first provide a calculation method callednodeinformationentropyto update the matrix. Then, we perform the search strategy in a continuous space and introduce the Deep Deterministic Policy Gradient (DDPG) method to overcome the drawback of the discrete space search. Finally, we integrate the GCN and reinforcement learning into an end-to-end framework. Our method can automatically define the Adjacency Matrix without prior knowledge. At the same time, the proposed approach can deal with any size of the matrix and provide a better AM for network. Four popular datasets are selected to evaluate the capability of our algorithm. The method in this paper achieves the state-of-the-art performance onCoraandPubmeddatasets, with the accuracy of 84.6% and 81.6% respectively. Yuhang Zhang 0010, Hongshuai Ren, Jiexia Ye, Yang Wang 0006, Kejiang Ye, Cheng-Zhong Xu 0001 |
ICPR | 5 |
| 2020 | Data Caching Based Transfer Optimization in Large Scale Networks
Xinxin Han, Guichen Gao, Yang Wang 0006, Hing-Fung Ting, Yong Zhang 0001 |
PDCAT | 3 |
| 2020 | Approximation Algorithm for the Offloading Problem in Edge Computing
Xinxin Han, Guichen Gao, Li Ning 0001, Yang Wang 0006, Yong Zhang 0001 |
WASA (1) | 4 |
| 2020 | IMCI: an efficient fingerprint retrieval approach based on 3D stacked memory
Wen Cheng 0003, Ran Cai, Lingfang Zeng, Dan Feng 0001, André Brinkmann, Yang Wang 0006 |
Sci. China Inf. Sci. | 6 |
| 2020 | Towards cost-effective service migration in mobile edge: A Q-learning approach
Yang Wang 0006, Hongshuai Ren, Kejiang Ye, Cheng-Zhong Xu 0001 |
J. Parallel Distributed Comput. | 1 |
| 2020 | Improving LSM-trie performance by parallel searchabstractSummary LSM‐trie‐based key‐value (KV) store is often used to manage an ultralarge dataset in reality by introducing a number of sublevels at each level, its linear growth pattern can fairly reduce the write amplification in store operations. Although this design is effective for the write operation, the last level holds a large proportion of KV items, leading to the extreme imbalance of data distribution. Therefore, to support efficient read, we need to carefully consider this imbalance. On the other hand, to ensure that acquired data is latest, the LSM‐trie needs to search the dataset at different levels one by one, and this search method may take a lot of unnecessary time. When the number of items is ultralarge, the random lookup performance may be poor due to the imbalance data distribution. To address this issue, we improve the read performance of the LSM‐trie by changing its serial search to parallel search, using two threads to simultaneously search at the last level and other levels, respectively. Our experiment results show that the read performance of the LSM‐trie can be improved up to 98.35% and on average 71.55%. Wen Cheng 0003, Lingfang Zeng, Yang Wang 0006, Lars Nagel 0001, Tim Süß, André Brinkmann |
Softw. Pract. Exp. | 4 |
| 2020 | Algorithmics of Cost-Driven Computation Offloading in the Edge-Cloud EnvironmentabstractComputation offloading between the edge and the cloud is an effective way for deployed service to fully utilize the resources at both sides for its QoS improvement and overall cost reduction. Although the offloading problem has been intensively studied in the context of mobile computing, existing algorithms in most cases cannot be effectively migrated to the edge-cloud environment because their inter-partition communication costs are always deemed as symmetric, and their intra-partition communication costs are often ignored, which, though reasonable to the traditional case, are not valid to our settings anymore. In this article, we propose a new algorithmic approach to the offloading problem in the edge-cloud environment, where a heterogeneous model is advocated to incorporate the communication cost between co-resident tasks while considering the asymmetry of communication costs between non-coresident tasks. We prove the offloading problem with respect to this model is NP-hard, and thereby designing an efficient algorithm to obtain a sub-optimal solution. Additionally, we also show that in a homogeneous case when the intra-partition and inter-partition communication costs between any pair of interactive tasks are symmetric, an optimal offloading algorithm can be devised by transforming the problem into a classical min-cut problem. We implemented and evaluated the algorithms by offloading a PageRank-based application in a controlled edge-cloud setting. Our empirical results show that the proposed algorithm for the heterogeneous case is always efficient to find a better offloading scheme, compared with the selected existing algorithms, while for the homogeneous case, the proposed solution can efficiently achieve the optimal strategy. Mingzhe Du, Yang Wang 0006, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Computers | 2 |
| 2020 | SMig-RL: An Evolutionary Migration Framework for Cloud Services Based on Deep Reinforcement LearningabstractService migration is an often-used approach in cloud computing to minimize the access cost by moving the service close to most users. Although it is effective in a certain sense, the service migration in existing research still suffers from some deficiencies in its evolutionary abilities in scalability , sensitivity , and adaptability to effectively react to the dynamically changing environments. This article proposes an evolutionary framework based on deep reinforcement learning for virtual service migration in large-scale mobile cloud centers. To enhance the spatio-temporal sensitivity of the algorithm, we design a scalable reward function for virtual service migration, redefine the input state, and add a Recurrent Neural Network ( RNN ) to the learning framework. Additionally, in order to enhance the adaptability of the algorithm, we also decompose the action space and exploit the network cost to adjust the number of virtual machine (VMs). The experimental results show that, compared with the existing results, the migration strategy generated by the algorithm can not only significantly reduce the total service cost and achieve the load balancing at the same time, but also address the burst situations with low cost in dynamic environments. Hongshuai Ren, Yang Wang 0006, Cheng-Zhong Xu 0001 |
ACM Trans. Internet Techn. | 2 |
| 2019 | On Cost-Driven Computation Offloading in the Edge: A New Model ApproachabstractComputation offloading is an often-used optimization method that exploits servers with powerful and plentiful resources to maximize computation efficiency with minimum cost. In this method, a client application is usually modeled as a weighted directed acyclic graph (DAG), which is typically split into two distinct parts - one running on client device and the other on server machine. To simplify the model, the inter-part communication costs are always assumed to be symmetric and the intra-part communication costs are commonly ignored. Although these assumptions are reasonable to the offloading in traditional mobile computing, they are not valid anymore when considering the problem in the edgecloud environment, especially with the development of microservice, where a provisioned multi-machine cluster at each side is involved. To address this problem, we propose a new offloading model in this paper, where both the intra-part communication costs as well as the asymmetry of inter-part communication costs are incorporated to carry out the client application, which are not a part of previous approaches. Given this model, we first prove the offloading problem is NP-hard, then design an efficient greedy algorithm to obtain a sub-optimal solution. Our numerical results show that our algorithm for the new model is always efficient to find a better offloading scheme, compared with other existing algorithms that lack the notion of communication costs between tasks co-located at the same side and the asymmetry of communication costs crossing sides. Mingzhe Du, Yang Wang 0006, Cheng-Zhong Xu 0001 |
CCGRID | 2 |
| 2019 | DP_Greedy: A Two-Phase Caching Algorithm for Mobile Cloud ServicesabstractIn this paper, we study the data caching problem in mobile cloud environment where multiple correlated data items could be packed and migrated to serve a predefined sequence of requests. By leveraging the spatial and temporal trajectory of requests, we propose a two-phase caching algorithm. We first investigate the correlation between data items to determine whether or not two data items could be packed to transfer, and then combine an existing dynamic programming (DP)-based algorithm and a greedy strategy to design a two-phase algorithm, named DP_Greedy, for effectively caching these shared data items to serve a predefined sequence of requests. Under homogeneous cost model, we prove the proposed algorithm is at most 2/α times worse than the optimal one in terms of the total service cost, where α is the defined discount factor, and also show that the algorithm can achieve this results within O(mn2) time and O(mn) space complexity for m caches to serve a n-length sequence. We evaluate our algorithm by effectively implementing it and comparing it with the non-packing case, the result show the proposed DP_Greedy algorithm not only presents excellent performances but is also more in line with the actual situation. Xiaopeng Fan 0002, Yang Wang 0006, Shuibing He, Cheng-Zhong Xu 0001 |
CLUSTER | 3 |
| 2019 | Multiscale Directional Fusion for Depth Map Super Resolution with DenoisingabstractTo tackle three main problems in depth map super resolution (SR) process, which are texture copy artifacts, blurred edge artifacts and jagged edge artifacts, we propose a depth map super resolution with denoising method based on multiscale directional fusion via nonsubsampled contourlet transform (NCST). We first transform low resolution depth maps of multiple views via NSCT. Then NSCT coefficients are denoised by a BayesShrink threshold in nonsubsampled directional filter banks (NSDFB) domain and fused by the max coefficient absolute value (mCAV) rule respectively within each scale and direction. Finally the fused coefficients are synthesized and upscaled to a high resolution depth map utilizing a modified edge-guided joint bilateral filter. Experimental results demonstrate that our method significantly outperforms the state-of-the-art super resolution algorithms quantitatively and visually while mitigating the corrupted noise. Dan Xu 0008, Xiaopeng Fan 0001, Yang Wang 0006, Debin Zhao, Wen Gao 0001 |
ICASSP | 4 |
| 2019 | Towards Cluster-wide Deduplication Based on CephabstractIn this paper, we design an efficient deduplication algorithm based on the distributed storage architecture of Ceph. The algorithm uses on-line block-level data deduplication technology to complete data slicing, which neither affects the data storage process in Ceph nor alter other interfaces and functions in Ceph. Without relying on any central node, the algorithm maintains the characteristics of Ceph by designing a special hash object to store the data fingerprint, and uses the CRUSH algorithm to judge the data duplication based on calculation, instead of global search. The algorithm replaces the duplicate data with the deduplicated objects, which storage their fingerprints with less storage space. We compare the effects of different block sizes with respect to the performance and deduplication rates through experimental studies, and select the most appropriate block size in our prototype implementation. The experimental results show that the algorithm can not only effectively save the storage space but also improve the bandwidth utilization when reading and writing the duplicate data. Yang Wang 0006, Hekang Wang, Kejiang Ye, Cheng-Zhong Xu 0001, Shuibing He, Lingfang Zeng |
NAS | 2 |
| 2019 | On Cost-Driven Collaborative Data Caching: A New Model ApproachabstractIn this paper we consider a new caching model that enables data sharing for network services in a cost-effective way. The proposed caching algorithms are characterized by using monetary cost and access information to control the cache replacements, instead of exploiting capacity-oriented strategies as in traditional approaches. In particular, given a stream of requests to a shared data item with respect to a homogeneous cost model, we first propose a fast off-line algorithm using dynamic programming techniques, which can generate an optimal schedule within$O(mn)$time-space complexity by using cache, migration as well as replication to serve a$n$-length request sequence in a$m$-node network, substantially improving the previous results. Furthermore, we also study the online form of this problem, and present an 3-competitive online algorithm by leveraging an idea of anticipatory caching. The algorithm can serve an online request in constant time and is space efficient in$O(m)$as well, rendering it more practical in reality. We evaluate our algorithms, together with some variants, by conducting extensive simulation studies. Our results show that the optimal cost of the off-line algorithm is changed in a parabolic form as the ratio of caching cost to transfer cost is increased, and the online algorithm is less than 2 times worse in most cases than its optimal off-line counterpart. Yang Wang 0006, Shuibing He, Xiaopeng Fan 0002, Cheng-Zhong Xu 0001, Xian-He Sun |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | How Does the Workload Look Like in Production Cloud? Analysis and Clustering of Workloads on Alibaba Cluster TraceabstractCloud computing technology is widely used in today's datacenters due to the benefits such as high scalability, on-demand services and low cost. An in-depth understanding of the characteristics of workloads running in production cloud environments is very important for improving the resource management efficiency. In this paper, we make a detailed analysis with visualization techniques and clustering methods on the trace dataset released by Alibaba which contains 11089 online services and 12951 batch jobs running on 1313 machines. Our methodology for clustering workloads contains: i) Select effective feature vectors as the dimensions of clustering; ii) Identify the cluster boundaries of each dimension using K-Means algorithm; iii) Classify jobs by combining the feature vectors which uses the results from previous step; iv) Analyze the characteristics of workload groups at runtime. Our analysis reveals several insights which previous work has not found on Alibaba cluster trace. For batch jobs: a) Average CPU cores of all batch jobs show bimodal-distribution obviously. b) At a random sampling time, more than 50 % machines only run one group of jobs with a short duration, medium CPU cores and small memory utilization, the remaining machines run mixed groups of jobs. For online instances: a) The resource usage (CPU, Memory, and Disk) of most online instances is low; b) There are up to six groups running on the same machine according to our clustering method at a random sampling time. Wenyan Chen 0001, Kejiang Ye, Yang Wang 0006, Guoyao Xu, Cheng-Zhong Xu 0001 |
ICPADS | 3 |
| 2018 | Modeling Application Performance in Docker Containers Using Machine Learning TechniquesabstractDocker container is experiencing a rapid development with the support from industry like Google and is being widely used in large scale production cloud environments. However the performance of applications running in Docker containers is still not clear due to the complex relationship between container resource allocation and application performance. In this paper, we first study the impact of key parameters in container resource allocation that affect the performance of containerized applications. Then, we present modeling techniques over CPU, memory and I/O resources to characterize the performance of applications running in containers. To address this multi-dimensional modeling problem, we propose three machine learning techniques, i.e. Linear Regression (LR), Support Vector Machine (SVM) and Artificial Neural Network (ANN). We implement and evaluate the modeling techniques for four complex benchmark workloads from Spark. Experimental results demonstrate the proposed models can achieve as low as 2.27% prediction error, with an average of 10.13% for most applications. Furthermore, the prediction accuracy of SVM and ANN models are substantially better than LR based approaches, with 48.13% and 29.30% improvement. Kejiang Ye, Yanmin Kou, Chengzhi Lu, Yang Wang 0006, Cheng-Zhong Xu 0001 |
ICPADS | 4 |
| 2018 | A Road-Aware Spatial Mapping for Moving ObjectsabstractThe Internet-of-Things (IoT) attracts great attention in the past few years. With millions of devices connected to the network, data are generated at an unprecedented speed and the data must be stored efficiently in the database to serve spatial queries. In existing spatial databases that use space-filling curves to organize the data, they store spatial data without considering on-road data distribution. This will introduce unnecessary computation and I/O cost in the service of users' queries about data on the roads. In this paper, we present a Road-Aware Spatial Mapping of data to the storage, or RASM for short, which can be applied in spatial databases for highly efficient storage and query services for moving objects. Usually, a space-filling curve, such as the Hilbert curve, is used to map data in a cell of a geographical area to a segment of linear storage space. However, in a road-network system where data are most distributed and queried along the roads, using a generic square cell as a mapping unit to aggregate data is in conflict with the data use pattern. In RASM, road segment, instead of the cell, is used as the unit of space mapping and data storage so that data requested in a road query can be stored together to enable efficient I/O. Furthermore, a substantial computation may be required to identify mapping units covered in a query in a geometric space. As RASM has grouped data in the road-segment units, one can efficiently found the units covered in a road query, which is usually concerned only about data on a few segments of roads. We implemented a prototype query-serving system using RASM to map data on road segments to a linear space enabled by LevelDB, a widely-used key-value store. Experiment results with real-world traffic data show that with RASM, the road query time can be reduced by up to 43%, and the I/O traffic can be reduced by up to 70%. In the meantime, other queries about geographical regions are well supported in RASM with minimal performance impacts. Xingsheng Zhao, Jingwen Shi, Mingzhe Du, Fan Ni, Song Jiang 0001, Yang Wang 0006 |
IPCCC | 6 |
| 2018 | A Migratory Heterogeneity-Aware Data Layout Scheme for Parallel File SystemsabstractParallel file systems (PFSs) are widely deployed to speed up the performance of high-performance computing (HPC) applications. In recent years, hybrid PFSs that consist of HDD-SSD servers, have attracted much attention in HPC community. However, existing data layout schemes do not well consider the characteristics of heterogeneous servers and heterogeneous access patterns, thus may experience considerable inefficiencies. In this study, we propose MHA, a migratory heterogeneity-aware data layout scheme to improve the data distribution of hybrid PFS. More specifically, to accommodate heterogeneous access patterns, MHA first migrates file data into several regions, each with similar access patterns. Then, by leveraging a data access cost model, MHA determines the appropriate stripe sizes on heterogeneous servers to get the best performance on each region. We have implemented MHA under MPI-IO library on top of OrangeFS file system. Experimental results show that MHA can significantly improve the hybrid PFS I/O system performance compared to existing data layout schemes. Shuibing He, Xian-He Sun, Yang Wang 0006, Cheng-Zhong Xu 0001 |
IPDPS | 3 |
| 2018 | A Deep Learning Approach for Network Anomaly Detection Based on AMF-LSTM
Mingyi Zhu, Kejiang Ye, Yang Wang 0006, Cheng-Zhong Xu 0001 |
NPC | 3 |
| 2017 | Data Caching in Next Generation Mobile Cloud Services, Online vs. Off-LineabstractIn this paper we consider the data caching problem in next generation data services in the cloud, which is characterized by using monetary cost and access trajectory information to control cache replacements, instead of exploiting capacityoriented strategies as in traditional research. In particular, given a stream of requests to a shared data item with respect to a homogeneous cost model, we first propose a fast off-line algorithm using dynamic programming techniques. The proposed algorithm can generate optimal schedule within O(mn) timespace complexity to cache, migrate as well as replicate the shared data item to serve an n-length request sequence with minimum cost in a fully connected m-node network, substantially improving the previous results. Additionally, we also study this problem in its online form, and present a 3-competitive online algorithm by leveraging a speculative caching idea. The algorithm can serve an online request in constant time, and is space efficient in O(m) as well, rendering it to be more practical in reality. Our research complements the shortage of similar research in literature on this problem. Yang Wang 0006, Shuibing He, Xiaopeng Fan 0002, Cheng-Zhong Xu 0001, Joseph C. Culberson, Joseph Horton |
ICPP | 1 |
| 2017 | Service Migrations in the Cloud for Mobile Accesses: A Reinforcement Learning ApproachabstractMigrating service to certain vantage locations that are close to its clients can not only reduce the service access latency,but also minimize the network costs for its service provider. As such, this problem is particularly important for time-bounded services to achieve both enhanced QoS and cost effectiveness as well. However, the service migration is not free, coming at costs of bulk-data transfer and likely service disruption, as a result, increasing the overall service costs. To gain the benefits of service migration while minimizing service costs, in this paper, we leverage reinforecement learning (RL) methods to propose an efficient algorithm, called Mig- RL, for the service migration in a cloud environment. The Mig-RL utilizes an agent to learn the optimal policy that determines service migration status by using a typical RL algorithm, called Q-learning. Specifically, the agent learns from the historical access information to decide when and to where the service should be migrated, without requiring any prior information regarding the service accesses. Therefore, the agent can dynamically adapt to the environment and achieve online migration in real time. Experimental results on the real and synthesized access sequences from cloud networks show that Mig-RL can minimize the service costs, and in the meantime, improve the quality of service (QoS) by adapting to the changes of mobile access patterns. Yang Wang 0006, Cheng-Zhong Xu 0001 |
NAS | 2 |
| 2017 | A Hash-Based Space-Efficient Page-Level FTL for Large-Capacity SSDsabstractWith increasing demands on high-performance and large-capacity SSDs in the enterprise-scale storage, the concern about the inefficient use of the DRAM space in SSDs rises, especially for those using page-level FTL (Flash Translation Layer). In such an FTL, the address mapping scheme allows a logical page address (LPA) to be mapped to any physical page address (PPA) in the disk. Though it provides flexible address management and minimizes internal data movements, it requires a large address mapping table whose size is proportional to the capacity of the disk. With the increase of SSD's capacity, the table can be too large to be held entirely in the DRAM buffer of the SSD, causing constantly accessing to the flash for the address translation. This performance penalty due to the buffer misses is particularly high with workloads of weak access locality and large working sets. In this paper, we propose a space- efficient page- level FTL using hash functions in the address translation, named Hash-based Page- level FTL, or HP-FTL in short, to address the concern. HP-FTL trades mapping flexibility with limited performance impact for high space efficiency allowing the entire table to fit in the buffer and eliminating translation misses. The experiment results show that HP-FTL can provide up to 2.6X throughput compared to DFTL, a representative page-level FTL, using the same amount of DRAM for buffering the table. Meanwhile, HP-FTL reduces the mapping table size to about 25% of the table space required by page- level mapping schemes, including DFTL, without having any buffer misses. Fan Ni, Chunyi Liu, Yang Wang 0006, Cheng-Zhong Xu 0001, Xiao Zhang 0014, Song Jiang 0001 |
NAS | 3 |
| 2017 | A Region-Based Approach to Pipeline Parallelism in Java Programs on MulticoresabstractAs multicore architectures dominate mainstream computing platforms, migrating legacy applications into their parallel representation becomes a viable approach to reaping the benefits of multicore computing. In this paper we present a dataflow analysis tool that assists programmers to exploit the coarse-grained pipeline parallelism in stream-like Java applications on multicores. With this tool, programmers can partition a source Java program into a set of regions, which as pipeline stages, are connected via data channels to execute on multicores. To this end, we propose a simple yet effective framework that leverages JVMTI (JVM Tool Interface) and Java agent techniques to track the data communication patterns among different regions, whereby a stream graph of the program is constructed. The graph is further used by the framework and programmers to re-factor the Java application into a pipelined program so that the potential of the multicores can be fully utilized. This procedure can be repeated in several rounds to progressively improve the performance. By applying this tool to several selected benchmarks, we demonstrate the effectiveness of the approach in terms of the performance improvements of some stream-like Java applications. Yang Wang 0006, Kenneth B. Kent |
PDP | 1 |
| 2017 | Naplus: a software distributed shared memory for virtual clusters in the cloudabstractSummary Virtual clusters (VCs) have exhibited various advantages over traditional cluster computing platforms by virtue of their extensibility, reconfigurability, and maintainability. As such, they have become a major execution environment for cloud‐based cluster applications. However, compared with traditional clusters, their distributed‐memory programming paradigm still remains largely unchanged, which implies that cluster applications cannot be efficiently deployed in VCs, especially when virtual machines (VMs) are running in different physical hosts. Recently, some efforts have been made to improve inter‐VM communication, resulting in many studies on how cluster applications could take advantages of VCs. However, most of them mainly focus on the situation that the VMs are all coresident on the same physical machine where the message passing mechanism is usually optimized away by exploiting the host's shared memory. In this paper, we present a design and implementation of Naplus, a kernel‐based virtual machine approach to the inter‐VM communications that are across different physical hosts. Naplus is based on Nahanni, a mechanism for shared‐memory communication in virtual environments. As such, it not only inherits the major merits of Nahanni with respect to flexible data structures and efficient synchronization but also achieves a shared‐memory paradigm among VMs. With Naplus, we enable the size of shared space to be maximized as large as the sum of each machine's local memory to accommodate cluster applications with large memory footprints. We prototype Naplus in a dual‐host system where an empirical study is conducted to show the effectiveness of the Naplus approach in achieving distributed shared memory for VCs in data centers. Copyright © 2017 John Wiley & Sons, Ltd. Lingfang Zeng, Yang Wang 0006, Kenneth B. Kent, Ziliang Xiao |
Softw. Pract. Exp. | 2 |
| 2017 | Toward cost-effective replica placements in cloud storage systems with QoS-awarenessabstractSummary In this paper, we propose a simulation model to study real‐world replication workflows for cloud storage systems. With this model, we present three new methods to maximize the storage space usage during replica creation, and two novel QoS aware greedy algorithms for replica placement optimization. By using a simulation method, our algorithms are evaluated, through a comparison with the existing placement algorithms, to show that (i) a more evenly distributed replicas for a data set can be achieved by using round‐robin methods in replica creation phase and (ii) the two proposed greedy algorithms, namedGS_QoSandGS_QoS_C1, not only have more economical results than those from Chenet al., but also guarantee the QoS for clients. Copyright © 2016 John Wiley & Sons, Ltd. Lingfang Zeng, Yang Wang 0006, Kenneth B. Kent, David Bremner, Cheng-Zhong Xu 0001 |
Softw. Pract. Exp. | 3 |
| 2017 | On Service Migrations in the Cloud for Mobile Accesses: A Distributed ApproachabstractWe study the problem of dynamically migrating a service in the cloud to satisfy an online sequence of mobile batch-request demands in a cost-effective way. The service may have single or multiple replicas, each running on a virtual machine. As the origin of mobile accesses frequently changes over time, this problem is particularly important for time-bounded services to achieve enhanced Quality of Service and cost effectiveness. Moving the service closer to the client locations not only reduces the service access latency but also minimizes the network costs for service providers. However, these benefits are not free. The migration comes at a cost of bulk-data transfer and service disruption, and hence, increasing the overall service costs. To gain the benefits of service migration while minimizing the caused monetary costs, we propose an efficient search-based algorithm Dmig to migrate a single server, and then extend it as a scalable algorithm, called mDmig , to the multi-server situation, a more general case in the cloud. Both algorithms are fully distributed, symmetric, and characterized by the effective use of historical access information to conduct virtual migration so that the limitations of local search in the cost reduction can be overcome. To evaluate the algorithms, we compared them with some existing algorithms and an off-line algorithm. Our simulation results showed that the proposed algorithms exhibit better performance in service migration by adapting to the changes of mobile access patterns in a cost-effective way. Yang Wang 0006, Bharadwaj Veeravalli, Chen-Khong Tham, Shuibing He, Cheng-Zhong Xu 0001 |
ACM Trans. Auton. Adapt. Syst. | 1 |
| 2017 | Heterogeneity-Aware Collective I/O for Parallel I/O Systems with Hybrid HDD/SSD ServersabstractCollective I/O is a widely used middleware technique that exploits I/O access correlation among multiple processes to improve I/O system performance. However, most existing implementations of collective I/O strategies are designed and optimized for homogeneous I/O systems. In practice, the homogeneity assumptions do not hold in heterogeneous parallel I/O systems, which consist of multiple HDD and SSD-based servers and become increasingly promising. In this paper, we propose a heterogeneity-aware collective-I/O (HACIO) strategy to enhance the performance of conventional collective I/O operations. HACIO reorganizes the order of I/O requests for each aggregator with awareness of the storage performance of heterogeneous servers, so that the hardware of the systems can be better utilized. We have implemented HACIO in ROMIO, a widely used MPI-IO library. Experimental results show that HACIO can significantly increase the I/O throughputs of heterogeneous I/O systems. Shuibing He, Yang Wang 0006, Xian-He Sun, Chuanhe Huang, Cheng-Zhong Xu 0001 |
IEEE Trans. Computers | 2 |
| 2017 | HARL: Optimizing Parallel File Systems with Heterogeneity-Aware Region-Level Data LayoutabstractParallel file system (PFS) is commonly used in high-end computing systems. With the emergence of solid state drives (SSDs), hybrid PFS, which consists of both HDD and SSD servers, provides a practical I/O system solution for data-intensive applications. However, most existing data layout schemes are inefficient for hybrid PFS due to their unawareness of server heterogeneities and workload changes in different parts of a file. In this study, we propose a heterogeneity-aware region-level data layout scheme, HARL, to improve the data distribution of a hybrid PFS. HARL first divides a file into fine-grained, varying sized regions according to the workload features of an application, then determines appropriate file stripe sizes on servers for each region based on the performance of heterogeneous servers. Furthermore, to further improve the performance of a hybrid PFS, we propose a dynamic region-level layout scheme, HARL-D, which creates multiple replicas for each region and redirects file requests to the proper replicas with the lowest access costs at the runtime. Experimental results of representative benchmarks and a real application show that HARL can greatly improve I/O system performance, and demonstrate the advantages of HARL-D over HARL. Shuibing He, Yang Wang 0006, Xian-He Sun, Cheng-Zhong Xu 0001 |
IEEE Trans. Computers | 2 |
| 2017 | Adaptive Scheduling of Task Graphs with Dynamic ResilienceabstractThis paper studies a scheduling problem of task graphs on a nondedicated networked computing platform. The networked platform is characterized by a set of fully connected processors such as a multiprocessor system that can be shared by multiple tasks. Therefore, the computation and communication capacities of the computing platform dynamically fluctuate. To deal with this fluctuations for high performance task graph computing, we propose an online dynamic resilience scheduling algorithm called Adaptive Scheduling Algorithm (ASA) that bears certain distinct features compared to existing algorithms. First, the proposed algorithm deliberately assigns tasks to idle processors in multiple rounds to prevent any unfavorable decisions and also to avoid inefficient assignments of certain key tasks to slow processors. Second, the algorithm adopts task duplication as an attempt to minimize serious increase of schedule length due to unexpected processor slowdown. Finally, a look-ahead message transmission policy is applied to save communication time and further improve the overall performance. Performance evaluation results are presented to demonstrate the effectiveness and competitiveness of our approaches when compared with the existing algorithms. Menglan Hu, Jun Luo 0001, Yang Wang 0006, Bharadwaj Veeravalli |
IEEE Trans. Computers | 3 |
| 2017 | CosaFS: A Cooperative Shingle-Aware File SystemabstractIn this article, we design and implement a cooperative shingle-aware file system, called CosaFS , on heterogeneous storage devices that mix solid-state drives (SSDs) and shingled magnetic recording (SMR) technology to improve the overall performance of storage systems. The basic idea of CosaFS is to classify objects as hot or cold objects based on a proposed Lookahead with Recency Weight scheme. If an object is identified as a hot (small) object, then it will be served by SSD. Otherwise, cold (large) objects are stored on SMR. For an SMR, large objects can be accessed in large sequential blocks, rendering the performance of their accesses comparable with that of accessing the same large sequential blocks as if they were stored on a hard drive. Small objects, such as inodes and directories, are stored on the SSD where “seeks” for such objects are nearly free. With thorough empirical studies, we demonstrate that CosaFS, as a cooperative shingle-aware file system, with metadata separation and cache-assistance, is a very effective way to handle the disk-based data demanded by the shingled writes and outperforms the device- and host-side shingle-aware file systems in terms of throughput, IOPS, and access latency as well. Lingfang Zeng, Zehao Zhang, Yang Wang 0006, Dan Feng 0001, Kenneth B. Kent |
ACM Trans. Storage | 3 |
| 2017 | Cost-Aware Region-Level Data Placement in Multi-Tiered Parallel I/O SystemsabstractMulti-tiered Parallel I/O systems that combine traditional HDDs with emerging SSDs mitigate the cost burden of SSDs while benefiting from their superior I/O performance. While a multi-tiered parallel I/O system is promising for data-intensive applications in high-performance (HPC) domains, placing data on each tier of the system to achieve high I/O performance remains a challenge. In this paper, we propose a cost-aware region-level (CARL) data placement scheme in multi-tiered parallel I/O systems. CARL divides a large file into several small regions, and then places regions on different types of servers based on region access costs. CARL includes a static policy S-CARL and a dynamic policy D-CARL. For applications whose I/O access patterns are completely known, S-CARL calculates the region costs within the entire workload duration, and uses a static data placement scheme to selectively place regions on the proper servers. To adapt to applications whose access patterns are unknown in advance, D-CARL uses a dynamic data placement scheme which migrates data among different servers within each time window. We have implemented CARL under MPI-IO library and OrangeFS parallel file system environment. Our evaluation with representative benchmarks and an application shows that CARL is both feasible and able to improve I/O performance significantly. Shuibing He, Yang Wang 0006, Zheng Li 0006, Xian-He Sun, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Using MinMax-Memory Claims to Improve In-Memory Workflow Computations in the CloudabstractIn this paper, we consider to improve scientific workflows in cloud environments where data transfers between tasks are performed via provisioned in-memory caching as a service, instead of relying entirely on slower disk-based file systems. However, this improvement is not free since services in the cloud are usually charged in a “pay-as-you-go” model. As a consequence, the workflow tenants have to estimate the amount of memory that they would like to pay. Given the intrinsic complexity of the workflows, it would be very hard to make an accurate prediction, which would lead to either oversubscription or undersubscription, resulting in unproductive spending or performance degradation. To address this problem, we propose a concept of minmax memory claim (MMC) to achieve cost-effective workflow computations in in-memory cloud computing environments. The minmax-memory claim is defined as the minimum amount of memory required to finish the workflow without compromising its maximum concurrency. With the concept of MMC, the workflow tenants can achieve the best performance via in-memory computing while minimizing the cost. In this paper, we present the procedure of how to find the MMCs for those workflows with arbitrary graphs in general and develop optimal efficient algorithms for some well-structured workflows in particular. To further show the values of this concept, we also implement these algorithms and apply them, through a simulation study, to improve deadlock resolutions in workflow-based workloads when memory resources are constrained. Shuibing He, Yang Wang 0006, Xian-He Sun, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | Raccoon: A Novel Network I/O Allocation Framework for Workload-Aware VM Scheduling in Virtual EnvironmentsabstractWe present a network I/O allocation framework, called Raccoon, for workload-aware VM scheduling algorithm to facilitate hybrid I/O workloads in virtual environments. Raccoon combines the strengths of paravirtual I/O and SR-IOV techniques to not only minimize the network latency, but also optimize the bandwidth utilization for workload-aware VM scheduling. In Raccoon, a limited number of VFs in SR-IOV are granted to I/O-intensive VMs while the paravirtual Network Interface Cards (vNICs) are allocated to other non-I/O-intensive VMs as the default resources. With this design, Raccoon provides latency reduction and bandwidth guarantee under the premise that I/O-intensive VMs will always be granted the VFs to facilitate their I/O operations. The types of workloads in each VM are identified at runtime by modified XenMon. By leveraging the ACPI Hotplug technique, Raccoon can adaptively plugin and plugout the SR-IOV VFs upon the changes of VM requirements so that an efficient I/O workload-aware VM scheduling algorithm can be implemented based on the bonding driver technique. The experimental results reveal that Raccoon can combine the benefits of para-virtual I/O and SR-IOV techniques to improve the overall performance of virtualized platforms with VMs that have diverse I/O workloads. Lingfang Zeng, Yang Wang 0006, Xiaopeng Fan 0002, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | A Low Disk-Bound Transaction Logging System for In-memory Distributed Data StoresabstractTransaction logging and snapshotting are techniques used to deliver durability to the data in in-memory data stores. Absolute durability guarantees are delivered to a system by sequentially recording the transaction logs and snapshots to a non-volatile disk. Recent advancements in database restoration techniques have given rise to lock-free fuzzy snapshots. Still the transaction log that completes the fuzzy snapshots is not lock-free. In addition to locking, the major overhead behind the transaction logging technique is the bottleneck involved in storing the logs to a persistent but slower disk. This paper concentrates on implementing an in-memory transaction logging system with a lesser disk dependency. This logging system mainly targets the distributed in-memory data stores that are transaction replicated, eventually consistent and fault tolerant to crash failures. By making logging in-memory, the performance will be improved, but during the crash fails, the state may be lost. On recovery, we restore the current state partially from the locally available fuzzy snapshot and the remaining from the non-failed nodes in the distributed replica. ZooKeeper, a distributed data store that offers distributed coordination as its major service is used to implement and test our research. On average, a 30 times write performance improvement has been achieved with this approach guaranteeing sufficient durability in replicated mode. Dayal Dilli, Kenneth B. Kent, Yang Wang 0006, Cheng-Zhong Xu 0001 |
CLUSTER | 3 |
| 2016 | TaskMe: multi-task allocation in mobile crowd sensingabstractTask allocation or participant selection is a key issue in Mobile Crowd Sensing (MCS). While previous participant selection approaches mainly focus on selecting a proper subset of users for a single MCS task, multi-task-oriented participant selection is essential and useful for the efficiency of large-scale MCS platforms. This paper proposes TaskMe, a participant selection framework for multi-task MCS environments. In particular, two typical multi-task allocation situations with bi-objective optimization goals are studied: (1) For FPMT (few participants, more tasks), each participant is required to complete multiple tasks and the optimization goal is to maximize the total number of accomplished tasks while minimizing the total movement distance. (2) For MPFT (more participants, few tasks), each participant is selected to perform one task based on pre-registered working areas in view of privacy, and the optimization objective is to minimize total incentive payments while minimizing the total traveling distance. Two optimal algorithms based on the Minimum Cost Maximum Flow theory are proposed for FPMT, and two algorithms based on the multi-objective optimization theory are proposed for MPFT. Experiments verify that the proposed algorithms outperform baselines based on a large-scale real-word dataset under different experiment settings (the number of tasks, various task distributions, etc.). Yan Liu 0045, Bin Guo 0001, Yang Wang 0006, Wenle Wu, Zhiwen Yu 0001, Daqing Zhang 0001 |
UbiComp | 3 |
| 2016 | On MinMax-Memory Claims for Scientific Workflows in the In-memory Cloud ComputingabstractWe propose a new concept of minmax memory claim (MMC) to achieve cost-effective workflow computations in in-memory cloud computing environments. The minmax-memory claim is defined as the minimum amount of memory required to finish the workflow without compromising its maximum concurrency. With MMC, the workflow tenants can achieve the best performance via the maximum concurrency while minimizing the cost to use the memory resources. In this paper, we present the algorithms to find the MMC for workflow computation and evaluate its value by applying it to deadlock avoidance algorithms. Yang Wang 0006, Cheng-Zhong Xu 0001, Shuibing He, Xian-He Sun |
ICDCS | 1 |
| 2016 | On Autonomous Service Migrations in the Cloud for Mobile AccessesabstractWe study the problem of autonomous service migration in the cloud to satisfy an online sequence of mobile batch-request demands in a cost-effective way. As the origins of the mobile accesses frequently change over time, this problem is particularly important for time-bounded services to achieve enhanced QoS and cost effectiveness. Moving the service closer to its client locations not only reduces the service access latency but also minimizes the network costs for service providers. However, the migration comes at costs of bulk-data transfer and service disruption, as a result, increasing the overall service costs. To gain the benefits of service migration while minimizing the service costs, we propose an efficient search-based algorithm Dmig the service migration in an autonomous way. Compared with existing algorithms, the proposed algorithm is fully distributed, symmetric, and characterized by the effective use of historical access information to perform virtual migration that overcomes the limitation of traditional local search in cost reduction. To evaluate the algorithm, we compared it with some existing algorithms, and show that the proposed algorithm exhibits better performance by adapting to the changes of mobile access patterns in a cost effective way. Yang Wang 0006, Shuibing He, Fuji Ren, Lujia Wang 0001, Cheng-Zhong Xu 0001 |
ICPADS | 1 |
| 2016 | VMBackup: an efficient framework for online virtual machine image backup and recoveryabstractSummary Although deduplication can reduce data volume for backup, it pauses the running system for the purpose of data consistency. This problem becomes severe when the target data are Virtual Machine Image (VMI), the volume of which can scale up to several gigabytes. In this paper, we propose an online framework for VM image backup and recovery, called VMBackup, which comprises three major components: (1) Similarity Retrieval that indexes chunks' fingerprints by its segment id for fast identification, (2) one‐level File‐Index that efficiently tracks file id to its content chunks in a correct order, and (3) Adjacent Storage model that places adjacent chunks of an image in the same disk partition to maximize chunk locality. The experimental results show that (1) the images of one OS serial and the same custom can share high percentage of duplicated contents, (2) variable‐length chunk partitioning is superior to fixed‐length chunk partitioning for deduplication, and (3) VMBackup, in our environment, can provide 8M/s backup throughput and 9.5M/s recovery throughput, which are only 15% and 4% less than storage systems without deduplication. Copyright © 2015 John Wiley & Sons, Ltd. Lingfang Zeng, Yang Wang 0006 |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | Black hole search in computer networks: State-of-the-art, challenges and future directions
Mengfei Peng, Wei Shi 0001, Jean-Pierre Corriveau, Richard Werner Nelem Pazzi, Yang Wang 0006 |
J. Parallel Distributed Comput. | 5 |
| 2016 | Boosting Parallel File System Performance via Heterogeneity-Aware Selective Data LayoutabstractHybrid parallel file systems (PFS) that combine HDD servers with SSD servers provide a promising solution for data intensive applications. The efficiency of a hybrid PFS relies on the data layout schemes. However, most current layout strategies are designed for homogeneous servers, which neither address the heterogeneity of servers nor the varying access patterns of applications. In this paper, we propose HAS, a novel heterogeneity-aware selective data layout scheme for hybrid PFSs. HAS alleviates inter-server load imbalance through skewing data distribution on heterogeneous servers based on their storage performance. Furthermore, to obtain the optimal performance for a specific access pattern, HAS selects one static data layout policy with lowest access cost from three typical layout candidates as the final file data layout method. To adapt to the mixed access patterns within an application, HAS uses a dynamic data layout scheme, which stores file with multiple copies, each using a different data layout policy, and then selects the copy with the lowest access cost to serve file requests. We have implemented HAS within MPICH2 and OrangeFS. Experimental results show that HAS can significantly increase the I/O throughput of hybrid PFSs, compared to existing data layout optimization methods. Shuibing He, Yang Wang 0006, Xian-He Sun |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Improving Performance of Parallel I/O Systems through Selective and Layout-Aware SSD CacheabstractParallel file systems (PFS) are widely-used to ease the I/O bottleneck of modern high-performance computing systems. However, PFSs do not work well for small requests, especially small random requests. Newer Solid State Drives (SSD) have excellent performance on small random data accesses, but also incur a high monetary cost. In this study, we propose SLA-Cache, a Selective and Layout-Aware Cache system that employs a small set of SSD-based file servers as a cache of conventional HDD-based file servers. SLA-Cache uses a novel scheme to identify performance-critical data, and conducts a selective cache admission (SCA) policy to fully utilize SSD-based file servers. Moreover, since data layout of the cache system can also largely influence its access performance, SLA-Cache applies a layout-aware cache placement scheme (LCP) to store data on SSD-based file servers. By storing data with an optimal layout requiring the lowest access cost among three typical layout candidates, LCP can further improve system performance. We have implemented SLA-Cache under the MPICH2 I/O library. Experimental results show that SLA-Cache can significantly improve I/O throughput, and is a promising approach for parallel applications. Shuibing He, Yang Wang 0006, Xian-He Sun |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | CloudSky: A Controllable Data Self-Destruction System for Untrusted Cloud Storage NetworksabstractIn cloud services, users may frequently be required to reveal their personal private information which could be stored in the cloud to used by different parts for different purposes. However, in a cloud-wide storage network, the servers are easily under strong attacks and also commonly experience software/hardware faults. As such, the private information could be under great risk in such an untrusted environment. Given that the presented personal sensitive information is usually out of user's controlin most cloud-based services, ensuring data security and privacy protection with respect to untrusted storage network has become a formidable challenge in research. To address these challenges, in this paper we propose a self-destruction system, named CloudSky, which is able to enforce the security of user privacy over the untrusted cloud in a controllable way. CloudSky exploits a key control mechanism based on the attribute-based encryption (ABE) and takes advantage of active storage networks to allow the user to control the subjective life-cycle and the access control polices of the private data whose integrity is ensured by using HMAC to cope with untrusted environments. %and thereby adapting it to the cloud in terms of both performance and security requirements. The feasibility of the system in terms of its performance and scalability is demonstrated by experiments on a real large-scale storage network. Lingfang Zeng, Yang Wang 0006, Dan Feng 0001 |
CCGRID | 2 |
| 2015 | A Heterogeneity-Aware Region-Level Data Layout for Hybrid Parallel File SystemsabstractParallel file systems (PFS) are commonly used in high-end computing systems. With the emergence of solid state drives (SSD), hybrid PFSs, which consist of both HDD and SSD servers, provide a practical I/O system solution for data-intensive applications. However, most existing PFS layout schemes are inefficient for hybrid PFSs due to their lack of awareness of the performance differences between heterogeneous servers and the workload changes between different parts of a file. This lack of recognition can result in severe I/O performance degradation. In this study, we propose a heterogeneity-aware region-level (HARL) data layout scheme to improve the data distribution of a hybrid PFS. HARL first divides a file into fine-grained, varying sized regions according to the changes of an application's I/O workload, then chooses appropriate file stripe sizes on heterogeneous servers based on the server performance for each file region. Experimental results of representative benchmarks show that HARL can greatly improve the I/O system performance. Shuibing He, Xian-He Sun, Yang Wang 0006, Antonios Kougkas, Adnan Haider |
ICPP | 3 |
| 2015 | Reusing Garbage Data for Efficient Workflow ComputationabstractHigh-performance computing (HPC) systems, including Clusters, Grids and the most recent Clouds, have emerged as attractive platforms to tackle various applications. One significant type of applications in the HPC systems is workflow computation, which has been applied in various scientific and engineering domains. The workflow computation frequently produces intermediate result files, which become garbage after being used and are usually cleaned up without making any contribution to future computation. In this paper, we argue that such garbage data could be useful in the future computation and should not be immediately cleaned up. This is because workflow computation usually contains multiple instances that may share some common data products produced in the past. This sharing scheme provides opportunities to reuse the historical data to speed-up subsequent computation and simplify re-computation due to faulty or crashed runs. To this end, we propose a novel approach, referred to as garbage data manager (GDM), for the workflow computation in HPC systems. The GDM organizes and manages the garbage data for batch schedulers to enhance the performance of subsequent computation. The essence of the GDM is to record the history of computation by constructing a dataflow graph on per instance (run) basis and set up inheritance relationships between the different instances of the same workflow, called run-tree, to achieve the data reuse. Our simulation results demonstrate that exploiting the garbage data is an effective way of improving the workflow computation. Yang Wang 0006, Menglan Hu |
Comput. J. | 1 |
| 2015 | Dataflow-Based Scheduling for Scientific Workflows in HPC with Storage ConstraintsabstractIn high-performance computing (HPC), workflow-based workloads are usually data intensive for exploratory analysis of a scientific computation problem that may involve a large parameter space. To achieve the best performance, storage resource constraint is always a pragmatic concern in reality as the potential problem space scale, especially in big data science, as well as its required dataset are ever growing to outpace any increasing rate of storage capacity. Therefore, the workflow computation in a HPC environment with finite storage resources is still a practical topic that is worthwhile studying. To this end, we propose a novel scheduling framework that enhances the scheduling policies of Versioned Name Space and Overwrite-Safe Concurrency, introduced in our earlier work, with abilities to handle the deadlock problem in workflow computation with finite storage constraints. We achieve this goal by leveraging the data dependency information of the workflow to integrate a collection of deadlock resolution algorithms into the workflow scheduler. With such integration, after extensive simulation-based studies we conclude that the enhanced scheduling policies can solve the deadlock problem introduced by the storage constraints caused by big data overflow. More interestingly, we demonstrate that our enhanced scheduling policies perform better than the cases where only pure deadlock algorithms are applied when storage is highly constrained in terms of makespan performance. Yang Wang 0006, Wei Shi 0001 |
Comput. J. | 1 |
| 2015 | Improving J9 virtual machine with LTTng for efficient and effective tracingabstractSummary The ability to observe the internal operation of the J9 virtual machine is essential for effective performance tuning. To this end, tracing is an important method, which is the action of recording events from a running system with minimum performance overhead for online or off‐line analysis. In this paper, we propose the integration of LTTng, an effective open‐source tracing toolset, with J9 to improve its tracing functions. With this integration, the tracing component is not only decoupled from the virtual machine but also performed efficiently at both user and kernel levels to achieve a high‐throughput result. To validate the integration and its impact performance, some empirical study results based on SpecJBB2005 and SQLBenchmark (supported by instrumented MariaDB) are also presented. Copyright © 2014 John Wiley & Sons, Ltd. Yang Wang 0006, Kenneth B. Kent, Graeme Johnson |
Softw. Pract. Exp. | 1 |
| 2015 | WaFS: A Workflow-Aware File System for Effective Storage Utilization in the CloudabstractWe present WaFS, a user-level file system, and a related scheduling algorithm for scientific workflow computation in the cloud. WaFS’s primary design goal is to automatically detect and gather the explicit and implicit data dependencies between workflow jobs, rather than high-performance file access. Using WaFS’s data, a workflow scheduler can either make effective cost-performance tradeoffs or improve storage utilization. Proper resource provisioning and storage utilization on pay-as-you-go clouds can be more cost effective than the uses of resources in traditional HPC systems. WaFS and the scheduler controls the number of concurrent workflow instances at runtime so that the storage is well used, while the total makespan (i.e., turnaround time for a workload) is not severely compromised. We describe the design and implementation of WaFS and the new workflow scheduling algorithm based on our previous work. We present empirical evidence of the acceptable overheads of our prototype WaFS and describe a simulation-based study, using representative workflows, to show the makespan benefits of our WaFS-enabled scheduling algorithm. Yang Wang 0006, Paul Lu, Kenneth B. Kent |
IEEE Trans. Computers | 1 |
| 2015 | Virtual Servers Co-Migration for Mobile Accesses: Online versus Off-LineabstractIn this paper, we study the problem of co-migrating a set of service replicas residing on one or more redundant virtual servers in clouds in order to satisfy a sequence of mobile batch-request demands in a cost effective way. With such a migration, we can not only reduce the service access latency for end users but also minimize the network costs for service providers. The co-migration can be achieved at the cost of bulk-data transfer and increases the overall monetary costs for the service providers. To gain the benefits of service migration while minimizing the overall costs, we propose a co-migration algorithmMigkfor multiple servers, each hosting a service replicas.Migkis a randomized algorithm with a competitive cost of$O(\frac{\gamma\, \log \,n}{\min \lbrace \frac{1}{\kappa },\frac{\mu }{\lambda \,+\,\mu }\rbrace })$to migrate$\kappa$services in a static$n$-node network where$\gamma$is the maximal ratio of the migration costs between any pair of neighbor nodes in the network, and where$\lambda$and$\mu$represent the maximum wired transmission cost and the wireless link cost respectively. For comparison, we also study this problem in its static off-line form by proposing a parallel dynamic programming (hereafter DP) based algorithm that integrates the branch&bound strategy with sampling techniques in order to approximate the optimal DP results. We validate the advantage of the proposed algorithms via extensive simulation studies using various requests patterns and cloud network topologies. Our simulation results show that the proposed algorithms can effectively adapt to mobile access patterns to satisfy the service request sequences in a cost-effective way. Yang Wang 0006, Wei Shi 0001, Menglan Hu |
IEEE Trans. Mob. Comput. | 1 |
| 2015 | XCollOpts: A Novel Improvement of Network Virtualizations in Xen for I/O-Latency Sensitive Applications on MulticoresabstractIt has long been recognized that the Credit scheduler selectively favors CPU-bound applications whereas for I/O-latency sensitive workloads, such as those related to stream-based audio/video services, it only exhibits tolerable, or even worse, unacceptable performance. The reasons behind this phenomenon are the poor understanding (to some degree) of the virtual machine scheduling as well as the network I/O virtualizations. In order to address these problems and make the system more responsive to the I/O-latency sensitive applications, in this paper, we present XCollOpts which performs a collection of novel optimizations to improve the Credit scheduler and the underlying I/O virtualizations in multicore environments, each from two perspectives. To optimize the schedule, in XCollOpts, we first pinpoint the Imbalanced Multi-Boosting problem among the cores thereby minimizing the system response time by load balancing the BOOST VCPUs. Then, we describe the Premature Preemption problem and address it by monitoring the received network packets in the driver domain and deliberately preventing it from being prematurely preempted during the packet delivery. However, these optimizations on the scheduling strategies cannot be fully exploited if the performance issues of the underlying supportive communication mechanisms are not considered. To this end, we make two further optimizations for the network I/O virtualizations, namely, Multi-Tasklet Pairs and Optimized Small Data Packet. Our empirical studies show that with XCollOpts, we can significantly improve the performance of the latency-sensitive applications at a cost of relatively small system overhead. Lingfang Zeng, Yang Wang 0006, Dan Feng 0001, Kenneth B. Kent |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2014 | Monetary-and-QoS Aware Replica Placements in Cloud-Based Storage SystemsabstractThis paper proposes a replication cost model and two greedy algorithms, named GS QoS and GS QoS C1, for replication placements in cloud-based storage systems. The model aims to minimize replication cost with full consideration of quality of user access to storage nodes. Our two algorithms employ a utility measurement to guide placement procedures. Our final experimental results show that 1) GS QoS outperforms GS QoS C1, 2) both algorithms have more economical results than those from existing greedy replica placement algorithm. Lingfang Zeng, Yang Wang 0006, Xiang Cui, Tan Wee Kiat, David Bremner, Kenneth B. Kent |
CloudCom | 3 |
| 2014 | Optimization on content service with local search in cloud of clouds
Lingfang Zeng, Yang Wang 0006 |
J. Netw. Comput. Appl. | 2 |
| 2014 | Budget-Driven Scheduling Algorithms for Batches of MapReduce Jobs in Heterogeneous CloudsabstractIn this paper, we consider task-level scheduling algorithms with respect to budget and deadline constraints for a batch of MapReduce jobs on a set of provisioned heterogeneous (virtual) machines in cloud platforms. The heterogeneity is manifested in the popular “pay-as-you-go” charging model where the service machines with different performance would have different service rates. We organize the batch of jobs as a k-stage workflow and study two related optimization problems, depending on whether the constraints are on monetary budget or on scheduling length of the workflow. First, given a total monetary budget B, by combining an in-stage local greedy algorithm (whose optimality is also proven) and dynamic programming (DP) techniques, we propose a global optimal scheduling algorithm to achieve minimum scheduling length of the workflow within O(kB2). Although the optimal algorithm is efficient when B is polynomially bounded by the number of tasks in the MapReduce jobs, the quadratic time complexity is still high. To improve the efficiency, we further develop two greedy algorithms, called Global Greedy Budget (GGB) and Gradual Refinement (GR), each adopting different greedy strategies. In GGB we extend the idea of the local greedy algorithm to the efficient global distribution of the budget with minimum scheduling length as a goal whilst in GR we iteratively apply the DP algorithm to the distribution of exponentially reduced budget so that the solutions are gradually refined. Second, we consider the optimization problem of minimizing cost when the (time) deadline of the computation D is fixed. We convert this problem into the standard Multiple-Choice Knapsack Problem via a parallel transformation. Our empirical studies verify the proposed optimal algorithms and show the efficiencies of the greedy algorithms in cost-effectiveness to distribute the budget for performance optimizations of the MapReduce workflows. Yang Wang 0006, Wei Shi 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2014 | Holistic Scheduling of Real-Time Applications in Time-Triggered In-Vehicle NetworksabstractAs time-triggered communication protocols [e.g., time-triggered controller area network (TTCAN), time-triggered protocol (TTP), and FlexRay] are widely used on vehicles, the scheduling of tasks and messages on in-vehicle networks becomes a critical issue for offering quality-of-service (QoS) guarantees to time-critical applications on vehicles. This paper studies a holistic scheduling problem for handling real-time applications in time-triggered in-vehicle networks where practical aspects in system design and integration are captured. The contributions of this paper are multifold. First, it designs a novel scheduling algorithm, referred to asUnfixed Start Time(UST) algorithm, which schedules tasks and messages in a flexible way to enhance schedulability. In addition, to tolerate assignment conflicts and further improve schedulability, it proposes two rescheduling and backtracking methods, namely,Rescheduling with Offset Modification(ROM) andBacktracking and Priority Promotion(BPP) procedures. Extensive performance evaluation studies are conducted to quantify the performance of the proposed algorithm under a variety of scenarios. Menglan Hu, Jun Luo 0001, Yang Wang 0006, Martin Lukasiewycz, Zeng Zeng |
IEEE Trans. Ind. Informatics | 3 |
| 2014 | Practical Resource Provisioning and Caching with Dynamic Resilience for Cloud-Based Content Distribution NetworksabstractContent distribution networks (CDNs) built on clouds have recently started to emerge. Compared to conventional CDNs, cloud-based CDNs have the benefit of cost efficient hosting services without owning infrastructure. However, resource provisioning and replica placement in cloud CDNs involve a number of challenging issues, mainly due to the dynamic nature of demand patterns. To deal with this dynamic nature, this paper proposes a set of novel algorithms to solve the joint problem of resource provisioning and caching (i.e., replica placement) for cloud-based CDNs with an emphasis on handling the dynamic demand patterns. Firstly, we propose a provisioning and caching algorithm framework called Differential Provisioning and Caching (DPC) algorithm, which aims to rent cloud resources to build CDNs and whereby to cache contents so that the total rental cost can be minimized while all demands are served. DPC consists of 2 steps. Step 1 first maximizes total demands supported by unexpired resources. Then, step 2 minimizes the total rental cost for new resources to serve all remaining demands. For each step we design both greedy and iterative heuristics, each with different advantages over the existing approaches. Moreover, to dynamically adjusts the placement of contents and route maps, we further propose the Caching and Request Balancing (CRB) algorithm, which is light-weight and thus can be frequently executed as a companion of DPC to maximize the total demands. Performance evaluation results are presented to demonstrate the effectiveness and competitiveness of our approaches when compared to existing algorithms. Menglan Hu, Jun Luo 0001, Yang Wang 0006, Bharadwaj Veeravalli |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2013 | On service migration in the cloud to facilitate mobile accessesabstractUsing service migration in Clouds to satisfy a sequence of mobile batch-request demands is a popular solution to enhanced QoS and cost effectiveness. As the origins of the mobile accesses are frequently changed over time, moving services closer to client locations not only reduces the service access latency but also minimizes the network cost for service providers. However, these benefits do not come without compromise. The migration comes at cost of bulk-data transfer and service disruption, as a result, increasing the overall service costs. In this paper, we study the problem of dynamically migrating a service in Clouds to satisfy a sequence of mobile batch-request demands in a cost effective way. More specifically, to gain the benefits of service migration while minimizing the increased monetary costs, we propose a search-based dynamic migration algorithm that can effectively migrate a single or multiple servers to adapt to the changes of access patterns with minimum service costs. The algorithm is characterized by effective uses of historical access information to conduct virtual moves of a set of servers as a whole under a certain condition so as to overcome the limitations of local search in cost reduction. Yang Wang 0006, Wei Shi 0001 |
CLUSTER | 1 |
| 2013 | On Scheduling Algorithms for MapReduce Jobs in Heterogeneous Clouds with Budget Constraints
Yang Wang 0006, Wei Shi 0001 |
OPODIS | 1 |
| 2013 | DDS: A deadlock detection-based scheduling algorithm for workflow computations in HPC systems with storage constraints
Yang Wang 0006, Paul Lu |
Parallel Comput. | 1 |
| 2013 | Maximizing Active Storage Resources with Deadlock Avoidance in Workflow-Based ComputationsabstractWorkflow-based workloads usually consist of multiple instances of the same workflow, which are jobs with control or data dependencies to carry out a well-defined scientific computation task, with each instance acting on its own input data. To maximize the performance, a high degree of concurrency is always achieved by running multiple instances simultaneously. However, since the amount of storage is limited on most systems, deadlock due to oversubscribed storage requests is a potential problem. To address this problem, we integrate two novel concepts with the traditional problem of deadlock avoidance by proposing two algorithms that can maximize active (not just allocated) resource utilization and minimize makespan. Our approach is based on the well-known banker's algorithm, but our algorithms make the important distinction between active and inactive resources, which is not a part of previous approaches. The central idea is to leverage the data-flow information to dynamically approximate localized maximum claim (i.e., the resource requirements of the remaining jobs of the instance) to improve either interinstance or intrainstance concurrency and still avoid deadlock. Through simulation-based studies, we show how our proposed algorithms are better than the classic banker's algorithm and the more recent Lang's algorithm in terms of makespan and active storage resource utilization. Yang Wang 0006, Paul Lu |
IEEE Trans. Computers | 1 |
| 2013 | On Data Staging Algorithms for Shared Data Accesses in CloudsabstractIn this paper, we study the strategies for efficiently achieving data staging and caching on a set of vantage sites in a cloud system with a minimum cost. Unlike the traditional research, we do not intend to identify the access patterns to facilitate the future requests. Instead, with such a kind of information presumably known in advance, our goal is to efficiently stage the shared data items to predetermined sites at advocated time instants to align with the patterns while minimizing the monetary costs for caching and transmitting the requested data items. To this end, we follow the cost and network models in [1] and extend the analysis to multiple data items, each with single or multiple copies. Our results show that under homogeneous cost model, when the ratio of transmission cost and caching cost is low, a single copy of each data item can efficiently serve all the user requests. While in multicopy situation, we also consider the tradeoff between the transmission cost and caching cost by controlling the upper bounds of transmissions and copies. The upper bound can be given either on per-item basis or on all-item basis. We present efficient optimal solutions based on dynamic programming techniques to all these cases provided that the upper bound is polynomially bounded by the number of service requests and the number of distinct data items. In addition to the homogeneous cost model, we also briefly discuss this problem under a heterogeneous cost model with some simple yet practical restrictions and present a 2-approximation algorithm to the general case. We validate our findings by implementing a data staging solver, whereby conducting extensive simulation studies on the behaviors of the algorithms. Yang Wang 0006, Bharadwaj Veeravalli, Chen-Khong Tham |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2011 | Dataflow detection and applications to workflow schedulingabstractAbstract In high‐performance computing (HPC)textitworkloads (i.e. the set of computations to be completed), the same computationalworkflowof jobs (e.g. a Pipeline, a Fork&Join, or a Lattice graph) may be applied to different input files and parameters. Each of theseworkflow instanceshas the same workflow shape, but accesses (possibly) separate input, intermediate, and output files. Therefore, the selective isolation of each workflow instance can be important for maximizing scheduling flexibility and performance. However, in practice, realizing this benefit is not obvious due to a variety of problems and constraints. For example, the unmediated interaction of different workflow instances can lead to a problem offilename conflictsbetween concurrent workflow instances overwriting common files, which, for a control‐flow driven batch scheduler, may result in either unsafe computation of the multiple instances in the same sub‐directory or storage overheads when multiple directories are used. We propose a novel approach of selectively coupling and integrating job schedulers and file systems, known as aWorkflow‐aware File System(WaFS), with two major benefits. First, separate namespaces can be constructed on a per‐instance basis to maximize the concurrency of workflow instances, despite filename conflicts, while minimizing storage overhead. Second, exploiting inferred dataflow information, trade‐offs can be made between makespan and storage overhead while maintaining correctness. Through a simulation‐based study, we have shown the potential benefits of WaFS to job concurrency and we have characterized the trade‐offs that can be made between storage overhead and performance. New scheduling policies,Versioned Namespace (VNS),Overwrite‐Safe Concurrency (OSC)and hybrids, are made possible by WaFS, with different advantages and disadvantages. Copyright © 2011 John Wiley & Sons, Ltd. Yang Wang 0006, Paul Lu |
Concurr. Comput. Pract. Exp. | 1 |
| 2009 | A distributed Key Message algorithm to optimize the communication in clusters
Yang Wang 0006 |
Parallel Comput. | 1 |
| 2005 | Towards a workflow-aware distributed versioning file system for metacomputing systemsabstractIn this paper, a novel workflow-aware distributed versioning file system, WAD-VFS is presented to overcome the shortcoming of traditional DFS and facilitate the high performance computing. Our preliminary simulation results are impressive and can hence serve as a supporting evidence of deploying WAD-VFS to our ongoing metacomputing project, Trellis system. Yang Wang 0006 |
HPDC | 1 |
| 2005 | Using Dataflow Information to Improve Inter-Workflow Instance ConcurrencyabstractThe control-flow-based design of traditional batch scheduling systems (i.e., Job A must finish before Job B is started) can constrain the concurrency in schedulers for high-performance computing (HPC) workloads. There are two main problems. First, the control-flow graph, representing the workflow, may be inherently limited in its degree of concurrency. Second, if the naming strategy of the input and output files of jobs is simplistic, there may be a filename conflict problem when multiple instances of the same workflow run concurrently. Yang Wang 0006, Paul Lu |
PDCAT | 1 |
| 2004 | Dynamic Key Messaging for Cluster ComputingabstractOver the past decades, distributed computing has been gaining popularity. It provides more computing power and memory space for parallel applications. On the other hand, such applications fight back and challenge the architectures of the distributed systems for more efficiency. To face the challenge, a key messaging (KM) scheme was proposed to realize the optimization of communication at a system architecture level in our previous papers. The contribution of KM is that it performs the optimization in both the underlying communication system and high level application model. The performance of an application is always determined by its critical path. Currently, messages along the critical path can be easily blocked by non-critical path messages, which degrade the performance. To solve this problem, KM provides an algorithm to identify the critical-path messages and optimizes them by introducing a prioritized protocol layer. Thus, these messages are served first before any low priority messages. Shorter processing time for the messages results in faster completion time of the critical path. Although KM was proved to be effective by a prototype system on the IBM SP2, there are limitations in its optimization procedures where dynamic natures of an underlying network should be considered when running parallel applications on cluster/distributed computing environments. To address the problem, a dynamic key message (DKM) algorithm is introduced. DKM takes into considerations the changes of network traffic load while tasks are running and dynamically updates the critical path on the fly. A comprehensive simulation method is adopted to evaluate the performance of this algorithm and the results show that under the same workload, DKM exhibits much more stability than the static key message (SKM) algorithm when the network traffic load changes. Yang Wang 0006, Constantine Katsinis |
NCA | 2 |
| 2004 | A space-efficient algorithm for sequence alignment with inversions and reversals
Zhi-Zhong Chen, Guohui Lin, Robert Niewiadomski, Yang Wang 0006 |
Theor. Comput. Sci. | 5 |
| 2003 | A Space Efficient Algorithm for Sequence Alignment with Inversions
Robert Niewiadomski, Yang Wang 0006, Zhi-Zhong Chen, Guohui Lin |
COCOON | 4 |