Shu Yin 0001

dblp:88/7556-1 · DBLP profile ↗
← Back
47ranked-venue papers
8as first author
14since 2021 · last 2025
0000-0001-6500-1790ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 35 · 4 first-author · 12 since 2021Computer networks · 5 · 2 first-author · 1 since 2021Security and privacy · 2 · 1 first-authorArtificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Optimizing Data Acquisitions in Multi-Robot Systems
abstract
We present ROSfs, a novel user-level file system designed to address critical data query inefficiencies in multi-robot systems (MRS). ROSfs introduces an innovative file organization model where robot data is structured as labeled sub-files, coupled with a time-indexed architecture that enables efficient querying of actively modified data. This design enables real-time cross-robot data acquisition and collaboration capabilities previously unattainable in MRS deployments. Our implementation integrates seamlessly with the Robot Operating System (ROS) and has been extensively evaluated using both physical UAV/UGV platforms and data servers. Experimental results demonstrate that ROSfs achieves a 7x reduction in online data query latency under wireless network conditions compared to conventional ROS storage methods, while simultaneously improving data freshness (Age of Information) by up to 271x. These advancements position ROSfs as a transformative solution for high-performance robotic data management in distributed systems.
Yanhao Li, Xuanjun Wen, Guancheng Li, Shu Yin 0001
SC6
2025 Optimizing parallel I/O performance in NVMe SSDs by Dynamic cache partitioning
Shu Yin 0001, Xiaojun Ruan
Perform. Evaluation2
2024 A Data Optimizer for Region-Aware Self-describing Files in Scientific Computing
abstract
Acquiring data from scientific simulations for analytical purposes is inherently challenging due to the complex and irregularly shaped regions within which the data resides, particularly when using self-describing data formats. The process of region-based data distillation becomes even more arduous when employing persistent memory or parallel file systems. To tackle this challenge, we introduce RASTER (Region-Aware Self-describing daTa optimizER), a lightweight middleware designed for region-aware data preprocessing. RASTER dynamically reorganizes data into variable groups based on regional identifiers during runtime, thereby eliminating the need for sequential searches to locate the required data. We have developed a prototype of RASTER and successfully integrated it into three distinct computing environments: a single-node server equipped with Intel® Optane™ DC persistent memory, the Huawei® OceanStor cloud storage platform, and the Sunway TaihuLight supercomputer. We then conducted a thorough evaluation of the RASTER prototype on the latter two platforms using a real-world scientific application, CESM (Community Earth System Model). Our experimental results demonstrate that RASTER enhances data acquisition performance by up to 2.83× and achieves a 2.36× speedup over conventional netCDF and the state-of-the-art ADIOS2. Additionally, RASTER significantly reduces memory usage by up to 400%, showcasing its scalability potential.
Tianyuan Wu, Guancheng Li, Shu Yin 0001, Wei Xue 0003
SoCC6
2024 Portus: Efficient DNN Checkpointing to Persistent Memory with Zero-Copy
abstract
We introduce Portus, an efficient checkpointing system for DNN models. The core of Portus is a three-level index structure and a direct RDMA datapath that enables fast check-points between GPUs and persistent memory in a serialization-free way. Portus offers a zero-copy approach between GPU and persistent memory without involving main memory and kernel crossings to underlying file systems. Portus also applies an asynchronous mechanism to hide the checkpointing overhead in the model training procedures. We integrated a Portus prototype into a high-performance AI cluster with NVIDIA®V100 and A40 GPUs and Intel®Optane™persistent memory, then evaluated its performance in both single-GPU and multi-GPU large model training scenarios. Experiment results show that compared to a state-of-the-art checkpointing system, Portus achieves up to 9.23× and 7.0× speedup in checkpointing and restoring, respectively. Portus achieves up to 2.6× higher throughput and 8× faster checkpointing operation on a large language model, GPT-22B.
Tianyuan Wu, Guancheng Li, Shu Yin 0001
ICDCS5
2024 Optimize Metadata Operations of Key-Value Store via A Flat Indexing LSM-Tree
abstract
The effectiveness of applying key-value store mechanisms to manage metadata of file systems has been demonstrated recently. However, traditional indirect metadata indexing schemes are not in concert with key-value data structures, which could degrade the performance of a KV-embedded file system due to the overhead of hierarchical path queries. In this paper, we propose FILT (Flat Indexing LSM-Tree), a lightweight file system middleware that can solve this issue by employing flat indexing. We introduce a new range rename mechanism over LSM-tree to provide fast directory renames. FILT exploits the benefits of both flat indexing and LSM-tree structure to eliminate redundant path look-ups. Our extensive performance evaluation studies show that FILT can offer up to 2.3X performance gain compared with TableFS and 56x to ordinary cloud file systems.
Chen Chen 0124, Shu Yin 0001
ISPA3
2024 Dynamic Cache Partitioning for Enhancing Parallel I/O Performance in NVMe SSDs
abstract
Solid State Drive cache, implemented as on-board shared DRAM memory, can significantly enhance 110 performance by caching frequently accessed data. Although SSD caching strategies for single 110 data flows have been extensively explored, studies on cache partitioning to optimize parallel 110 in an SSD are scarce. In this paper, we present a novel dynamic cache partitioning approach designed to improve overall performance of multi-parallel 110 data flows by minimizing per-formance degradation of cache pollution and resource contention. By dynamically adjusting cache partition sizes for each data flow by considering cache sensitivity on performance, our strategy seeks to determine the optimal cache partition sizes to maximize overall 110 throughput. We implemented the strategy in the SSD simulator MQSim and evaluated its performance using various synthetic and real-world workloads. Our experimental results indicate that our dynamic cache partitioning strategy achieves an overall throughput increase of up to 33.22 % compared to shared cache methods and outperforms static cache partitioning strategies by up to 21.19%.
Guancheng Li, Songhui Cao, Shu Yin 0001, Xiaojun Ruan
NAS4
2023 RBC: A bandwidth controller to reduce write-stalls and tail latency
abstract
We present RBC (Request Bandwidth Controller for LSM-tree), a controller that manages LSM-tree read and write operations and optimizes the requests for flash-based storage devices. RBC minimizes the number of write-stalls and reduces tail latency by precisely measuring the requests for solid-state disk bandwidth from various components within the LSM-tree. We have successfully implemented RBC on the state-of-the-art LSM-tree storage system, RocksDB. The experimental results show that using YSCB to evaluate RBC leads to a significant decrease in write-stalls and tail latency. Additionally, RBC provides higher write throughput (up to 1.52x) and high read throughput (up to 1.96x), while reducing read and write amplification by 55% and 37%, respectively. Furthermore, we observed a reduction of approximately 50% in tail latency with the use of RBC.
Zepeng Wang 0007, Shu Yin 0001
ICPP2
2023 Critique of "A Parallel Framework for Constraint-Based Bayesian Network Learning via Markov Blanket Discovery" by SCC Team From ShanghaiTech University
abstract
In SC20, (Srivastava et al. 2020) proposed a Parallel Framework forBayesianLearning, or ramBLe, for short, which is a highly parallel and efficient framework for learning the structure of Bayesian Networks (BNs) from samples,There was a discrepancy in Bibliography in the PDF and the source file. We have followed the source file. ?> particularly large genome-scale networks. As part of our participation in the SC21 Student Cluster Competition, our task was to verify conclusions from the original work (Srivastava et al. 2020). Here we present the outcome of our experiments, which were performed on a four-node cluster from the Oracle Cloud HPC platform. We reproduce the numerical results from (Srivastava et al. 2020), namely the algorithm's performance and scaling behavior using MPI and different Python and Boost libraries on the Oracle cloud.
Guancheng Li, Songhui Cao, Chuyi Zhao, Siyuan Zhang 0001, Yuchen Ji, Haotian Jing, Yiwei Yang 0002, Shu Yin 0001
IEEE Trans. Parallel Distributed Syst.10
2022 User-level parallel file system: Case studies and performance optimizations
abstract
Abstract User‐level file systems are usually adopted to bridge the gap between efficacy and efficiency of file system developments for new applications' I/O demands. And the widely known user‐space file system framework, FUSE, is commonly utilized to deployed user‐level file systems. This article first uses a popular stack‐able file system as a case study to exam how FUSE affects I/O performance. Based on the testing and analytical results, this article then presents SHC, an implementation method to implement a user‐level file system without FUSE intervention. Experimental results indicate that SHC improves write bandwidth by up to 5.6x compared with that of FUSE and present leading superiority on read cases.
Yanliang Zou, Chen Chen 0124, Tongliang Deng, Jian Zhang 0070, Xiaomin Zhu 0001, Si Chen 0009, Shu Yin 0001
Concurr. Comput. Pract. Exp.7
2022 ARSpy: Breaking Location-Based Multi-Player Augmented Reality Application for User Location Tracking
abstract
Augmented reality (AR) applications that overlay the perception of the real world with digitally generated information are on the cusp of commercial viability. AR has appeared in several commercial platforms like Microsoft HoloLens and smartphones. They extend the user experience beyond two dimensions and supplement the normal 3D world of a user. A typical location-based multi-player AR application works through a three-step process, wherein the system collects sensory data from the real world, identifies objects based on their context, and finally, renders information on top of senses of a user. However, because these AR applications frequently exchange data with users, they have exposed new individual and public safety issues. In this paper, we develop ARSpy, a user location tracking system solely based on network traffic information of the user, and we test it on location-based multi-player AR applications. We demonstrate the effectiveness and efficiency of the proposed scheme via real-world experiments on 12 volunteers and show that we could obtain the geolocation of any target with high accuracy. We also propose three mitigation methods to mitigate these side channel attacks. Our results reveal a potential security threat in current location-based multi-player AR applications and serve as a critical security reminder to a vast number of AR users.
Jiacheng Shang, Si Chen 0009, Jie Wu 0001, Shu Yin 0001
IEEE Trans. Mob. Comput.4
2022 Reproducibility: Performance Evaluation of MemXCT on Azure CycleCloud Platform
abstract
Memory-Centric X-ray Computational Tomography(CT) is an iterative reconstruction technique that trades compute simplifications with higher memory accesses. MemXCT implements a sparse matrix-vector multiplication(SpMV) with multi-stage buffering and two-level pseudo-Hilbert ordering for optimization. Motivated by the need to validate conclusions from previous work, we reproduce the numerical results, the algorithm’s performance, and the scaling behavior of the algorithms as the number of MPI processes increases on Azure. Digital artifacts from these experiments are available at: 10.5281/zenodo.5598108
Yixuan Meng, Tianyuan Wu, Yiwei Yang 0002, Shu Yin 0001
IEEE Trans. Parallel Distributed Syst.7
2021 ADA: An Application-Conscious Data Acquirer for Visual Molecular Dynamics
abstract
Visual molecular dynamics (VMD) has been widely used by numerous molecular dynamics (MD) applications to animate and analyze the trajectory of an MD simulation. One challenge faced by domain scientists, however, is how to filter out inactive data (i.e., data irrelevant to the subject) from the enormous output of an MD simulation. To solve it, we propose ADA (application-conscious data acquirer), a light-weight file system middleware that can perform an application-conscious data pre-processing. It provides host CPUs with only the data needed instead of an entire raw dataset. Next, we implement an ADA prototype, which is then integrated into three computing platforms: an SSD server, a nine-node OrangeFS storage cluster, and a fat-node server with 1 TB memory. Further, we evaluate ADA by running a computational biology application on the three platforms. Our experimental results show that compared to a traditional file system an ADA-assisted file system improves data processing turnaround time by up to 13.4x and reduces memory usage for data rendering by up to 2.5x. Besides, ADA allows the 1TB memory server to render more than 2x the VMD graphs while cutting energy consumption by 3x.
Hanpei Wu, Tongliang Deng, Yanliang Zou, Shu Yin 0001, Si Chen 0009, Tao Xie 0004
ICPP4
2021 Improving bioinformatics applications performance via active storage systems
Zhiyang Ding, Xiao Qin 0001, Shu Yin 0001
CCF Trans. High Perform. Comput.3
2021 Refocusable Gigapixel Panoramas for Immersive VR Experiences
abstract
There have been significant advances in capturing gigapixel panoramas (GPP). However, solutions for viewing GPPs on head-mounted displays (HMDs) are lagging: an immersive experience requires ultra-fast rendering while directly loading a GPP onto the GPU is infeasible due to limited texture memory capacity. In this paper, we present a novel out-of-core rendering technique that supports not only classic panning, tilting, and zooming but also dynamic refocusing for viewing a GPP on HMD. Inspired by the network package transmission mechanisms in distributed visualization, our approach employs hierarchical image tiling and on-demand data updates across the main and the GPU memory. We further present a multi-resolution rendering scheme and a refocused light field rendering technique based on RGBD GPPs with minimal memory overhead. Comprehensive experiments demonstrate that our technique is highly efficient and reliable, able to achieve ultra-high frame rates ( fps) even on low-end GPUs. With an embedded gaze tracker, our technique enables immersive panorama viewing experiences with unprecedented resolutions, field-of-view, and focus variations while maintaining smooth spatial, angular, and focal transitions.
Wentao Lyu, Yingliang Zhang, Anpei Chen, Minye Wu, Shu Yin 0001, Jingyi Yu 0001
IEEE Trans. Vis. Comput. Graph.6
2020 FILT: Optimizing KV-Embedded File Systems through Flat Indexing
abstract
The effectiveness of applying key-value store mechanisms to manage metadata of file systems has been demonstrated recently. However, traditional indirect metadata indexing schemes are not in concert with modern key-value data structures, which could degrade the performance of a KV-embedded file system due to the overhead of hierarchical path queries. In this paper, we propose FILT, a proof-of-concept file system middleware that can solve this problem by employing flat indexing. FILT exploits the benefits of both flat indexing and LSM-tree structure to eliminate redundant path lookups. Our extensive performance evaluation studies show that FILT can offer up to 5.8x performance gain compared with sophisticated local file systems.
Chen Chen 0124, Tongliang Deng, Jian Zhang 0070, Yanliang Zou, Xiaomin Zhu 0001, Shu Yin 0001
ICDCS6
2020 BORA: a bag optimizer for robotic analysis
abstract
We present BORA (Bag Optimizer for Robotic Analysis), a file system middleware that optimizes the acquisition of bags, which are specially formatted files used to store timestamped ROS (robot operating system) messages. BORA sits between ROS and an existing file system to conduct semantic-aware data pre-processing. In particular, it categorizes ROS bag data into multiple groups with each having a distinct label. BORA predigests data index constructions and reduces file open time via a hash-based label management scheme. It is also capable of providing ROS analytic applications with only data needed without a sequence of data searching and locating operations. We implement a BORA prototype, which is then integrated into three computing platforms: a single-node server, a four-node PVFS storage cluster, and a Tianhe-1A Supercomputer storage subsystem. Next, we evaluate the BORA prototype on the three platforms using four real-world ROS applications. Our experimental results show that compared to a traditional bag management scheme BORA improves data acquisition performance by up to 11x. In addition, it offers up to 10x data acquisition performance improvement and 3,100x bags open improvement under a swarm robotics data analysis scenario where data is retrieved across multiple bags simultaneously.
Jian Zhang 0070, Tao Xie 0004, Yuzhuo Jing, Guanzhou Hu, Si Chen 0009, Shu Yin 0001
SC7
2019 Virtual machine allocation and migration based on performance-to-power ratio in energy-efficient clouds
Xiaojun Ruan, Haiquan Chen 0001, Yun Tian 0004, Shu Yin 0001
Future Gener. Comput. Syst.4
2018 PEA: Parallel Evolutionary Algorithm by Separating Convergence and Diversity for Large-Scale Multi-Objective Optimization
abstract
Running evolutionary algorithms in parallel is an intuitive way to speed up the process of solving large-scale multi-objective optimization problems, which have hundreds or thousands of decision variables. However, the framework of the existing multi-objective evolutionary algorithms seriously limits their parallelization. During each iteration, the environmental selection operators present in the existing framework need to collect and compare all the candidate solutions to balance the convergence and diversity, thus dividing the whole evolutionary process into a series of dependent sub-processes and resulting in frequent data transmission. To address this issue, we propose a novel parallel framework that separates the environmental selection operator from the entire evolutionary process, evidently removing the dependencies among sub-processes and reducing the data transmission. On the basis of the parallel framework, a new parallel evolutionary algorithm, namely PEA, is designed. In PEA, the convergence is achieved by a series of independent sub-populations, and the diversity is merely emphasized at the converged solutions from each subpopulation, which is helpful for avoiding that the environmental selection operator limits the parallelization of the algorithm. Moreover, a new environmental selection strategy is proposed to improve the diversity without considering the convergence. To assess the performance of the proposed PEA, we compare it with five representative multi-objective evolutionary algorithms in terms of both the convergence and diversity. The performance of the parallel framework is also analyzed by comparing with two existing parallel models. The experimental results demonstrate the superiority of the proposed parallel algorithms in terms of the convergence, diversity, and speedup.
Huangke Chen, Xiaomin Zhu 0001, Witold Pedrycz, Shu Yin 0001, Guohua Wu 0001
ICDCS4
2018 DuoFS: A Hybrid Storage System Balancing Energy-Efficiency, Reliability, and Performance
abstract
As the Energy Wall and the Reliability Wall become unavoidable, it is a demanding and challenging task to reduce energy consumption in large-scale storage systems in modern data centers while retaining acceptable systems reliability. We propose a reliable energy-efficient storage system called DuoFS, which aims at balancing the energy efficiency, the reliability and the performance of parallel storage systems by seamlessly integrating one HDD-based file system and one SSD-based file system. At the heart of the DuoFS is a transformative middleware layer that dispatches files to the one of the two independent parallel file systems based on the files' I/O access popularity. By replicating popular files to the SSD-based file system and pushing the HDD-based file system into the low-power mode under light workload conditions, DuoFS can reduce significant energy consumption, avoid major factors that harm the storage systems reliability, and extract SSDs good I/O performance. Experimental results show that the DuoFS system saves up to 40% of energy, achieves up to 50% better I/O performance while only sacrificing less than 15% of the system's reliability.
Shu Yin 0001, Bing Jiao, Xiaomin Zhu 0001, Xiaojun Ruan, Si Chen 0009, Zhuo Tang
PDP1
2017 DuoFS: An Attempt at Energy-Saving and Retaining Reliability of Storage Systems
abstract
As issues of the Energy Wall and the Reliability Wall become unavoidable, it is a demanding and challenging task to reduce energy consumption in large-scale storage systems in modern data centres while retaining acceptable systems reliability. Most energy conservation techniques inevitably have adverse impacts on the parallel disk systems. To address the reliability issues of energy-efficient parallel storage systems, we propose a reliable energy-efficient storage system called DuoFS, which aims at improving both energy efficiency and reliability of parallel storage systems by seamlessly integrating HDDs and SSDs. With the help of the middleware layer, DuoFS can distribute popular data to SSD-based nodes and put HDD-based nodes into the low-power mode under light workload conditions without modification of the parallel systems.
Bing Jiao, Xiaomin Zhu 0001, Xiaojun Ruan, Xiao Qin 0001, Shu Yin 0001
ICDCS5
2016 Unequal Failure Protection Coding Technology for Cloud Storage Systems
abstract
In recent years, erasure codes have become the de facto standard for data protection of large scale distributed cloud storage systems at the cost of an affordable storage overhead. While traditional erasure coding schemes, such as Reed-Solomon codes, suffer from high reconstruction cost and I/Os. The recent past has seen a plethora of efforts to optimize the tradeoff between the reconstruction cost, I/Os and storage overhead. Quietly different from all prior studies, in this paper, our erasure coding technology makes the first attempt to take advantage of the unequal failure rates across the disks/nodes to optimize the reconstruction performance and system reliability. Specifically, our proposed technology, the Unequal Failure Protection based Local Reconstruction Code (UFP-LRC) divides the data blocks into several unequal-sized groups with local parities, assigning the data blocks stored on more failure-prone disks/nodes into the smaller-sized group, so as to provide unequal failure protection for each group. In this way, by exploiting the nonuniform local parity degrees, the proposed UFP-LRC enables the data blocks that are stored on more failure-prone disks/nodes to tolerate a greater number of failures while suffer from less repair cost than others, leading to a substantial improvement of overall repair performance and reliability for cloud storage system. We perform numerical analysis and build a prototype storage system to verify our approach. The analytical results show that the UFPLRC technique gradually outperforms LRC along the increase of failure rate ratio. Also, extensive experiments show that, when compared to LRC, UFP-LRC is able to achieve a 10% to 13% improvement in throughput, and a 8% to 12% reduction in decoding latency, while retaining a comparable overall reliability.
Yupeng Hu 0004, Yonghe Liu, Wenjia Li, Nong Xiao 0001, Zheng Qin 0001, Shu Yin 0001
CLUSTER6
2016 RESS: A Reliable Energy-Efficient Storage System
abstract
Extracting high I/O performance from parallel file systems is no longer the only goal in modern data centres. As issues of the Energy Wall and the Reliability Wall become unavoidable, it is a demanding and challenging task to reduce energy consumption in large-scale storage systems in modern data centres while retaining acceptable systems reliability. Most energy conservation techniques inevitably have adverse impacts on the parallel disk systems. To address the reliability issues of energy-efficient parallel storage systems, we propose a reliable energy-efficient storage system called RESS, which aims at improving both energy efficiency and reliability of parallel storage systems by seamlessly integrating HDDs and SSDs. At the heart of the RESS is a transformative middleware layer, which reorganizes the I/O workload for the underlying parallel file systems. With the help of the middleware layer, RESS can distribute popular data to SSDs and put HDDs into the low-power mode under light workload conditions without modification of the parallel systems.
Shu Yin 0001, Zhaoyu Xiao, Kenli Li 0001, Jianzhong Huang 0001, Xiaojun Ruan, Xiaomin Zhu 0001, Xiao Qin 0001
ICPADS1
2015 REED: A Reliable Energy-Efficient RAID
abstract
Recent studies indicate that the energy cost and carbon footprint of data centers have become exorbitant. It is a demanding and challenging task to reduce energy consumption in large-scale storage systems in modern data centers. Most energy conservation techniques inevitably have adverse impacts on parallel disk systems. To address the reliability issues of energy-efficient parallel disks, we propose a reliable energy-efficient RAID system called REED, which aims at improving both energy efficiency and reliability of RAID systems by seamlessly integrating HDDs and SSDs. At the heart of REED is a high-performance cache mechanism powered by SSDs, which are serving popular data. Under light workload conditions, REED spins down HDDs into the low-power mode, thereby offering energy conservation. Importantly, during an I/O access turbulence (i.e., I/O load is dynamically and frequently changing), REED is conducive to reducing the number of disk power-state transitions by keeping HDDs in the low-power mode while serving requests with SSDs. We build a model to quantitatively show that REED is capable of improving the reliability of energy-efficient RAIDs. We implement the REED prototype in a real-world RAID-0 system. Our experimental results demonstrate that REED improves the energy-efficiency of conventional RAID-0 by up to 73% while maintaining good reliability.
Shu Yin 0001, Xuewu Li, Kenli Li 0001, Jianzhong Huang 0001, Xiaojun Ruan, Xiaomin Zhu 0001, Wei Cao 0006, Xiao Qin 0001
ICPP1
2015 A cost-optimal parallel algorithm for the 0-1 knapsack problem and its performance on multicore CPU and GPU implementations
Kenli Li 0001, Jing Liu 0032, Lanjun Wan, Shu Yin 0001, Keqin Li 0001
Parallel Comput.4
2015 Exploiting Pipelined Encoding Process to Boost Erasure-Coded Data Archival
abstract
This paper addresses an issue of erasure-coded data archival, where (k + r; k) erasure codes are employed to archive rarely accessed replicas. The traditional synchronous encodingprocess neither leverages the existence of replicas, nor handles encoding operations in a decentralized manner. To overcome these drawbacks, we exploit pipelined encoding processes to boost the data archival performance on storage clusters. First, we propose two data layouts called [D + P]cdand [3X]cdby applying a chained-declustering mechanism to both Mirrored RAID-5 and triplication redundancy groups. Second, in light of the [D + P]cdand [3X]cdlayouts, we design two archiving schemes named DP and 3X, which exhibit the following three salient features: (i) exploiting data locality-two or three local blocks are read by each involved node for encoding; (ii) decentralized computation load-encoding operations are distributed among k nodes; and (iii) parallel archival processing-two or three encoding pipelines are simultaneously deployed to generate parity blocks. We implement both the DPand 3X schemes and three existing solutions (i.e., SynE, DE, and RapidRAID) in a real-world storage cluster. Experimental results show that our archival schemes outperform the other three solutions in terms of archiving time by a factor of at least 3.41 in a nine-node storage cluster. The experiments strongly indicate that the performance bottleneck of SynE lies in its block-receiving stage; it is disk I/O rather than network traffic that dominates archiving time for both the DE and RapidRAID schemes.
Jianzhong Huang 0001, Yanqun Wang, Xiao Qin 0001, Xianhai Liang, Shu Yin 0001, Changsheng Xie 0001
IEEE Trans. Parallel Distributed Syst.5
2014 Real-Time Tasks Oriented Energy-Aware Scheduling in Virtualized Clouds
abstract
Energy conservation is a major concern in cloud computing systems because it can bring several important benefits such as reducing operating costs, increasing system reliability, and prompting environmental protection. Meanwhile, power-aware scheduling approach is a promising way to achieve that goal. At the same time, many real-time applications, e.g., signal processing, scientific computing have been deployed in clouds. Unfortunately, existing energy-aware scheduling algorithms developed for clouds are not real-time task oriented, thus lacking the ability of guaranteeing system schedulability. To address this issue, we first propose in this paper a novel rolling-horizon scheduling architecture for real-time task scheduling in virtualized clouds. Then a task-oriented energy consumption model is given and analyzed. Based on our scheduling architecture, we develop a novel energy-aware scheduling algorithm named EARH for real-time, aperiodic, independent tasks. The EARH employs a rolling-horizon optimization policy and can also be extended to integrate other energy-aware scheduling algorithms. Furthermore, we propose two strategies in terms of resource scaling up and scaling down to make a good trade-off between task’s schedulability and energy conservation. Extensive simulation experiments injecting random synthetic tasks as well as tasks following the last version of the Google cloud tracelogs are conducted to validate the superiority of our EARH by comparing it with some baselines. The experimental results show that EARH significantly improves the scheduling quality of others and it is suitable for real-time task scheduling in virtualized clouds.
Xiaomin Zhu 0001, Laurence T. Yang, Huangke Chen, Ji Wang 0002, Shu Yin 0001, Xiaocheng Liu
IEEE Trans. Cloud Comput.5
2014 MINT: A Reliability Modeling Frameworkfor Energy-Efficient Parallel Disk Systems
abstract
The Popular Disk Concentration (PDC) technique and the Massive Array of Idle Disks (MAID) technique are two effective energy conservation schemes for parallel disk systems. The goal of PDC and MAID is to skew I/O load toward a few disks so that other disks can be transitioned to low power states to conserve energy. I/O load skewing techniques like PDC and MAID inherently affect reliability of parallel disks, because disks storing popular data tend to have high failure rates than disks storing cold data. To study reliability impacts of energy-saving techniques on parallel disk systems, we develop a mathematical modeling framework called MINT. We first model the behaviors of parallel disks coupled with power management optimization policies. We make use of data access patterns as input parameters to estimate each disk's utilization and power-state transitions. Then, we derive each disk's reliability in terms of annual failure rate from the disk's utilization, age, operating temperature, and power-state transition frequency. Next, we calculate the reliability of PDC and MAID parallel disk systems in accordance with the annual failure rate of each disk in the systems. Finally, we use real-world trace to validate out MINT model. Validation result shows that the behaviors of PDC and MAID which are modeled by MINT have a similar trend as that in the real-world.
Shu Yin 0001, Xiaojun Ruan, Adam Manzanares, Xiao Qin 0001, Kenli Li 0001
IEEE Trans. Dependable Secur. Comput.1
2013 Interrelation analysis of celestial spectra data using constrained frequent pattern trees
Jifu Zhang, Xujun Zhao, Sulan Zhang, Shu Yin 0001, Xiao Qin 0001
Knowl. Based Syst.4
2012 Multicore-Enabled Smart Storage for Clusters
abstract
We present a multicore-enabled smart storage for clusters in general and MapReduce clusters in particular. The goal of this research is to improve performance of data-intensive parallel applications on clusters by offloading data processing to multicore processors in storage nodes. Compared with traditional storage devices, next-generation disks will have computing capability to reduce computational load of host processors or CPUs. With the advance of processor and memory technologies, smart storage systems are promising devices to perform complex on-disk operations. The proposed smart storage system can avoid moving a huge amount of data back and forth between storage nodes and computing nodes in a cluster. To enhance the performance of data-intensive applications, we have designed a smart storage system called Multicore-enabled Smart Storage (McSD), in which a multicore processor is integrated in storage nodes. We have implemented a programming framework for data-intensive applications running on a computing system coupled with McSD. The programming framework aims at balancing load between computing nodes and multicore-enabled smart storage nodes. To fully utilize multicore processors in smart storage nodes, we have implemented the MapReduce model for McSDs to handle parallel computing on a cluster. A prototype of McSD has been implemented in a cluster connected by Gigabit Ethernet. Experimental results show that McSD can significantly reduce the execution times of three real-world applications - word count, string matching, and matrix multiplication. We demonstrate that the integration of multicore-enabled smart storage with MapReduce clusters is a promising approach to improving overall performance of data-intensive applications on clusters.
Zhiyang Ding, Xunfei Jiang, Shu Yin 0001, Xiao Qin 0001, Kai-Hsiung Chang, Xiaojun Ruan, Mohammed I. Alghamdi, Meikang Qiu
CLUSTER3
2012 ES-MPICH2: A Message Passing Interface with Enhanced Security
abstract
An increasing number of commodity clusters are connected to each other by public networks, which have become a potential threat to security sensitive parallel applications running on the clusters. To address this security issue, we developed a Message Passing Interface (MPI) implementation to preserve confidentiality of messages communicated among nodes of clusters in an unsecured network. We focus on M PI rather than other protocols, because M PI is one of the most popular communication protocols for parallel computing on clusters. Our MPI implementation-called ES-MPICH2-was built based on MPICH2 developed by the Argonne National Laboratory. Like MPICH2, ES-MPICH2 aims at supporting a large variety of computation and communication platforms like commodity clusters and high-speed networks. We integrated encryption and decryption algorithms into the MPICH2 library with the standard MPI interface and; thus, data confidentiality of MPI applications can be readily preserved without a need to change the source codes of the MPI applications. MPI-application programmers can fully configure any confidentiality services in MPICHI2, because a secured configuration file in ES-MPICH2 offers the programmers flexibility in choosing any cryptographic schemes and keys seamlessly incorporated in ES-MPICH2. We used the Sandia Micro Benchmark and Intel MPI Benchmark suites to evaluate and compare the performance of ES-MPICH2 with the original MPICH2 version. Our experiments show that overhead incurred by the confidentiality services in ES-MPICH2 is marginal for small messages. The security overhead in ES-MPICH2 becomes more pronounced with larger messages. Our results also show that security overhead can be significantly reduced in ES-MPICH2 by high-performance clusters. The executable binaries and source code of the ES-MPICH2 implementation are freely available at http:// www.eng.auburn.edu/~xqin/software/es-mpich2/.
Xiaojun Ruan, Qing Yang 0003, Mohammed I. Alghamdi, Shu Yin 0001, Xiao Qin 0001
IEEE Trans. Dependable Secur. Comput.4
2011 Reliability analysis of an energy-aware RAID system
abstract
We develop a mathematical model - MREED - to quantitatively evaluate the failure rate of energy-efficient parallel storage systems. The Power-Aware Redundant Array of Inexpensive Disk (PARAID) aims to reduce energy use of commodity server-class disks without specialized hardware. The goal of PARAID is to skewed striping pattern to adapt to the system load by changing the number of powered disks. By spinning down disks during light workloads, PARAID can reduce power consumption, while still meeting performance demands. We show that MREED can be used to estimate a five-disk PARAID-0 system. We validate the accuracy of MREED using the DiskSim simulator. Our approach shows that MREED can rely on file access pattern to estimate system utilization correctly. Furthermore, even thought PARAID may achieve reasonable reliability, our model shows that PARAID's reliability is affected by data locality.
Shu Yin 0001, Yun Tian 0004, Jiong Xie, Xiao Qin 0001, Mohammed I. Alghamdi, Xiaojun Ruan, Meikang Qiu
IPCCC1
2011 Secure Fragment Allocation in a Distributed Storage System with Heterogeneous Vulnerabilities
abstract
There is a growing demand for large-scale distributed storage systems to support resource sharing and fault tolerance. Although heterogeneity issues of distributed systems have been widely investigated, little attention has yet been paid to security solutions designed for distributed storage systems with heterogeneous vulnerabilities. This fact motivates us to investigate a fragment allocation scheme called S-FAS to improve security of a distributed system where storage sites have a wide variety of vulnerabilities. In the S-FAS approach, we integrate file fragmentation with the secret sharing technique in a distributed storage system with heterogeneous vulnerabilities. Storage sites in a distributed systems are classified into a variety of different server types based on vulnerability characteristics. Given a file and a distributed system, S-FAS allocates fragments of the file to as many different types of nodes as possible in the system. Data confidentiality is preserved because fragments of a file are allocated to multiple storage nodes. We develop storage assurance and dynamic assurance models to evaluate the quality of security offered by S-FAS. Analysis results show that fragment allocations made by S-FAS lead to enhanced security because of the consideration of heterogeneous vulnerabilities in distributed storage systems.
Yun Tian 0004, Shu Yin 0001, Jiong Xie, Ji Zhang 0002, Xiao Qin 0001, Mohammed I. Alghamdi, Meikang Qiu
NAS2
2011 Quality of security adaptation in parallel disk systems
Mais Nijim, Ziliang Zong, Shu Yin 0001, Kiranmai Bellam, Xiao Qin 0001
J. Parallel Distributed Comput.3
2011 PRE-BUD: Prefetching for energy-efficient parallel I/O systems with buffer disks
abstract
A critical problem with parallel I/O systems is the fact that disks consume a significant amount of energy. To design economically attractive and environmentally friendly parallel I/O systems, we propose an energy-aware prefetching strategy (PRE-BUD) for parallel I/O systems with disk buffers. We introduce a new architecture that provides significant energy savings for parallel I/O systems using buffer disks while maintaining high performance. There are two buffer disk configurations: (1) adding an extra buffer disk to accommodate prefetched data, and (2) utilizing an existing disk as the buffer disk. PRE-BUD is not only able to reduce the number of power-state transitions, but also to increase the length and number of standby periods. As such, PRE-BUD conserves energy by keeping data disks in the standby state for increased periods of time. Compared with the first prefetching configuration, the second configuration lowers the capacity of the parallel disk system. However, the second configuration is more cost-effective and energy-efficient than the first one. Finally, we quantitatively compare PRE-BUD with both disk configurations against three existing strategies. Empirical results show that PRE-BUD is able to reduce energy dissipation in parallel disk systems by up to 50 percent when compared against a non-energy aware approach. Similarly, our strategy is capable of conserving up to 30 percent energy when compared to the dynamic power management technique.
Adam Manzanares, Xiao Qin 0001, Xiaojun Ruan, Shu Yin 0001
ACM Trans. Storage4
2011 A Message-Scheduling Scheme for Energy Conservation in Multimedia Wireless Systems
abstract
Reducing power consumption of wireless networks has become a major goal in designing modern multimedia wireless systems. In an effort to reduce power consumption, this paper addresses the issue of scheduling real-time messages in multimedia wireless networks subject to both timing and power constraints. A power-consumption model is introduced to calculate power-consumption rates in accordance with message-transmission rates. Next, a new message-scheduling scheme called Power-aware Real-time Message (PARM) is developed to generate message-transmission schedules that minimize power consumption of multimedia wireless-network interfaces and the probability of missing deadlines for real-time messages. With a power-aware scheduling policy in place, the proposed PARM scheme is very energy-efficient. Experimental results based on a wide variety of synthetic workloads and eight real-world applications show that PARM significantly reduces energy dissipation while maintaining low missed rates. PARM reduces power consumption of data transmissions by up to 99.4% (with an average of 86.7%) for synthetic network traffic and saves energy by up to 60.0% (with an average of 34.1%) in the eight real-world applications.
Xiaojun Ruan, Shu Yin 0001, Adam Manzanares, Mohammed I. Alghamdi, Xiao Qin 0001
IEEE Trans. Syst. Man Cybern. Part A2
2010 Improving Energy Efficiency and Security for Disk Systems
abstract
Improving security and minimizing power consumption are crucial for large-scale data storage systems. Although a handful of studies have been focused on data security and energy efficiency, most of the existing approaches have concentrated on only one of these two metrics. In this paper, we present a new approach to integrating power optimization with security services to enhance the security of energy-efficient large-scale storage systems. In our approach, we make use of the dynamic speed control for power management technique, or DRPM, to conserve energy in secure storage systems. In this study we develop two ways of integrating confidentiality services with the dynamic disk speed control technique. The first strategy - security aggressive in nature - is focused on the improvement of storage system security with less emphasis on energy conservation. The second strategy gives higher priority to energy conservation as opposed to the security optimization. Our experimental results show that the energy-aggressive approach provides better energy savings than the security-aggressive approach. However, the quality of security achieved by the security-aggressive scheme is higher than that of the energy-aggressive approach. Moreover, the empirical results show that energy savings yielded by the two approaches become more pronounced when the data size is increased. The findings illustrate that the response time of the security-aggressive approach is more sensitive to data size than that of the energy-aggressive scheme.
Shu Yin 0001, Mohammed I. Alghamdi, Xiaojun Ruan, Mais Nijim, Ashwin Tamilarasan, Ziliang Zong, Xiao Qin 0001
HPCC1
2010 Energy Efficient Prefetching with Buffer Disks for Cluster File Systems
abstract
Energy efficient computing is becoming increasingly important as the scale of parallel computing systems is expanding. As the processing power of parallel computing systems has been incremented there has been an increased demand for large scale storage systems to store the output of these parallel computing systems. Data centers are growing at an enormous pace and it is important to investigate a means of managing the energy efficiency of large scale parallel storage systems. To address these issues we introduce EEVFS (Energy Efficient Virtual File System), which is able to manage data placement and disk states to help improve the energy efficiency of a parallel disk system. EEVFS places data on the storage disks in an energy efficient layout and attempts to predict when each disk will be idle for a large period of time, facilitating a state transition into the standby state. EEVFS should also maintain relatively high performance, so we have built a load balancing policy into the data partitioning of EEVFS. The implementation architecture and measured results are presented to demonstrate the energy efficiency and performance characteristics of EEVFS.
Adam Manzanares, Xiaojun Ruan, Shu Yin 0001, Jiong Xie, Zhiyang Ding, Yun Tian 0004, James Majors, Xiao Qin 0001
ICPP3
2010 ES-MPICH2: A Message Passing Interface with enhanced security
abstract
In largely distributed clusters, computing nodes are geographically deployed in various computing sites. Information processed in a distributed cluster is shared among a group of distributed processes or users by virtue of messages passing protocols (e.g. message passing interface - MPI) running on the Internet. Because of the open accessible nature of the Internet, data encryption for these large-scale distributed clusters becomes a non-trivial and challenging problem. To address this issue, we enhanced the security of the MPI (Message Passing Interface) protocol by encrypting and decrypting messages sent and received among computing nodes. In this study we focused on MPI rather than other protocols because MPI is one of the most popular communication protocols for cluster computing environments. From among a variety of MPI implementations, we picked MPICH2 developed by the Argonne National Laboratory. The design goal of MPICH2 - a widely used MPI implementation - is to combine portability with high performance. We integrated encryption algorithms into the MPICH2 library so that data confidentiality of MPI applications could be readily preserved without a need to change the source codes of the MPI applications. since we provide a security enhanced MPI-library with the standard MPI interfact, data communications of a conventional MPI program can be secured without converting the program into the corresponding secure version. We used Sandia Micro Benchmark and Intel MPI Benchmarks to evaluate and compared the performance of original MPICH2 and Enhanced Security MPICH2. According to the performance evaluation, ES-MPICH2 provides secured Message Passing Interface by sacrificing reasonable system performance.
Xiaojun Ruan, Qing Yang 0003, Mohammed I. Alghamdi, Shu Yin 0001, Zhiyang Ding, Jiong Xie, Joshua Lewis, Xiao Qin 0001
IPCCC4
2010 Communication-Aware Load Balancing for Parallel Applications on Clusters
abstract
Cluster computing has emerged as a primary and cost-effective platform for running parallel applications, including communication-intensive applications that transfer a large amount of data among the nodes of a cluster via the interconnection network. Conventional load balancers have proven effective in increasing the utilization of CPU, memory, and disk I/O resources in a cluster. However, most of the existing load-balancing schemes ignore network resources, leaving an opportunity to improve the effective bandwidth of networks on clusters running parallel applications. For this reason, we propose a communication-aware load-balancing technique that is capable of improving the performance of communication-intensive applications by increasing the effective utilization of networks in cluster environments. To facilitate the proposed load-balancing scheme, we introduce a behavior model for parallel applications with large requirements of network, CPU, memory, and disk I/O resources. Our load-balancing scheme can make full use of this model to quickly and accurately determine the load induced by a variety of parallel applications. Simulation results generated from a diverse set of both synthetic bulk synchronous and real parallel applications on a cluster show that our scheme significantly improves the performance, in terms of slowdown and turn-around time, over existing schemes by up to 206 percent (with an average of 74 percent) and 235 percent (with an average of 82 percent), respectively.
Xiao Qin 0001, Hong Jiang 0001, Adam Manzanares, Xiaojun Ruan, Shu Yin 0001
IEEE Trans. Computers5
2010 Conserving energy in real-time storage systems with I/O burstiness
abstract
Energy conservation has become a critical problem for real-time embedded storage systems. Although a variety of approaches for reducing energy consumption have been extensively studied, energy conservation for real-time embedded storage systems is still an open problem. In this article, we propose an energy management strategy, I/O Burstiness for Energy Conservation (IBEC), exploiting the burstiness of real-time embedded storage systems applications. Our approach aims at combining the IBEC energy-management strategy with a Linux-based disk block-scheduling mechanism to conserve the energy of storage systems. Extensive experiments are conducted involving a number of synthetic disk traces as well as real-world data-intensive traces. To evaluate the energy efficiency of IBEC, we compare the performance of IBEC against three existing strategies, namely, PA-EDF, DP-EDF, and EDF. Compared with the alternative strategies, IBEC reduces the power consumption of real-time embedded disks system by up to 60%.
Adam Manzanares, Xiaojun Ruan, Shu Yin 0001, Xiao Qin 0001, Adam Roth, Mais Nijim
ACM Trans. Embed. Comput. Syst.3
2009 How reliable are parallel disk systems when energy-saving schemes are involved?
abstract
Many energy conservation techniques have been proposed to achieve high energy efficiency in disk systems. Unfortunately, growing evidence shows that energy-saving schemes in disk drives usually have negative impacts on storage systems. Existing reliability models are inadequate to estimate reliability of parallel disk systems equipped with energy conservation techniques. To solve this problem, we propose a mathematical model - called MINT - to evaluate the reliability of a parallel disk system where energy-saving mechanisms are implemented. In this paper, we focus on modeling the reliability impacts of two well-known energy-saving techniques - the Popular Disk Concentration technique (PDC) and the Massive Array of Idle Disks (MAID). We started this research by investigating how PDC and MAID affect the utilization and power-state transition frequency of each disk in a parallel disk system. We then model the annual failure rate of each disk as a function of the disk's utilization, power state transition frequency as well as operating temperature, because these parameters are key reliability-affecting factors in addition to disk ages. Next, the reliability of a parallel disk system can be derived from the annual failure rate of each disk in the parallel disk system. Finally, we used MINT to study the reliability of a parallel disk system equipped with the PDC and MAID techniques. Experimental results show that PDC is more reliable than MAID when disk workload is low. In contrast, the reliability of MAID is higher than that of PDC under relatively high I/O load.
Shu Yin 0001, Xiaojun Ruan, Adam Manzanares, Xiao Qin 0001
CLUSTER1
2009 Performance Evaluation of Energy-Efficient Parallel I/O Systems with Write Buffer Disks
abstract
In the past decade, parallel disk systems have been developed to address the problem of I/O performance. A critical challenge with modern parallel I/O systems is that parallel disks consume a significant amount of energy in servers and high performance computers. To conserve energy consumption in parallel I/O systems, one can immediately spin down disks when disk are idle; however, spinning down disks might not be able to produce energy savings due to penalties of spinning operations. Unlike powering up CPUs, spinning down and up disks need physical movements. Therefore, energy savings provided by spinning down operations must offset energy penalties of the disk spinning operations. To substantially reduce the penalties incurred by disk spinning operations, we developed a novel approach to conserving energy of parallel I/O systems with write buffer disks, which are used to accumulate small writes using a log file system. Data sets buffered in the log file system can be transferred to target data disks in a batch way. Thus, buffer disks aim to serve a majority of incoming write requests, attempting to reduce the large number of disk spinning operations by keeping data disks in standby for long period times. Interestingly, the write buffer disks not only can achieve high energy efficiency in parallel I/O systems, but also can shorten response times of write requests. To evaluate the performance and energy efficiency of our parallel I/O systems with buffer disks, we implemented a prototype using a cluster storage system as a testbed. Experimental results show that under light and moderate I/O load, buffer disks can be employed to significantly reduce energy dissipation in parallel I/O systems without adverse impacts on I/O performance.
Xiaojun Ruan, Adam Manzanares, Shu Yin 0001, Ziliang Zong, Xiao Qin 0001
ICPP3
2009 ECOS: An energy-efficient cluster storage system
abstract
Cluster storage systems are essential building blocks for many high-end computing infrastructures. Although energy conservation techniques have been intensively studied in the context of clusters and disk arrays, improving energy efficiency of cluster storage systems remains an open issue. To address this problem, we describe in this paper an approach to implementing an energy-efficient cluster storage system or ECOS for short. ECOS relies on the architecture of cluster storage systems in which each I/O node manages multiple disks - one buffer disk and several data disks. Given an I/O node, the key idea behind ECOS is to redirect disk requests from data disks to the buffer disk. To balance I/O load among I/O nodes, ECOS might redirect requests from one I/O node into the others. Redirecting requests is a driving force of energy saving, and the reason is two-fold. First, ECOS makes an effort to keep buffer disks active while placing data disks into standby in a long time period to conserve energy. Second, ECOS reduces the number of disk spin downs/ups in I/O nodes. The idea of ECOS was implemented in a Linux cluster, where each I/O node contains one buffer disk and two data disks. Experimental results show that ECOS improves the energy efficiency of traditional cluster storage systems where buffer disks are not employed. Adding one extra buffer disk into each I/O node seemingly has negative impact on energy saving. Interestingly, our results indicate that ECOS equipped with extra buffer disks is more energy efficient than the same cluster storage system without the buffer disks. The implication of the experiments is that using existing data disks in I/O nodes to perform as buffer disks can achieve even higher energy efficiency.
Xiaojun Ruan, Shu Yin 0001, Adam Manzanares, Jiong Xie, Zhiyang Ding, James Majors, Xiao Qin 0001
IPCCC2
2009 Improving reliability of energy-efficient parallel storage systems by disk swapping
abstract
The Popular Disk Concentration (PDC) technique and the Massive Array of Idle Disks (MAID) technique are two effective energy saving schemes for parallel disk systems. The goal of PDC and MAID is to skew I/O load towards a few disks so that other disks can be transitioned to low power states to conserve energy. I/O load skewing techniques like PDC and MAID inherently affect reliability of parallel disks because disks storing popular data tend to have high failure rates than disks storing cold data. To achieve good tradeoffs between energy efficiency and disk reliability, we first present a reliability model to quantitatively study the reliability of energy-efficient parallel disk systems equipped with the PDC and MAID schemes. Then, we propose a novel strategy—disk swapping—to improve disk reliability by alternating disks storing hot data with disks holding cold data. We demonstrate that our disk-swapping strategies not only can increase the lifetime of cache disks in MAID-based parallel disk systems, but also can improve reliability of PDC-based parallel disk systems.
Shu Yin 0001, Xiaojun Ruan, Adam Manzanares, Zhiyang Ding, Jiong Xie, James Majors, Xiao Qin 0001
IPCCC1
2009 Can We Improve Energy Efficiency of Secure Disk Systems without Modifying Security Mechanisms?
abstract
Improving energy efficiency of security-aware storage systems is challenging, because security and energy efficiency are often two conflicting goals. The first step toward making the best tradeoffs between high security and energy efficiency is to profile encryption algorithms to decide if storage systems would be able to produce energy savings for security mechanisms. We are focused on encryption algorithms rather than other types of security services, because encryption algorithms are usually computation-intensive. In this study, we used the XySSL libraries and profiled operations of several test problems using Conky - a lightweight system monitor that is highly configurable. Using our profiling techniques we concluded that although 3DES is much slower than AES encryption,it more likely to save energy in security-aware storage systems using 3DES than AES. The CPU is the bottleneck in 3DES, allowing us to take advantage of dynamic power management schemes to conserve energy at the disk level.After profiling several hash functions, we noticed that the CPU is not the bottleneck for any of these functions,indicating that it is difficult to leverage the dynamic power management technique to conserve energy of a single disk where hash functions are implemented for integrity checking.
Xiaojun Ruan, Adam Manzanares, Shu Yin 0001, Mais Nijim, Xiao Qin 0001
NAS3
2009 Energy-Aware Prefetching for Parallel Disk Systems: Algorithms, Models, and Evaluation
abstract
Parallel disk systems consume a significant amount of energy due to the large number of disks. To design economically attractive and environmentally friendly parallel disk systems, in this paper we design and evaluate an energy-aware prefetching strategy for parallel disk systems consisting of a small number of buffer disks and large number of data disks. Using buffer disks to temporarily handle requests for data disks, we can keep data disks in the low-power mode as long as possible. Our prefetching algorithm aims to group many small idle periods in data disks to form large idle periods, which in turn allow data disks to remain in the standby state to save energy. To achieve this goal, we utilize buffer disks to aggressively fetch popular data from regular data disks into buffer disks, thereby putting data disks into the standby state for longer time intervals. A centrepiece in the prefetching mechanism is an energy-saving prediction model, based on which we implement the energy-saving calculation module that is invoked in the prefetching algorithm. We quantitatively compare our energy-aware prefetching mechanism against existing solutions, including the dynamic power management strategy. Experimental results confirm that the buffer-disk-based prefetching can significantly reduce energy consumption in parallel disk systems by up to 50 percent. In addition, we systematically investigate the energy efficiency impact that varying disk power parameters has on our prefetching algorithm.
Adam Manzanares, Xiaojun Ruan, Shu Yin 0001, Mais Nijim, Xiao Qin 0001
NCA3
2009 Dynamic load balancing for I/O-intensive applications on clusters
abstract
Load balancing for clusters has been investigated extensively, mainly focusing on the effective usage of global CPU and memory resources. However, previous CPU- or memory-centric load balancing schemes suffer significant performance drop under I/O-intensive workloads due to the imbalance of I/O load. To solve this problem, we propose two simple yet effective I/O-aware load-balancing schemes for two types of clusters: (1) homogeneous clusters where nodes are identical and (2) heterogeneous clusters, which are comprised of a variety of nodes with different performance characteristics in computing power, memory capacity, and disk speed. In addition to assigning I/O-intensive sequential and parallel jobs to nodes with light I/O loads, the proposed schemes judiciously take into account both CPU and memory load sharing in the system. Therefore, our schemes are able to maintain high performance for a wide spectrum of workloads. We develop analytic models to study mean slowdowns, task arrival, and transfer processes in system levels. Using a set of real I/O-intensive parallel applications and synthetic parallel jobs with various I/O characteristics, we show that our proposed schemes consistently improve the performance over existing non-I/O-aware load-balancing schemes, including CPU- and Memory-aware schemes and a PBS-like batch scheduler for parallel and sequential jobs, for a diverse set of workload conditions. Importantly, this performance improvement becomes much more pronounced when the applications are I/O-intensive. For example, the proposed approaches deliver 23.6--88.0 % performance improvements for I/O-intensive applications such as LU decomposition, Sparse Cholesky, Titan, Parallel text searching, and Data Mining. When I/O load is low or well balanced, the proposed schemes are capable of maintaining the same level of performance as the existing non-I/O-aware schemes.
Xiao Qin 0001, Hong Jiang 0001, Adam Manzanares, Xiaojun Ruan, Shu Yin 0001
ACM Trans. Storage5