EDBT 2026 Demo / reviewers in the wild / expert
Yanlong Yin
dblp:23/9789
· DBLP profile ↗
22ranked-venue papers
2as first author
12since 2021 · last 2025
0009-0004-1505-4295ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 2 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GoPIM: GCN-Oriented Pipeline Optimization for PIM AcceleratorsabstractGraph convolutional networks (GCNs) are popular for a variety of graph learning tasks. ReRAM-based processing-in-memory (PIM) accelerators are promising to expedite GCN training owing to their in-situ computing capability. However, existing accelerators can be severely underutilized even with pipelines, due to the oversight of the skewed execution times of various GCN stages and the ignorance of skewed degrees of graph vertices. In this work, we propose GOPIM, a GCN-oriented pipeline optimization for PIM accelerators to expedite GCN training. First, GOPIM proposes an ML-based scheme that allocates crossbar resources to the most needed stages to streamline the overall pipeline. Second, GOPIM utilizes a selective vertex updating technique that evenly distributes vertices on crossbars by interleaved mapping. These techniques collectively reduce the overall execution time without losing much accuracy. We also provide a practical architecture design for GOPIM. Our experimental results show that, GoPIM achieves up to 191 × speedup and 16.1 × energy saving, compared to the state-of-the-art work. Siling Yang, Shuibing He, Wenjiong Wang, Yanlong Yin, Weijian Chen 0002, Xuechen Zhang 0001, Xian-He Sun |
HPCA | 4 |
| 2025 | Advanced Maximal Biclique Enumeration on GPUs Using BitmapsabstractMaximal biclique enumeration (MBE) in bipartite graphs is an important problem in data mining with many real-world applications. Parallel MBE algorithms for GPUs are needed for MBE acceleration leveraging its many computing cores. However, enumerating maximal bicliques using GPUs has three main challenges including large memory requirement, thread divergence, and load imbalance. In this paper, we propose GMBE+, an advanced GPU solution for the MBE problem. To overcome the challenges, we design (1) a node-reuse approach to reduce GPU memory usage with advanced node pruning, (2) a bitmap-based set intersection approach to minimize thread divergence, and (3) a load-aware task scheduling framework to achieve load balance among threads within GPU warps, facilitated by a novel set union approach. Our experiments reveal that GMBE+ is 1.2× faster than the latest GPU-based MBE algorithm GMBE on average when running on the same NVIDIA A100 GPU. Zhe Pan 0001, Shuibing He, Xu Li 0026, Xuechen Zhang 0001, Rui Wang 0076, Yanlong Yin, Gang Chen 0001 |
IEEE Trans. Computers | 6 |
| 2024 | AUTOHET: An Automated Heterogeneous ReRAM-Based Accelerator for DNN InferenceabstractReRAM-based accelerators have become prevalent in accelerating deep neural network inference owing to their in-situ computing capability of ReRAM crossbars. However, most existing ReRAM-based accelerators are designed with homogeneous crossbars, leading to either low resource utilization or sub-optimal energy efficiency. In this paper, we propose AutoHet, an automated heterogeneous ReRAM-based accelerator with varied-size crossbars for different DNN layers. To achieve both high crossbar utilization and energy efficiency, AutoHet uses a reinforcement learning algorithm to automatically determine the proper crossbar configuration for each DNN layer. Additionally, AutoHet introduces rectangle crossbars and a tile-shared crossbar allocation scheme to reduce crossbar wastage and energy consumption. Experiment results show that AutoHet effectively improves crossbar utilization by up to 3.1 × and reduces energy consumption by up to 94.6%, compared to approaches with homogeneous ReRAM crossbars. Shuibing He, Weijian Chen 0002, Siling Yang, Yanlong Yin, Xuechen Zhang 0001, Xian-He Sun, Gang Chen 0001 |
ICPP | 7 |
| 2024 | IOWA: An I/O-Aware Adaptive Sampling Framework for Deep LearningabstractTraining deep DNN models is time-consuming, especially when using large datasets. In the standard model training process, data instances are sampled uniformly and fed into the neural networks. However, not all instances contribute equally to the resulting model, and even the same data instance may affect the model differently in different training iterations. In addition to computational costs, I/O overhead can significantly impact the training speed, particularly for I/O-intensive processes. Given these observations, we propose an I/O-aware sampling metric in this paper. Building on this, we introduce an I/O-Aware Adaptive Sampling Framework (IOWA), which includes data profiling, adaptive data sampling, and redundant data instance replacement to accelerate the training process. Extensive exper-iments demonstrate that, compared to traditional DNN training processes, our approach can achieve up to a 3 x speedup without compromising the resulting model. Weijian Chen 0002, Yanlong Yin, Shuibing He |
NAS | 3 |
| 2024 | Enumeration of Billions of Maximal Bicliques in Bipartite Graphs without Using GPUsabstractMaximal biclique enumeration (MBE) is crucial in bipartite graph analysis. Recent studies rely on extensive set intersections on static bipartite graphs to solve the MBE problem. However, the computational subgraphs dynamically change during enumeration, leading to redundant memory accesses and degraded set intersection performance. To overcome this limitation, we propose an AdaMBE algorithm. First, we redesign its core operations using local neighborhood information derived from computational subgraphs to minimize redundant memory accesses. Second, we dynamically create computational subgraphs using bitmaps leveraging its fast bitwise operations to accelerate set intersections. Finally, we integrate them in AdaMBE. Our experimental results show that AdaMBE is $1.6 \times-49.7 \times$ faster than its closest CPU-based competitor and successfully enumerates all 19 billion maximal bicliques on the TVTropes dataset, a large task beyond the capabilities of existing algorithms. Notably, on certain datasets, our parallel version, ParAdaMBE, on CPUs even outperforms GMBE on GPUs by up to $5.07 \times$. Zhe Pan 0001, Shuibing He, Xu Li 0026, Xuechen Zhang 0001, Yanlong Yin, Rui Wang 0076, Lidan Shou, Mingli Song, Xian-He Sun, Gang Chen 0001 |
SC | 5 |
| 2023 | Textile pattern recommendations with convolutional neural networks and autoencoderabstractAbstract Textile pattern design is a time‐consuming and tedious work. Mihui, our ongoing developing system, employs deep‐learning techniques to automatically generate huge volumes of patterns with the help of human guidance. However, trained as a black box, Mihui cannot provide customized service for each individual designer who shows unique aesthetics preferences. In this article, we introduce the recommendation module of Mihui. The module forwards all generated pattern images to a deep encoding network, where images are mapped into 128‐dimension vectors. For each user of Mihui, we create a profile by his/her purchased or downloaded history. A novel encoder network is proposed to learn a personal taste vector for each user, based on which, we recommend new patterns to him/her. Our records in Mihui show that the recommendation module effectively improve users' experience on Mihui. Kuang Mao, Sai Wu, Jiajia He, Haichao Huang, Yanlong Yin, Zujie Ren |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | APQ: Automated DNN Pruning and Quantization for ReRAM-Based AcceleratorsabstractEmerging ReRAM-based accelerators support in-memory computation to accelerate deep neural network (DNN) inference. Weight matrix pruning is a widely used technique to reduce the size of DNN models, thereby reducing the resource and energy consumption of ReRAM-based accelerators. However, existing pruning works for ReRAM-based accelerators have three major issues. First, they use heuristics or rules from domain experts to prune the weights, leading to sub-optimal pruning policies. Second, they use row or column-level coarse-granularity methods to prune weights, resulting in poor compression rates with model accuracy constraints. Third, they only apply the weight pruning technique individually, losing the compression opportunity of both pruning and quantization. In this article, we propose an Automated DNN Pruning and Quantization framework, namedAPQ, for ReRAM-based accelerators. First,APQadopts reinforcement learning (RL) to automatically determine the pruning policy for DNN layers for a global optimum. Second, it prunes and maps weight matrices to a ReRAM-based accelerator in a finer granularity of column-vector, which improves the compression rates with the accuracy constraints. To address the dislocation problem, it uses a new data path in ReRAM-based accelerators to correctly index and feed input to matrix-vector computation. Third, to further reduce resource consumption,APQalso leverages reinforcement learning to automatically determine the quantization bitwidth of each layer of the pruned DNN model. Experimental results show that,APQachieves up to 4.52X compression rate, 4.11X area efficiency, and 4.51X energy efficiency with similar or even higher model accuracy, compared to the state-of-the-art work. Siling Yang, Shuibing He, Hexiao Duan, Weijian Chen 0002, Xuechen Zhang 0001, Yanlong Yin |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2022 | FPL Demo: SERVE: Agile Hardware Development Platform with Cloud IDE and Cloud FPGAsabstractWe introduce SERVE, a cloud platform for agile hardware software co-design, with cloud IDE and cloud FPGAs integrated. SERVE enables users to focus on logic designs, without facing the hassle of setting up FPGA tools and development environment. Users can write and simulate hardware logic in the cloud IDE and then generate bitstream files through a Continuous Integration (CI) pipeline. Finally, the bitstream files are deployed on an FPGA board. A great amount of testbenches will be executed to ensure the correctness of the hardware logic. We will demo a workflow of modifying a RISC- V processor and getting the design change quickly evaluated using SERVE. Ke Zhang 0017, Yisong Chang, Yanlong Yin, Yuxiao Chen 0009, Songyue Wang, Mingyu Chen 0001, Yungang Bao |
FPL | 4 |
| 2022 | Accelerating Tensor Swapping in GPUs With Self-Tuning CompressionabstractData swapping between CPUs and GPUs is widely used to address the GPU memory shortage issue when training deep neural networks (DNNs) requiring a larger amount of memory than that a GPU may have. Data swapping may become a bottleneck when its latency is longer than the latency of DNN computations. Tensor compression in GPUs can reduce the data swapping time. However, existing works on compressing tensors in the virtual memory of GPUs have three major issues: lack of portability because its implementation requires additional (de)compression units in memory controllers, sub-optimal compression performance for varying tensor compression ratios and sizes, and poor adaptation to dense tensors because they only focus on sparse tensors. We propose a self-tuning tensor compression framework, namedCSwap+, for improving the virtual memory management of GPUs. It uses GPUs for (de)compression directly and thus has high portability and is minimally dependent on GPU architecture features. Furthermore, it only applies compression on tensors that are deemed to be cost-effective considering their compression ratio, size, and the characteristics of compression algorithms at runtime. Finally, to adapt to DNN models with dense tensors, it also supports cost-effective lossy compression for dense tensors with nearly no model training accuracy degradation. We conduct the experiments through six representative memory-intensive DNN models. Compared to vDNN,CSwap+reduces tensor swapping latency by up to 50.9% and 46.1% with NVIDIA V100 GPU, for DNN models with sparse and dense tensors, respectively. Shuibing He, Xuechen Zhang 0001, Shuaiben Chen, Peiyi Hong, Yanlong Yin, Xian-He Sun |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2021 | CSWAP: A Self-Tuning Compression Framework for Accelerating Tensor Swapping in GPUsabstractGraphic Processing Units (GPUs) have limited memory capacity. Training popular deep neural networks (DNNs) often requires a larger amount of memory than that a GPU may have. Consequently, training data needs to be swapped between CPUs and GPUs. Data swapping may become a bottleneck when its latency is longer than the latency of DNN computations. Tensor compression in GPUs can reduce the data swapping time. However, existing works on compressing tensors in the virtual memory of GPUs have two major issues: sub-optimal compression performance for varying tensor sparsity and sizes and lack of portability because its implementation requires additional (de)compression units in memory controllers. We propose a self-tuning tensor compression framework, named CSWAP, for improving the virtual memory management of GPUs. It has high portability and is minimally dependent on GPU architecture features. Furthermore, its runtime only applies compression on tensors that are deemed to be cost-effective considering their sparsity and size and the characteristics of compression algorithms. Finally, our framework is fully automated and can customize the compression policy for different neural network architectures and GPU architectures. Our experimental results using six representative memory-intensive DNN models show that CSWAP reduces tensor swapping latency by up to 50.9% and reduces the DNN training time by 20.7% on average with NVIDIA V100 GPUs compared to vDNN. Shuibing He, Xuechen Zhang 0001, Shuaiben Chen, Peiyi Hong, Yanlong Yin, Xian-He Sun, Gang Chen 0001 |
CLUSTER | 6 |
| 2021 | A Novel Multi-CPU/GPU Collaborative Computing Framework for SGD-based Matrix FactorizationabstractThis paper presents a heterogeneous collaborative computing framework for SGD-based Matrix Factorization, named HCC-MF. HCC-MF can train the feature matrix efficiently using multiple CPUs and GPUs. It performs collaborative computing with data parallelism, where a server CPU is in charge of management and synchronization and other heterogeneous worker CPUs and worker GPUs performs calculation with their data assignments. HCC-MF adopts two data partition strategies, “data partition with heterogeneous load balance” and “data partition with hidden synchronization.” We build a time cost model to guide the data distribution among multiple workers and we design several communication optimization techniques with consideration of datasets’ and processors’ characteristics. Experimental results indicate that HCC-MF can utilize more than 88% of the platform’s computing power, yielding a speedup of 2.9 compared with advanced SGD-based MF, CuMF_SGD, on large-scale data sets. Yanlong Yin, Yan Liu 0032, Shuibing He, Yang Bai 0007, Renfa Li |
ICPP | 2 |
| 2021 | AUTO-PRUNE: automated DNN pruning and mapping for ReRAM-based acceleratorabstractEmergent ReRAM-based accelerators support in-memory computation to accelerate deep neural network (DNN) inference. Weight matrix pruning of DNNs is a widely used technique to reduce the size of DNN models, thereby reducing the resource and energy consumption of ReRAM-based accelerators. However, conventional works on weight matrix pruning for ReRAM-based accelerators have three major issues. First, they use heuristics or rules from domain experts to prune the weights, leading to suboptimal pruning policies. Second, they mostly focus on improving compression ratio, thus may not meet accuracy constraints. Third, they ignore direct feedback of hardware. In this paper, we introduce an automated DNN pruning and mapping framework, named AUTO-PRUNE. It leverages reinforcement learning (RL) to automatically determine the pruning policy considering the constraint of accuracy loss. The reward function of RL agents is designed using hardware’s direct feedback (i.e., accuracy and compression rate of occupied crossbars). The function directs the search of the pruning ratio of each layer for a global optimum considering the characteristics of individual layers of DNN models. Then AUTO-PRUNE maps the pruned weight matrices to crossbars to store only nontrivial elements. Finally, to avoid the dislocation problem, we design a new data-path in ReRAM-based accelerators to correctly index and feed input to matrix-vector computation leveraging the mechanism of operation units. Experimental results show that, compared to the state-of-the-art work, AUTO-PRUNE achieves up to 3.3X compression rate, 3.1X area efficiency, and 3.3X energy efficiency with a similar or even higher accuracy. Siling Yang, Weijian Chen 0002, Xuechen Zhang 0001, Shuibing He, Yanlong Yin, Xian-He Sun |
ICS | 5 |
| 2020 | Optimizing Parallel I/O Accesses through Pattern-Directed and Layout-Aware ReplicationabstractAs the performance gap between processors and storage devices keeps increasing, I/O performance becomes a critical bottleneck of modern high-performance computing systems. In this paper, we propose a pattern-directed and layout-aware data replication design, named PDLA, to improve the performance of parallel I/O systems. PDLA includes an HDD-based scheme H-PDLA and an SSD-based scheme S-PDLA. For applications with relatively low I/O concurrency, H-PDLA identifies access patterns of applications and makes a reorganized data replica for each access pattern on HDD-based servers with an optimized data layout. Moreover, to accommodate applications with high I/O concurrency, S-PDLA replicates critical access patterns that can bring performance benefits on SSD-based servers or on HDD-based and SSD-based servers. We have implemented the proposed replication scheme under MPICH2 library on top of OrangeFS file system. Experimental results show that H-PDLA can significantly improve the original parallel I/O system performance and demonstrate the advantages of S-PDLA over H-PDLA. Shuibing He, Yanlong Yin, Xian-He Sun, Xuechen Zhang 0001, Zongpeng Li |
IEEE Trans. Computers | 2 |
| 2020 | A Holistic Heterogeneity-Aware Data Placement Scheme for Hybrid Parallel I/O SystemsabstractWe presentH2DP, a holistic heterogeneity-aware data placement scheme for hybrid parallel I/O systems, which consist of HDD servers and SSD servers. Most of the existing approaches focus on server performance or application I/O pattern heterogeneity in data placement.H2DPconsiders three axes of heterogeneity: server performance, server space, and application I/O pattern. More specifically,H2DPdetermines the optimized stripe sizes on servers based on server performance, keeps only critical data on all hybrid servers and the rest data on HDD servers, and dynamically migrates data among different types of servers at run-time. This holistic heterogeneity-awareness enablesH2DPto achieve high performance by alleviating server load imbalance, efficiently utilizing SSD space, and accommodating application pattern variation. We have implemented a prototype ofH2DPunder MPICH2 atop OrangeFS. Extensive experimental results demonstrate thatH2DPsignificantly improve I/O system performance compared to existing data placement schemes. Shuibing He, Zheng Li 0006, Yanlong Yin, Xiaohua Xu 0002, Yong Chen 0001, Xian-He Sun |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2015 | IOSIG+: On the Role of I/O Tracing and Analysis for Hadoop SystemsabstractHadoop, as one of the most widely accepted MapReduce frameworks, is naturally data-intensive. Its several dependent projects, such as Mahout and Hive, inherent this characteristic. Meanwhile I/O optimization becomes a daunting work, since applications' source code is not always available. I/O traces for Hadoop and its dependents are increasingly important, because it can faithfully reveal intrinsic I/O behaviors without knowing the source code. This method can not only help to diagnose system bottlenecks but also further optimize performance. To achieve this goal, we propose a transparent tracing and analysis tool suite, namely IOSIG+, which can be plugged into Hadoop system. We make several contributions: 1) we describe our approach of tracing, 2) we release the tracer, which can trace I/O operations without modifying targets' source code, 3) this work adopts several techniques to mitigate the introduced execution overhead at runtime, 4) we create an analyzer, which helps to discover new approaches to address I/O problems according to access patterns. The experimental results and analysis confirm its effectiveness and the observed overhead can be as low as 1.97%. Xi Yang 0002, Yanlong Yin, Xian-He Sun |
CLUSTER | 4 |
| 2014 | SCALER: Scalable parallel file write in HDFSabstractTwo camps of file systems exist: parallel file systems designed for conventional high performance computing (HPC) and distributed file systems designed for newly emerged data-intensive applications. Addressing the big data challenge requires an approach that utilizes both high performance computing and data-intensive computing power. Thus, HPC applications may need to interact with distributed file systems, such as HDFS. The N-1 (N-to-1) parallel file write is a critical technical challenge, because it is very common for HPC applications but HDFS does not allow it. This study introduces a system solution, named SCALER, which allows MPI based applications to directly access HDFS without extra data movement. SCALER supports N-1 file write at both the inter-block level and intra-block level. Experimental results confirm that SCALER achieves the design goal efficiently. Xi Yang 0002, Yanlong Yin, Hui Jin 0001, Xian-He Sun |
CLUSTER | 2 |
| 2013 | Runtime system design of decoupled execution paradigm for data-intensive high-end computingabstractHigh performance computing are widely used for scientific discoveries by running scientific computation programs. Many of these applications are getting more and more data intensive [1]. They generate or access huge amount of data during some execution phases. However, traditional supercomputers are designed for computing-intensive tasks. They usually have highdensity clusters of processing cores and their storage systems are placed remotely and connected to the computing clusters with networks. This separation of the computing system and the storage system causes the data Input/Output performance bottleneck, especially for the data-intensive phases of HPC applications. This bottleneck degrades the HPC system's efficiency. Yanlong Yin, Hassan Eslami, Xian-He Sun, Yong Chen 0001, Rajeev Thakur, William Gropp |
CLUSTER | 2 |
| 2013 | Pattern-Direct and Layout-Aware Replication Scheme for Parallel I/O SystemsabstractThe performance gap between computing power and the I/O system is ever increasing, and in the meantime more and more High Performance Computing (HPC) applications are becoming data intensive. This study describes an I/O data replication scheme, named Pattern-Direct and Layout-Aware (PDLA) data replication scheme, to alleviate this performance gap. The basic idea of PDLA is replicating identified data access pattern, and saving these reorganized replications with optimized data layouts based on access cost analysis. A runtime system is designed and developed to integrate the PDLA replication scheme and existing parallel I/O system; a prototype of PDLA is implemented under the MPICH2 and PVFS2 environments. Experimental results show that PDLA is effective in improving data access performance of parallel I/O systems. Yanlong Yin, Jibing Li, Xian-He Sun, Rajeev Thakur |
IPDPS | 1 |
| 2012 | Boosting Application-Specific Parallel I/O Optimization Using IOSIGabstractMany scientific applications spend a significant portion of their execution time in accessing data from files. Various optimization techniques exist to improve data access performance, such as data prefetching and data layout optimization. However, optimization process is usually a difficult task due to the complexity involved in understanding I/O behavior. Tools that can help simplify the optimization process have a significant importance. In this paper, we introduce a tool, called IOSIG, for providing a better understanding of parallel I/O accesses and information to be used for optimization techniques. The tool enables tracing parallel I/O calls of an application and analyzing the collected information to provide a clear understanding of I/O behavior of the application. We show that performance overheads of the tool in trace collection and analysis are negligible. The analysis step creates I/O signatures that various optimizations can use for improving I/O performance. I/O signatures are compact, easy-to-understand, and parameterized representations containing data access pattern information such as size, strides between consecutive accesses, repetition, timing, etc. The signatures include local I/O behavior for each process and global behavior for an overall application. We illustrate the usage of the IOSIG tool in data prefetching and data layout optimizations. Yanlong Yin, Surendra Byna, Huaiming Song, Xian-He Sun, Rajeev Thakur |
CCGRID | 1 |
| 2011 | A Segment-Level Adaptive Data Layout Scheme for Improved Load Balance in Parallel File SystemsabstractParallel file systems are designed to mask the ever-increasing gap between CPU and disk speeds via parallel I/O processing. While they have become an indispensable component of modern high-end computing systems, their inadequate performance is a critical issue facing the HPC community today. Conventionally, a parallel file system stripes a file across multiple file servers with a fixed stripe size. The stripe size is a vital performance parameter, but the optimal value for it is often application dependent. How to determine the optimal stripe size is a difficult research problem. Based on the observation that many applications have different data-access clusters in one file, with each cluster having a distinguished data access pattern, we propose in this paper a segmented data layout scheme for parallel file systems. The basic idea behind the segmented approach is to divide a file logically into segments such that an optimal stripe size can be identified for each segment. A five-step method is introduced to conduct the segmentation, to identify the appropriate stripe size for each segment, and to carry out the segmented data layout scheme automatically. Experimental results show that the proposed layout scheme is feasible and effective, and it improves performance up to 163% for writing and 132% for reading on the widely used IOR and IOzone benchmarks. Huaiming Song, Yanlong Yin, Xian-He Sun, Rajeev Thakur, Samuel Lang |
CCGRID | 2 |
| 2011 | A cost-intelligent application-specific data layout scheme for parallel file systemsabstractI/O data access is a recognized performance bottleneck of high-end computing. Several commercial and research parallel file systems have been developed in recent years to ease the performance bottleneck. These advanced file systems perform well on some applications but may not perform well on others. They have not reached their full potential in mitigating the I/O-wall problem. Data access is application dependent. Based on the application-specific optimization principle, in this study we propose a cost-intelligent data access strategy to improve the performance of parallel file systems. We first present a novel model to estimate data access cost of different data layout policies. Next, we extend the cost model to calculate the overall I/O cost of any given application and choose an appropriate layout policy for the application. A complex application may consist of different data access patterns. Averaging the data access patterns may not be the best solution for those complex applications that do not have a dominant pattern. We then further propose a hybrid data replication strategy for those applications, so that a file can have replications with different layout policies for the best performance. Theoretical analysis and experimental testing have been conducted to verify the newly proposed cost-intelligent layout approach. Analytical and experimental results show that the proposed cost model is effective and the application-specific data layout approach achieved up to 74% performance improvement for data-intensive applications. Huaiming Song, Yanlong Yin, Yong Chen 0001, Xian-He Sun |
HPDC | 2 |
| 2011 | Server-side I/O coordination for parallel file systemsabstractParallel file systems have become a common component of modern high-end computers to mask the ever-increasing gap between disk data access speed and CPU computing power. However, while working well for certain applications, current parallel file systems lack the ability to effectively handle concurrent I/O requests with data synchronization needs, whereas concurrent I/O is the norm in data-intensive applications. Recognizing that an I/O request will not complete until all involved file servers in the parallel file system have completed their parts, in this paper we propose a serverside I/O coordination scheme for parallel file systems. The basic idea is to coordinate file servers to serve one application at a time in order to reduce the completion time, and in the meantime maintain the server utilization and fairness. A window-wide coordination concept is introduced to serve our purpose. We present the proposed I/O coordination algorithm and its corresponding analysis of average completion time in this study. We also implement a prototype of the proposed scheme under the PVFS2 file system and MPI-IO environment. Experimental results demonstrate that the proposed scheme can reduce average completion time by 8% to 46%, and provide higher I/O bandwidth than that of default data access strategies adopted by PVFS2 for heavy I/O workloads. Experimental results also show that the server-side I/O coordination scheme has good scalability. Huaiming Song, Yanlong Yin, Xian-He Sun, Rajeev Thakur, Samuel Lang |
SC | 2 |