VLDB 2026 Research / reviewers in the wild / expert
Yong Chen 0001
dblp:67/6351-1
· DBLP profile ↗
133ranked-venue papers
13as first author
26since 2021 · last 2026
0000-0002-9961-9051ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 110 · 13 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 4 since 2021Artificial intelligence and machine learning · 12 · 4 since 2021Databases, data management, data science and information retrieval · 11 · 3 since 2021Security and privacy · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TokenPowerBench: Benchmarking the Power Consumption of LLM InferenceabstractLarge language model (LLM) services now answer billions of queries per day, and industry reports show that inference, not training, accounts for more than 90% of total power consumption. However, existing benchmarks focus on either training/fine-tuning or performance of inference and provide little support for power consumption measurement and analysis of inference. We introduce TokenPowerBench, the first lightweight and extensible benchmark designed for LLM-inference power consumption studies. The benchmark combines a declarative configuration interface covering model choice, prompt set, and inference engine, a measurement layer that captures GPU-, node-, and system-level power without specialized power meters, and a phase-aligned metrics pipeline that attributes energy to the prefill and decode stages of every request. These elements make it straightforward to explore the power consumed by an LLM inference run; furthermore, by varying batch size, context length, parallelism strategy and quantization, users can quickly assess how each setting affects joules per token and other energy-efficiency metrics. We evaluate TokenPowerBench on four of the most widely used model series (Llama, Falcon, Qwen, and Mistral). Our experiments cover from 1 billion parameters up to the frontier-scale Llama3-405B model. Furthermore, we release TokenPowerBench as open source to help users to measure power consumption, forecast operating expenses, and meet sustainability targets when deploying LLM services. Chenxu Niu 0001, Wei Zhang 0097, Jie Li 0057, Tongyang Wang, Xi Wang 0009, Yong Chen 0001 |
AAAI | 7 |
| 2026 | FIXME: Towards End-to-End Benchmarking of LLM-Aided Design VerificationabstractDespite the transformative potential of Large Language Models (LLMs) in hardware design, a comprehensive evaluation of their capabilities in design verification remains underexplored. Current efforts predominantly focus on RTL generation and basic debugging, overlooking the critical domain of functional verification, which is the primary bottleneck in modern design methodologies due to the rapid escalation of hardware complexity. We present FIXME, the first end-to-end, multi-model, and open-source evaluation framework for assessing LLM performance in hardware functional verification (FV) to address this crucial gap. FIXME introduces a structured three-level difficulty hierarchy spanning six verification sub-domains and 180 diverse tasks, enabling in-depth analysis across the design lifecycle. Leveraging a collaborative AI-human approach, we construct a high-quality dataset using 100% silicon-proven designs, ensuring comprehensive coverage of real-world challenges. Furthermore, we enhance the functional coverage by 45.57% through expert-guided optimization. By rigorously evaluating state-of-the-art LLMs such as GPT-4, Claude3, and LlaMA3, we identify key areas for improvement and outline promising research directions to unlock the full potential of LLM-driven automation in hardware design verification. The benchmark is available at https://github.com/ChatDesignVerification/FIXME. Gwok-Waa Wan, Sam-Zaak Wong, Shengchu Su, Chenxu Niu 0001, Ning Wang 0071, Xinlai Wan, Qixiang Chen, Mengnv Xing, Jianmin Ye, Rongchang Song, Qiang Xu 0001, Nan Guan, Zhe Jiang 0004, Xi Wang 0009, Yong Chen 0001, Jun Yang 0006 |
AAAI | 18 |
| 2025 | ICEAGE: Intelligent Contextual Exploration and Answer Generation Engine for Scientific Data Discovery
Chenxu Niu 0001, Wei Zhang 0097, Mert Side, Yong Chen 0001 |
SSDBM | 4 |
| 2025 | LaAeb: A comprehensive log-text analysis based approach for insider threat detection
Kexiong Fei, Yucan Zhou, Xiaoyan Gu 0001, Haihui Fan, Bo Li 0063, Weiping Wang 0005, Yong Chen 0001 |
Comput. Secur. | 8 |
| 2025 | Log2Graph: A graph convolution neural network based method for insider threat detectionabstractWith the advancement of network security equipment, insider threats gradually replace external threats and become a critical contributing factor for cluster security threats. When detecting and combating insider threats, existing methods often concentrate on users’ behavior and analyze logs recording their operations in an information system. Traditional sequence-based method considers temporal relationships for user actions, but cannot represent complex logical relationships well between various entities and different behaviors. Current machine learning-based approaches, such as graph-based methods, can establish connections among log entries but have limitations in terms of complexity and identifying malicious behavior of user’s inherent intention. In this paper, we propose Log2Graph, a novel insider threat detection method based on graph convolution neural network. To achieve efficient anomaly detection, Log2Graph first retrieves logs and corresponding features from log files through feature extraction. Specifically, we use an auxiliary feature of anomaly index to describe the relationship between entities, such as users and hosts, instead of establishing complex connections between them. Second, these logs and features are augmented through a combination of oversampling and downsampling, to prepare for the next-stage supervised learning process. Third, we use three elaborated rules to construct the graph of each user by connecting the logs according to chronological and logical relationships. At last, the dedicated built graph convolution neural network is used to detect insider threats. Our validation and extensive evaluation results confirm that Log2Graph can greatly improve the performance of insider threat detection compared to existing state-of-the-art methods. Kexiong Fei, Weiping Wang 0005, Yong Chen 0001 |
J. Comput. Secur. | 5 |
| 2025 | RuYi: Optimizing Burst Buffer Through Automated, Fine-Grained Process-to-BB MappingabstractCurrent supercomputers use an SSD-based storage layer called Burst Buffer (BB) to provide I/O-intensive applications with accelerated storage access. However, efficiently utilizing this limited and expensive storage remains a critical issue, creating an urgent need for implementing Quality of Service (QoS) in BB. To address this, we propose RuYi, a QoS-aware method to provide applications with bandwidth guarantees in the BB file system. RuYi tackles two main issues. First, it quantitatively profiles available bandwidth resources in BB to ensure reliable QoS, a crucial aspect seldom studied in the literature. Second, RuYi offers fine-grained process-level QoS via an innovative process-to-BB mapping, maximizing resource utilization—something not achievable with conventional coarse-grained compute-to-BB mapping. We evaluated RuYi on a subsystem of the leading exascale supercomputer Sunway, consisting of 4,000 compute nodes and 200 BB nodes. The experimental results demonstrate that RuYi achieves an impressive end-to-end bandwidth control accuracy of 97%, while improving BB utilization by up to 116% compared to conventional coarse-grained compute-to-BB mapping. Yusheng Hua, Xuanhua Shi, Ligang He, Teng Zhang 0001, Hai Jin 0001, Yong Chen 0001 |
IEEE Trans. Computers | 7 |
| 2024 | Job Scheduling in High Performance Computing Systems with Disaggregated Memory ResourcesabstractDisaggregated memory promises to meet growing memory requirements of applications while improving system resource utilization in high-performance computing (HPC) systems. Compared to traditional systems-where expensive resources such as CPUs, GPUs, and memory, are assigned to jobs in units of nodes-systems with disaggregated memory introduce memory pools that can be shared among jobs; this introduces new optimization metrics to the job scheduler. In this paper, we propose a data-driven approach to evaluate job scheduling and resource configuration in HPC systems with disaggregated memory. To incorporate the memory requirements of jobs for both local and disaggregated memory resources and improve system efficiency in open-science HPC systems, we introduce a novel job scheduling algorithm called FM (Fair Memory). Our simulation results show that FM outperforms commonly-used job schedulers in terms of jobs' bounded slowdown when the shared memory pool capacity is limited, and in terms of fairness under all conditions. Jie Li 0057, George Michelogiannakis, Samuel A. Maloney, Brandon Cook 0001, Estela Suarez, John Shalf, Yong Chen 0001 |
CLUSTER | 7 |
| 2024 | Revisiting Erasure Codes: A Configuration PerspectiveabstractErasure coding (EC) plays a crucial role in the fault tolerance of modern distributed storage systems (DSS). Inspired by recent research on storage configuration, we study the configuration sensitivity of EC in real DSS in this paper. We systematically inject faults to trigger EC recovery under various configurations, and measure the impact on recovery time and storage overhead quantitatively. Our results show that configurations may affect the EC recovery time significantly (e.g., up to 426%). More interestingly, theoretically superior codes may perform worse in DSS under certain configurations. Also, there is a system checking period before EC recovery that accounts for 41% to 58% of the overall system recovery time, which has been largely ignored in previous studies. Finally, in terms of storage overhead, EC may introduce 32.3% to 72.0% more write amplification (WA) than the theoretical expectation, and we derive a formula to help estimate WA more precisely. Our work suggests the importance of considering the context of real DSS for EC research, and we hope the methodology and findings can contribute to a firmer footing for EC optimization in practice. Runzhou Han, Tabassum Mahmud, Zeren Yang, Vladislav Esaulov, Lipeng Wan 0001, Yong Chen 0001, Jim Wayda, Matthew Wolf, Mai Zheng |
HotStorage | 7 |
| 2024 | PROV-IO$^+$+: A Cross-Platform Provenance Framework for Scientific Data on HPC SystemsabstractData provenance, or data lineage, describes the life cycle of data. In scientific workflows on HPC systems, scientists often seek diverse provenance (e.g., origins of data products, usage patterns of datasets). Unfortunately, existing provenance solutions cannot address the challenges due to their incompatible provenance models and/or system implementations. In this paper, we analyze four representative scientific workflows in collaboration with the domain scientists to identify concrete provenance needs. Based on the first-hand analysis, we propose a provenance framework called PROV-IO$^+$, which includes an I/O-centric provenance model for describing scientific data and the associated I/O operations and environments precisely. Moreover, we build a prototype of PROV-IO$^+$to enable end-to-end provenance support on real HPC systems with little manual effort. The PROV-IO$^+$framework can support both containerized and non-containerized workflows on different HPC platforms with flexibility in selecting various classes of provenance. Our experiments with realistic workflows show that PROV-IO$^+$can address the provenance needs of the domain scientists effectively with reasonable performance (e.g., less than 3.5% tracking overhead for most experiments). Moreover, PROV-IO$^+$outperforms a state-of-the-art system (i.e., ProvLake) in our experiments. Runzhou Han, Mai Zheng, Surendra Byna, Houjun Tang, Bin Dong 0002, Dong Dai 0001, Yong Chen 0001, Dongkyun Kim, Joseph Hassoun, David Thorsley |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2023 | Effective Management of Time Series DataabstractCloud computing systems, consisting of numerous nodes and components, require constant monitoring to satisfy the Quality-of-Service (QoS), making the management of large-scale time series data challenging. To address this issue, age threshold retention policies have been implemented to remove historical data, but this eliminates valuable information from older periods. In this paper, we proposed an alternative approach that applies time series deduplication with metric-based tolerance to discard readings that stabilize within a calculated tolerance window. This approach can reduce the data volume by 70.38% on average. Once the data-reduced interval is queried, the readings can be reconstructed to retrieve the original granularity with low query runtime overhead and a Mean Absolute Percentage Error of 0.74%. Cristiano E. Caon, Jie Li 0057, Yong Chen 0001 |
CLOUD | 3 |
| 2023 | Workload Failure Prediction for Data CentersabstractFailed workloads that consumed significant computational resources in time and space affect the efficiency of HPC data centers significantly and thus limit the amount of scientific work that can be achieved. While the computational power has increased significantly over the years, detection and prediction of workload failures have lagged far behind and will become increasingly critical as the system scale and complexity further increase. In this study, we analyze workload traces collected from a production cluster and train machine learning models on a large amount of data sets to predict workload failures. Our prediction models consist of a queue-time model that estimates the probability of workload failures before execution and a runtime model that predicts failures at runtime. Evaluation results show that the queue-time model and runtime model can predict workload failures with a maximum precision score of 90.61% and 97.75%, respectively. By integrating the runtime model with the job scheduler, it helps reduce CPU time, and memory usage by up to 16.7% and 14.53%, respectively. Jie Li 0057, Ghazanfar Ali, Tommy Dang, Alan Sill, Yong Chen 0001 |
CLOUD | 6 |
| 2023 | PSQS: Parallel Semantic Querying Service for Self-describing File FormatsabstractFinding relevant datasets can be a time-consuming and challenging task, especially for self-describing file formats. Current solutions use either exact or partial keyword matching approaches to extract and process metadata queries, but they fail to capture semantic relationships between the metadata content and query keywords. To address this challenge, we introduce PSQS, a novel parallel semantic search method for self-describing files. The method leverages parallel processing and kv2vec semantic similarity measures to retrieve semantically relevant data efficiently. Our evaluation against existing metadata search solutions shows that PSQS offers a new, efficient and effective semantic search functionality for various fields where large self-describing files are used, such as scientific data management, leading to more accurate and efficient data retrieval. Chenxu Niu 0001, Wei Zhang 0097, Surendra Byna, Yong Chen 0001 |
IEEE Big Data | 4 |
| 2023 | Performance-Aware Energy-Efficient GPU Frequency Selection using DNN-based ModelsabstractEnergy efficiency will be important in future accelerator-based HPC systems for sustainability and to improve overall performance. This study proposes a deep neural network (DNN)-based learning model for execution time and power consumption of workloads across GPUs DVFS design space. Micro-architectural data obtained by running SPEC-ACCEL, DGEMM, and STREAM benchmarks are used for model training. These features are consistent for a workload unaffected by frequency and input size reducing the data required significantly. For real-world applications - LAMMPS, NAMD, GROMACS, LSTM, BERT, and ResNet50 power and time models show 89% – 98% accuracy on NVIDIA Ampere. Multi-objective functions help select optimal frequencies that lower power and minimize performance impact showing maximum energy savings of 27% at a performance loss of 1.8%. The same models trained on Ampere showed an accuracy of greater than 93% on an NVIDIA Volta, thereby demonstrating model portability across architectures. Ghazanfar Ali, Mert Side, Sridutt Bhalachandra, Nicholas J. Wright, Yong Chen 0001 |
ICPP | 5 |
| 2023 | An automated and portable method for selecting an optimal GPU frequency
Ghazanfar Ali, Mert Side, Sridutt Bhalachandra, Nicholas J. Wright, Yong Chen 0001 |
Future Gener. Comput. Syst. | 5 |
| 2023 | Data Distribution for Heterogeneous Storage SystemsabstractThe exponential growth of data in many science and engineering domains poses significant challenges to storage systems. Data distribution is a critical component in large-scale distributed storage systems and plays a vital role in placing petabytes of data and beyond, among tens to hundreds of thousands of storage devices. Meantime, heterogeneous storage systems, such as those having devices with hard disk drives (HDDs) and storage class memories (SCMs), have become increasingly popular for massive data storage due to their distinct and complement characteristics. This paper presents a new data distribution algorithm called SUORA (Scalable and Uniform storage via Optimally-adaptive and Random number Addressing) specifically for heterogeneous devices to maximize the benefits of them. SUORA provides a fully symmetric, highly efficient methodology to distribute data across a hybrid and tiered storage cluster. It divides heterogeneous devices into different buckets and segments, and adopts pseudo-random functions to map data onto them with the balanced consideration of capacity, performance and life-time. By analyzing hotness and access patterns, SUORA gradually moves hot data from HDDs to SCMs to optimize the throughput, and moves cold data reversely for load balance. It combines data replication with migration to significantly reduce movement overhead while making data placement more adaptive to different workloads. Extensive evaluations on simulation and Sheepdog storage system show that, with considering distinct characteristics of various devices thoroughly, SUORA improves the overall performance efficiency of heterogeneous storage systems. Yong Chen 0001, Mai Zheng, Weiping Wang 0005 |
IEEE Trans. Computers | 2 |
| 2022 | JobViewer: Graph-based Visualization for Monitoring High-Performance Computing SystemabstractVisualization aims to strengthen data exploration and analysis, especially for complex and high-dimensional data. High-performance computing (HPC) systems are typically large and complicated instruments that generate massive performance and operation time series. Monitoring HPC systems’ performance is a daunting task for HPC admins and researchers due to their dynamic natures. This work proposes a visual design using the bipartite graph’s idea to visualize HPC clusters’ structure, metrics, and job scheduling data. We built a web-based prototype, called JobViewer, that integrates advanced methods in visualization and human-computer interaction (HCI) to demonstrate the benefits of visualization in real-time monitoring HPC centers. We also showed real use cases and a user study to validate the efficiency and highlight the current approach’s drawbacks. Tommy Dang, Ngan V. T. Nguyen, Jie Li 0057, Alan Sill, Jon R. Hass, Yong Chen 0001 |
BDCAT | 6 |
| 2022 | Automating CPU Dynamic Thermal Control for High Performance ComputingabstractIn a production high-performance computing (HPC) data center, numerous factors, including workload compute in-tensity, cooling infrastructure failure, and the use of economized cooling can substantially increase the CPU temperature. CPU thermal design-related studies have shown that slight variances in the operational temperature can significantly impact the lifetime, durability, and performance of a CPU. Therefore, it is critical to monitor and control the operating temperature of the CPU. In this study, we design an automated and continuous CPU thermal monitoring and control methodology to maintain and control a healthy CPU thermal state. This research utilizes the Redfish protocol to monitor the CPU temperature and dynamic voltage frequency scaling to control the temperature. We developed a reference implementation and evaluated our methodology using a cluster of 150 Raspberry Pi3 nodes. We performed extensive CPU thermal analyses in different scenarios. We analyzed how quickly a CPU can attain the maximum temperature under 100% load at room temperature. Based on our experiments, the temperature of a CPU with 100% load can increase to ~72°C (161.6°F) and ~86°C (186.8°F) with the lowest and highest CPU frequency configurations, respectively. We analyzed the impact of applying thermal control at eight temperature configurations on the thermal and frequency scaling behavior of a CPU. We observed that applying thermal control at lower temperature configurations (e.g., 70°C (158°F)) is a better configuration for healing an overheated CPU. As a result of the proposed model, the CPU operating at normal temperature will consume comparatively less energy, deliver higher performance, and augment its durability. Ghazanfar Ali, Lowell Wofford, Yong Chen 0001 |
CCGRID | 4 |
| 2022 | On the Reproducibility of Bugs in File-System Aware Storage ApplicationsabstractMany storage applications such as file system checkers, defragmentation tools, etc. require a detailed understanding of file systems. Such file-system aware applications play an essential role today, but unfortunately they are error-prone. To better understand the challenges as well as the opportunities to address the issues, this paper presents an empirical study of real world bugs in file-system aware storage applications. By analyzing 59 bug cases from 4 representative applications in depth, we derive multiple insights in terms of general bug patterns, triggering conditions, and implications for building effective tools to address the issues. We hope that our study and the resulting dataset could contribute to the development of reliability tools for building robust file-system aware storage applications in general. Tabassum Mahmud, Om Rameshwar Gatla, Runzhou Han, Yong Chen 0001, Mai Zheng |
NAS | 5 |
| 2022 | A Study of Failure Recovery and Logging of High-Performance Parallel File SystemsabstractLarge-scale parallel file systems (PFSs) play an essential role in high-performance computing (HPC). However, despite their importance, their reliability is much less studied or understood compared with that of local storage systems or cloud storage systems. Recent failure incidents at real HPC centers have exposed the latent defects in PFS clusters as well as the urgent need for a systematic analysis. To address the challenge, we perform a study of the failure recovery and logging mechanisms of PFSs in this article. First, to trigger the failure recovery and logging operations of the target PFS, we introduce a black-box fault injection tool called PFault , which is transparent to PFSs and easy to deploy in practice. PFault emulates the failure state of individual storage nodes in the PFS based on a set of pre-defined fault models and enables examining the PFS behavior under fault systematically. Next, we apply PFault to study two widely used PFSs: Lustre and BeeGFS. Our analysis reveals the unique failure recovery and logging patterns of the target PFSs and identifies multiple cases where the PFSs are imperfect in terms of failure handling. For example, Lustre includes a recovery component called LFSCK to detect and fix PFS-level inconsistencies, but we find that LFSCK itself may hang or trigger kernel panics when scanning a corrupted Lustre. Even after the recovery attempt of LFSCK, the subsequent workloads applied to Lustre may still behave abnormally (e.g., hang or report I/O errors). Similar issues have also been observed in BeeGFS and its recovery component BeeGFS-FSCK. We analyze the root causes of the abnormal symptoms observed in depth, which has led to a new patch set to be merged into the coming Lustre release. In addition, we characterize the extensive logs generated in the experiments in detail and identify the unique patterns and limitations of PFSs in terms of failure logging. We hope this study and the resulting tool and dataset can facilitate follow-up research in the communities and help improve PFSs for reliable high-performance computing. Runzhou Han, Om Rameshwar Gatla, Mai Zheng, Jinrui Cao, Di Zhang 0015, Dong Dai 0001, Yong Chen 0001, Jonathan E. Cook 0001 |
ACM Trans. Storage | 7 |
| 2022 | LoomIO: Object-Level Coordination in Distributed File SystemsabstractDevice-level interference is recognized as a major cause of the performance degradation in distributed file systems. Although the approaches of mitigating interference through coordination at application-level, middleware-level, and server-level have shown beneficial results in previous studies, we find their effectiveness is largely reduced since I/O requests are re-arranged by underlying object file systems. In this research study, we prove that object-level coordination is critical and often the key to address the interference issue, as the scheduling of object requests determines the device-level accesses and thus determines the actual I/O bandwidth and latency. This article proposes an object-level coordination system, LoomIO, which uses an OBOP (One-Broadcast-One-Propagate) method and a time-limited coordination process to deliver highly efficient coordination service. Specifically, LoomIO enables object requests to achieve an optimized scheduling decision within a few milliseconds and largely mitigates the device-level interference. We have implemented a LoomIO prototye and integrated it into Ceph file system. The evaluation results show that LoomIO achieved the considerable improvements in resource utilization (by up to 35%), in I/O throughput (by up to 31%), and in 99th percentile latency (by up to 54%) compared to the K-optimal method which uses the same scheduling algorithm as LoomIO but does not have the coordination support. Yusheng Hua, Xuanhua Shi, Hai Jin 0001, Wei Xie 0017, Ligang He, Yong Chen 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2021 | xBGAS: A Global Address Space Extension on RISC-V for High Performance ComputingabstractThe tremendous expansion of data volume has driven the transition from monolithic architectures towards systems integrated with discrete and distributed subcomponents in modern scalable high performance computing (HPC) systems. As such, multi-layered software infrastructures have become essential to bridge the gap between heterogeneous commodity devices. However, operations across synthesized components with divergent interfaces inevitably lead to redundant software footprints and undesired latency. Therefore, a scalable and unified computing platform, capable of supporting efficient interactions between individual components, is desirable for largescale data-intensive applications. In this work, we introduce the Extended Base Global Address Space, or xBGAS, microarchitecture extension to the RISC-V instruction set architecture (ISA) for scalable high performance computing. The xBGAS extension provides native ISA-level support for direct accesses to remote shared memory by mapping remote data objects into a system's extended address space. We perform both software and hardware evaluations of the xBGAS design. The results show that xBGAS reduces instruction count generated by interprocess communication by 69.26% on average. Overall, xBGAS achieves an average performance gain of 21.96% (up to 37.29%) across the tested workloads. Xi Wang 0009, John D. Leidel, Brody Williams, Alan Ehret, Miguel Mark, Michel A. Kinsy, Yong Chen 0001 |
IPDPS | 7 |
| 2021 | Exploiting user activeness for data retention in HPC systemsabstractHPC systems typically rely on the fixed-lifetime (FLT) data retention strategy, which only considers temporal locality of data accesses to parallel file systems. However, our extensive analysis based on the leadership-class HPC system traces suggests that the FLT approach often fails to capture the dynamics in users' behavior and leads to undesired data purge. In this study, we propose an activeness-based data retention (ActiveDR) solution, which advocates considering the data retention approach from a holistic activeness-based perspective. By evaluating the frequency and impact of users' activities, ActiveDR prioritizes the file purge process for inactive users and rewards active users with extended file lifetime on parallel storage. Our extensive evaluations based on the traces of the prior Titan supercomputer show that, when reaching the same purge target, ActiveDR achieves up to 37% file miss reduction as compared to the current FLT retention methodology. Wei Zhang 0097, Surendra Byna, Hyogi Sim, Sankeun Lee 0001, Sudharshan S. Vazhkudai, Yong Chen 0001 |
SC | 6 |
| 2021 | I/O characteristic discovery for storage system optimizations
Yong Chen 0001, Dong Dai 0001, Weiping Wang 0005 |
J. Parallel Distributed Comput. | 2 |
| 2021 | HAM: Hotspot-Aware Manager for Improving Communications With 3D-Stacked MemoryabstractEmerging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, and big data science, are data-intensive. Data-intensive workloads usually present fine-grained memory accesses with limited or no data locality, and thus incur frequent cache misses and low utilization of memory bandwidth. 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) can provide significantly higher bandwidth than conventional memory modules. However, the traditional interfaces and optimization methods for JEDEC DDR devices do not allow to fully exploit the potential performance of 3D-stacked memory with the massive amount of irregular memory accesses of data-intensive applications. In this article, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices capable of optimizing memory access streams via request aggregation, hotspot detection, and in-memory prefetching. We present the HAM design and implementation, and simulate it on a system using RISC-V embedded cores with attached HMC devices. We extensively evaluate HAM with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results show that, on average, HAM reduces redundant requests by 37.51 percent and increases the prefetch buffer hit rate by 4.2 times, compared to a baseline streaming prefetcher. On the selected benchmark set, HAM provides performance gains of 21.81 percent in average (up to 34.28 percent), and power savings of 35.07 percent over a standard 3D-stacked memory. Xi Wang 0009, Antonino Tumeo, John D. Leidel, Jie Li 0057, Yong Chen 0001 |
IEEE Trans. Computers | 5 |
| 2021 | Trigger-Based Incremental Data Processing with Unified Sync and Async ModelabstractIn recent years, more and more applications in the cloud have needs to process large-scale on-line datasets, which evolve over time as new entries are added and existing entries are modified. Several programming frameworks, such as Percolator and Oolong, are proposed for such incremental data processing and can achieve efficient processing with an event-driven abstraction. However, these frameworks are inherently asynchronous, leaving the heavy burden of managing synchronization to applications' developers, which further significantly restricts their usabilities. In this study, we propose a trigger-based incremental computing framework in the cloud, called Domino, with both synchronous and asynchronous mechanisms to coordinate parallel triggers. With this new framework, both synchronous and asynchronous applications can be seamlessly developed. Use cases and extensive evaluation results confirm that it can deliver sufficient performance, and also is easy to use for incremental applications in large-scale distributed computing. Dong Dai 0001, Yong Chen 0001, Dries Kimpe, Robert B. Ross |
IEEE Trans. Cloud Comput. | 2 |
| 2021 | HiperView: real-time monitoring of dynamic behaviors of high-performance computing centers
Tommy Dang, Ngan V. T. Nguyen, Yong Chen 0001 |
J. Supercomput. | 3 |
| 2020 | Remote Atomic Extension (RAE) for Scalable High Performance ComputingabstractEmerging data-intensive applications such as graph analytics, machine learning, and data-driven scientific computing are driving the evolution of high-performance computing (HPC) systems from monolithic to scaled-out, heterogeneous, and complex architectures. In these systems, enormous data sets are mapped to discrete nodes to improve the performance of the system by using distributed storage and computing resources. As such, these data distributions induce frequent cross-node data transactions which challenge the performance of large-scale systems. Global atomic operations are one emerging class of the remote data operations that enable lock-free remote shared data operations. However, the cross-node read-modify-write operations consist of multiple distinct data operations and specific atomicity management, which induces a large amount of overhead. As such, these global atomic operations require an efficient communication methodology Existing advanced compo-nents, such as network interface controllers, network fabrics, network-on-chip (NoC) interconnects, are architected together to improve the system performance. However, complex software infrastructures are needed to provide integration between each discrete component. As a result, the redundant software routines across distinct devices induce a large amount of overhead that causes performance degradationIn this paper, we propose a remote atomic extension (RAE) design that provides inherent ISA-level instructions and micro-architecture support for remote atomic operations based on the RISC-V instruction set architecture (ISA). We design a toolchain and evaluate the RAE infrastructure via simulation. Our experiment results show that RAE eliminates 89.71% of the redundant software instructions used for remote atomic accesses and improves the performance by 17.61% on average (up to 23.35%), compared with the OpenSHMEM. Xi Wang 0009, Brody Williams, John D. Leidel, Alan Ehret, Michel A. Kinsy, Yong Chen 0001 |
DAC | 6 |
| 2020 | PAC: Paged Adaptive Coalescer for 3D-Stacked MemoryabstractMany contemporary data-intensive applications exhibit irregular and highly concurrent memory access patterns and thus challenge the performance of conventional memory systems. Driven by an expanding need for high-bandwidth memory featuring low access latency, 3D-stacked memory devices, such as the Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM), were designed to provide significantly higher throughput as compared to standard JEDEC DDR devices. However, existing memory interfaces and coalescing models, designed for conventional DDR devices, are unable to fully exploit the bandwidth potential inherent in these new 3D-stacked memory devices. In order to remedy this disparity, we introduce in this work a novel paged adaptive coalescer (PAC) infrastructure with a scalable coalescing network for 3D-stacked memory. We present the design and simulated implementation of this approach on RISC-V embedded cores with attached HMC devices. We have carried out extensive evaluations and the results show that the proposed PAC methodology yields an average coalescing efficiency of 56.01%. Further, our evaluation results also show that the PAC reduces bank conflicts and the power consumption by 85.16% and 59.21%, respectively. Overall, PAC achieves an average performance gain of 14.35% (and up to 26.06%) across 14 test suites. These results showcase the potential of the PAC methodology as applied to architecture design for increasingly critical data-intensive algorithms and applications. Xi Wang 0009, John D. Leidel, Brody Williams, Yong Chen 0001 |
HPDC | 4 |
| 2020 | Optimizing the SSD Burst Buffer by Traffic DetectionabstractCurrently, HPC storage systems still use hard disk drive (HDD) as their dominant storage device. Solid state drive (SSD) is widely deployed as the buffer to HDDs. Burst buffer has also been proposed to manage the SSD buffering of bursty write requests. Although burst buffer can improve I/O performance in many cases, we find that it has some limitations such as requiring large SSD capacity and harmonious overlapping between computation phase and data flushing phase. In this article, we propose a scheme, called SSDUP+. 1 SSDUP+ aims to improve the burst buffer by addressing the above limitations. First, to reduce the demand for the SSD capacity, we develop a novel method to detect and quantify the data randomness in the write traffic. Further, an adaptive algorithm is proposed to classify the random writes dynamically. By doing so, much less SSD capacity is required to achieve the similar performance as other burst buffer schemes. Next, to overcome the difficulty of perfectly overlapping the computation phase and the flushing phase, we propose a pipeline mechanism for the SSD buffer, in which data buffering and flushing are performed in pipeline. In addition, to improve the I/O throughput, we adopt a traffic-aware flushing strategy to reduce the I/O interference in HDD. Finally, to further improve the performance of buffering random writes in SSD, SSDUP+ transforms the random writes to sequential writes in SSD by storing the data with a log structure. Further, SSDUP+ uses the AVL tree structure to store the sequence information of the data. We have implemented a prototype of SSDUP+ based on OrangeFS and conducted extensive experiments. The experimental results show that our proposed SSDUP+ can save an average of 50% SSD space while delivering almost the same performance as other common burst buffer schemes. In addition, SSDUP+ can save about 20% SSD space compared with the previous version of this work, SSDUP, while achieving 20–30% higher I/O throughput than SSDUP. Xuanhua Shi, Wei Liu 0004, Ligang He, Hai Jin 0001, Yong Chen 0001 |
ACM Trans. Archit. Code Optim. | 6 |
| 2020 | Automated Performance Modeling of HPC Applications Using Machine LearningabstractAutomated performance modeling and performance prediction of parallel programs are highly valuable in many use cases, such as in guiding task management and job scheduling, offering insights of application behaviors, and assisting resource requirement estimation. The performance of parallel programs is affected by numerous factors, including but not limited to hardware, applications, algorithms, and input parameters, thus an accurate performance prediction is often a challenging and daunting task. In this article, we focus on automatically predicting the execution time of parallel programs (more specifically, MPI programs) with different inputs, at different scales, and without domain knowledge. We model the correlation between the execution time and domain-independent runtime features. These features include values of variables, counters of branches, loops, and MPI communications. Through automatically instrumenting an MPI program, each execution of the program will output a feature vector and its corresponding execution time. After collecting data from executions with different inputs, a random forest machine learning approach is used to build an empirical performance model, which can predict the execution time of the program given a new input. A transfer learning method is used to reuse an existing performance model and improve the prediction accuracy on a new platform that lacks historical execution data. Our experiments and analyses of three parallel applications, Graph500, GalaxSee, and SMG2000, on three different systems confirm that our method performs well, with less than 20 percent prediction error on average. Jingwei Sun 0001, Guangzhong Sun, Shiyan Zhan, Jiepeng Zhang, Yong Chen 0001 |
IEEE Trans. Computers | 5 |
| 2020 | PRS: A Pattern-Directed Replication Scheme for Heterogeneous Object-Based StorageabstractData replication is a key technique to achieve high data availability, reliability, and optimized performance in distributed storage systems. In recent years, with emerged new storage devices, heterogeneous object-based storage systems, such as a storage system with a mix of hard disk drives, solid state drives, and other non-volatile memory devices have become increasingly attractive since they combine the merits of different storage devices to deliver better promises. However, existing data replication schemes do not well consider distinct characteristics of heterogeneous storage devices yet, which could lead to suboptimal performance. This article introduces a new data replication scheme called Pattern-directed Replication Scheme (PRS) to achieve efficient data replication for heterogeneous storage systems. Different from traditional schemes, the PRS selectively replicates data objects and distributes replicas to various storage devices based on their characteristics. It aggregates objects that have I/O correlation into object groups by calculating object distance and makes replication for grouped objects according to application's data access pattern identified. In addition, the PRS uses a pseudo random algorithm to optimize replica placement by considering the storage device performance and capacity features. We have evaluated the pattern-directed replication scheme with extensive tests in Sheepdog, a typical object-based storage system. The experimental results confirm that it is a highly efficient replication scheme for heterogeneous storage systems. For instance, the read performance was improved by 105 percent to nearly 10x compared with existing replication schemes. Yong Chen 0001, Wei Xie 0017, Dong Dai 0001, Shuibing He, Weiping Wang 0005 |
IEEE Trans. Computers | 2 |
| 2020 | Segmented In-Advance Data Analytics for Fast Scientific DiscoveryabstractScientific discovery usually involves data generation, data preprocessing, data storage and data analysis. As the data volume exceeds a few terabytes (TB) in a single simulation run, the data movement, which happens during each cycle of the scientific discovery, continues to be the bottleneck in most scientific big data applications. A lot of research works have been conducted on reducing the data movement. Among the existing efforts and based on our previous research, reusing the analysis results shows a significant potential in optimizing the data movement between analysis operations. In this work, we propose the Segmented In-Advance (SIA) data analytics approach for optimizing the data movement and we also provide a cloud-based elastic distributed in-memory database to manage the intermediate analysis results. The fundamental idea of this Segmented In-Advance approach is to analyze the history operations and to predict the future interesting analytics operations. The predicted analysis operation is in-advance performed on the finer segmented dataset and the segmented results are distributed in an in-memory key-value store for future reuse. The evaluation shows that the segmented in-advance data analytics approach achieves 1.2X-6.1X speedup. The evaluation also shows a good scalability of the in-memory distributed data store. The proposed Segmented In-Advance data analytics approach is a promising data movement reduction solution for scientific big data applications and fast scientific discovery. Jialin Liu 0002, Yong Chen 0001 |
IEEE Trans. Cloud Comput. | 2 |
| 2020 | A Holistic Heterogeneity-Aware Data Placement Scheme for Hybrid Parallel I/O SystemsabstractWe presentH2DP, a holistic heterogeneity-aware data placement scheme for hybrid parallel I/O systems, which consist of HDD servers and SSD servers. Most of the existing approaches focus on server performance or application I/O pattern heterogeneity in data placement.H2DPconsiders three axes of heterogeneity: server performance, server space, and application I/O pattern. More specifically,H2DPdetermines the optimized stripe sizes on servers based on server performance, keeps only critical data on all hybrid servers and the rest data on HDD servers, and dynamically migrates data among different types of servers at run-time. This holistic heterogeneity-awareness enablesH2DPto achieve high performance by alleviating server load imbalance, efficiently utilizing SSD space, and accommodating application pattern variation. We have implemented a prototype ofH2DPunder MPICH2 atop OrangeFS. Extensive experimental results demonstrate thatH2DPsignificantly improve I/O system performance compared to existing data placement schemes. Shuibing He, Zheng Li 0006, Yanlong Yin, Xiaohua Xu 0002, Yong Chen 0001, Xian-He Sun |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2020 | A Highly Reliable Metadata Service for Large-Scale Distributed File SystemsabstractMany massive data processing applications nowadays often need long, continuous, and uninterrupted data accesses. Distributed file systems are used as the back-end storage to provide the global namespace management and reliability guarantee. Due to increasing hardware failures and software issues with the growing system scale, metadata service reliability has become a critical issue as it has a direct impact on file and directory operations. Existing metadata management mechanisms can provide fault tolerance capability to some level but are inadequate. They often have limitations in system availability, state consistence, and performance overhead and lack an effective mechanism to offer metadata reliability. This paper introduces a novel highly reliable metadata service to address these issues in large-scale file systems. Different from traditional strategies, this proposed reliable metadata service adopts a new active-standby architecture for fault tolerance and uses a holistic approach to improve file system availability. A new shared storage pool (SSP) is designed for transparent metadata synchronization and replication between active and standby servers. Based on the SSP, a new policy called multiple actives multiple standbys (MAMS) is presented to perform metadata service recovery in case of failures. A new global state recovery strategy and a smart client fault tolerance mechanism are achieved to maintain the continuity of metadata service. We have implemented such highly reliable metadata service in a prototype file system CFS (Clover file system) and conducted extensive tests to evaluate it. Experimental results confirm that it can significantly improve file system reliability with fast failover under different failure scenarios while having negligible influence on performance. Compared with typical reliability designs in Hadoop Avatar, Hadoop HA, and Boom-FS file systems, the mean-time-to-recovery (MTTR) with the highly reliable metadata service was reduced by 80.23, 65.46 and 28.13 percent, respectively. Yong Chen 0001, Weiping Wang 0005, Shuibing He, Dan Meng 0002 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | POSTER: Memory Hotspot Optimization for Data-Intensive ApplicationsabstractEmerging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, big data science, are data-intensive. The data-intensive workloads usually present irregular memory footprints with limited data locality, and thus incur frequent cache misses and a growing desire for memory bandwidth. Driven by this need, 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) are introduced to yield significantly higher throughput. However, the traditional interfaces and optimization methods for JEDEC DDR devices cannot fully exploit the potential performance of 3D-stacked memory to handle massive irregular memory accesses accompanied with data-intensive applications. In this paper, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices that is capable of optimizing memory access streams via request aggregation, hotspot detection, prefetching, and an associated hotspot-aware page policy. We present the HAM design and simulation implementation on RISC-V embedded cores with attached HMC devices. We have conducted extensive evaluations with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results reveal that HAM reduces redundant memory accesses by 37.51% and achieves a 4.19X enhancement on the prefetch buffer hit rate on average. Overall, HAM exhibits an average of 21.81% performance gain (up to 34.28%) and 35.07% power saving over the standard 3D-stacked memory. Xi Wang 0009, Jie Li 0057, Antonino Tumeo, John D. Leidel, Yong Chen 0001 |
PACT | 5 |
| 2019 | MTSAD: Multivariate Time Series Abnormality Detection and VisualizationabstractDetecting outliers is one of the fundamental tasks in visual analytics and valuable in many application domains, such as suspicious network cyberattack recognition. This paper introduces an approach to analyzing and visualizing high-dimensional time series, focusing on identifying multivariate observations that are significantly different from the others. We also propose a prototype, called MTSAD, to guide users when interactively exploring abnormalities in large time series. The prototype contains two views: the main window provides an overview of identified outliers overtime, the detail window investigates and explores the ranked temporal data entries based on their outlying contributions to the overall plots. The visual interface supports a full range of interactions, such as lensing, brushing and linking, ranking, and filtering. To validate the benefits and usefulness of our approach, we demonstrate MTSAD on real-world datasets of different numbers of attributes. Vung Pham, Ngan V. T. Nguyen, Jie Li 0057, Jon R. Hass, Yong Chen 0001, Tommy Dang |
IEEE BigData | 5 |
| 2019 | RE-Store: Reliable and Efficient KV-Store with Erasure Coding and ReplicationabstractIn-memory key/value stores (KV-stores) are a key building block for numerous applications running on a cluster. As cluster scales have grown, efficiency and availability have become increasingly critical characteristics. Traditional replication provides redundancy, but is inefficient due to its high storage overhead. Erasure coding can provide data reliability with significantly lower storage requirements, but is primarily used for long-term archival data due to the limitation of its write performance. Recent studies have attempted to combine these two techniques by using replication for frequently-updated metadata, and erasure coding for large, read-only data. In this study, we propose RE-Store, an in-memory key/value store system which utilizes a novel hybrid replication/erasure coding scheme to achieve both efficiency and reliability. RE-Store introduces replication into erasure coding by making one copy of each encoded datum and replacing partial parity with replicas for improved storage-efficiency. When failures occur, it uses these replicas to ensure data availability and thus avoids the inefficiencies of erasure coding during repair. RE-Store provides fault tolerance through fast, online recovery during different failure scenarios with little performance degradation. We have implemented RE-Store on a real key/value system and conducted extensive evaluations to validate its design and to study its performance, efficiency, and reliability. Experimental results show that RE-Store performs similarly to erasure coding and replication under normal operations while saving 18% to 34% of the memory used by replication when tolerating 2 to 4 failures. Yuzhe Li 0001, Weiping Wang 0005, Yong Chen 0001 |
CLUSTER | 4 |
| 2019 | Exploring Metadata Search Essentials for Scientific Data ManagementabstractScientific experiments and observations store massive amounts of data in various scientific file formats. Metadata, which describes the characteristics of the data, is commonly used to sift through massive datasets in order to locate data of interest to scientists. Several indexing data structures (such as hash tables, trie, self-balancing search trees, sparse array, etc.) have been developed as part of efforts to provide an efficient method for locating target data. However, efficient determination of an indexing data structure remains unclear in the context of scientific data management, due to the lack of investigation on metadata, metadata queries, and corresponding data structures. In this study, we perform a systematic study of the metadata search essentials in the context of scientific data management. We study a real-world astronomy observation dataset and explore the characteristics of the metadata in the dataset. We also study possible metadata queries based on the discovery of the metadata characteristics and evaluate different data structures for various types of metadata attributes. Our evaluation on real-world dataset suggests that trie is a suitable data structure when prefix/suffix query is required, otherwise hash table should be used. We conclude our study with a summary of our findings. These findings provide a guideline and offers insights in developing metadata indexing methodologies for scientific applications. Wei Zhang 0097, Surendra Byna, Chenxu Niu 0001, Yong Chen 0001 |
HiPC | 4 |
| 2019 | MAC: Memory Access Coalescer for 3D-Stacked MemoryabstractEmerging data-intensive applications, such as graph analytics and data mining, exhibit irregular memory access patterns. Research has shown that with these memory-bound applications, traditional cache-based processor architectures, which exploit locality and regular patterns to mitigate the memory-wall issue, are inefficient. Meantime, novel 3D-stacked memory devices, such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM), promise significant increases in bandwidth that appear extremely appealing for memory-bound applications. However, conventional memory interfaces designed for cache-based architectures and JEDEC DDR devices fit poorly with the 3D-stacked memory, which leads to significant under-utilization of the promised high bandwidth. Xi Wang 0009, Antonino Tumeo, John D. Leidel, Jie Li 0057, Yong Chen 0001 |
ICPP | 5 |
| 2019 | MIQS: metadata indexing and querying service for self-describing file formatsabstractScientific applications often store datasets in self-describing data file formats, such as HDF5 and netCDF. Regrettably, to efficiently search the metadata within these files remains challenging due to the sheer size of the datasets. Existing solutions extract the metadata and store it in external database management systems (DBMS) to locate desired data. However, this practice introduces significant overhead and complexity in extraction and querying. In this research, we propose a novel Metadata Indexing and Querying Service (MIQS), which removes the external DBMS and utilizes in-memory index to achieve efficient metadata searching. MIQS follows the self-contained data management paradigm and provides portable and schema-free metadata indexing and querying functionalities for self-describing file formats. We have evaluated MIQS with the state-of-the-art MongoDB-based metadata indexing solution. MIQS achieved up to 99% time reduction in index construction and up to 172kx search performance improvement with up to 75% reduction in memory footprint. Wei Zhang 0097, Surendra Byna, Houjun Tang, Brody Williams, Yong Chen 0001 |
SC | 5 |
| 2019 | Software-defined QoS for I/O in exascale computing
Yusheng Hua, Xuanhua Shi, Hai Jin 0001, Wei Liu 0004, Yong Chen 0001, Ligang He |
CCF Trans. High Perform. Comput. | 6 |
| 2019 | Foreword to the special issue for the Workshop on Parallel Programming Models and Systems Software for High-End Computing (P2S2 2017)
Pavan Balaji, Abhinav Vishnu, Yong Chen 0001 |
Parallel Comput. | 3 |
| 2019 | Vectorizing disks blocks for efficient storage system via deep learning
Dong Dai 0001, Forrest Sheng Bao, Xuanhua Shi, Yong Chen 0001 |
Parallel Comput. | 5 |
| 2019 | CARS: A contention-aware scheduler for efficient resource management of HPC storage systems
Weihao Liang, Yong Chen 0001, Jialin Liu 0002, Hong An |
Parallel Comput. | 2 |
| 2019 | Parallel programming models and systems software for high-end computing (P2S2 2018)
Min Si, Abhinav Vishnu, Yong Chen 0001 |
Parallel Comput. | 3 |
| 2019 | Client-side straggler-aware I/O scheduler for object-based parallel file systems
Neda Tavakoli, Dong Dai 0001, Yong Chen 0001 |
Parallel Comput. | 3 |
| 2019 | Guest Editor's Introduction: P2S2: SI 2016
Abhinav Vishnu, Pavan Balaji, Yong Chen 0001 |
Parallel Comput. | 3 |
| 2019 | Improving Nighttime Light Imagery With Location-Based Social Media DataabstractLocation-based social media have been extensively utilized in the concept of “social sensing” to exploit dynamic information about human activities, yet joint uses of social sensing and remote sensing images are underdeveloped at present. In this paper, the close relationship between the number of Twitter users and brightness of nighttime lights (NTL) over the contiguous United States is calculated and geotagged tweets are then used to upsample a stable light image for 2013. An associated outcome of the upsampling process is the solution of two major problems existing in the NTL image, pixel saturation, and blooming effects. Compared with the original stable light image, digital number (DN) values of the upsampled stable light image have larger correlation coefficients with gridded population (0.47 versus 0.09) and DN values of the new generation NTL image product (0.56 versus 0.52), i.e., the Visible Infrared Imaging Radiometer Suite day/night band image composite. In addition, total personal incomes of states are disaggregated to each pixel in proportion to the DN value of the pixel in the NTL images and then aggregate by counties. Personal incomes distributed by the upsampled NTL image are closer to the official demographic data than those distributed by the original stable light image. All of these results explore the potential of geotagged tweets to improve the quality of NTL images for more accurately estimating or mapping socioeconomic factors. Naizhuo Zhao, Wei Zhang 0097, Eric L. Samson, Yong Chen 0001, Guofeng Cao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | Managing Rich Metadata in High-Performance Computing Systems Using a Graph ModelabstractHigh-performance computing (HPC) systems generate huge amounts of metadata about different entities such as jobs, users, and files. Existing systems can efficiently record and manage part of these metadata, mainly the POSIX metadata of data files (e.g., file size, name, and permissions mode). But another important set of metadata, referred to as “rich” metadata in this study, which record not only wider range of entities (e.g., running processes and jobs) but also more complex relationships between them, are mostly missing in current HPC systems. Yet such rich metadata are critical for supporting many advanced data management functions such as identifying data sources and parameters behind a given result; auditing data usage; or understanding details about how inputs are transformed into outputs. To uniformly and efficiently manage the rich metadata generated in HPC systems, We propose to utilize a graph model in this study. We identify the key challenges of implementing such a graph-based HPC rich metadata management system and present GraphMeta, a graph-based rich metadata management system designed and optimized for HPC platforms, to tackle these challenges. Extensive evaluations on both synthetic and real HPC metadata workloads show its advantages in both performance and scalability compared with existing solutions. Dong Dai 0001, Yong Chen 0001, Philip H. Carns, John Jenkins, Wei Zhang 0097, Robert B. Ross |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | I/O Characteristics Discovery in Cloud Storage SystemsabstractThe data growth from many applications in clouds poses significant challenges to cloud storage systems. To deliver the best storage and I/O performance possible, it is often required to understand and leverage the I/O characteristics based on data accesses. A number of research studies have been carried out on this topic. However, most of them either utilize a limited number of data-access attributes, restricting the general applicability of the method for different applications, or heavily rely on the domain knowledge or expertise about applications' I/O behaviors to select the best representative features, introducing bias for certain workloads. To overcome these limitations, in this study, we present a new I/O characteristic discovery methodology. This method enables capturing data-access features as many as possible to eliminate human bias. It utilizes a machine-learning based strategy to derive the most important set of features automatically, and groups data objects with a clustering algorithm (DBSCAN) to reveal I/O characteristics discovered. These I/O characteristics revealed can direct I/O performance optimizations in numerous scenarios, such as in data prefeteching and data reorganization optimizations in cloud storage systems. Dong Dai 0001, Yong Chen 0001 |
IEEE CLOUD | 6 |
| 2018 | DART: distributed adaptive radix tree for efficient affix-based keyword search on HPC systemsabstractAffix-based search is a fundamental functionality for storage systems. It allows users to find desired datasets, where attributes of a dataset match an affix. While building inverted index to facilitate efficient affix-based keyword search is a common practice for standalone databases and for desktop file systems, building local indexes or adopting indexing techniques used in a standalone data store is insufficient for high-performance computing (HPC) systems due to the massive amount of data and distributed nature of the storage devices within a system. In this paper, we propose Distributed Adaptive Radix Tree (DART), to address the challenge of distributed affix-based keyword search on HPC systems. This trie-based approach is scalable in achieving efficient affix-based search and alleviating imbalanced keyword distribution and excessive requests on keywords at scale. Our evaluation at different scales shows that, comparing with the "full string hashing" use case of the most popular distributed indexing technique - Distributed Hash Table (DHT), DART achieves up to 55× better throughput with prefix search and with suffix search, while achieving comparable throughput with exact and infix searches. Also, comparing to the "initial hashing" use case of DHT, DART maintains a balanced keyword distribution on distributed nodes and alleviates excessive query workload against popular keywords. Wei Zhang 0097, Houjun Tang, Surendra Byna, Yong Chen 0001 |
PACT | 4 |
| 2018 | Atributed consistent hashing for heterogeneous storage systemsabstractStorage systems are critical building blocks of high-end computing systems and data centers. They demand the flexibility to distribute data effectively and provide high I/O performance. The consistent hashing algorithm is widely used in parallel/distributed file systems due to its decentralized design, scalability, and adaptability to node changes. However, it lacks efficiency in a heterogeneous environment where different storage devices, e.g. hard disk drives and solid state drives, co-exist. In this study, we propose an attributed consistent hashing (attributedCH), to overcome this deficiency. AttributedCH manages heterogeneous nodes on a consistent hashing ring and maintains attributes for each node to characterize distinct node features. It divides the hash ring into sectors and selects nodes from the sector with a comprehensive selection strategy. By considering different attributes, attributedCH achieves adaptive and efficient data placement for heterogeneous storage systems. We have carried out extensive evaluations and the evaluation results confirm that the attributedCH overcomes the deficiency of existing consistent hashing algorithms well and is particularly suitable for heterogeneous storage systems. Yong Chen 0001, Weiping Wang 0005 |
PACT | 2 |
| 2018 | AKIN: A Streaming Graph Partitioning Algorithm for Distributed Graph Storage SystemsabstractMany graph-related applications face the challenge of managing excessive and ever-growing graph data in a distributed environment. Therefore, it is necessary to consider a graph partitioning algorithm to distribute graph data onto multiple machines as the data comes in. Balancing data distribution and minimizing edge-cut ratio are two basic pursuits of the graph partitioning problem. While achieving balanced partitions for streaming graphs is easy, existing graph partitioning algorithms either fail to work on streaming workloads, or leave edge-cut ratio to be further improved. Our research aims to provide a better solution that fits the need of streaming graph partitioning in a distributed system, which further reduces the edge-cut ratio while maintaining rough balance among all partitions. We exploit the similarity measure on the degree of vertices to gather structuralrelated vertices in the same partition as much as possible, this reduces the edge-cut ratio even further as compared to the state-of-the-art streaming graph partitioning algorithm - FENNEL. Our evaluation shows that our streaming graph partitioning algorithm is able to achieve better partitioning quality in terms of edge-cut ratio (up to 20% reduction as compared to FENNEL) while maintaining decent balance between all partitions, and such improvement applies to various real-life graphs. Wei Zhang 0097, Yong Chen 0001, Dong Dai 0001 |
CCGrid | 2 |
| 2018 | Memory Coalescing for Hybrid Memory CubeabstractArguably, many data-intensive applications pose significant challenges to conventional architectures and memory systems, especially when applications exhibit non-contiguous, irregular, and small memory access patterns. The long memory access latency can dramatically slow down the overall performance of applications. The growing desire of high memory bandwidth and low latency access stimulate the advent of novel 3D-staked memory devices such as the Hybrid Memory Cube (HMC), which provides significantly higher bandwidth compared with the conventional JEDEC DDR devices. Even though many existing studies have been devoted to achieving high bandwidth throughput of HMC, the bandwidth potential cannot be fully exploited due to the lack of highly efficient memory coalescing and interfacing methodology for HMC devices. In this research, we introduce a novel memory coalescer methodology that facilitates memory bandwidth efficiency and the overall performance through an efficient and scalable memory request coalescing interface for HMC. We present the design and implementation of this approach on RISC-V embedded cores with attached HMC devices. Our evaluation results show that the new memory coalescer eliminates 47.47% memory accesses to HMC and improves the overall performance by 13.14% on average. Xi Wang 0009, John D. Leidel, Yong Chen 0001 |
ICPP | 3 |
| 2018 | PFault: A General Framework for Analyzing the Reliability of High-Performance Parallel File SystemsabstractHigh-performance parallel file systems (PFSes) are of prime importance today. However, despite the importance, their reliability is much less studied compared with that of local storage systems, largely due to the lack of an effective analysis methodology. Jinrui Cao, Om Rameshwar Gatla, Mai Zheng, Dong Dai 0001, Vidya Eswarappa, Yan Mu, Yong Chen 0001 |
ICS | 7 |
| 2018 | GRAM: A GPU-Based Property Graph Traversal and Query for HPC Rich Metadata Management
Wenke Li, Xuanhua Shi, Hong Huang 0001, Hai Jin 0001, Dong Dai 0001, Yong Chen 0001 |
NPC | 7 |
| 2018 | Exploiting Internal Parallelism for Address Translation in Solid-State DrivesabstractSolid-state Drives (SSDs) have changed the landscape of storage systems and present a promising storage solution for data-intensive applications due to their low latency, high bandwidth, and low power consumption compared to traditional hard disk drives. SSDs achieve these desirable characteristics using internal parallelism —parallel access to multiple internal flash memory chips—and a Flash Translation Layer (FTL) that determines where data are stored on those chips so that they do not wear out prematurely. However, current state-of-the-art cache-based FTLs like the Demand-based Flash Translation Layer (DFTL) do not allow IO schedulers to take full advantage of internal parallelism, because they impose a tight coupling between the logical-to-physical address translation and the data access. To address this limitation, we introduce a new FTL design called Parallel-DFTL that works with the DFTL to decouple address translation operations from data accesses. Parallel-DFTL separates address translation and data access operations into different queues, allowing the SSD to use concurrent flash accesses for both types of operations. We also present a Parallel-LRU cache replacement algorithm to improve the concurrency of address translation operations. To compare Parallel-DFTL against existing FTL approaches, we present a Parallel-DFTL performance model and compare its predictions against those for DFTL and an ideal page-mapping approach. We also implemented the Parallel-DFTL approach in an SSD simulator using real device parameters, and used trace-driven simulation to evaluate Parallel-DFTL’s efficacy. Our evaluation results show that Parallel-DFTL improved the overall performance by up to 32% for the real IO workloads we tested, and by up to two orders of magnitude with synthetic test workloads. We also found that Parallel-DFTL is able to achieve reasonable performance with a very small cache size and that it provides the best benefit for those workloads with large request size or with high write ratio. Wei Xie 0017, Yong Chen 0001, Philip C. Roth |
ACM Trans. Storage | 2 |
| 2017 | Lightweight Provenance Service for High-Performance ComputingabstractProvenance describes detailed information about the history of a piece of data, containing the relationships among elements such as users, processes, jobs, and workflows that contribute to the existence of data. Provenance is key to supporting many data management functionalities that are increasingly important in operations such as identifying data sources, parameters, or assumptions behind a given result; auditing data usage; or understanding details about how inputs are transformed into outputs. Despite its importance, however, provenance support is largely underdeveloped in highly parallel architectures and systems. One major challenge is the demanding requirements of providing provenance service in situ. The need to remain lightweight and to be always on often conflicts with the need to be transparent and offer an accurate catalog of details regarding the applications and systems. To tackle this challenge, we introduce a lightweight provenance service, called LPS, for high-performance computing (HPC) systems. LPS leverages a kernel instrument mechanism to achieve transparency and introduces representative execution and flexible granularity to capture comprehensive provenance with controllable overhead. Extensive evaluations and use cases have confirmed its efficiency and usability. We believe that LPS can be integrated into current and future HPC systems to support a variety of data management needs. Dong Dai 0001, Yong Chen 0001, Philip H. Carns, John Jenkins, Robert B. Ross |
PACT | 2 |
| 2017 | Pattern-Directed Replication Scheme for Heterogeneous Object-based StorageabstractData replication is a key technique to achieve data availability, reliability, and optimized performance in distributed storage systems and data centers. In recent years, with the emergence of new storage devices, heterogeneous object-based storage system, such as a storage system with the co-existence of hard disk drives and solid state drives, have become increasingly attractive as they combine merits of different storage devices to deliver better promise. However, existing data replication schemes do not place data based on heterogeneous device characteristics as well as considering distinct data access patterns. In this paper, we introduce a novel data replication scheme PRS to achieve efficient data replication for heterogeneous storage systems. Different from traditional schemes, the PRS groups objects according to data access patterns and distributes replicas to heterogeneous devices with their features. It uses a pseudo random algorithm to optimize replica layout by considering storage device performance and capacity. The experimental results confirm that PRS is a highly efficient replication scheme for heterogeneous storage systems. Wei Xie 0017, Dong Dai 0001, Yong Chen 0001 |
CCGrid | 4 |
| 2017 | OpenSoC system architect: An open toolkit for building soft-cores on FPGAsabstractGiven the recent difficulty in continuing the classic CMOS manufacturing density and power scaling curves, also known as Moore's Law and Dennard Scaling, respectively, we find that modern complex system architectures are increasingly relying upon accelerators in order to optimize the placement of specific computational workloads. In addition, large-scale computing infrastructures utilized in HPC, data intensive computing, and cloud computing must rely almost exclusively upon commodity device architectures provided by third-party manufacturers. The end result being a final system architecture that lacks specificity for the target software workload. At the same time, there is a trend in the FPGA space of much larger FPGAs with a lot more resources and hardened IP blocks, making this type of architecture design space exploration much easier. The OpenSoC System Architect infrastructure combines several open source design tools and methodologies into a central infrastructure for designing, developing, and verifying the necessary hardware and software modules required to implement application-specific processors for use in FPGAs. The end result is an infrastructure that permits rapid development and deployment of application-specific accelerators and softcores, including a fully functional software development tool chain. Farzad Fatollahi-Fard, David Donofrio, John Shalf, John D. Leidel, Xi Wang 0009, Yong Chen 0001 |
FPL | 6 |
| 2017 | IOGP: An Incremental Online Graph Partitioning Algorithm for Distributed Graph DatabasesabstractGraphs have become increasingly important in many applications and domains such as querying relationships in social networks or managing rich metadata generated in scientific computing. Many of these use cases require high-performance distributed graph databases for serving continuous updates from clients and, at the same time, answering complex queries regarding the current graph. These operations in graph databases, also referred to as online transaction processing (OLTP) operations, have specific design and implementation requirements for graph partitioning algorithms. In this research, we argue it is necessary to consider the connectivity and the vertex degree changes during graph partitioning. Based on this idea, we designed an Incremental Online Graph Partitioning (IOGP) algorithm that responds accordingly to the incremental changes of vertex degree. IOGP helps achieve better locality, generate balanced partitions, and increase the parallelism for accessing high-degree vertices of the graph. Over both real-world and synthetic graphs, IOGP demonstrates as much as 2x better query performance with a less than 10% overhead when compared against state-of-the-art graph partitioning algorithms. Dong Dai 0001, Wei Zhang 0097, Yong Chen 0001 |
HPDC | 3 |
| 2017 | Pipelining Computation and Optimization Strategies for Scaling GROMACS on the Sunway Many-Core Processor
Hong An, Junshi Chen 0003, Weihao Liang, Qingqing Xu, Yong Chen 0001 |
ICA3PP | 6 |
| 2017 | SSDUP: a traffic-aware ssd burst buffer for HPC systemsabstractMany high performance computing (HPC) applications are highly data intensive. Current HPC storage systems still use hard disk drives (HDDs) as their dominant storage devices, which suffer from disk head thrashing when accessing random data. New storage devices such as solid state drives (SSDs), which can handle random data access much more efficiently, have been widely deployed as the buffer to HDDs in many production HPC systems. Burst buffer has also been proposed to manage the SSD buffering of bursty write requests. Although burst buffer can improve I/O performance in many cases, we find that it has some limitations such as requiring large SSD capacity and harmonious overlapping between computation phase and data flushing stage. Xuanhua Shi, Wei Liu 0004, Hai Jin 0001, Chen Yu 0003, Yong Chen 0001 |
ICS | 6 |
| 2017 | Elastic Consistent Hashing for Distributed Storage SystemsabstractElastic distributed storage systems have been increasingly studied in recent years because power consumption has become a major problem in data centers. Much progress has been made in improving the agility of resizing small- and large-scale distributed storage systems. However, most of these studies focus on metadata based distributed storage systems. On the other hand, emerging consistent hashing based distributed storage systems are considered to allow better scalability and are highly attractive. We identify challenges in achieving elasticity in consistent hashing based distributed storage. These challenges cannot be easily solved by techniques used in current studies. In this paper, we propose an elastic consistent hashing based distributed storage to solve two problems. First, in order to allow a distributed storage to resize quickly, we modify the data placement algorithm using a primary server design and achieve an equal-work data layout. Second, we propose a selective data re-integration technique to reduce the performance impact when resizing a cluster. Our experimental and trace analysis results confirm that our proposed elastic consistent hashing works effectively and allows significantly better elasticity. Wei Xie 0017, Yong Chen 0001 |
IPDPS | 2 |
| 2017 | POSTER: IOGP: An Incremental Online Graph Partitioning for Large-Scale Distributed Graph DatabasesabstractLarge-scale graphs are becoming critical in various domains such as social network, scientific application, knowledge discovery, and even system software, etc. Many of those use cases require large-scale high-performance graph databases, which are designed for serving continuous updates from the clients, and at the same time, answering complex queries towards the current graph in an on-line manner. Those operations in graph databases, also referred as OLTP (online transaction processing) operations, need specific design and implementation in graph partitioning algorithms. In this study, we designed an incremental online graph partitioning (IOGP), optimized for OLTP workloads. It is designed to achieve better locality, generate balanced partitions, and increase the parallelism for accessing hotspots of the graph. Our evaluation results on both real world and synthetic graphs in both simulation and real system confirm a better performance on graph queries (as much as 2X) with small overheads during graph insertion (less than 10%). Dong Dai 0001, Wei Zhang 0097, Yong Chen 0001 |
PPoPP | 3 |
| 2017 | HMC-Sim-2.0: A co-design infrastructure for exploring custom memory cube operations
John D. Leidel, Yong Chen 0001 |
Parallel Comput. | 2 |
| 2017 | ASA-FTL: An adaptive separation aware flash translation layer for solid state drives
Wei Xie 0017, Yong Chen 0001, Philip C. Roth |
Parallel Comput. | 2 |
| 2016 | GraphMeta: A Graph-Based Engine for Managing Large-Scale HPC Rich MetadataabstractHigh-performance computing (HPC) systems face increasingly critical metadata management challenges, especially in the approaching exascale era. These challenges arise not only from exploding metadata volumes but also from increasingly diverse metadata, which contains data provenance and user-defined attributes in addition to traditional POSIX metadata. This "rich" metadata is critical to support many advanced data management functionality such as data auditing and validation. In our prior work, we presented a graph-based model that could be a promising solution to uniformly manage such rich metadata because of its flexibility and generality. At the same time, however, graph-based rich metadata management introduces significant challenges. In this study, we first identify the challenges presented by the underlying infrastructure in supporting scalable, high-performance rich metadata management. To tackle these challenges, we then present GraphMeta, a graph-based engine designed for managing large-scale rich metadata. We also utilize a series of optimizations designed for rich metadata graphs. We evaluate GraphMeta with both synthetic and real HPC metadata workloads and compare it with other approaches. The results show that its advantages in terms of rich metadata management in HPC systems, including better performance and scalability compared with existing solutions. Dong Dai 0001, Yong Chen 0001, Philip H. Carns, John Jenkins, Wei Zhang 0097, Robert B. Ross |
CLUSTER | 2 |
| 2016 | SSDUP: An Efficient SSD Write Buffer Using PipelineabstractHigh performance computing (HPC) applications are becoming more data-intensive and produce increasingly large I/O demands on storage systems. New storage devices such as SSD which has nearly no seek latency and high throughput have been widely used together with HDD to serve as a hybrid storage system. To solve the I/O bottleneck problem, existing hybrid storage solutions such as Burst Buffer have been proposed as intermediate layer between clients and disks to absorb burst I/O requests and improve write performance. However Burst Buffer needs sufficient SSD space to meet the maximum burst I/O requests which is still a costly solution. In this paper, we propose a hybrid architecture called SSDUP (an SSD write buffer Using Pipeline) for HPC storage systems, which uses NAND flash based SSD as a write-back buffer for HDD. With our efforts, SSDUP can achieve a good performance by using limited SSD space. Xuanhua Shi, Wei Liu 0004, Hai Jin 0001, Yong Chen 0001 |
CLUSTER | 5 |
| 2016 | Active Burst-Buffer: In-Transit Processing Integrated into Hierarchical StorageabstractThe data volume of many scientific applications has substantially increased in the past decade and continues to increase due to the rising needs of high-resolution and fine- granularity scientific discovery. The data movement between storage and compute nodes has become a critical performance factor and has attracted intense research and development attention in recent years. In this paper, we propose a novel solution, named Active burst-buffer, to reduce the unnecessary data movement and to speed up scientific workflow. Active burst-buffer enhances the existing burst-buffer concept with analysis capabilities by reconstructing the cached data to a logic file and providing a MapReduce-like computing framework for programming and executing the analysis codes. An extensive set of experiments were conducted to evaluate the performance of Active burst-buffer by comparing it against existing mainstream schemes, and more than 30% improvements were observed. The evaluations confirm that Active burst-buffer is capable of enabling efficient data analysis in-transit on burst-buffer nodes and is a promising solution to scientific discoveries with large-scale data sets. Michael Lang 0003, Latchesar Ionkov, Yong Chen 0001 |
NAS | 4 |
| 2016 | Parallel-DFTL: A Flash Translation Layer That Exploits Internal Parallelism in Solid State DrivesabstractSolid State Drives (SSDs) using flash memory storage technology present a promising storage solution for data-intensive applications due to their low latency, high bandwidth, and low power consumption compared to traditional hard disk drives. SSDs achieve these desirable characteristics using internal parallelism - parallel access to multiple internal flash memory chips - and a Flash Translation Layer (FTL) that determines where data is stored on those chips so that they do not wear out prematurely. Unfortunately, current state-of- the-art cache-based FTLs like the Demand-based Flash Translation Layer (DFTL) do not allow IO schedulers to take full advantage of internal parallelism because they impose a tight coupling between the logical-to-physical address translation and the data access. In this work, we propose an innovative IO scheduling policy called Parallel-DFTL that works with the DFTL to break the coupled address translation operations from data accesses. Parallel-DFTL schedules address translation and data access operations separately, allowing the SSD to use its flash access channel resources concurrently and fully for both types of operations. We present a performance model of FTL schemes that predicts the benefit of Parallel-DFTL against DFTL. We implemented our approach in an SSD simulator using real SSD device parameters, and used trace-driven simulation to evaluate its efficacy. Parallel-DFTL improved overall performance by up to 32% for the real IO workloads we tested, and up to two orders of magnitude for our synthetic test workloads. It is also found that Parallel-DFTL is able to achieve reasonable performance with a very small cache size. Wei Xie 0017, Yong Chen 0001, Philip C. Roth |
NAS | 2 |
| 2016 | SUORA: A Scalable and Uniform Data Distribution Algorithm for Heterogeneous Storage SystemsabstractThe data scale in many data centers is growing explosively with emerging applications and usages of big data technologies. Data distribution is a key issue in large-scale distributed storage systems to place petabytes of data or even beyond, among tens or hundreds of thousands of storage devices. In the meantime, heterogeneous storage systems, such as those having devices with hard disk drives (HDDs) and storage class memories (SCMs), have become increasingly popular for massive data storage due to balanced performance, capacity, and cost. Current data distribution algorithms can achieve efficient, scalable, and balanced mapping, but do not distinguish different characteristics of heterogeneous devices well. This paper presents a novel data distribution algorithm called SUORA (Scalable and Uniform storage via Optimally-adaptive and Random number Addressing), to take full advantage of heterogeneous devices. SUORA is a pseudo-random algorithm that uniformly distributes data cross a hybrid and tiered storage cluster. It divides heterogeneous devices, maps them onto different buckets and assigns them to various segments in each bucket. A pseudo-random and deterministic number sequence is generated to map data among segments and devices. Data movement is performed for achieving better read throughput while keeping load balance according to data hotness and bucket threshold. With considering distinct characteristics of heterogeneous storage devices well, the SUORA algorithm achieves a highly efficient adaptive data distribution for data centers and heterogeneous storage systems. Wei Xie 0017, Jason Noble, Kace Echo, Yong Chen 0001 |
NAS | 5 |
| 2016 | Special Issue on Parallel Programming Models and Systems Software for High-End Computing
Pavan Balaji, Abhinav Vishnu, Yong Chen 0001 |
Parallel Comput. | 3 |
| 2016 | An asynchronous traversal engine for graph-based rich metadata management
Dong Dai 0001, Philip H. Carns, Robert B. Ross, John Jenkins, Nicholas Muirhead, Yong Chen 0001 |
Parallel Comput. | 6 |
| 2015 | Two-mode data distribution scheme for heterogeneous storage in data centersabstractFast growing "Big Data" demands present new challenges to the traditional distributed storage system solutions. In order to support cloud-scale data centers, new types of distributed storage systems are emerging. They are designed to scale to thousands of nodes, maintain petabytes of data and be highly reliable. The support for virtual machines is also becoming essential as it is one of the most important technology that supports cloud computing. To meet these needs, these distributed storage systems are implemented with advanced data distribution schemes. Data are striped and distributed across the storage cluster based on distribution algorithms instead of mapping tables. The existing algorithms usually balance the data distribution across nodes proportional to their capacity. However, they overlook distinct performance characteristics across different nodes and devices in the emerging heterogeneous storage environment. We propose a two-mode data distribution scheme in this study to maximize the overall performance and keep data balanced across the storage cluster at the same time. The working principle of the two-mode data distribution scheme is provided. We also present a new data read and write strategy to work with the two-mode scheme. We evaluate the computation time for data distribution using two-mode scheme and analyze its implication on the overall IO performance. We expect significant performance improvement while it still needs more analytical and experimental evaluation to further examine the details. Wei Xie 0017, Mark Reyes, Jason Noble, Yong Chen 0001 |
IEEE BigData | 5 |
| 2015 | GraphTrek: Asynchronous Graph Traversal for Property Graph-Based Metadata ManagementabstractProperty graphs are a promising data model for rich metadata management in high-performance computing (HPC) systems because of their ability to represent not only metadata attributes but also the relationships between them. A property graph can be used to record the relationships between users, jobs, and data, for example, with unique annotations for each entity. This high-volume, power-law distributed use case is a natural fit for an out-of-core distributed property graph database. Such a system must support live updates (to ingest production information in real time), low-latency point queries (for frequent metadata operations such as permission checking), and large-scale traversals (for provenance data mining). Large-scale property graph traversals are particularly challenging for distributed graph databases, however. Most existing graph databases implement a "level-synchronous" breadth-first search algorithm that relies on global synchronization in each traversal step. This traversal model performs well in many problem domains, but a rich metadata management system is characterized by imbalanced graphs, long traversal lengths, and concurrent workloads, each of which has the potential to introduce or exacerbate stragglers. We define stragglers as abnormally slow steps (or servers) in a graph traversal that lead to low overall throughput for synchronous traversal algorithms. The straggler problem can be mitigated by the use of asynchronous traversal algorithms. Asynchronous traversal has been successfully demonstrated in graph processing frameworks, but such systems require the graph to be loaded into a separate batch-processing framework. In this work, we propose GraphTrek, a general asynchronous graph traversal engine working with graph databases for processing rich metadata management in their native format. We also outline a traversal-aware query language and key optimizations (traversal-affiliate caching and execution merging) necessary for efficient performance. Our experiments show that the asynchronous graph traversal engine is more efficient than its synchronous counterpart in the case of HPC rich metadata processing, where more servers are involved and larger traversals are needed. Dong Dai 0001, Philip H. Carns, Robert B. Ross, John Jenkins, Kyle Blauer, Yong Chen 0001 |
CLUSTER | 6 |
| 2015 | A Cache Management Scheme for Hiding Garbage Collection Latency in Flash-Based Solid State DrivesabstractRecent advancements in flash-based solid state drive (SSD) make it a highly desirable storage device, especially for data-intensive applications. There are significant more SSDs used in data centers and high performance computing systems. SSDs perform one or two orders better than traditional hard disk drives generally. However, the performance of random writes on SSDs, especially small random writes, is still largely limited due to the garbage collection (GC) process. Existing work tried to utilize the on-device RAM as a write cache to improve the write performance, however directly utilizing it as a normal write cache under utilizes the RAM cache. In this poster, we present our initial study of a cache management scheme that hides the GC latency. Wei Xie 0017, Yong Chen 0001 |
CLUSTER | 2 |
| 2015 | MAMS: A Highly Reliable Policy for Metadata ServiceabstractMost mass data processing applications nowadays often need long, continuous, and uninterrupted data access. Parallel/distributed file systems often use multiple metadata servers to manage the global namespace and provide a reliability guarantee. With the rapid increase of data amount and system scale, the probability of hardware or software failures keeps increasing, which easily leads to multiple points of failures. Metadata service reliability has become a crucial issue as it affects file and directory operations in the event of failures. Existing reliable metadata management mechanisms can provide fault tolerance but have disadvantages in system availability, state consistence, and performance overhead. This paper introduces a new highly reliable policy called MAMS (multiple actives multiple standbys) to ensure multiple metadata service reliability in file systems. Different from traditional strategies, the MAMS divides metadata servers into different replica groups and maintains more than one standby node for failover in each group. Combining the global view with distributed protocols, the MAMS achieves an automatic state transition and service takeover. We have implemented the MAMS policy in a prototyping file system and conducted extensive tests to validate and evaluate it. The experimental results confirm that the MAMS policy can achieve a faster transparent fault tolerance in different error scenarios with less influence on metadata operations. Compared with typical designs in Hadoop Avatar, Hadoop HA, and Boom-FS file systems, the mean time to recovery (MTTR) with the MAMS was reduced by 80.23%, 65.46% and 28.13%, respectively. Yong Chen 0001, Weiping Wang 0005, Dan Meng 0002 |
ICPP | 2 |
| 2015 | A virtual shared metadata storage for HDFSabstractHadoop is a popular open-source framework that allows distributed analysis of large datasets using the MapReduce programming model. A distributed file system HDFS is implemented to provide high-throughput access to datasets. HDFS can achieve high performance metadata service but has two disadvantages. First, when the metadata server stores metadata on persistent devices, it is restricted to read and write operations of local disks. Second, it also lacks effective methods for metadata synchronization and replication, which is critical for metadata availability and reliability. In this research, we introduce a novel Virtual Shared Storage Pool (VSSP) concept and design for storing and sharing metadata in HDFS. The VSSP is a virtual storage device which is built on existing servers and transparent to upper layers. Two strategies, a journal synchronization based on the 2PC protocol and a fine-grained image replication, are introduced in the VSSP according to different metadata access features. The VSSP not only reduces the overhead on metadata modification operations, but also improves the I/O performance for namespace storage. Experimental results show that the VSSP improved the average performance by 40.51% and 23.46% when writing logs compared with the BookKeeper and Hadoop QJM. The average image read and write throughput was nearly 5 times and 2.4 times better than NFS and the original approach. These results confirm that the proposed VSSP solution significantly improves the metadata access performance, scalability, and reliability for HDFS. Yong Chen 0001, Xiaoyan Gu 0001, Weiping Wang 0005, Dan Meng 0002 |
NAS | 2 |
| 2015 | Performance model-directed data sieving for high-performance I/O
Yong Chen 0001, Yin Lu, Prathamesh Amritkar, Rajeev Thakur |
J. Supercomput. | 1 |
| 2015 | Mammoth: Gearing Hadoop Towards Memory-Intensive MapReduce ApplicationsabstractThe MapReduce platform has been widely used for large-scale data processing and analysis recently. It works well if the hardware of a cluster is well configured. However, our survey has indicated that common hardware configurations in small- and medium-size enterprises may not be suitable for such tasks. This situation is more challenging for memory-constrained systems, in which the memory is a bottleneck resource compared with the CPU power and thus does not meet the needs of large-scale data processing. The traditional high performance computing (HPC) system is an example of the memory-constrained system according to our survey. In this paper, we have developed Mammoth, a new MapReduce system, which aims to improve MapReduce performance using global memory management. In Mammoth, we design a novel rule-based heuristic to prioritize memory allocation and revocation among execution units (mapper, shuffler, reducer, etc.), to maximize the holistic benefits of the Map/Reduce job when scheduling each memory unit. We have also developed a multi-threaded execution engine, which is based on Hadoop but runs in a single JVM on a node. In the execution engine, we have implemented the algorithm of memory scheduling to realize global memory management, based on which we further developed the techniques such as sequential disk accessing, multi-cache and shuffling from memory, and solved the problem of full garbage collection in the JVM. We have conducted extensive experiments to compare Mammoth against the native Hadoop platform. The results show that the Mammoth system can reduce the job execution time by more than 40 percent in typical cases, without requiring any modifications of the Hadoop programs. When a system is short of memory, Mammoth can improve the performance by up to 5.19 times, as observed for I/O intensive applications, such as PageRank. We also compared Mammoth with Spark. Although Spark can achieve better performance than Mammoth for interactive and iterative applications when the memory is sufficient, our experimental results show that for batch processing applications, Mammoth can adapt better to various memory environments and outperform Spark when the memory is insufficient, and can obtain similar performance as Spark when the memory is sufficient. Given the growing importance of supporting large-scale data processing and analysis and the proven success of the MapReduce platform, the Mammoth system can have a promising potential and impact. Xuanhua Shi, Ligang He, Lu Lu 0006, Hai Jin 0001, Yong Chen 0001, Song Wu 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2014 | Provenance-based object storage prediction scheme for scientific big data applicationsabstractObject storage has been increasingly adopted in high-performance computing for scientific, big data applications. With object storage, applications usually use object IDs, queries, or collections to identify the data instead of using files. Since the object store changes the way data is accessed in applications, it introduces new challenges for I/O prediction, which used to work based on interfile or intrafile pattern detection. The key challenge is that the inputs of object-based applications are no longer expressed as static file names: they become much more dynamic and unstable, hidden inside application logic. Traditional prediction strategies do not work well in such conditions. In this paper, we introduce the use of provenance information, which was collected for data management in high-performance computing systems, in order to build an accurate coarse-grained (object-level) input prediction. The prediction results can be preloaded into a burst buffer to accelerate future reads. To our best knowledge, this study is the first to use provenance information in object stores to predict application inputs. Evaluation results confirm the effectiveness and accuracy of our provenance-based prediction and show that the proposed prediction system is feasible for real-work deployment. Dong Dai 0001, Yong Chen 0001, Dries Kimpe, Robert B. Ross |
IEEE BigData | 2 |
| 2014 | In-advance data analytics for reducing time to discoveryabstractScientific workflow involves data generation, data analysis, and knowledge discovery. As the data volume exceeds a few terabytes (TB) in a single simulation run, the data movement, which happens among data generation, data analysis, and knowledge discovery, becomes a bottleneck in most scientific big data applications. Our previous work shows that reusing the analysis results can have a significant potential in reducing the overlap between data movement among compute nodes and storage nodes. In this work, we propose a new in-advance data analytics method to augment the result reuse. The fundamental idea of this in-advance data analytics method and its prototyping system is to predict the potential useful analytics operations by studying the users' analysis pattern. The predicted analysis operation is pro-actively performed on existing data and the analysis results are stored in an in-memory database for result reuse. The evaluation shows that the in-advance data analytics method and its prototyping system gains 1.2X-6.1X speedup in I/O performance improvement with 50% data overlapping and 10%-100% operation recommendation hit rate. The proposed in-advance data analytics method brings a new promising data reduction solution for big data applications. Jialin Liu 0002, Yin Lu, Yong Chen 0001 |
IEEE BigData | 3 |
| 2014 | Provenance-Based Prediction Scheme for Object Storage System in HPCabstractObject-based storage model is recently widely adopted both in industry and academia to support growingly data intensive applications in high-performance computing. However, the I/O prediction strategies which have been proven effective in traditional parallel file systems, have not been thoroughly studied under this new object-based storage model. There are new challenges introduced from object storage that make traditional prediction systems not work properly. In this paper, we propose a new I/O access prediction system based on provenance analysis on both applications and objects. We argue that the provenance, which contains metadata that describes the history of data, reveals the detailed information about applications and data sets, which can be used to capture the system status and provide accurate I/O prediction efficiently. Our current evaluations based on real-world trace data (Darshan datasets) simulation also confirm that provenance-based prediction system is able to provide accurate predictions for object storage systems. Dong Dai 0001, Yong Chen 0001, Dries Kimpe, Robert B. Ross |
CCGRID | 2 |
| 2014 | Iteration Based Collective I/O Strategy for Parallel I/O SystemsabstractMPI collective I/O is a widely used I/O method that helps data-intensive scientific applications gain better I/O performance. However, it has been observed that existing collective I/O strategies do not perform well due to the access contention problem. Existing collective I/O optimization strategies mainly focus on the I/O phase efficiency and ignore the shuffle cost that may limit the potential of their performance improvement. We observe that as the size of I/O becomes larger, one I/O operation from the upper application would be separated into several iterations to complete. So, I/O requests in each file domain do not necessarily issue to the parallel file system simultaneously unless they are carried out within the same iteration step. Based on that observation, this paper proposes a new collective I/O strategy that reorganizes I/O requests within each file domain instead of coordinating requests across file domains, such that we can eliminate access contentions without introducing extra shuffle cost between aggregators and computing processes. Using benchmark workloads IOR, we evaluate our new strategy and compare with the conventional one. The proposed strategy achieves up to 47%-63% I/O bandwidth improvement compared to the existing ROMIO collective I/O strategy. Xuanhua Shi, Hai Jin 0001, Song Wu 0001, Yong Chen 0001 |
CCGRID | 5 |
| 2014 | An Adaptive Separation-Aware FTL for Improving the Efficiency of Garbage Collection in SSDsabstractHot/cold data separation in flash-based solid state drives has been considered important to the overall performance due to the costly garbage collection overhead. This work proposes a method that accurately and naturally identifies and separates hot/cold data while only incurs minimal overhead. The proposed method only requires minimal on-device RAM space. Simulation results have shown that the proposed ASA-FTL reduce the GC overhead by up to 33% and improve the overall response time by 9% against the most advanced existing FTL in both real workloads and synthetic workloads. Wei Xie 0017, Yong Chen 0001 |
CCGRID | 2 |
| 2014 | Domino: an incremental computing framework in cloud with eventual synchronizationabstractIn recent years, more and more applications in cloud have needed to process large-scale on-line data sets that evolve over time as entries are added or modified. Several programming frameworks, such as Percolator and Oolong, are proposed for such incremental data processing and can achieve efficient updates with an event-driven abstraction. However, these frameworks are inherently asynchronous, leaving the heavy burden of managing synchronization to applications developers. Such a limitation significantly restricts their usability. In this paper, we introduce a trigger-based incremental computing framework, called Domino, with a flexible synchronization mechanism and runtime optimizations to coordinate parallel triggers efficiently. With this new framework, both synchronous and asynchronous applications can be seamlessly developed. Use cases and current evaluation results confirm that the new Domino programming model delivers sufficient performance and is easy to use in large-scale distributed computing. Dong Dai 0001, Yong Chen 0001, Dries Kimpe, Robert B. Ross, Xuehai Zhou |
HPDC | 2 |
| 2014 | Revealing applications' access pattern in collective I/O for cache managementabstractCollective I/O is a critical I/O strategy on high-performance parallel computing systems that enables programmers to reveal parallel processes' I/O accesses collectively and makes possible for the parallel I/O middleware to carry out I/O requests in a highly efficient manner. Collective I/O has been proven as a core parallel I/O optimization technique. However, due to the collective nature of collective I/O, the access pattern of each individual process can be lost after I/O requests are aggregated at the parallel I/O middleware layer. In this study, we analyze this issue in detail. We show that such lost access pattern can have a negative impact on underlying caching algorithms' view of locality and can result in many unnecessary cache misses in low level buffer caches and additional disk accesses. To address this issue, we propose to reveal unseen access patterns - performing collective I/O but more importantly retaining applications' access patterns to underlying cache management. With such an idea, we have prototyped a new collective I/O aware cache management methodology. The evaluations with various cache management algorithms have confirmed clear advantages over the existing collective I/O strategy that throws away applications' original access pattern. Yin Lu, Yong Chen 0001, Robert Latham |
ICS | 2 |
| 2014 | Two-Choice Randomized Dynamic I/O Scheduler for Object Storage SystemsabstractObject storage is considered a promising solution for next-generation (exascale) high-performance computing platform because of its flexible and high-performance object interface. However, delivering high burst-write throughput is still a critical challenge. Although deploying more storage servers can potentially provide higher throughput, it can be ineffective because the burst-write throughput can be limited by a small number of stragglers (storage servers that are occasionally slower than others). In this paper, we propose a two-choice randomized dynamic I/O scheduler that schedules the concurrent burst-write operations in a balanced way to avoid stragglers and hence achieve high throughput. The contributions in this study are threefold. First, we propose a two-choice randomized dynamic I/O scheduler with collaborative probe and preassign strategies. Second, we design and implement a redirect table and metadata maintainer to address the metadata management challenge introduced by dynamic I/O scheduling. Third, we evaluate the proposed scheduler with both simulation tests and experimental tests in an HPC cluster. The evaluation results confirm the scalability and performance benefits of the proposed I/O scheduler. Dong Dai 0001, Yong Chen 0001, Dries Kimpe, Robert B. Ross |
SC | 2 |
| 2014 | Guest Editors' introduction to the special issue on "DISCS-2013"
Philip C. Roth, Yong Chen 0001 |
Parallel Comput. | 2 |
| 2013 | Multilevel Active Storage for big data applications in high performance computingabstractGiven the growing importance of supporting dataintensive sciences and big data applications, an effective HPC I/O solution has become a key issue and has attracted intensive attention in recent years. Active storage has been shown effective in reducing data movement and network traffic as a potential new I/O solution. Existing prototypes and systems, however, are primarily designed for read-intensive applications. In addition, they generally assume that offloaded processing kernels have small computational demands, which makes this solution a poor fit for data-intensive operations that have significant computational demands, including write-intensive operations. In this research, we propose a new Multilevel Active Storage (MAS) solution. The new MAS design can support and handle both read- and write-intensive operations, as well as complex operations that have considerable computational demands. Experimental tests have been carried out and confirmed that the MAS approach is feasible and outperformed existing approaches. The new multilevel active storage design has a potential to deliver a high performance I/O solution for big data applications in HPC. Michael Lang 0003, Yong Chen 0001 |
IEEE BigData | 3 |
| 2013 | Using pattern-models to guide SSD deployment for Big Data applications in HPC systemsabstractFlash-memory based Solid State Drives (SSDs) embrace higher performance and lower power consumption compared to traditional storage devices (HDDs). These benefits are needed in HPC systems, especially with the growing demand of supporting Big Data applications. In this paper, we study placement and deployment strategies of SSDs in HPC systems to maximize the performance improvement, given a practical fixed hardware budget constraint. We propose a pattern-model approach to guide SSD deployment for HPC systems through two steps; characterizing workload and mapping deployment strategy. The first step is responsible for characterizing the access patterns of the workload and the second step contributes the actual deployment recommendation for Parallel File System (PFS) configuration combining with an analytical model. We have carried out initial experimental tests and the results confirmed that the proposed approach can guide placement of SSDs in HPC systems for accelerating data accesses. Our research will be helpful in guiding designs and developments for Big Data applications in current and projected HPC systems including exascale systems. Philip C. Roth, Yong Chen 0001 |
IEEE BigData | 3 |
| 2013 | Segmented analysis for reducing data movementabstractMany scientific applications nowadays generate a few terabytes (TB) of data in a single run and the data sizes are expected to reach petabytes (PB) in the near future. Enabling fast extraction of knowledge through analyzing these large datasets holds the key to faster scientific discoveries. However, reading data from traditional storage subsystem is a slow process as the I/O performance lags far behind computational performance. Reducing data movement from the storage subsystem is widely considered a viable option for improving performance of data analysis. In this paper, we propose Segmented Analysis, a data movement reduction strategy through reusing results, where multiple similar analysis tasks process the same segments of data. The basic idea is to segment the data accessed in an analysis task, to process the data segments with a given analysis task, and to store the results of segments in a cache for future use. In future, when an analysis task needs to perform the same process on the data segments for which the results are available in the cache, the task can avoid moving data and performing computation for the available results. The Segmented Analysis framework contains modules for computation and I/O access overlap detection, in situ segmentation, and segment result caching. We evaluate the Segmented Analysis strategy by varying factors like the overlap rate among analysis tasks, the request size and the granularity of segmentation. We observed 2X to 13X I/O and to 2X to 8X computation speedups when the overlap is above 50%. Jialin Liu 0002, Surendra Byna, Yong Chen 0001 |
IEEE BigData | 3 |
| 2013 | Locality-driven high-level I/O aggregation for processing scientific datasetsabstractScientific I/O libraries, like PnetCDF, ADIOS, and HDF5, have been commonly used to facilitate the array-based scientific dataset processing. The underlying physical data layout information, however, is usually hidden from the upper layer's logical access. Such mismatching can lead to poor I/O. In this research, we have observed performance degradation in the case of concurrent sub-array accesses, where overlaps among calls that access sub-arrays led to high contention on storage servers due to the logical-physical mismatching. We propose a locality-driven high-level I/O aggregation approach to address these issues in this work. By designing a logical-physical mapping scheme, we try to utilize the scientific dataset's structured formats and the file systems' data distribution to resolve the mismatching issue. Therefore the I/O can be carried out in a locality-driven fashion. The proposed approach is effective and complements the existing I/O strategies, such as the independent I/O and collective I/O strategy. We have also carried out experimental tests and the results confirm the performance improvement compared to existing I/O strategies. The proposed locality-driven highlevel I/O aggregation approach holds a promise for efficiently processing scientific datasets, which is critical for the data intensive or big data computing era. Jialin Liu 0002, Bradly Crysler, Yin Lu, Yong Chen 0001 |
IEEE BigData | 4 |
| 2013 | Hierarchical I/O Scheduling for Collective I/OabstractThe non-contiguous access pattern of many scientific applications results in a large number of I/O requests, which can seriously limit the data-access performance. Collective I/O has been widely used to address this issue. However, the performance of collective I/O could be dramatically degraded in today's high-performance computing system due to the increasing shuffle cost caused by highly concurrent data accesses. This situation tends to be even worse as many applications become more and more data intensive. Previous research has primarily focused on optimizing I/O access cost in collective I/O but largely ignored the shuffle cost involved. In this study, we propose a new hierarchical I/O scheduling (HIO) algorithm to address the increasing shuffle cost in collective I/O. The fundamental idea is to schedule applications' I/O requests based on a shuffle cost analysis to achieve the optimal overall performance, instead of achieving optimal I/O accesses only. The algorithm is currently evaluated with the MPICH2 andPVFS2. Both theoretical analysis and experimental tests show that the proposed hierarchical I/O scheduling has a potential in addressing the degraded performance issue of collective I/O with highly concurrent accesses. Jialin Liu 0002, Yong Chen 0001 |
CCGRID | 2 |
| 2013 | Cost-Aware Client-Side File Caching for Data-Intensive ApplicationsabstractParallel and distributed file systems are widely used to provide high throughput in high-performance computing and Cloud computing systems. To increase the parallelism, I/O requests are partitioned into multiple sub-requests (or `flows') and distributed across different data nodes. The performance of file systems is extremely poor if data nodes have highly unbalanced response time. Client-side caching offers a promising direction for addressing this issue. However, current work has primarily used client-side memory as a read cache and employed a write-through policy which requires synchronous update for every write and significantly under-utilizes the client-side cache when the applications are write-intensive. Realizing that the cost of an I/O request depends on the struggler sub-requests, we propose a cost-aware client-side file caching (CCFC) strategy, that is designed to cache the sub-requests with high I/O cost on the client end. This caching policy enables a new trade-off across write performance, consistency guarantee and cache size dimensions. Using benchmark workloads MADbench2, we evaluate our new cache policy alongside conventional write-through. We find that the proposed CCFC strategy can achieve up to 110% throughput improvement compared to the conventional write-through policies with the same cache size on an 85-node cluster. Yaning Huang, Hai Jin 0001, Xuanhua Shi, Song Wu 0001, Yong Chen 0001 |
CloudCom (2) | 5 |
| 2013 | Unified and efficient HEC storage system with a working-set based reorganization schemeabstractHigh-end computing (HEC) applications and simulations have become increasingly data intensive. The pressure on the storage system capability has substantially increased in recent years. Traditional hard disk drives (HDD) are dominant storage devices in HEC, but suffer seek time delays and rotational latencies. The emerged storage class memory such as Solid State Drive (SSD) provides a new promising high bandwidth and low latency storage solution, but with inherent limitations of small capacity, limited writing cycles, and high cost. SSDs and HDDs have complement characteristics in nature and there is a desire to combine and unify them to best serve HEC workloads. In this paper, we introduce our initial study of a novel Working-Set based Reorganization Scheme (WS-ROS), to manage and leverage the merits of both SSDs and HDDs and to provide a highly efficient storage system. The goal of the proposed scheme is to leverage SSDs service more read requests through keeping active data and to leverage HDDs to keep large non-active data and extend SSDs writing cycles. The preliminary evaluations have shown that the proposed scheme has significant advantages and holds a promise as a unified storage management scheme for HEC workloads. Yong Chen 0001 |
CLUSTER | 2 |
| 2013 | Runtime system design of decoupled execution paradigm for data-intensive high-end computingabstractHigh performance computing are widely used for scientific discoveries by running scientific computation programs. Many of these applications are getting more and more data intensive [1]. They generate or access huge amount of data during some execution phases. However, traditional supercomputers are designed for computing-intensive tasks. They usually have highdensity clusters of processing cores and their storage systems are placed remotely and connected to the computing clusters with networks. This separation of the computing system and the storage system causes the data Input/Output performance bottleneck, especially for the data-intensive phases of HPC applications. This bottleneck degrades the HPC system's efficiency. Yanlong Yin, Hassan Eslami, Xian-He Sun, Yong Chen 0001, Rajeev Thakur, William Gropp |
CLUSTER | 6 |
| 2013 | Fast data analysis with integrated statistical metadata in scientific datasetsabstractScientific datasets and libraries, such as HDF5, ADIOS, and NetCDF, have been used widely in many data-intensive applications. These libraries have their special file formats and I/O functions to provide efficient access to large datasets. Recent studies have started to utilize indexing, subsetting, and data reorganization to manage the increasingly large datasets. In this work, we present an approach to boost the data analysis performance, namely Fast Analysis with Statistical Metadata (FASM), via data subsetting and integrating a small amount of statistics into the original datasets. The added statistical information illustrates the data shape and provides knowledge of the data distribution; therefore the original I/O libraries can utilize these statistical metadata to perform fast queries and analyses. Various subsetting schemes can affect the access pattern and the I/O performance. We present a comparison study of different subsetting schemes by focusing on three dominant factors, the shape, the concurrency, and the locality. The added statistical metadata slightly increases the original data size, and we evaluate the cost and trade-off as well. This work is the first study that utilizes statistical metadata with various subsetting schemes to perform fast queries and analyses on large datasets. The proposed FASM approach is currently evaluated with the PnetCDF on Lustre file systems, but can also be implemented with other scientific libraries. The FASM can potentially lead to a new dataset design and can have an impact on big data analysis. Jialin Liu 0002, Yong Chen 0001 |
CLUSTER | 2 |
| 2013 | Special issue on programming models, systems software, and tools for High-End Computing
Yong Chen 0001, Pavan Balaji, Abhinav Vishnu |
Parallel Comput. | 1 |
| 2013 | Guest Editors' introduction
Abhinav Vishnu, Pavan Balaji, Yong Chen 0001 |
J. Supercomput. | 3 |
| 2012 | Checkpointing Orchestration: Toward a Scalable HPC Fault-Tolerant EnvironmentabstractCheck pointing is widely used in technical computing. However, the overhead of check pointing is a subject of increasing in concern in recent years, especially for large-scale parallel computer systems. In these systems, check pointing generates a huge number of concurrent I/O writes. The burst of writes plus the worsening I/O-wall problem often leads to network and I/O congestion, and makes the overall system performance painfully slow. Recognizing contention as a dominant performance factor, in this paper we propose a systematic approach named check pointing orchestration to reduce write contention, which combines the marshaling of concurrent checkpoint requests and the adopting of vertical data access in coordination. A prototype of the proposed check pointing orchestration approach has been implemented at the system-level under Open MPI over the PVFS2 file system. Extensive experiments based on NPB benchmarks have been conducted to verify the design and implementation. Experimental results show that check pointing orchestration reduced the check pointing cost at a degree of more than 30%. Check pointing cost was halved for 4 out of 5 the C class NPB benchmarks. Hui Jin 0001, Tao Ke, Yong Chen 0001, Xian-He Sun |
CCGRID | 3 |
| 2012 | DOSAS: Mitigating the Resource Contention in Active Storage SystemsabstractActive storage provides an effective method to mitigate the I/O bottleneck problem of data intensive high performance computing applications. It can reduce the amount of data transferred as the application runs by moving appropriate computations close to the data. Prior research has achieved considerable progress in developing several active storage prototypes. However, existing studies have neglected the impact of resource contention when concurrent processes request IOoperations from the same storage node simultaneously, which happens frequently in practice. In this paper, we analyze the impact of resource contention on active storage systems. Motivated by our analysis, we propose a novel Dynamic Operation Scheduling Active Storage architecture to address the resource contention issue. It offloads the active processing operations dynamically between storage nodes and compute nodes according to the system environment. By evaluating our architecture, we observed that: (1) resource contention is a critical problem for active storage systems, (2) the proposed dynamic operation scheduling method mitigates the problem, and (3) the new active storage architecture outperforms existing active storage systems. Yong Chen 0001, Philip C. Roth |
CLUSTER | 2 |
| 2012 | A Decoupled Execution Paradigm for Data-Intensive High-End ComputingabstractHigh-end computing (HEC) applications in critical areas of science and technology tend to be more and more data intensive. I/O has become a vital performance bottleneck of modern HEC practice. Conventional HEC execution paradigms, however, are computing-centric for computation intensive applications. They are designed to utilize memory and CPU performance and have inherent limitations in addressing the critical I/O bottleneck issues of HEC. In this study, we propose a decoupled execution paradigm (DEP) to address the challenging I/O bottleneck issues. DEP is the first paradigm enabling users to identify and handle data-intensive operations separately. It can significantly reduce costly data movement and is better than the existing execution paradigms for data-intensive applications. The initial experimental tests have confirmed its promising potential. Its data-centric architecture could have an impact in future HEC systems, programming models, and algorithms design and development. Yong Chen 0001, Xian-He Sun, William Gropp, Rajeev Thakur |
CLUSTER | 1 |
| 2012 | Dynamic Active Storage for High Performance I/OabstractMany high-end computing applications in critical areas of science and technology are becoming more and more data intensive. These applications transfer large amounts of data from storage nodes to compute nodes for processing, which is costly and bandwidth consuming. The data movement often dominates the applications' run time. Active storage provides a promising solution for these applications by moving appropriate computations from compute nodes to storage nodes. The prior research has achieved considerable progress and developed several active storage models. However, the existing studies have neglected the influence of data dependence on the performance of active storage systems. This study shows that the data dependence has a critical impact on active storage, and the ignorance of dependence can lead to waste of the precious bandwidth. To address this issue in active storage, this paper also presents a new Dynamic Active Storage (DAS) architecture that analyzes the bandwidth requirement of operations, determines the applicability for active storage requests, and optimizes data layout on servers to minimize the bandwidth requirement. Experimental tests have been conducted, and the results have confirmed that the proposed DAS architecture outperforms existing active storage systems. The DAS architecture reduces the data movement caused by data dependence and improves applications' performance over existing schemes. It holds a promise for high performance I/O system in high-end computing. Yong Chen 0001 |
ICPP | 2 |
| 2012 | CHAIO: Enabling HPC Applications on Data-Intensive File SystemsabstractThe computing paradigm of "HPC in the Cloud" has gained a surging interest in recent years, due to its merits of cost-efficiency, flexibility, and scalability. Cloud is designed on top of distributed file systems such as Google file system (GFS). The capability of running HPC applications on top of data-intensive file systems is a critical catalyst in promoting Clouds for HPC. However, the semantic gap between data-intensive file systems and HPC imposes numerous challenges. For example, N-1 (N to 1) is a widely used data access pattern for HPC applications such as check pointing, but cannot perform well on data-intensive file systems. In this study, we propose the CHunk-Aware I/O (CHAIO) strategy to enable efficient N-1 data access on data-intensive distributed file systems. CHAIO reorganizes I/O requests to favor data-intensive file systems and avoid possible access contention. It balances the workload distribution and promotes data locality. We have tested the CHAIO design over the Kosmos file system (KFS). Experimental results show that CHAIO achieves a more than two-fold improvement in I/O bandwidth for both write and read operations. Experiments in large-scale environment confirm the potential of CHAIO for small and irregular requests. The aggregator selection algorithm works well to balance the workload distribution. CHAIO is a critical and necessary step to enable HPC in the Cloud. Hui Jin 0001, Jiayu Ji, Xian-He Sun, Yong Chen 0001, Rajeev Thakur |
ICPP | 4 |
| 2012 | Algorithm-level Feedback-controlled Adaptive data prefetcher: Accelerating data access for high-performance processors
Yong Chen 0001, Huaiyu Zhu 0002, Hui Jin 0001, Xian-He Sun |
Parallel Comput. | 1 |
| 2011 | A Hybrid Shared-Nothing/Shared-Data Storage Architecture for Large Scale DatabasesabstractShared-nothing and shared-disk are two widely-used storage architectures in current parallel database systems, and each of them has its own merits for different query patterns. However, there is no much effort in investigating the integration of these two architectures and exploiting their merits together. In this study, we propose a novel hybrid shared-nothing/shared-data storage scheme for large-scale databases, to leverage the benefits of both shared-nothing and shared-disk architectures. We adopt a shared-nothing architecture as the hardware layer and leverage a parallel file system as the storage layer. The proposed hybrid storage scheme can provide a high degree of parallelism in both I/O and computing, like that in a shared-nothing system. In the meantime, it can achieve convenient and high-speed data sharing across multiple database nodes, like that in a shared-disk system. The hybrid scheme is more appropriate for large-scale and data-intensive applications than each of the two individual types of systems. Huaiming Song, Xian-He Sun, Yong Chen 0001 |
CCGRID | 3 |
| 2011 | PAC-PLRU: A Cache Replacement Policy to Salvage Discarded Predictions from Hardware PrefetchersabstractCache replacement policy plays an important role in guaranteeing the availability of cache blocks, reducing miss rates, and improving applications' overall performance. However, recent research efforts on improving replacement policies require either significant additional hardware or major modifications to the organization of the existing cache. In this study, we propose the PAC-PLRU cache replacement policy. PAC-PLRU not only utilizes but also judiciously salvages the prediction information discarded from a widely-adopted stride prefetcher. The main idea behind PAC-PLRU is utilizing the prediction results generated by the existing stride prefetcher and preventing these predicted cache blocks from being replaced in the near future. Experimental results show that leveraging the PAC-PLRU with a stride prefetcher reduces the average L2 cache miss rate by 91% over a baseline system with only PLRU policy, and by 22% over a system using PLRU with an unconnected stride prefetcher. Most importantly, PAC-PLRU only requires minor modifications to existing cache architecture to get these benefits. The proposed PAC-PLRU policy is promising in fostering the connection between prefetching and replacement policies, and have a lasting impact on improving the overall cache performance. Zhensong Wang, Yong Chen 0001, Huaiyu Zhu 0002, Xian-He Sun |
CCGRID | 3 |
| 2011 | A cost-intelligent application-specific data layout scheme for parallel file systemsabstractI/O data access is a recognized performance bottleneck of high-end computing. Several commercial and research parallel file systems have been developed in recent years to ease the performance bottleneck. These advanced file systems perform well on some applications but may not perform well on others. They have not reached their full potential in mitigating the I/O-wall problem. Data access is application dependent. Based on the application-specific optimization principle, in this study we propose a cost-intelligent data access strategy to improve the performance of parallel file systems. We first present a novel model to estimate data access cost of different data layout policies. Next, we extend the cost model to calculate the overall I/O cost of any given application and choose an appropriate layout policy for the application. A complex application may consist of different data access patterns. Averaging the data access patterns may not be the best solution for those complex applications that do not have a dominant pattern. We then further propose a hybrid data replication strategy for those applications, so that a file can have replications with different layout policies for the best performance. Theoretical analysis and experimental testing have been conducted to verify the newly proposed cost-intelligent layout approach. Analytical and experimental results show that the proposed cost model is effective and the application-specific data layout approach achieved up to 74% performance improvement for data-intensive applications. Huaiming Song, Yanlong Yin, Yong Chen 0001, Xian-He Sun |
HPDC | 3 |
| 2011 | LACIO: A New Collective I/O Strategy for Parallel I/O SystemsabstractParallel applications benefit considerably from the rapid advance of processor architectures and the available massive computational capability, but their performance suffers from large latency of I/O accesses. The poor I/O performance has been attributed as a critical cause of the low sustained performance of parallel systems. Collective I/O is widely considered a critical solution that exploits the correlation among I/O accesses from multiple processes of a parallel application and optimizes the I/O performance. However, the conventional collective I/O strategy makes the optimization decision based on the logical file layout to avoid multiple file system calls and does not take the physical data layout into consideration. On the other hand, the physical data layout in fact decides the actual I/O access locality and concurrency. In this study, we propose a new collective I/O strategy that is aware of the underlying physical data layout. We confirm that the new Layout-Aware Collective I/O (LACIO) improves the performance of current parallel I/O systems effectively with the help of noncontiguous file system calls. It holds promise in improving the I/O performance for parallel systems. Yong Chen 0001, Xian-He Sun, Rajeev Thakur, Philip C. Roth, William Gropp |
IPDPS | 1 |
| 2011 | A Hybrid Shared-Nothing/Shared-Data Storage Scheme for Large-Scale Data ProcessingabstractShared-nothing and shared-disk are the two most common storage architectures of parallel databases in the past two decades. Both two types of systems have their own merits for different applications. However, there are no much efforts in investigating the integration of these two architectures and exploiting their merits together. In this paper, we propose a novel hybrid storage architecture for large-scale data processing, to leverage the benefits of both shared-nothing and shared-disk architectures. In the proposed hybrid system, we adopt a shared-nothing architecture as the hardware layer and leverage a parallel file system as the storage layer to combine the scattered disks on all database nodes. We present an overall design of the new scheme, including data and storage organization, data access modes, and query processing methods. The proposed hybrid scheme can achieve both high I/O performance as a shared-nothing system, and high-speed data sharing across all server nodes as a share-disk system. Preliminary experimental results demonstrate that the hybrid scheme is promising and more appropriate for large-scale and data-intensive applications than each of the two individual types of systems. Huaiming Song, Xian-He Sun, Yong Chen 0001 |
ISPA | 3 |
| 2010 | An Adaptive Data Prefetcher for High-Performance ProcessorsabstractWhile computing speed continues increasing rapidly, data-access technology is lagging behind. Data-access delay, not the processor speed, becomes the leading performance bottleneck of high-end/high-performance computing. Prefetching is an effective solution to masking the gap between computing speed and data-access speed. Existing works of prefetching, however, are very conservative in general, due to the computing power consumption concern of the past. They suffer in effectiveness especially when applications' access pattern changes. In this study, we propose an Algorithm-level Feedback-controlled Adaptive (AFA) data prefetcher to address these issues. The AFA prefetcher is based on the Data-Access History Cache, a hardware structure that is specifically designed for data prefetching. It provides an algorithm-level adaptation and is capable of dynamically adapting to appropriate prefetching algorithms at runtime. We have conducted extensive simulation testing with Simple Scalar simulator to validate the design and to illustrate the performance gain. The simulation results show that AFA prefetcher is effective and achieves considerable IPC (Instructions Per Cycle) improvement in average. Yong Chen 0001, Huaiyu Zhu 0002, Xian-He Sun |
CCGRID | 1 |
| 2010 | REMEM: REmote MEMory as Checkpointing StorageabstractCheck pointing is a widely used mechanism for supporting fault tolerance, but notorious in its high-cost disk access. The idea of memory-based check pointing has been extensively studied in research but made little success in practice due to its complexity and potential reliability concerns. In this study we present the design and implementation of REMEM, a Remote Memory check pointing system to extend the check pointing storage from disk to remote memory. A unique feature of REMEM is that it can be integrated into existing disk-based check pointing systems seamlessly. A user can flexibly switch between REMEM and disk as check pointing storage to balance the efficiency and reliability. The implementation of REMEM on Open MPI is also introduced. The experimental results confirm that REMEM and the proposed adaptive check pointing storage selection are promising in both performance, reliability and scalability. Hui Jin 0001, Xian-He Sun, Yong Chen 0001, Tao Ke |
CloudCom | 3 |
| 2010 | Improving Parallel I/O Performance with Data Layout AwarenessabstractParallel applications can benefit greatly from massive computational capability, but their performance suffers from large latency of I/O accesses. The poor I/O performance has been attributed as a critical cause of the low sustained performance of parallel computing systems. In this study, we propose a data layout-aware optimization strategy to promote a better integration of the parallel I/O middleware and parallel file systems, two major components of the current parallel I/O systems, and to improve the data access performance. We explore the layout-aware optimization in both independent I/O and collective I/O, two primary forms of I/O in parallel applications. We illustrate that the layout-aware I/O optimization could improve the performance of current parallel I/O strategy effectively. The experimental results verify that the proposed strategy could improve parallel I/O performance by nearly 40% on average. The proposed layout-aware parallel I/O has a promising potential in improving the I/O performance of parallel systems. Yong Chen 0001, Xian-He Sun, Rajeev Thakur, Huaiming Song, Hui Jin 0001 |
CLUSTER | 1 |
| 2010 | A layout-aware optimization strategy for collective I/OabstractIn this study, we propose an optimization strategy to promote a better integration of the parallel I/O middleware and parallel file systems. We illustrate that a layout-aware optimization strategy can improve the performance of current collective I/O in parallel I/O system. We present the motivation, prototype design and initial verification of the proposed layout-aware optimization strategy. The analytical and initial experimental testing results demonstrate that the proposed strategy has a potential in improving the parallel I/O system performance. Yong Chen 0001, Huaiming Song, Rajeev Thakur, Xian-He Sun |
HPDC | 1 |
| 2010 | Optimizing HPC Fault-Tolerant Environment: An Analytical ApproachabstractThe increasingly large ensemble size of modern High-Performance Computing (HPC) systems has drastically increased the possibility of failures. Performance under failures and its optimization become timely important issues facing the HPC community. In this study, we propose an analytical model to predict the application performance. The model characterizes the impact of coordinated checkpointing and system failures on application performance, considering all the factors including workload, the number of nodes, failure arrival rate, recovery cost, and checkpointing interval and overhead. Based on the model, we gauge three parameters, the number of compute nodes, checkpointing interval, and the number of spare nodes to conduct a comprehensive study of performance optimization under failures. Performance scalability under failures is also studied to explore the performance improvement space for different parameters. Experimental results from both synthetic and actual system failure logs confirm that the proposed model and optimization methodologies are effective and feasible. Hui Jin 0001, Yong Chen 0001, Huaiyu Zhu 0002, Xian-He Sun |
ICPP | 2 |
| 2010 | Timing local streams: improving timeliness in data prefetchingabstractData prefetching technique is widely used to bridge the growing performance gap between processor and memory. Numerous prefetching techniques have been proposed to exploit data patterns and correlations in the miss address stream. In general, the miss addresses are grouped by some common characteristics, such as program counter or memory region they belong to, into localized streams to improve prefetch accuracy and coverage. However, the existing stream localization technique lacks the timing information of misses. This drawback can lead to a large fraction of untimely prefetches, which in turn limits the effectiveness of prefetching, wastes precious bandwidth and leads to high cache pollution potentially. This paper proposes a novel mechanism named stream timing technique that can largely reduce untimely prefetches and in turn increase the overall performance. Based on the proposed stream timing technique, we extend the conventional stride prefetcher and propose a new stride prefetcher called Time-Aware Stride (TAS) prefetcher. We have carried out extensive simulation experiments to verify the design of the stream timing technique and the TAS prefetcher. The simulation results show that the proposed stream timing technique is promising in reducing untimely prefetches and the IPC improvement of TAS prefetcher outperforms the existing stride prefetcher by 11%. Huaiyu Zhu 0002, Yong Chen 0001, Xian-He Sun |
ICS | 2 |
| 2010 | Reevaluating Amdahl's law in the multicore era
Xian-He Sun, Yong Chen 0001 |
J. Parallel Distributed Comput. | 2 |
| 2009 | V-MCS: A configuration system for virtual machinesabstractVitual Machine (VM) technology encapsulates shared computing resources into secure, stable, isolated and customizable private computing environments. While service-oriented computing becomes more and more a norm of computing, VM becomes a must-have common structure. However, creating and customizing a VM system on different hardware/software environments to meet versatile demands is a state-of-the-art task, especially for casual users working in new computing environments. In addition, VM configuration without system support is tedious, time consuming, and error prone. In this study, we propose a Virtual Machine Configuration System (V-MCS) for tackling this issue. V-MCS takes a systematic approach to enhance the flexibility and usability of VM. It provides an easy-to-use web interface to users to create their preferred configurations, and to convert the configurations into PAN documents for human-computer interaction and XML documents for machine automation. The underlying definition component parses the configurations and the spawn component generates customized VMs on the fly. V-MCS maintains and deploys these two-level documents when users login in the future. With the help of V-MCS, users can generate their customized VMs easily and swiftly. V-MCS has been implemented and tested. Experimental results match the design goal well. Xian-He Sun, Hongbo Zou, Yong Chen 0001, Prerak Shukla |
CLUSTER | 4 |
| 2009 | Performance under Failure of Multi-tier Web ServicesabstractPerformance issues of multi-tier Web services have been studied extensively in recent years. Performance modeling and prediction under failure of multi-tier architectures, however, is not well addressed yet. We propose a novel model named Performance under failure of multi-tier architecture, or PerFAMA in short, to address this issue. We first show that the multi-tier architecture with failure considerations is a product-form network, and then analyze and model the failure impact. By applying the PerFAMA model, we are able to predict the end-to-end response time of multi-tier Web services under failures. We have simulated two representative Web services architectures and various failure scenarios to verify the proposed PerFAMA model. The experimental results show that the proposed model works well and the prediction accuracy is up to 98%. Yong Chen 0001, Xian-He Sun, Hui Jin 0001 |
ICPADS | 2 |
| 2009 | Core-aware memory access scheduling schemesabstractMulti-core processors have changed the conventional hardware structure and require a rethinking of system scheduling and resource management to utilize them efficiently. However, current multi-core systems are still using conventional single-core memory scheduling. In this study, we investigate and evaluate traditional memory access scheduling techniques, and propose a core-aware memory scheduling for multi-core environments. Since memory requests from the same source exhibit better locality, it is reasonable to schedule the requests by taking the source of the requests into consideration. Motivated from this principle of locality, we propose two core-aware policies based on traditional bank-first and row-first schemes. Simulation results show that the core-aware policies can effectively improve the performance. Compared with the bank-first and row-first policies, the proposed core-aware policies reduce the execution time of certain NAS Parallel Benchmarks by up to 20% in running the benchmarks separately, and by 11% in running them concurrently. Zhibin Fang, Xian-He Sun, Yong Chen 0001, Surendra Byna |
IPDPS | 3 |
| 2009 | Taxonomy of Data Prefetching for Multicore Processors
Surendra Byna, Yong Chen 0001, Xian-He Sun |
J. Comput. Sci. Technol. | 2 |
| 2008 | 2008 International Conference on Parallel Processing September 8-12, 2008 Portland, Oregon Exploring Parallel I/O Concurrency with Speculative PrefetchingabstractParallel applications can benefit greatly from massive computational capability, but their performance usually suffers due to large latency in I/O accesses. Conventional I/O prefetching techniques are conservative and are limited by low accuracy and coverage. As the processor performance has been increasing rapidly and the computing power is virtually free, we introduce a novel speculative approach for comprehensive and aggressive parallel I/O prefetching in this study. We present the design of our approach as well as challenges, solutions, and our prototype implementation. The experiments have shown promising results in reducing I/O access latency. Yong Chen 0001, Surendra Byna, Xian-He Sun, Rajeev Thakur, William Gropp |
ICPP | 1 |
| 2008 | Parallel I/O prefetching using MPI file caching and I/O signaturesabstractParallel I/O prefetching is considered to be effective in improving I/O performance. However, the effectiveness depends on determining patterns among future I/O accesses swiftly and fetching data in time, which is difficult to achieve in general. In this study, we propose an I/O signature-based prefetching strategy. The idea is to use a predetermined I/O signature of an application to guide prefetching. To put this idea to work, we first derived a classification of patterns and introduced a simple and effective signature notation to represent patterns. We then developed a toolkit to trace and generate I/O signatures automatically. Finally, we designed and implemented a thread-based client-side collective prefetching cache layer for MPI-IO library to support prefetching. A prefetching thread reads I/O signatures of an application and adjusts them by observing I/O accesses at runtime. Experimental results show that the proposed prefetching method improves I/O performance significantly for applications with complex patterns. Surendra Byna, Yong Chen 0001, Xian-He Sun, Rajeev Thakur, William Gropp |
SC | 2 |
| 2008 | Hiding I/O latency with pre-execution prefetching for parallel applicationsabstractParallel applications are usually able to achieve high computational performance but suffer from large latency in I/O accesses. I/O prefetching is an effective solution for masking the latency. Most of existing I/O prefetching techniques, however, are conservative and their effectiveness is limited by low accuracy and coverage. As the processor-I/O performance gap has been increasing rapidly, data-access delay has become a dominant performance bottleneck. We argue that it is time to revisit the ldquoI/O wallrdquo problem and trade the excessive computing power with data-access speed. We propose a novel pre-execution approach for masking I/O latency. We describe the pre-execution I/O prefetching framework, the pre-execution thread construction methodology, the underlying library support, and the prototype implementation in the ROMIO MPI-IO implementation in MPICH2. Preliminary experiments show that the pre-execution approach is promising in reducing I/O access latency and has real potential. Yong Chen 0001, Surendra Byna, Xian-He Sun, Rajeev Thakur, William Gropp |
SC | 1 |
| 2008 | Algorithm-system scalability of heterogeneous computing
Yong Chen 0001, Xian-He Sun, Ming Wu 0006 |
J. Parallel Distributed Comput. | 1 |
| 2007 | Improving Data Access Performance with Server Push ArchitectureabstractData prefetching, where data is fetched before CPU demands for it, has been considered as an effective solution to mask data access latency. However, the current client-initiated prefetching strategies do not work well for applications with complex, non-contiguous data access patterns. While technology advances continue to enlarge the gap between computing and data access performance, trading computing power for data access delay has become a natural choice. We propose a server-based data-push approach. In this server-push architecture, a dedicated server named data push server (DPS) initiates and proactively pushes data closer to the client in time. We present the DPS architecture and study the issues such as what data to fetch, when to fetch, how to push, and data access modeling. Xian-He Sun, Surendra Byna, Yong Chen 0001 |
IPDPS | 3 |
| 2007 | Data access history cache and associated data prefetching mechanismsabstractData prefetching is an effective way to bridge the increasing performance gap between processor and memory. As computing power is increasing much faster than memory performance, we suggest that it is time to have a dedicated cache to store data access histories and to serve prefetching to mask data access latency effectively. We thus propose a new cache structure, named Data Access History Cache (DAHC), and study its associated prefetching mechanisms. The DAHC behaves as a cache for recent reference information instead of as a traditional cache for instructions or data. Theoretically, it is capable of supporting many well known history-based prefetching algorithms, especially adaptive and aggressive approaches. We have carried out simulation experiments to validate DAHC design and DAHC-based data prefetching methodologies and to demonstrate performance gains. The DAHC provides a practical approach to reaping data prefetching benefits and its associated prefetching mechanisms are proven more effective than traditional approaches. Yong Chen 0001, Surendra Byna, Xian-He Sun |
SC | 1 |
| 2007 | Server-Based Data Push Architecture for Multi-Processor Environments
Xian-He Sun, Surendra Byna, Yong Chen 0001 |
J. Comput. Sci. Technol. | 3 |
| 2006 | QoS Oriented Resource Reservation in Shared EnvironmentsabstractResource sharing across different computers and organizations makes it possible to support diverse, dynamic changing resource requirements of distributed applications. Reservation mechanisms have been used to reserve resources for external applications through service level agreements between local resource organizations and external applications. However, the effects of resource reservation on local applications, and therefore the trustfulness of the successful fulfillment of the service agreement, have been ignored. In this paper, we investigate the effect of resource reservation on external applications as well as local jobs, and design efficient task scheduling algorithms considering the tolerance of local jobs to resource reservation. Extensive simulations and implementation experiments have been carried out to confirm our analysis results. Experimental results show that the relative slowdown metric and the failureminimization scheduling algorithms proposed in this study are practically effective and have a real potential. Ming Wu 0006, Xian-He Sun, Yong Chen 0001 |
CCGRID | 3 |
| 2006 | STAS: A Scalability Testing and Analysis SystemabstractScalability is a crucial factor in performance evaluation and analysis of parallel and distributed systems. Much effort has been devoted to scalability research and several metrics are proposed. However, the lacking of an effective scalability analysis toolkit is still a major barrier for researchers to measure and analyze scalabilities. Isospeed scalability is a known metric and has been extended for general computing systems recently. This paper proposes an effective scalability testing and analysis system, called STAS, and presents its implementation with isospeed-e scalability metric. STAS provides the facility to conduct automated isospeed-e scalability measure and analysis. It reduces the burden for users to evaluate the performance of algorithms and systems. Experiments have been conducted to verify the design and implementation Yong Chen 0001, Xian-He Sun |
CLUSTER | 1 |
| 2005 | Scalability of Heterogeneous ComputingabstractScalability is a key factor of the design of distributed systems and parallel algorithms and machines. However, conventional scalabilities are designed for homogeneous parallel processing. There is no suitable and commonly accepted definition of scalability metric for heterogeneous systems. Isospeed scalability is a well-defined metric for homogeneous computing. This study extends the isospeed scalability metric to general heterogeneous computing systems. The proposed isospeed-efficiency metric is suitable for both homogeneous and heterogeneous computing. Through theoretical analysis, we derive methodologies of scalability measurement and prediction for heterogeneous systems. Experimental results verify the analytical results and confirm that the proposed isospeed-efficiency scalability works well in both homogeneous and heterogeneous environments. Xian-He Sun, Yong Chen 0001, Ming Wu 0006 |
ICPP | 2 |