EDBT 2026 Demo / reviewers in the wild / expert
Qing Liu 0002
dblp:53/4481-2
· DBLP profile ↗
70ranked-venue papers
9as first author
27since 2021 · last 2026
0000-0002-7600-7976ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 49 · 2 first-author · 18 since 2021Computer networks · 11 · 7 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QProR: An Efficient Framework for Quantity-of-Interest Based Progressive Retrieval with Guaranteed Error ControlabstractScientific applications generate an unprecedented volume of data, overwhelming the network and file systems’ bandwidth and posing challenges for efficient and scalable data retrieval and analysis. Progressive data compression offers a promising solution by enabling on-demand retrieval at reduced size. However, existing progressive methods either fail to bound the errors in essential quantities of interest (QoIs) derived from raw data or suffer from suboptimal retrieval efficiency. In this work, we propose QProR, an efficient QoI-based progressive framework that optimizes progressive retrieval for target QoIs. Our key contributions include: (1) a systematic framework that integrates error-controlled lossy compressors with bitplane encoding while decoupling the two processes for high flexibility and adaptability; (2) a novel weighted bitplane encoding method which incorperates QoI knowledge into data refactoring to enhance retrieval efficiency; (3) an optimized retrieval strategy that accounts for the varying impacts of different variables on multivariate QoIs; (4) comprehensive evaluations using six real-world datasets from multiple scientific applications and thorough comparisons against state of the arts. Experimental results demonstrate that QProR achieves up to \(80.38\%\) reduction in the retrieval size under the same requested QoI error tolerance, when compared with the best-performing existing methods. When transferring 384 GB of scientific data to remote sites, QProR delivers up to 1.68 × speedup in the end-to-end data transfer performance. Qian Gong, Jieyang Chen, Qing Liu 0002, Xubin He, Norbert Podhorszki, Scott Klasky, Xin Liang 0001 |
HPDC | 5 |
| 2026 | Anchored Maximum Communities over Large Directed Graphs
Xu Zhou 0001, Yan Ding 0004, Qing Liu 0002, Haoxian Xu, Kenli Li 0001 |
Proc. VLDB Endow. | 4 |
| 2026 | fPIM: A Holistic Design to Optimize PIM Data Flow for High Execution EfficiencyabstractAs applications demand more bandwidth, the the “memory wall” problem becomes increasingly severe. Therefore, the processing-in-memory (PIM) architecture has attracted significant research interest due to its ability to execute instructions offloaded by the processor. Existing works on PIM architectures are classified into two categories: regional offloading, where all instructions within a programmer-specified code region are offloaded, and selective offloading, where only instructions of interest are offloaded via hardware support. However, PIM architectures pose the amplified in-PIM traffic overhead challenge that endangers the performance of PIM and degrades the performance of the entire system. To address the challenge, we propose a PIM architecture, called fast PIM (fPIM), which integrates the PIM cache within each Channel Controller to optimize the data flow within the PIM. This design cooperates with theProcessing Unit Load-balancerandBehavior-based Offloaderto achieve high execution efficiency. To evaluatefPIM, we perform extensive experiments, and the results show thatfPIMreduces the workload finish time by up to$88.6\%$,$87.5\%$, and$79.6\%$(with an average of$68.7\%$,$66.2\%$, and$59.8\%$), compared to three state-of-the-art PIM designs, PEI, Fafnir, and SpaceA, respectively. Wenjie Liu 0002, Qing Liu 0002, Xubin He |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | Understanding and Estimating Error Propagation in Neural Networks for Scientific Data AnalysisabstractNeural networks are increasingly integrated into scientific discovery, where input data reduction and model quantization play a key role in accelerating inference. However, understanding and mitigating the impact of these techniques on output error is critical for ensuring reliable results, particularly in tasks demanding high numerical precision. This paper introduces a comprehensive framework for optimizing neural network inference in scientific computing by combining data reduction and model quantization while maintaining error-controlled outcomes. We develop theoretical analyses to bound error propagation under these techniques and propose a framework that balances computational performance with error constraints. Evaluation on real-world learning-based combustion simulations and satellite image classification shows that our derived error bounds accurately predict observed errors while enabling significant computational speedup under our framework. This work highlights the potential for further leveraging advancements in modern lossy compression algorithms and hardware accelerators that support lower-precision formats. Weiming He, Qian Gong, Jing Li 0025, Qing Liu 0002, Norbert Podhorszki, Scott Klasky, Ki Sung Jung, Cristian Lacey, Jackie Chen, Hongjian Zhu |
ICDE | 5 |
| 2025 | HPDR: High-Performance Portable Scientific Data Reduction FrameworkabstractThe rapid growth in scientific data generation is outpacing advancements in computing systems necessary for efficient storage, transfer, and analysis, particularly in the context of exascale computing. With the deployment of first-generation exascale computing systems and next-generation experimental facilities, this gap is widening and necessitates effective data reduction techniques to manage enormous data volumes. Over the past decade, various data reduction methods, including lossless compression, error-controlled lossy compression, and data refactoring, have been developed to accelerate I/O in scientific workflows. Despite significant reductions in data volume, these methods introduce considerable computational overhead, which can become the new bottleneck in data processing. To mitigate this, GPU-accelerated data reduction algorithms have been introduced. However, challenges remain in their integration into exascale workflows, including limited portability across different GPU architectures, substantial memory transfer overhead, and reduced scalability on dense multi-GPU systems. To address these challenges, we propose HPDR, a high-performance and portable data reduction framework. HPDR is designed to enable the execution of state-of-the-art reduction algorithms across diverse processor architectures while reducing memory transfer overhead to 2.3 % of the original, resulting in up to$3.5 \times$faster throughput compared to existing solutions. It also achieves up to 96% of the theoretical speedup in multi-GPU settings. In addition, evaluations on accelerating I/O operations at scale up to 1,024 nodes of the Frontier supercomputer demonstrate that HPDR can achieve up to$103\ \text{TB} / \mathrm{s}$reduction throughput, providing up to$4 \times$acceleration in parallel I/O performance compared to existing data reduction routines. This work highlights the potential of HPDR to significantly enhance data reduction efficiency in exascale computing environments. Jieyang Chen, Qian Gong, Yanliang Li, Xin Liang 0001, Lipeng Wan 0001, Qing Liu 0002, Norbert Podhorszki, Scott Klasky |
IPDPS | 6 |
| 2025 | Stability-preserving Lossy Compression for Large-scale Partial Differential EquationsabstractCheckpoint/Restart (C/R) strategies are vital for fault tolerance in PDE-based scientific simulations, yet traditional checkpointing incurs significant I/O overhead. Lossy compression offers a scalable solution by reducing checkpoint data size, but conventional methods often lack control over physical invariants (e.g., energy), leading to instability such as oscillations or divergence in Partial Differential Equations (PDE) systems. This paper introduces a stability-preserving compression approach tailored for PDE simulations by explicitly controlling kinetic and potential energy perturbations to ensure stable restarts. Extensive experiments conducted across diverse PDE configurations demonstrate that our method maintains numerical stability with minimal error magnification—even across multiple checkpoint-restart cycles—outperforming state-of-the-art lossy compressors. Parallel evaluations on the Frontier supercomputer show up to 8.4× improvement in checkpoint write performance and 6.3× in read performance, while maintaining relative L2 errors ∼ 2e-6 throughout continued simulation. These results provide practical guidance for balancing compression accuracy, stability, and computational efficiency in large-scale PDE applications. Qian Gong, Mark Ainsworth, Jieyang Chen, Xin Liang 0001, Liangji Zhu, Ethan Klasky, Tushar M. Athawale, Qing Liu 0002, Anand Rangarajan 0001, Sanjay Ranka, Scott Klasky |
SC | 8 |
| 2025 | HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUsabstractScientific applications produce vast amounts of data, posing grand challenges in the underlying data management and analytic tasks. Progressive compression is a promising way to address this problem, as it allows for on-demand data retrieval with significantly reduced data movement cost. However, most existing progressive methods are designed for CPUs, leaving a gap for them to unleash the power of today’s heterogeneous computing systems with GPUs.In this work, we propose HP-MDR, a high-performance and portable data refactoring and progressive retrieval framework for GPUs. Our contributions are four-fold: (1) We carefully optimize the bitplane encoding and lossless encoding, two key stages in progressive methods, to achieve high performance on GPUs; (2) We propose pipeline optimization and incorporate it with data refactoring and progressive retrieval workflows to further enhance the performance for large data process; (3) We leverage our framework to enable high-performance data retrieval with guaranteed error control for common Quantities of Interest; (4) We evaluate HP-MDR and compare it with state of the arts using five real-world datasets. Experimental results demonstrate that HP-MDR delivers an average 13.68 × and 6.31 × throughput in data refactoring and progressive retrieval tasks, respectively. It also leads to 11.22 × throughput for recomposing required data representations under Quantity-of-Interest error control and 6.04 × performance for the corresponding end-to-end data retrieval, when compared with state-of-the-art solutions. Yanliang Li, Qian Gong, Qing Liu 0002, Norbert Podhorszki, Scott Klasky, Xin Liang 0001, Jieyang Chen |
SC | 4 |
| 2024 | A General Framework for Error-controlled Unstructured Scientific Data CompressionabstractData compression plays a key role in reducing storage and I/O costs. Traditional lossy methods primarily target data on rectilinear grids and cannot leverage the spatial coherence in unstructured mesh data, leading to suboptimal compression ratios. We present a multi-component, error-bounded compression framework designed to enhance the compression of floating-point unstructured mesh data, which is common in scientific applications. Our approach involves interpolating mesh data onto a rectilinear grid and then separately compressing the grid interpolation and the interpolation residuals. This method is general, independent of mesh types and typologies, and can be seamlessly integrated with existing lossy compressors for improved performance. We evaluated our framework across twelve variables from two synthetic datasets and two real-world simulation datasets. The results indicate that the multi-component framework consistently outperforms state-of-the-art lossy compressors on unstructured data, achieving, on average, a 2.3 − 3.5× improvement in compression ratios, with error bounds ranging from 1 × 10 the−6to 1×10−2. We further investigate impact of hyperparameters, such as grid spacing and error allocation, to deliver optimal compression ratios in diverse datasets. Qian Gong, Zhe Wang 0059, Viktor Reshniak, Xin Liang 0001, Jieyang Chen, Qing Liu 0002, Tushar M. Athawale, Yi Ju, Anand Rangarajan 0001, Sanjay Ranka, Norbert Podhorszki, Rick Archibald, Scott Klasky |
e-Science | 6 |
| 2024 | Tango: A Cross-layer Approach to Managing I/O Interference over Local Ephemeral StorageabstractAs simulation-based scientific discovery advances to exascale, a major question that the community is striving to answer is how to co-design data storage and complex physicsrich analytics in a way that the time to knowledge can be minimized for post-processing. A particular challenge is how to accommodate a broad spectrum of data analytics needsparticularly those that become clear only until very late during the post-processing, a scenario where existing methods, such as in situ processing, are unable or less effective in supporting data analytics. As HPC storage systems have become deeper and more complex with the recent addition of NVMe, die-stacked memory, and burst buffer, it requires fundamentally rethinking new paradigms and methods for data storage and analysis. This paper aims to address the issue of I/O interference for data analytics over local ephemeral storage, which is shared by multiple applications in a non-exclusive node usage scenario-often configured for small- to medium-sized clusters. At the core of this work is a coordinated cross-layer approach that reacts to storage interference from both storage and application layers. By decomposing and distributing analysis data across the storage hierarchy, data analytics can adapt to the interference by reducing or completely avoiding access to lower tiers whenever there is a high interference, while maintaining a prescribed error bound to limit the information loss. Meanwhile, proper actions are also taken at the storage layer to ensure sufficient bandwidth is allocated for retrieving an augmentation, which is based upon the cardinality and accuracy of the augmentation as well as the nature of an application. We evaluate three realworld data analytics, XGC, GenASiS, and CFD, on Chameleon, and quantitatively demonstrate that the I/O performance can be vastly improved, e.g., by 52% versus no adaptivity and 36% versus single-layer adaptivity, while maintaining acceptable outcomes of data analysis. Zhenbo Qiao, Qirui Tian, Zhenlu Qin, Jinzhen Wang, Qing Liu 0002, Norbert Podhorszki, Scott Klasky, Hongjian Zhu |
SC | 5 |
| 2024 | Error-controlled Progressive Retrieval of Scientific Data under Derivable Quantities of InterestabstractThe unprecedented amount of scientific data has introduced heavy pressure on the current data storage and transmission systems. Progressive compression has been proposed to mitigate this problem, which offers data access with on-demand precision. However, existing approaches only consider precision control on primary data, leaving uncertainties on the quantities of interest (QoIs) derived from it. In this work, we present a progressive data retrieval framework with guaranteed error control on derivable QoIs. Our contributions are three-fold. (1) We carefully derive the theories to strictly control QoI errors during progressive retrieval. Our theory is generic and can be applied to any QoIs that can be composited by the basis of derivable QoIs proved in the paper. (2) We design and develop a generic progressive retrieval framework based on the proposed theories, and optimize it by exploring feasible progressive representations. (3) We evaluate our framework using five real-world datasets with a diverse set of QoIs. Experiments demonstrate that our framework can faithfully respect any user-specified QoI error bounds in the evaluated applications. This leads to over $2.02 \times$ performance gain in data transfer tasks compared to transferring the primary data while guaranteeing a QoI error that is less than 1E-5. Qian Gong, Jieyang Chen, Qing Liu 0002, Norbert Podhorszki, Xin Liang 0001, Scott Klasky |
SC | 4 |
| 2023 | RAPIDS: Reconciling Availability, Accuracy, and Performance in Managing Geo-Distributed Scientific DataabstractIn modern science, big data plays an increasingly important role. Many scientific applications, such as running simulations on supercomputers or conducting experiments on advanced instruments, produce huge amount of data at unprecedented speed. Analyzing and understanding such big data is the key for scientists to make scientific breakthroughs. However, data might become unavailable for scientists to access when outages or maintenance of the storage system occur, which severely hinders scientific discovery. To improve the data availability, data duplication and erasure coding (EC) are often used. But as the scientific data gets larger, using these two methods can cause considerable storage and network overhead. Lipeng Wan 0001, Jieyang Chen, Xin Liang 0001, Ana Gainaru, Qian Gong, Qing Liu 0002, Ben Whitney, Joy Arulraj, Zhengchun Liu, Ian T. Foster, Scott Klasky |
HPDC | 6 |
| 2023 | Improving Progressive Retrieval for HPC Scientific Data using Deep Neural NetworkabstractAs the disparity between compute and I/O on high-performance computing systems has continued to widen, it has become increasingly difficult to perform post-hoc data analytics on full-resolution scientific simulation data due to the high I/O cost. Error-bounded data decomposition and progressive data retrieval framework has recently been developed to address such a challenge by performing data decomposition before storage and reading only part of the decomposed data when necessary. However, the performance of the progressive retrieval framework has been suffering from the over-pessimistic error control theory, such that the achieved maximum error of recomposed data is significantly lower than the required error. Therefore, more data than required is fetched for recomposition, incurring additional I/O overhead. In order to tackle this issue, we propose a DNN-based progressive retrieval framework that can better identify the minimum amount of data to be retrieved. Our contributions are as follows: 1) We provide an in-depth investigation of the recently developed progressive retrieval framework; 2) We propose two designs of prediction models (named D-MGARD and E-MGARD) to estimate the amount of retrieved data size based on error bounds. 3) We evaluate our proposed solutions using scientific datasets generated by real-world simulations from two domains. Evaluation results demonstrate the effectiveness of our solution in accurately predicting the amount of retrieval data size, as well as the advantages of our solution over the traditional approach to reducing the I/O overhead. Based on our evaluation, our solution is shown to read significantly less data (5% - 40% with D-MGARD, 20% - 80% with E-MGARD). Jinzhen Wang, Xin Liang 0001, Ben Whitney, Jieyang Chen, Qian Gong, Xubin He, Lipeng Wan 0001, Scott Klasky, Norbert Podhorszki, Qing Liu 0002 |
ICDE | 10 |
| 2023 | High-Ratio Lossy Compression: Exploring the Autoencoder to Compress Scientific DataabstractScientific simulations on high-performance computing (HPC) systems can generate large amounts of floating-point data per run. To mitigate the data storage bottleneck and lower the data volume, it is common for floating-point compressors to be employed. As compared to lossless compressors, lossy compressors, such as SZ and ZFP, can reduce data volume more aggressively while maintaining the usefulness of the data. However, a reduction ratio of more than two orders of magnitude is almost impossible without seriously distorting the data. In deep learning, the autoencoder technique has shown great potential for data compression, in particular with images. Whether the autoencoder can deliver similar performance on scientific data, however, is unknown. In this article, we for the first time conduct a comprehensive study on the use of autoencoders to compress real-world scientific data and illustrate several key findings on using autoencoders for scientific data reduction. We implement an autoencoder-based compression prototype to reduce floating-point data. Our study shows that the out-of-the-box implementation needs to be further tuned in order to achieve high compression ratios and satisfactory error bounds. Our evaluation results show that, for most of the test datasets, the tuned autoencoder outperforms SZ by up to 4X, and ZFP by up to 50X in compression ratios, respectively. Our practices and lessons learned in this work can direct future optimizations for using autoencoders to compress scientific data. Tong Liu 0030, Jinzhen Wang, Qing Liu 0002, Shakeel Alibhai, Tao Lu 0014, Xubin He |
IEEE Trans. Big Data | 3 |
| 2023 | A Data-driven Approach to Harvesting Latent Reduced Models to Precondition Lossy Compression for Scientific DataabstractIn this paper, we propose and evaluate the idea that data need to be preconditioned prior to compression, such that they can better match the design philosophies of lossy compressors for HPC scientific data. In particular, we aim to identify a reduced model that can be utilized to transform the original data into a more compressible form. We begin with two PDE applications as a proof of concept, in which we demonstrate that a reduced model can indeed reside in the full model output, and can be utilized to improve compression ratios. A mathematical proof is also presented to show how the compression ratio is improved by the reduced model. We further explore more general dimension reduction techniques to extract the reduced model, including principal component analysis, singular value decomposition, and discrete wavelet transform. After preconditioning, the reduced model in conjunction with difference between the reduced model and full model is stored, which results in higher compression ratios. We evaluate the reduced models on ten scientific datasets, and the results show the effectiveness of our approaches. Given that there is no single method that consistently achieves the best performance, we further propose a selection strategy that guides users to select the best reduced model prior to data reduction. Huizhang Luo, Junqi Wang 0002, Zhenlu Qin, Dan Huang 0001, Qing Liu 0002, MengChu Zhou, Hong Jiang 0001 |
IEEE Trans. Big Data | 5 |
| 2023 | zPerf: A Statistical Gray-Box Approach to Performance Modeling and Extrapolation for Scientific Lossy CompressionabstractWith the scaling up of simulation-based scientific discovery on high-performance computing systems, the disparity between compute and I/O has increased, forcing domain scientists to save only a small amount of simulation data to persistent storage. This can result in the loss of essential physics fields that are needed for data analysis. While error-bounded lossy compression has made tremendous progress in bridging the gap between compute and I/O, the lack of understanding of compression performance remains a key hurdle to its wide adoption. In this work, we present zPerf, a statistical gray-box performance modeling approach for scientific lossy compression. Our contributions are threefold: 1) We develop zPerf to estimate the performance of lossy compression techniques, based on in-depth understanding and statistical modeling for data features and core compression metrics; 2) We demonstrate the in-detailed implementation of zPerf using two case studies, where we derive the performance modeling for SZ and ZFP, two leading lossy compressors; 3) We evaluate the effectiveness of zPerf on real-world datasets across various domains. Based on the evaluation, we demonstrate the efficacy of the zPerf performance model; 4) We further discuss three case studies where zPerf is applied to extrapolate the compression ratio of SZ and ZFP with alternative encoding schemes as well as ZFP with an alternative transform scheme. Through the case studies, we demonstrate the potential of zPerf for exploring the design space of lossy compression, which has hardly been studied in the literature. Jinzhen Wang, Tong Liu 0030, Qing Liu 0002, Xubin He |
IEEE Trans. Computers | 4 |
| 2023 | Exploring Memory Access Similarity to Improve Irregular Application Performance for Distributed Hybrid Memory SystemsabstractWith the increasing problem complexity, more irregular applications are deployed on high-performance clusters due to the parallel working paradigm, and yield irregular memory access behaviors across nodes. However, the irregularity of memory access behaviors is not comprehensively studied, which results in low utilization of the integrated hybrid memory system compositing of stacked DRAM and off-chip DRAM. To address this problem, we devise a novel method calledSimilarity-Managed Hybrid Memory System(SM-HMS) to improve the hybrid memory system performance by leveraging the memory access similarity among nodes in a cluster. WithinSM-HMS, two techniques are proposed,Memory Access Similarity MeasuringandSimilarity-based Memory Access Behavior Sharing. To quantify the memory access similarity, memory access behaviors of each node are vectorized, and the distance between two vectors is used as the memory access similarity. The calculated memory access similarity is used to share memory access behaviors precisely across nodes. With the shared memory access behaviors,SM-HMSdivides the stacked DRAM into two sections, thesliding window sectionand theoutlier section. The shared memory access behaviors guide the replacement of thesliding window sectionwhile theoutlier sectionis managed in the LRU manner. Our evaluation results with a set of irregular applications on various clusters consisting of up to 256 nodes have shown thatSM-HMSoutperforms the state-of-the-art approaches,Cameo,Chameleon, andHyrbid2, on job finish time reduction by up to$58.6\%$,$56.7\%$, and$31.3\%$, with$46.1\%$,$41.6\%$, and$19.3\%$on average, respectively.SM-HMScan also achieve up to$98.6\%$($91.9\%$on average) of the ideal hybrid memory system performance. Wenjie Liu 0002, Xubin He, Qing Liu 0002 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | Region-adaptive, Error-controlled Scientific Data Compression using Multilevel DecompositionabstractThe increase of computer processing speed is significantly outpacing improvements in network and storage bandwidth, leading to the big data challenge in modern science, where scientific applications can quickly generate much more data than that can be transferred and stored. As a result, big scientific data must be reduced by a few orders of magnitude while the accuracy of the reduced data needs to be guaranteed for further scientific explorations. Moreover, scientists are often interested in some specific spatial/temporal regions in their data, where higher accuracy is required. The locations of the regions requiring high accuracy can sometimes be prescribed based on application knowledge, while other times they must be estimated based on general spatial/temporal variation. In this paper, we develop a novel multilevel approach which allows users to impose region-wise compression error bounds. Our method utilizes the byproduct of a multilevel compressor to detect regions where details are rich and we provide the theoretical underpinning for region-wise error control. With spatially varying precision preservation, our approach can achieve significantly higher compression ratios than single-error bounded compression approaches and control errors in the regions of interest. Qian Gong, Ben Whitney, Chengzhu Zhang, Xin Liang 0001, Anand Rangarajan 0001, Jieyang Chen, Lipeng Wan 0001, Paul Ullrich, Qing Liu 0002, Robert Jacob, Sanjay Ranka, Scott Klasky |
SSDBM | 9 |
| 2022 | Locality-based transfer learning on compression autoencoder for efficient scientific data lossy compression
Tong Liu 0030, Jinzhen Wang, Qing Liu 0002, Shakeel Alibhai, Xubin He |
J. Netw. Comput. Appl. | 4 |
| 2022 | Identifying challenges and opportunities of in-memory computing on large HPC systems
Dan Huang 0001, Zhenlu Qin, Qing Liu 0002, Norbert Podhorszki, Scott Klasky |
J. Parallel Distributed Comput. | 3 |
| 2022 | MGARD+: Optimizing Multilevel Methods for Error-Bounded Scientific Data ReductionabstractNowadays, data reduction is becoming increasingly important in dealing with the large amounts of scientific data. Existing multilevel compression algorithms offer a promising way to manage scientific data at scale, but may suffer from relatively low performance and reduction quality. In this paper, we propose MGARD+, a multilevel data reduction and refactoring framework drawing on previous multilevel methods, to achieve high-performance data decomposition and high-quality error-bounded lossy compression. Our contributions are four-fold: 1) We propose to leverage a level-wise coefficient quantization method, which uses different error tolerances to quantize the multilevel coefficients. 2) We propose an adaptive decomposition method which treats the multilevel decomposition as a preconditioner and terminates the decomposition process at an appropriate level. 3) We leverage a set of algorithmic optimization strategies to significantly improve the performance of multilevel decomposition/recomposition. 4) We evaluate our proposed method using four real-world scientific datasets and compare with several state-of-the-art lossy compressors. Experiments demonstrate that our optimizations improve the decomposition/recomposition performance of the existing multilevel method by up to$70 \times$, and the proposed compression method can improve compression ratio by up to$2 \times$compared with other state-of-the-art error-bounded lossy compressors under the same level of data distortion. Xin Liang 0001, Ben Whitney, Jieyang Chen, Lipeng Wan 0001, Qing Liu 0002, Dingwen Tao, James Kress, David Pugmire, Matthew Wolf, Norbert Podhorszki, Scott Klasky |
IEEE Trans. Computers | 5 |
| 2022 | zMesh: Theories and Methods to Exploring Application Characteristics to Improve Lossy Compression Ratio for Adaptive Mesh RefinementabstractScientific simulations on high-performance computing systems produce vast amounts of data that need to be stored and analyzed efficiently. Lossy compression significantly reduces the data volume by trading accuracy for performance. Despite the recent success of lossy compressions, such as ZFP and SZ, the compression performance is still far from being able to keep up with the exponential growth of data. This article aims to further take advantage of application characteristics, an area that is often under-explored, to improve the compression ratios of adaptive mesh refinement (AMR) - a widely used numerical solver that allows for an improved resolution in limited regions. We propose a level reordering techniquezMeshto reduce the storage footprint of AMR applications. In particular, we group the data points that are mapped to the same or adjacent geometric coordinates such that the dataset is smoother and more compressible. Unlike the prior work where the compression performance is affected by the overhead of metadata, this work re-generates the restore recipe using a chained tree structure, thus involving no extra storage overhead for compressed data, which substantially improves the compression ratios. We further derive a mathematical proof that lays the foundation for our method. The results demonstrate that zMesh can improve the smoothness of data by 67.9% and 71.3% for Z-ordering and Hilbert, respectively. Overall, zMesh improves the compression ratios by up to 16.5% and 133.7% for ZFP and SZ, respectively. Despite that zMesh involves additional compute overhead for tree and restore recipe construction, we show that the cost can be amortized as the number of quantities to be compressed increases. Huizhang Luo, Junqi Wang 0002, Qing Liu 0002, Jieyang Chen, Scott Klasky, Norbert Podhorszki |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Accelerating Multigrid-based Hierarchical Scientific Data Refactoring on GPUsabstractRapid growth in scientific data and a widening gap between computational speed and I/O bandwidth make it increasingly infeasible to store and share all data produced by scientific simulations. Instead, we need methods for reducing data volumes: ideally, methods that can scale data volumes adaptively so as to enable negotiation of performance and fidelity tradeoffs in different situations. Multigrid-based hierarchical data representations hold promise as a solution to this problem, allowing for flexible conversion between different fidelities so that, for example, data can be created at high fidelity and then transferred or stored at lower fidelity via logically simple and mathematically sound operations. However, the effective use of such representations has been hindered until now by the relatively high costs of creating, accessing, reducing, and otherwise operating on such representations. We describe here highly optimized data refactoring kernels for GPU accelerators that enable efficient creation and manipulation of data in multigrid-based hierarchical forms. We demonstrate that our optimized design can achieve up to 250 TB/s aggregated data refactoring throughput-83% of theoretical peak-on 1024 nodes of the Summit supercomputer. We showcase our optimized design by applying it to a large-scale scientific visualization workflow and the MGARD lossy compression software. Jieyang Chen, Lipeng Wan 0001, Xin Liang 0001, Ben Whitney, Qing Liu 0002, David Pugmire, Nicholas Thompson, Jong Choi 0001, Matthew Wolf, Todd S. Munson, Ian T. Foster, Scott Klasky |
IPDPS | 5 |
| 2021 | zMesh: Exploring Application Characteristics to Improve Lossy Compression Ratio for Adaptive Mesh RefinementabstractScientific simulations on high-performance computing systems produce vast amounts of data that need to be stored and analyzed efficiently. Lossy compression significantly reduces the data volume by trading accuracy for performance. Despite the recent success of lossy compression, such as ZFP and SZ, the compression performance is still far from being able to keep up with the exponential growth of data. This paper aims to further take advantage of application characteristics, an area that is often under-explored, to improve the compression ratios of adaptive mesh refinement (AMR) - a widely used numerical solver that allows for an improved resolution in limited regions. We propose a level reordering technique zMesh to reduce the storage footprint of AMR applications. In particular, we group the data points that are mapped to the same or adjacent geometric coordinates such that the dataset is smoother and more compressible. Unlike the prior work where the compression performance is affected by the overhead of metadata, this work re-generates restore recipe using a chained tree structure, thus involving no extra storage overhead for compressed data, which substantially improves the compression ratios. The results demonstrate that zMesh can improve the smoothness of data by 67.9% and 71.3% for Z-ordering and Hilbert, respectively. Overall, zMesh improves the compression ratios by up to 16.5% and 133.7% for ZFP and SZ, respectively. Despite that zMesh involves additional compute overhead for tree and restore recipe construction, we show that the cost can be amortized as the number of quantities to be compressed increases. Huizhang Luo, Junqi Wang 0002, Qing Liu 0002, Jieyang Chen, Scott Klasky, Norbert Podhorszki |
IPDPS | 3 |
| 2021 | Reducing the Training Overhead of the HPC Compression Autoencoder via Dataset ProportioningabstractAs the storage overhead of high-performance computing (HPC) data reaches into the petabyte or even exabyte scale, it could be useful to find new methods of compressing such data. The compression autoencoder (CAE) has recently been proposed to compress HPC data with a very high compression ratio. However, this machine learning-based method suffers from the major drawback of lengthy training time. In this paper, we attempt to mitigate this problem by proposing a proportioning scheme to reduce the amount of data that is used for training relative to the amount of data to be compressed. We show that this method drastically reduces the training time without, in most cases, significantly increasing the error. We further explain how this scheme can even improve the accuracy of the CAE on certain datasets. Finally, we provide some guidance on how to determine a suitable proportion of the training dataset to use in order to train the CAE for a given dataset. Tong Liu 0030, Shakeel Alibhai, Jinzhen Wang, Qing Liu 0002, Xubin He |
NAS | 4 |
| 2021 | Error-controlled, progressive, and adaptable retrieval of scientific data with multilevel decompositionabstractExtreme-scale simulations and high-resolution instruments have been generating an increasing amount of data, which poses significant challenges to not only data storage during the run, but also post-processing where data will be repeatedly retrieved and analyzed for a long period of time. The challenges in satisfying a wide range of post-hoc analysis needs while minimizing the I/O overhead caused by inappropriate and/or excessive data retrieval should never be left unmanaged. In this paper, we propose a data refactoring, compressing, and retrieval framework capable of 1) fine-grained data refactoring with regard to precision; 2) incrementally retrieving and recomposing the data in terms of various error bounds; and 3) adaptively retrieving data in multi-precision and multi-resolution with respect to different analysis. With the progressive data re-composition and the adaptable retrieval algorithms, our framework significantly reduces the amount of data retrieved when multiple incremental precision are requested and/or the downstream analysis time when coarse resolution is used. Experiments show that the amount of data retrieved under the same progressively requested error bound using our framework is 64% less than that using state-of-the-art single-error-bounded approaches. Parallel experiments with up to 1, 024 cores and ~ 600 GB data in total show that our approach yields 1.36× and 2.52× performance over existing approaches in writing to and reading from persistent storage systems, respectively. Xin Liang 0001, Qian Gong, Jieyang Chen, Ben Whitney, Lipeng Wan 0001, Qing Liu 0002, David Pugmire, Rick Archibald, Norbert Podhorszki, Scott Klasky |
SC | 6 |
| 2021 | Biobjective Task Scheduling for Distributed Green Data CentersabstractThe industry of data centers is the fifth largest energy consumer in the world. Distributed green data centers (DGDCs) consume 300 billion kWh per year to provide different types of heterogeneous services to global users. Users around the world bring revenue to DGDC providers according to actual quality of service (QoS) of their tasks. Their tasks are delivered to DGDCs through multiple Internet service providers (ISPs) with different bandwidth capacities and unit bandwidth price. In addition, prices of power grid, wind, and solar energy in different GDCs vary with their geographical locations. Therefore, it is highly challenging to schedule tasks among DGDCs in a high-profit and high-QoS way. This work designs a multiobjective optimization method for DGDCs to maximize the profit of DGDC providers and minimize the average task loss possibility of all applications by jointly determining the split of tasks among multiple ISPs and task service rates of each GDC. A problem is formulated and solved with a simulated-annealing-based biobjective differential evolution (SBDE) algorithm to obtain an approximate Pareto-optimal set. The method of minimum Manhattan distance is adopted to select a knee solution that specifies the Pareto-optimal task service rates and task split among ISPs for DGDCs in each time slot. Real-life data-based experiments demonstrate that the proposed method achieves lower task loss of all applications and larger profit than several existing scheduling algorithms. Note to Practitioners-This work aims to maximize the profit and minimize the task loss for DGDCs powered by renewable energy and smart grid by jointly determining the split of tasks among multiple ISPs. Existing task scheduling algorithms fail to jointly consider and optimize the profit of DGDC providers and QoS of tasks. Therefore, they fail to intelligently schedule tasks of heterogeneous applications and allocate infrastructure resources within their response time bounds. In this work, a new method that tackles drawbacks of existing algorithms is proposed. It is achieved by adopting the proposed SBDE algorithm that solves a multiobjective optimization problem. Simulation experiments demonstrate that compared with three typical task scheduling approaches, it increases profit and decreases task loss. It can be readily and easily integrated and implemented in real-life industrial DGDCs. The future work needs to investigate the real-time green energy prediction with historical data and further combine prediction and task scheduling together to achieve greener and even net-zero-energy data centers. Haitao Yuan 0001, Jing Bi 0001, MengChu Zhou, Qing Liu 0002, Ahmed Chiheb Ammari |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2021 | Enhancing Proportional IO Sharing on Containerized Big Data File SystemsabstractBig Data platforms recently employ resource management systems, such as YARN, Mesos, and Google Borg, to provision computational resources. These systems adopt containerization to share the computing resources in a multi-tenant setting with low performance overhead and interference. However, it may be observed that tenants often interfere with each other on the underlying Big Data File Systems (BDFS), e.g., Hadoop File System, which have been widely deployed as a persistent layer in current data centers. A solution with systematic generality is to containerize BDFS itself to isolate and allocate its IO sources to multiple tenants. To this end, we conduct analysis on the ineffectiveness of proportionally sharing BDFS IO resource via containerization. This ineffectiveness is due to the scheduler of containerization in “pseudo-starvation” status, in which most of IO requests are backlogged in BDFS rather than in containerization scheduler. Without enough backlogged IO requests, existing schedulers might have to maximize device utilization rather than enforce proportional sharing policy. To resolve this ineffectiveness issue, we develop a cross-layer system calledBDFS-Container, which containerizes BDFS at the Linux block IO level. Central to BDFS-Container, we propose and design a proactive IOPS throttling-based mechanism namedIOPS Regulator, which achieves a trade-off between maximizing IO utilization and accurately proportional IO sharing. The evaluation results show that our method can improve proportionally sharing BDFS IO resources by 74.4 percent on average. Dan Huang 0001, Jun Wang 0001, Qing Liu 0002, Nong Xiao 0001, Huafeng Wu, Jiangling Yin |
IEEE Trans. Computers | 3 |
| 2020 | A Comprehensive Study of In-Memory Computing on Large HPC SystemsabstractWith the increasing fidelity and resolution enabled by high-performance computing systems, simulation-based scientific discovery is able to model and understand microscopic physical phenomena at a level that was not possible in the past. A grand challenge that the HPC community is faced with is how to handle the large amounts of analysis data generated from simulations. In-memory computing, among others, is recognized to be a viable path forward and has experienced tremendous success in the past decade. Nevertheless, there has been a lack of a complete study and understanding of in-memory computing as a whole on HPC systems. This paper presents a comprehensive study, which goes well beyond the typical performance metrics. In particular, we assess the in-memory computing with regard to its usability, portability, robustness and internal design trade-offs, which are the key factors that of interest to domain scientists. We use two realistic scientific workflows, LAMMPS and Laplace, to conduct comprehensive studies on state-of-the-art in-memory computing libraries, including DataSpaces, DIMES, Flexpath and Decaf. We conduct cross-platform experiments at scale on two leading supercomputers, Titan at ORNL and Cori at NERSC, and summarize our key findings in this critical area. Dan Huang 0001, Zhenlu Qin, Qing Liu 0002, Norbert Podhorszki, Scott Klasky |
ICDCS | 3 |
| 2020 | Taming I/O variation on QoS-less HPC storage: what can applications do?abstractAs high-performance computing (HPC) is being scaled up to exascale to accommodate new modeling and simulation needs, I/O has continued to be a major bottleneck in the end-to-end scientific processes. Nevertheless, prior work in this area mostly aimed to maximize the average performance, and there has been a lack of study and solutions that can manage I/O performance variation on HPC systems. This work aims to take advantage of the storage characteristics and explore application level solutions that are interference-aware. In particular, we monitor the performance of data analytics and estimate the state of shared storage resources using discrete fourier transform (DFT). If heavy I/O interference is predicted to occur at a given timestep, data analytics can dynamically adapt to the environment by lowering the accuracy and performing partial or no augmentation from the shared storage, dictated by an augmentation-bandwidth plot. We evaluate three data analytics, XGC, GenASiS, and Jet, on Chameleon, and quantitatively demonstrate that both the average and variation of I/O performance can be vastly improved using our dynamic augmentation, with the mean and variance improved by as much as 67% and 96%, respectively, while maintaining acceptable outcome of data analysis. Zhenbo Qiao, Qing Liu 0002, Norbert Podhorszki, Scott Klasky, Jieyang Chen |
SC | 2 |
| 2020 | Compression Ratio Modeling and Estimation across Error Bounds for Lossy CompressionabstractScientific simulations on high-performance computing (HPC) systems generate vast amounts of floating-point data that need to be reduced in order to lower the storage and I/O cost. Lossy compressors trade data accuracy for reduction performance and have been demonstrated to be effective in reducing data volume. However, a key hurdle to wide adoption of lossy compressors is that the trade-off between data accuracy and compression performance, particularly the compression ratio, is not well understood. Consequently, domain scientists often need to exhaust many possible error bounds before they can figure out an appropriate setup. The current practice of using lossy compressors to reduce data volume is, therefore, through trial and error, which is not efficient for large datasets which take a tremendous amount of computational resources to compress. This paper aims to analyze and estimate the compression performance of lossy compressors on HPC datasets. In particular, we predict the compression ratios of two modern lossy compressors that achieve superior performance, SZ and ZFP, on HPC scientific datasets at various error bounds, based upon the compressors' intrinsic metrics collected under a given base error bound. We evaluate the estimation scheme using twenty real HPC datasets and the results confirm the effectiveness of our approach. Jinzhen Wang, Tong Liu 0030, Qing Liu 0002, Xubin He, Huizhang Luo, Weiming He |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | Identifying Latent Reduced Models to Precondition Lossy CompressionabstractWith the high volume and velocity of scientific data produced on high-performance computing systems, it has become increasingly critical to improve the compression performance. Leveraging the general tolerance of reduced accuracy in applications, lossy compressors can achieve much higher compression ratios with a user-prescribed error bound. However, they are still far from satisfying the reduction requirements from applications. In this paper, we propose and evaluate the idea that data need to be preconditioned prior to compression, such that they can better match the design philosophies of a compressor. In particular, we aim to identify a reduced model that can be utilized to transform the original data to a more compressible form. We begin with a case study of Heat3d as a proof of concept, in which we demonstrate that a reduced model can indeed reside in the full model output, and can be utilized to improve compression ratios. We further explore more general dimension reduction techniques to extract the reduced model, including principal component analysis, singular value decomposition, and discrete wavelet transform. After preconditioning, the reduced model in conjunction with difference between the reduced model and full model is stored, which results in higher compression ratios. We evaluate the reduced models on nine scientific datasets, and the results show the effectiveness of our approaches. Huizhang Luo, Dan Huang 0001, Qing Liu 0002, Zhenbo Qiao, Hong Jiang 0001, Jing Bi 0001, Haitao Yuan 0001, MengChu Zhou, Jinzhen Wang, Zhenlu Qin |
IPDPS | 3 |
| 2019 | Exploring Transfer Learning to Reduce Training Overhead of HPC Data in Machine LearningabstractNowadays, scientific simulations on high-performance computing (HPC) systems can generate large amounts of data (in the scale of terabytes or petabytes) per run. When this huge amount of HPC data is processed by machine learning applications, the training overhead will be significant. Typically, the training process for a neural network can take several hours to complete, if not longer. When machine learning is applied to HPC scientific data, the training time can take several days or even weeks. Transfer learning, an optimization usually used to save training time or achieve better performance, has potential for reducing this large training overhead. In this paper, we apply transfer learning to a machine learning HPC application. We find that transfer learning can reduce training time without, in most cases, significantly increasing the error. This indicates transfer learning can be very useful for working with HPC datasets in machine learning applications. Tong Liu 0030, Shakeel Alibhai, Jinzhen Wang, Qing Liu 0002, Xubin He, Chentao Wu |
NAS | 4 |
| 2019 | Load-aware Elastic Data Reduction and Re-computation for Adaptive Mesh RefinementabstractThe increasing performance gap between computation and I/O creates huge data management challenges for simulation-based scientific discovery. Data reduction, among others, is deemed to be a promising technique to bridge the gap through reducing the amount of data migrated to persistent storage. However, the reduction performance is still far from what is being demanded from production applications. To this end, we propose a new methodology that aggressively reduces data despite the substantial loss of information, and re-computes the original accuracy on-demand. As a result, our scheme creates an illusion of a fast and large storage medium with the availability of high-accuracy data. We further design a load-aware data reduction strategy that monitors the I/O overhead at runtime, and dynamically adjusts the reduction ratio. We verify the efficacy of our methodology through adaptive mesh refinement, a popular numerical technique for solving partial differential equations. We evaluate data reduction and selective data re-computation on Titan, using a real application in FLASH and mini-applications in Chombo. To clearly demonstrate the benefits of re-computation, we compare it with other state-of-the-art data reduction methods including SZ, ZFP, FPC and deduplication, and it is shown to be superior in both write and read speeds, particularly when a small amount of data (e.g., 1%) need to be retrieved, as well as reduction ratio. Our results confirm that data reduction and selective data re-computation can 1) reduce the performance gap between I/O and compute via aggressively reducing AMR levels, and more importantly 2) can recover the target accuracy efficiently for AMR through re-computation. Mengxiao Wang, Huizhang Luo, Qing Liu 0002, Hong Jiang 0001 |
NAS | 3 |
| 2019 | Can I/O Variability Be Reduced on QoS-Less HPC Storage Systems?abstractFor a production high-performance computing (HPC) system, where storage devices are shared between multiple applications and managed in a best effort manner, I/O contention is often a major problem. In this paper, we propose a balanced messaging-based re-routing in conjunction with throttling at the middleware level. This work tackles two key challenges that have not been fully resolved in the past: whether I/O variability can be reduced on a QoS-less HPC storage system, and how to design a runtime scheduling system that can scale up to a large amount of cores. The proposed scheme uses a two-level messaging system to re-route I/O requests to a less congested storage location so that write performance is improved, while limiting the impact on read by throttling re-routing. An analytical model is derived to guide the setup of optimal throttling factor. We thoroughly analyze the virtual messaging layer overhead and explore whether the in-transit buffering is effective in managing I/O variability. Contrary to the intuition, in-transit buffer cannot completely solve the problem. It can reduce the absolute variability but not the relative variability. The proposed scheme is verified against a synthetic benchmark as well as being used by production applications. Dan Huang 0001, Qing Liu 0002, Jong Choi 0001, Norbert Podhorszki, Scott Klasky, Jeremy Logan, George Ostrouchov, Xubin He, Matthew Wolf |
IEEE Trans. Computers | 2 |
| 2019 | Harnessing Data Movement in Virtual Clusters for In-Situ ExecutionabstractAs a result of increasing data volume and velocity, Big Data science at exascale has shifted towards the in-situ paradigm, where large scale simulations run concurrently alongside data analytics. With in-situ, data generated from simulations can be processed while still in memory, thereby avoiding the slow storage bottleneck. However, running simulations and analytics together on shared resources will likely result in substantial contention if left unmanaged, as demonstrated in this work, leading to much reduced efficiency of simulations and analytics. Recently, virtualization technologies such as Linux containers have been widely applied to data centers and physical clusters to provide highly efficient and elastic resource provisioning for consolidated workloads including scientific simulations and data analytics. In this paper, we investigate to facilitate network traffic manipulation and reduce mutual interference on the network for in-situ applications in virtual clusters. In order to dynamically allocate the network bandwidth when it is needed, we adopt SARIMA-based techniques to analyze and predict MPI traffic issued from simulations. Although this can be an effective technique, the naïve usage of network virtualization can lead to performance degradation for bursty asynchronous transmissions within an MPI job. We analyze and resolve this performance degradation in virtual clusters. Dan Huang 0001, Qing Liu 0002, Scott Klasky, Jun Wang 0001, Jong Choi 0001, Jeremy Logan, Norbert Podhorszki |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | Coupling Exascale Multiphysics Applications: Methods and Lessons LearnedabstractWith the growing computational complexity of science and the complexity of new and emerging hardware, it is time to re-evaluate the traditional monolithic design of computational codes. One new paradigm is constructing larger scientific computational experiments from the coupling of multiple individual scientific applications, each targeting their own physics, characteristic lengths, and/or scales. We present a framework constructed by leveraging capabilities such as in-memory communications, workflow scheduling on HPC resources, and continuous performance monitoring. This code coupling capability is demonstrated by a fusion science scenario, where differences between the plasma at the edges and at the core of a device have different physical descriptions. This infrastructure not only enables the coupling of the physics components, but it also connects in situ or online analysis, compression, and visualization that accelerate the time between a run and the analysis of the science content. Results from runs on Titan and Cori are presented as a demonstration. Jong Choi 0001, Choong-Seock Chang, Julien Dominski, Scott Klasky, Gabriele Merlo, Eric Suchyta, Mark Ainsworth, Bryce Allen, Franck Cappello, Michael Churchill, Philip E. Davis, Sheng Di, Greg Eisenhauer, Stéphane Ethier, Ian T. Foster, Berk Geveci, Hanqi Guo 0001, Kevin A. Huck, Frank Jenko, Mark Kim, James Kress, Seung-Hoe Ku, Qing Liu 0002, Jeremy Logan, Allen D. Malony, Kshitij Mehta, Kenneth Moreland, Todd S. Munson, Manish Parashar, Tom Peterka, Norbert Podhorszki, David Pugmire, Ozan Tugluk, Ben Whitney, Matthew Wolf, Chad Wood |
eScience | 23 |
| 2018 | A View from ORNL: Scientific Data Research Opportunities in the Big Data AgeabstractOne of the core issues across computer and computational science today is adapting to, managing, and learning from the influx of "Big Data". In the commercial space, this problem has led to a huge investment in new technologies and capabilities that are well adapted to dealing with the sorts of human-generated logs, videos, texts, and other large-data artifacts that are processed and resulted in an explosion of useful platforms and languages (Hadoop, Spark, Pandas, etc.). However, translating this work from the enterprise space to the computational science and HPC community has proven somewhat difficult, in part because of some of the fundamental differences in type and scale of data and timescales surrounding its generation and use. We describe a forward-looking research and development plan which centers around the concept of making Input/Output (I/O) intelligent for users in the scientific community, whether they are accessing scalable storage or performing in situ workflow tasks. Much of our work is based on our experience with the Adaptable I/O System (ADIOS 1.X), and our next generation version of the software ADIOS 2.X [1]. Scott Klasky, Matthew Wolf, Mark Ainsworth, Chuck Atkins, Jong Choi 0001, Greg Eisenhauer, Berk Geveci, William F. Godoy, Mark Kim, James Kress, Tahsin M. Kurç, Qing Liu 0002, Jeremy Logan, Arthur B. Maccabe, Kshitij Mehta, George Ostrouchov, Manish Parashar, Norbert Podhorszki, David Pugmire, Eric Suchyta, Lipeng Wan 0001 |
ICDCS | 12 |
| 2018 | A Cost-effective and Energy-efficient Architecture for Die-stacked DRAM/NVM Memory SystemsabstractTraditional DRAM-based memory systems are facing two major scalability issues. First, the memory wall problem becomes a major performance bottleneck. Second, conventional memory systems consume increasing power as the capacity increases, which could be as much as 40% of the total system power. These issues hinder the scaling of DRAM-based memory systems. Fortunately, emerging memory technologies, such as high bandwidth memory (HBM) and phase change memory (PCM), have the potential to solve these scalability issues. However, there is no single memory technology that can overcome these issues together. Therefore, a hybrid memory system could be a promising way to build a high-performance, large-capacity, and energy-efficient memory system. To achieve this goal, we propose a cost-effective and energy-efficient architecture for HBM/PCM memory systems, called Dual Role HBM (DR-HBM). In DR-HBM, the HBM plays two roles and is divided into two parts. A small portion of which, called HBM cache, is used as a cache for the PCM. The remaining HBM is used as a part of main memory. Furthermore, the HBM cache is also used to track page hotness without additional hardware support. Hot pages will be migrated to HBM when they are evicted from the HBM cache. The experimental results show DR-HBM outperforms two state-of-the-art hybrid memory systems, CAMEO [1] and RaPP [2]. Compared to the baseline in which both HBM and PCM are architected as a part of main memory without page migration, DR-HBM improves the performance by 63% on average. Yuhua Guo, Weijun Xiao, Qing Liu 0002, Xubin He |
IPCCC | 3 |
| 2018 | Understanding and Modeling Lossy Compression Schemes on HPC Scientific DataabstractScientific simulations generate large amounts of floating-point data, which are often not very compressible using the traditional reduction schemes, such as deduplication or lossless compression. The emergence of lossy floating-point compression holds promise to satisfy the data reduction demand from HPC applications; however, lossy compression has not been widely adopted in science production. We believe a fundamental reason is that there is a lack of understanding of the benefits, pitfalls, and performance of lossy compression on scientific data. In this paper, we conduct a comprehensive study on state-of-the-art lossy compression, including ZFP, SZ, and ISABELA, using real and representative HPC datasets. Our evaluation reveals the complex interplay between compressor design, data features and compression performance. The impact of reduced accuracy on data analytics is also examined through a case study of fusion blob detection, offering domain scientists with the insights of what to expect from fidelity loss. Furthermore, the trial and error approach to understanding compression performance involves substantial compute and storage overhead. To this end, we propose a sampling based estimation method that extrapolates the reduction ratio from data samples, to guide domain scientists to make more informed data reduction decisions. Tao Lu 0014, Qing Liu 0002, Xubin He, Huizhang Luo, Eric Suchyta, Jong Choi 0001, Norbert Podhorszki, Scott Klasky, Matthew Wolf, Tong Liu 0030, Zhenbo Qiao |
IPDPS | 2 |
| 2018 | Write Energy Reduction for PCM via Pumping Efficiency ImprovementabstractThe emerging Phase Change Memory (PCM) is considered to be a promising candidate to replace DRAM as the next generation main memory due to its higher scalability and lower leakage power. However, the high write power consumption has become a major challenge in adopting PCM as main memory. In addition to the fact that writing to PCM cells requires high write current and voltage, current loss in the charge pumps also contributes a large percentage of high power consumption. The pumping efficiency of a PCM chip is a concave function of the write current. Leveraging the characteristics of the concave function, the overall pumping efficiency can be improved if the write current is uniform. In this article, we propose a peak-to-average (PTA) write scheme, which smooths the write current fluctuation by regrouping write units. In particular, we calculate the current requirements for each write unit by their values when they are evicted from the last level cache (LLC). When the write units are waiting in the memory controller, we regroup the write units by LLC-assisted PTA to reach the current-uniform goal. Experimental results show that LLC-assisted PTA achieved 13.4% of overall energy saving compared to the baseline. Huizhang Luo, Qing Liu 0002, Jingtong Hu, Qiao Li 0001, Liang Shi 0001, Qingfeng Zhuge, Edwin H.-M. Sha |
ACM Trans. Storage | 2 |
| 2017 | TGE: Machine Learning Based Task Graph Embedding for Large-Scale Topology MappingabstractTask mapping is an important problem in parallel and distributed computing. The goal in task mapping is to find an optimal layout of the processes of an application (or a task) onto a given network topology. We target this problem in the context of staging applications. A staging application consists of two or more parallel applications (also referred to as staging tasks) which run concurrently and exchange data over the course of computation. Task mapping becomes a more challenging problem in staging applications, because not only data is exchanged between the staging tasks, but also the processes of a staging task may exchange data with each other. We propose a novel method, called Task Graph Embedding (TGE), that harnesses the observable graph structures of parallel applications and network topologies. TGE employs a machine learning based algorithm to find the best representation of a graph, called an embedding, onto a space in which the task-to-processor mapping problem can be solved. We evaluate and demonstrate the effectiveness of TGE experimentally with the communication patterns extracted from runs of XGC, a large-scale fusion simulation code, on Titan. Jong Choi 0001, Jeremy Logan, Matthew Wolf, George Ostrouchov, Tahsin M. Kurç, Qing Liu 0002, Norbert Podhorszki, Scott Klasky, Melissa Romanus, Manish Parashar, Michael Churchill, Choong-Seock Chang |
CLUSTER | 6 |
| 2017 | Canopus: A Paradigm Shift Towards Elastic Extreme-Scale Data Analytics on HPC StorageabstractScientific simulations on high performance computing (HPC) platforms generate large quantities of data. To bridge the widening gap between compute and I/O, and enable data to be more efficiently stored and analyzed, simulation outputs need to be refactored, reduced, and appropriately mapped to storage tiers. However, a systematic solution to support these steps has been lacking on the current HPC software ecosystem. To that end, this paper develops Canopus, a progressive JPEGlike data management scheme for storing and analyzing big scientific data. It co-designs the data decimation, compression and data storage, taking the hardware characteristics of each storage tier into considerations. With reasonably low overhead, our approach refactors simulation data into a much smaller, reduced-accuracy base dataset, and a series of deltas that is used to augment the accuracy if needed. The base dataset and deltas are compressed and written to multiple storage tiers. Data saved on different tiers can then be selectively retrieved to restore the level of accuracy that satisfies data analytics. Thus, Canopus provides a paradigm shift towards elastic data analytics and enables end users to make trade-offs between analysis speed and accuracy on-the-fly. We evaluate the impact of Canopus on unstructured triangular meshes, a pervasive data model used by scientific modeling and simulations. In particular, we demonstrate the progressive data exploration of Canopus using the “blob detection” use case on the fusion simulation data. Tao Lu 0014, Eric Suchyta, David Pugmire, Jong Choi 0001, Scott Klasky, Qing Liu 0002, Norbert Podhorszki, Mark Ainsworth, Matthew Wolf |
CLUSTER | 6 |
| 2017 | Computing Just What You Need: Online Data Analysis and Reduction at Extreme Scales
Ian T. Foster, Mark Ainsworth, Bryce Allen, Julie Bessac, Franck Cappello, Jong Choi 0001, Emil M. Constantinescu, Philip E. Davis, Sheng Di, Zichao Wendy Di, Hanqi Guo 0001, Scott Klasky, Kerstin Kleese van Dam, Tahsin M. Kurç, Qing Liu 0002, Abid Malik, Kshitij Mehta, Klaus Mueller 0001, Todd S. Munson, George Ostrouchov, Manish Parashar, Tom Peterka, Line C. Pouchard, Dingwen Tao, Ozan Tugluk, Stefan M. Wild, Matthew Wolf, Justin M. Wozniak, Wei Xu 0020, Shinjae Yoo |
Euro-Par | 15 |
| 2017 | Canopus: Enabling Extreme-Scale Data Analytics on Big HPC Storage via Progressive Refactoring
Tao Lu 0014, Eric Suchyta, Jong Choi 0001, Norbert Podhorszki, Scott Klasky, Qing Liu 0002, David Pugmire, Matthew Wolf, Mark Ainsworth |
HotStorage | 6 |
| 2017 | Exacution: Enhancing Scientific Data Management for ExascaleabstractAs we continue toward exascale, scientific data volume is continuing to scale and becoming more burdensome to manage. In this paper, we lay out opportunities to enhance state of the art data management techniques. We emphasize well-principled data compression, and using it to achieve progressive refinement. This can both accelerate I/O and afford the user increased flexibility when she interacts with the data. The formulation naturally maps onto enabling partitioning of the progressively improving-quality representations of a data quantity into different media-type destinations, to keep the highest priority information as close as possible to the computation, and take advantage of deepening memory/storage hierarchies in ways not previously possible. Careful monitoring is requisite to our vision, not only to verify that compression has not eliminated salient features in the data, but also to better understand the performance of massively parallel scientific applications. Increased mathematical rigor would be ideal,to help bring compression on a better-understood theoretical footing, closer to the relevant scientific theory, more aware of constraints imposed by the science, and more tightly error-controlled. Throughout, we highlight pathfinding research we have begun exploring related these topics, and comment toward future work that will be needed. Scott Klasky, Eric Suchyta, Mark Ainsworth, Qing Liu 0002, Ben Whitney, Matthew Wolf, Jong Choi 0001, Ian T. Foster, Mark Kim, Jeremy Logan, Kshitij Mehta, Todd S. Munson, George Ostrouchov, Manish Parashar, Norbert Podhorszki, David Pugmire, Lipeng Wan 0001 |
ICDCS | 4 |
| 2017 | SELF: A High Performance and Bandwidth Efficient Approach to Exploiting Die-Stacked DRAM as Part of MemoryabstractDie-stacked DRAM (a.k.a., on-chip DRAM) provides much higher bandwidth and lower latency than off-chip DRAM. It is a promising technology to break the "memory wall". Die-stacked DRAM can be used either as a cache (i.e., DRAM cache) or as a part of memory (PoM). A DRAM cache design would suffer from more page faults than a PoM design as the DRAM cache cannot contribute towards capacity of main memory. At the same time, obtaining high performance requires PoM systems to swap requested data to the die-stacked DRAM. Existing PoM designs fall into two categories – line-based and page-based. The former ensures low off-chip bandwidth utilization but suffers from a low hit ratio of on-chip memory due to limited temporal locality. In contrast, page-based designs achieve a high hit ratio of on-chip memory albeit at the cost of moving large amounts of data between on-chip and off-chip memories, leading to increased off-chip bandwidth utilization and significant system performance degradation.To achieve a similar high hit ratio of on-chip memory as page-based designs, and eliminate excessive off-chip traffic involved, we propose SELF, a high performance and bandwidth efficient approach. The key idea is to SElectively swap Lines in a requested page that are likely to be accessed according to page Footprint, instead of blindly swapping an entire page. In doing so, SELF allows incoming requests to be serviced from the on-chip memory as much as possible, while avoiding swapping unused lines to reduce memory bandwidth consumption. We evaluate a memory system which consists of 4GB on-chip DRAM and 12GB off-chip DRAM. Compared to a baseline system that has the same total capacity of 16GB off-chip DRAM, SELF improves the performance in terms of instructions per cycle by 26.9%, and reduces the energy consumption per memory access by 47.9% on average. In contrast, state-of-the-art line-based and page-based PoM designs can only improve the performance by 9.5% and 9.9%, respectively, against the same baseline system. Yuhua Guo, Qing Liu 0002, Weijun Xiao, Ping Huang 0001, Norbert Podhorszki, Scott Klasky, Xubin He |
MASCOTS | 2 |
| 2015 | Exploring Memory Hierarchy to Improve Scientific Data Read PerformanceabstractImproving read performance is one of the major challenges with speeding up scientific data analytic applications. Utilizing the memory hierarchy is one major line of researches to address the read performance bottleneck. Related methods usually combine solide-state-drives(SSDs) with dynamic random-access memory(DRAM) and/or parallel file system(PFS) to mitigate the speed and space gap between DRAM and PFS. However, these methods are unable to handle key performance issues plaguing SSDs, namely read contention that may cause up to 50% performance reduction. In this paper, we propose a framework that exploits the memory hierarchy resource to address the read contention issues involved with SSDs. The framework employs a general purpose online read algorithm that able to detect and utilize memory hierarchy resource to relieve the problem. To maintain a near optimal operating environment for SSDs, the framework is able to orchastrate data chunks across different memory layers to facilitate the read algorithm. Compared to existing tools, our framework achieves up to 50% read performance improvement when tested on datasets from real-world scientific simulations. Wenzhao Zhang, Houjun Tang, Xiaocheng Zou, Steve Harenberg, Qing Liu 0002, Scott Klasky, Nagiza F. Samatova |
CLUSTER | 5 |
| 2015 | Combining phase identification and statistic modeling for automated parallel benchmark generationabstractParallel application benchmarks are indispensable for evaluating/optimizing HPC software and hardware. However, it is very challenging and costly to obtain high-fidelity benchmarks reflecting the scale and complexity of state-of-the-art parallel applications. Hand-extracted synthetic benchmarks are time- and labor-intensive to create. Real applications themselves, while offering most accurate performance evaluation, are expensive to compile, port, recon- figure, and often plainly inaccessible due to security or ownership concerns. This work contributes APPRIME, a novel tool for trace-based automatic parallel benchmark generation. Taking as input standard communication-I/O traces of an application’s execution, it couples accurate automatic phase identification with statistical regeneration of event parameters to create compact, portable, and to some degree reconfigurable parallel application benchmarks. Experiments with four NAS Parallel Benchmarks (NPB) and three real scientific simulation codes confirm the fidelity of APPRIME benchmarks. They retain the original applications’ performance characteristics, in particular the relative performance across platforms. Xiaosong Ma, Qing Liu 0002, Jeremy Logan, Norbert Podhorszki, Jong Choi 0001, Scott Klasky |
PPoPP | 4 |
| 2015 | Combining Phase Identification and Statistic Modeling for Automated Parallel Benchmark GenerationabstractParallel application benchmarks are indispensable for evaluating/optimizing HPC software and hardware. However, it is very challenging and costly to obtain high-fidelity benchmarks reflecting the scale and complexity of state-of-the-art parallel applications. Hand-extracted synthetic benchmarks are time- and labor-intensive to create. Real applications themselves, while offering most accurate performance evaluation, are expensive to compile, port, reconfigure, and often plainly inaccessible due to security or ownership concerns. This work contributes APPrime, a novel tool for trace-based automatic parallel benchmark generation. Taking as input standard communication-I/O traces of an application's execution, it couples accurate automatic phase identification with statistical regeneration of event parameters to create compact, portable, and to some degree reconfigurable parallel application benchmarks. Experiments with four NAS Parallel Benchmarks (NPB) and three real scientific simulation codes confirm the fidelity of APPrime benchmarks. They retain the original applications' performance characteristics, in particular their relative performance across platforms. Also, the result benchmarks, already released online, are much more compact and easy-to-port compared to the original applications. Xiaosong Ma, Qing Liu 0002, Jeremy Logan, Norbert Podhorszki, Jong Choi 0001, Scott Klasky |
SIGMETRICS | 4 |
| 2014 | Transparent in Situ Data Transformations in ADIOSabstractThough an abundance of novel "data transformation" technologies have been developed (such as compression, level-of-detail, layout optimization, and indexing), there remains a notable gap in the adoption of such services by scientific applications. In response, we develop an in situ data transformation framework in the ADIOS I/O middleware with a "plug in" interface, thus greatly simplifying both the deployment and use of data transform services in scientific applications. Our approach ensures user-transparency, runtime-configurability, compatibility with existing I/O optimizations, and the potential for exploiting read-optimizing transforms (such as level-of-detail) to achieve I/O reduction. We demonstrate use of our framework with the QLG simulation at up to 8,192 cores on the leadership-class Titan supercomputer, showing negligible overhead. We also explore the read performance implications of data transforms with respect to parameters such as chunk size, access pattern, and the "opacity" of different transform methods including compression and level-of-detail. David A. Boyuka II, Sriram Lakshminarasimhan, Xiaocheng Zou, Zhenhuan Gong, John Jenkins, Eric R. Schendel, Norbert Podhorszki, Qing Liu 0002, Scott Klasky, Nagiza F. Samatova |
CCGRID | 8 |
| 2014 | Hello ADIOS: the challenges and lessons of developing leadership class I/O frameworksabstractSUMMARY Applications running on leadership platforms are more and more bottlenecked by storage input/output (I/O). In an effort to combat the increasing disparity between I/O throughput and compute capability, we created Adaptable IO System (ADIOS) in 2005. Focusing on putting users first with a service oriented architecture, we combined cutting edge research into new I/O techniques with a design effort to create near optimal I/O methods. As a result, ADIOS provides the highest level of synchronous I/O performance for a number of mission critical applications at various Department of Energy Leadership Computing Facilities. Meanwhile ADIOS is leading the push for next generation techniques including staging and data processing pipelines. In this paper, we describe the startling observations we have made in the last half decade of I/O research and development, and elaborate the lessons we have learned along this journey. We also detail some of the challenges that remain as we look toward the coming Exascale era. Copyright © 2013 John Wiley & Sons, Ltd. Qing Liu 0002, Jeremy Logan, Yuan Tian 0004, Hasan Abbasi, Norbert Podhorszki, Jong Choi 0001, Scott Klasky, Roselyne Tchoua, Jay F. Lofstead, Ron A. Oldfield, Manish Parashar, Nagiza F. Samatova, Karsten Schwan, Arie Shoshani, Matthew Wolf, Kesheng Wu, Weikuan Yu |
Concurr. Comput. Pract. Exp. | 1 |
| 2013 | PARLO: PArallel Run-Time Layout Optimization for Scientific Data Explorations with Heterogeneous Access PatternsabstractThe size and scope of cutting-edge scientific simulations are growing much faster than the I/O and storage capabilities of their run-time environments. The growing gap is exacerbated by exploratory, data-intensive analytics, such as querying simulation data with multivariate, spatio-temporal constraints, which induces heterogeneous access patterns that stress the performance of the underlying storage system. Previous work addresses data layout and indexing techniques to improve query performance for a single access pattern, which is not sufficient for complex analytics jobs. We present PARLO a parallel run-time layout optimization framework, to achieve multi-level data layout optimization for scientific applications at run-time before data is written to storage. The layout schemes optimize for heterogeneous access patterns with user-specified priorities. PARLO is integrated with ADIOS, a high-performance parallel I/O middleware for large-scale HPC applications, to achieve user-transparent, light-weight layout optimization for scientific datasets. It offers simple XML-based configuration for users to achieve flexible layout optimization without the need to modify or recompile application codes. Experiments show that PARLO improves performance by 2 to 26 times for queries with heterogeneous access patterns compared to state-of-the-art scientific database management systems. Compared to traditional post-processing approaches, its underlying run-time layout optimization achieves a 56% savings in processing time and a reduction in storage overhead of up to 50%. PARLO also exhibits a low run-time resource requirement, while also limiting the performance impact on running applications to a reasonable level. Zhenhuan Gong, David A. Boyuka II, Xiaocheng Zou, Qing Liu 0002, Norbert Podhorszki, Scott Klasky, Xiaosong Ma, Nagiza F. Samatova |
CCGRID | 4 |
| 2013 | FlexQuery: An online query system for interactive remote visual data exploration at large scaleabstractThe remote visual exploration of live data generated by scientific simulations is useful for scientific discovery, performance monitoring, and online validation for the simulation results. Online visualization methods are challenged, however, by the continued growth in the volume of simulation output data that has to be transferred from its source - the simulation running on the high end machine - to where it is analyzed, visualized, and displayed. A specific challenge in this context is limits in the communication bandwidth between data source(s) and sinks. Previous work places queries `near' data sources, exploiting their data reduction capabilities, but such work does not address the common scenario in which scientists make multiple different queries on the data being produced. This paper considers the general case in which science users are interested in different (sub)sets of the data produced by a high end simulation. We offer the FlexQuery online data query system that can deploy and execute data queries `along' the I/O and analytics pipelines. FlexQuery carefully extends such analytics pipelines, using online performance monitoring and data location tracking, to realize data queries in ways that minimize additional data movement and offer low latency in data query execution. Using a real-world scientific application - the Maya astrophysics code and its analytics workflow - we demonstrate FlexQuery's ability to dynamically deploy queries for low-latency remote data visualization. Hongbo Zou, Karsten Schwan, Magdalena Slawiñska, Matthew Wolf, Greg Eisenhauer, Fang Zheng 0003, Jai Dayal, Jeremy Logan, Qing Liu 0002, Scott Klasky, Tanja Bode, Michael Clark, Matthew Kinsey |
CLUSTER | 9 |
| 2013 | ADIOS Visualization Schema: A First Step Towards Improving Interdisciplinary Collaboration in High Performance ComputingabstractScientific communities have benefitted from a significant increase of available computing and storage resources in the last few decades. For science projects that have access to leadership scale computing resources, the capacity to produce data has been growing exponentially. Teams working on such projects must now include, in addition to the traditional application scientists, experts in various disciplines including applied mathematicians for development of algorithms, visualization specialists for large data, and I/O specialists. Sharing of knowledge and data is becoming a requirement for scientific discovery, providing useful mechanisms to facilitate this sharing is a key challenge for e-Science. Our hypothesis is that in order to decrease the time to solution for application scientists we need to lower the barrier of entry into related computing fields. We aim at improving users' experience when interacting with a vast software ecosystem and/or huge amount of data, while maintaining focus on their primary research field. In this context we present our approach to bridge the gap between the application scientists and the visualization experts through a visualization schema as a first step and proof of concept for a new way to look at interdisciplinary collaboration among scientists dealing with big data. The key to our approach is recognizing that our users are scientists who mostly work as islands. They tend to work in very specialized environment but occasionally have to collaborate with other researchers in order to take full advantage of computing innovations and get insight from big data. We present an example of identifying the connecting elements between one of such relationships and offer a liaison schema to facilitate their collaboration. Roselyne Tchoua, Jong Choi 0001, Scott Klasky, Qing Liu 0002, Jeremy Logan, Kenneth Moreland, Jingqing Mu, Manish Parashar, Norbert Podhorszki, David Pugmire, Matthew Wolf |
e-Science | 4 |
| 2013 | Runtime I/O Re-Routing + Throttling on HPC Storage
Qing Liu 0002, Norbert Podhorszki, Jeremy Logan, Scott Klasky |
HotStorage | 1 |
| 2012 | Understanding I/O Performance Using I/O Skeletal Applications
Jeremy Logan, Scott Klasky, Hasan Abbasi, Qing Liu 0002, George Ostrouchov, Manish Parashar, Norbert Podhorszki, Yuan Tian 0004, Matthew Wolf |
Euro-Par | 4 |
| 2012 | ISOBAR hybrid compression-I/O interleaving for large-scale parallel I/O optimizationabstractCurrent peta-scale data analytics frameworks suffer from a significant performance bottleneck due to an imbalance between their enormous computational power and limited I/O bandwidth. Using data compression schemes to reduce the amount of I/O activity is a promising approach to addressing this problem. In this paper, we propose a hybrid framework for interleaving I/O with data compression to achieve improved I/O throughput side-by-side with reduced dataset size. We evaluate several interleaving strategies, present theoretical models, and evaluate the efficiency and scalability of our approach through comparative analysis. With our theoretical model, considering 19 real-world scientific datasets both from the public domain and peta-scale simulations, we estimate that the hybrid method can result in a 12 to 46 increase in throughput on hard-to-compress scientific datasets. At the reported peak bandwidth of 60 GB/s of uncompressed data for a current, leadership-class parallel I/O system, this translates into an effective gain of 7 to 28 GB/s in aggregate throughput. Eric R. Schendel, Saurabh V. Pendse, John Jenkins, David A. Boyuka II, Zhenhuan Gong, Sriram Lakshminarasimhan, Qing Liu 0002, Hemanth Kolla, Jackie Chen, Scott Klasky, Robert B. Ross, Nagiza F. Samatova |
HPDC | 7 |
| 2011 | EDO: Improving Read Performance for Scientific Applications through Elastic Data OrganizationabstractLarge scale scientific applications are often bottlenecked due to the writing of checkpoint-restart data. Much work has been focused on improving their write performance. With the mounting needs of scientific discovery from these datasets, it is also important to provide good read performance for many common access patterns, which requires effective data organization. To address this issue, we introduce Elastic Data Organization (EDO), which can transparently enable different data organization strategies for scientific applications. Through its flexible data ordering algorithms, EDO harmonizes different access patterns with the underlying file system. Two levels of data ordering are introduced in EDO. One works at the level of data groups (a.k.a process groups). It uses Hilbert Space Filling Curves (SFC) to balance the distribution of data groups across storage targets. Another governs the ordering of data elements within a data group. It divides a data group into sub chunks and strikes a good balance between the size of sub chunks and the number of seek operations. Our experimental results demonstrate that EDO is able to achieve balanced data distribution across all dimensions and improve the read performance of multidimensional datasets in scientific applications. Yuan Tian 0004, Scott Klasky, Hasan Abbasi, Jay F. Lofstead, Ray W. Grout, Norbert Podhorszki, Qing Liu 0002, Yandong Wang 0001, Weikuan Yu |
CLUSTER | 7 |
| 2011 | Six degrees of scientific data: reading patterns for extreme scale science IOabstractPetascale science simulations generate 10s of TBs of application data per day, much of it devoted to their checkpoint/restart fault tolerance mechanisms. Previous work demonstrated the importance of carefully managing such output to prevent application slowdown due to IO blocking, resource contention negatively impacting simulation performance and to fully exploit the IO bandwidth available to the petascale machine. This paper takes a further step in understanding and managing extreme-scale IO. Specifically, its evaluations seek to understand how to efficiently read data for subsequent data analysis, visualization, checkpoint restart after a failure, and other read-intensive operations. In their entirety, these actions support the 'end-to-end' needs of scientists enabling the scientific processes being undertaken. Contributions include the following. First, working with application scientists, we define 'read' benchmarks that capture the common read patterns used by analysis codes. Second, these read patterns are used to evaluate different IO techniques at scale to understand the effects of alternative data sizes and organizations in relation to the performance seen by end users. Third, defining the novel notion of a 'data district' to characterize how data is organized for reads, we experimentally compare the read performance seen with the ADIOS middleware's log-based BP format to that seen by the logically contiguous NetCDF or HDF5 formats commonly used by analysis tools. Measurements assess the performance seen across patterns and with different data sizes, organizations, and read process counts. Outcomes demonstrate that high end-to-end IO performance requires data organizations that offer flexibility in data layout and placement on parallel storage targets, including in ways that can make tradeoffs in the performance of data writes vs. reads. Jay F. Lofstead, Milo Polte, Garth A. Gibson, Scott Klasky, Karsten Schwan, Ron A. Oldfield, Matthew Wolf, Qing Liu 0002 |
HPDC | 8 |
| 2010 | Enhanced Crankback Signaling for Multi-Domain Traffic EngineeringabstractMulti-domain traffic engineering is a major focus area for carriers and crankback signaling offers a very promising alternative. However, even though various crankback studies have been done, there remains significant latitude for improved multi-domain designs. To address these challenges, this work develops a novel solution for intra/inter-domain signaling crankback in IP/MPLS networks. Namely, dynamic intra-domain link-state routing information is coupled with inter-domain path/distance-vector routing state to improve the search process. Mechanisms are also added to limit setup signaling overheads and track crankback history from congested links. The performance of the proposed solution is analyzed using simulation and compared against other techniques including hierarchical inter-domain routing. Mostafa Esmaeili, Chongyang Xie, Nasir Ghani, Min Peng 0002, Qing Liu 0002 |
ICC | 6 |
| 2010 | PreDatA - preparatory data analytics on peta-scale machinesabstractPeta-scale scientific applications running on High End Computing (HEC) platforms can generate large volumes of data. For high performance storage and in order to be useful to science end users, such data must be organized in its layout, indexed, sorted, and otherwise manipulated for subsequent data presentation, visualization, and detailed analysis. In addition, scientists desire to gain insights into selected data characteristics `hidden' or `latent' in these massive datasets while data is being produced by simulations. PreDatA, short for Preparatory Data Analytics, is an approach to preparing and characterizing data while it is being produced by the large scale simulations running on peta-scale machines. By dedicating additional compute nodes on the machine as `staging' nodes and by staging simulations' output data through these nodes, PreDatA can exploit their computational power to perform select data manipulations with lower latency than attainable by first moving data into file systems and storage. Such intransit manipulations are supported by the PreDatA middleware through asynchronous data movement to reduce write latency, application-specific operations on streaming data that are able to discover latent data characteristics, and appropriate data reorganization and metadata annotation to speed up subsequent data access. PreDatA enhances the scalability and flexibility of the current I/O stack on HEC platforms and is useful for data pre-processing, runtime data analysis and inspection, as well as for data exchange between concurrently running simulations. Fang Zheng 0003, Hasan Abbasi, Ciprian Docan, Jay F. Lofstead, Qing Liu 0002, Scott Klasky, Manish Parashar, Norbert Podhorszki, Karsten Schwan, Matthew Wolf |
IPDPS | 5 |
| 2010 | Managing Variability in the IO Performance of Petascale Storage SystemsabstractSignificant challenges exist for achieving peak or even consistent levels of performance when using IO systems at scale. They stem from sharing IO system resources across the processes of single largescale applications and/or multiple simultaneous programs causing internal and external interference, which in turn, causes substantial reductions in IO performance. This paper presents interference effects measurements for two different file systems at multiple supercomputing sites. These measurements motivate developing a 'managed' IO approach using adaptive algorithms varying the IO system workload based on current levels and use areas. An implementation of these methods deployed for the shared, general scratch storage system on Oak Ridge National Laboratory machines achieves higher overall performance and less variability in both a typical usage environment and with artificially introduced levels of 'noise'. The latter serving to clearly delineate and illustrate potential problems arising from shared system usage and the advantages derived from actively managing it. Jay F. Lofstead, Fang Zheng 0003, Qing Liu 0002, Scott Klasky, Ron A. Oldfield, Todd Kordenbrock, Karsten Schwan, Matthew Wolf |
SC | 3 |
| 2009 | Distributed Grooming in Multi-Domain IP/MPLS-DWDM NetworksabstractThis paper studies distributed multi-domain, multi-layer provisioning (grooming) in IP/MPLS-DWDM networks. Although many multi-domain studies have emerged over the years, these have primarily considered "homogeneous" network layers. Meanwhile, most grooming studies have assumed idealized settings with "global" link state across all layers. Hence there is a critical need to develop practical distributed grooming schemes for real-world networks consisting of multiple domains and technology layers. Along these lines, a detailed hierarchical framework is proposed to implement inter-layer routing, distributed grooming, and setup signaling. The performance of this solution is analyzed using simulation studies and future directions high-lighted. Qing Liu 0002, Tannous Frangieh, Chongyang Xie, Nasir Ghani, Ashwin Gumaste, Tom Lehman, Chin Guok, Scott Klasky |
GLOBECOM | 1 |
| 2009 | Multi-Point Ethernet over Next-Generation SONET/SDHabstractAdvances in SONET/SDH technologies have introduced novel features for improved services mapping and provisioning, enabling many new avenues for new Carrier Ethernet support. However Ethernet-over-SONET studies have mostly focused on provisioning point-to-point Ethernet private line offerings. This paper considers the more challenging case of provisioning multi-point-to-multi-point Ethernet LAN services over advanced SONET/SDH networks and presents novel strategies based upon connection group overlays. Detailed simulation results are also presented along with directions for future work. Chongyang Xie, Nasir Ghani, Qing Liu 0002, Wei Wennie Shu, Ashwin Gumaste, Min-You Wu |
ICC | 3 |
| 2008 | Inter-Domain Routing Scalability in Optical DWDM NetworksabstractRecent studies on inter-domain DWDM networks have focused on topology abstraction for state summarization, i.e. transforming a physical topology to a virtual mesh, tree, or star network. Although these schemes give very good inter- domain blocking reduction, associated inter-domain routing overheads are significant, particularly as the number of domains and border OXC nodes increase. To address these scalability limitations, novel routing update triggering policies for multi-domain DWDM networks are developed. The performance of inter-domain lightpath RWA and signaling schemes in conjunction with these strategies is then studied in order to gauge the overall effectiveness of these approaches. Qing Liu 0002, Chongyang Xie, Tannous Frangieh, Nasir Ghani, Ashwin Gumaste, Nageswara S. V. Rao, Tom Lehman |
ICCCN | 1 |
| 2007 | Distributed inter-domain lightpath provisioning in the presence of wavelength conversion
Qing Liu 0002, Nasir Ghani, Nageswara S. V. Rao, Ashwin Gumaste, M. L. Garcia |
Comput. Commun. | 1 |
| 2006 | Inter-Domain Provisioning in DWDM NetworksabstractAs DWDM technology proliferates there is a growing need to address distributed inter-domain lightpath provisioning issues. Although inter-domain provisioning has been well-studied for packet/cell- switching networks, the wavelength dimension presents many additional challenges. This paper develops a novel hierarchical GMPLS-based framework for provisioning all-optical and opto-electronic multi-domain DWDM networks. The scheme adapts topology abstraction schemes to aggregate domain-level state to improve routing scalability and lower inter-domain blocking. Related inter-domain lightpath RWA and signaling schemes are also tabled. Performance analysis results are presented along with directions for future work. Qing Liu 0002, Mehmet A. Kok, Nasir Ghani, V. M. Muthalaly |
GLOBECOM | 1 |
| 2006 | Application of topology abstraction techniques in multi-domain optical networksabstractAs DWDM networks proliferate there is a growing need to address the issue of distributed interdomain lightpath provisioning. Although inter-domain provisioning has been well-studied for packet/cellswitching networks, the wavelength dimension presents many additional challenges. This paper develops a novel hierarchical GMPLS-based framework for provisioning all-optical and opto-electronic multi-domain DWDM networks. In particular, several topology abstraction schemes are proposed for aggregating domain-level state to improve routing scalability and lower inter-domain blocking. Inter-domain lightpath RWA and signaling schemes are also tabled. Performance analysis results are presented along with directions for future work. Qing Liu 0002, Nasir Ghani, Mehmet A. Kok |
ICCCN | 1 |
| 2006 | Hierarchical Inter-Domain Routing in Optical DWDM NetworksabstractAs DWDM technology proliferates there is a growing need to address distributed inter-domain lightpath provisioning issues. Although inter-domain provisioning has been well-studied for packet/cell- switching networks, the wavelength dimension presents many additional challenges. This paper develops a novel hierarchical GMPLS-based framework for provisioning all-optical and opto-electronic multi-domain DWDM networks. The scheme adapts topology abstraction schemes to aggregate domain-level state to improve routing scalability and lower inter-domain blocking. Related inter-domain lightpath RWA and signaling schemes are also tabled. Performance analysis results are presented along with directions for future work. Qing Liu 0002, Mehmet A. Kok, Nasir Ghani, V. M. Muthalaly |
INFOCOM | 1 |
| 2006 | Hierarchical routing in multi-domain optical networks
Qing Liu 0002, Mehmet A. Kok, Nasir Ghani, Ashwin Gumaste |
Comput. Commun. | 1 |