VLDB 2026 Research / reviewers in the wild / expert
Lipeng Wan 0001
dblp:124/2262-1
· DBLP profile ↗
32ranked-venue papers
10as first author
17since 2021 · last 2026
0000-0003-2347-8667ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 7 first-author · 13 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Computer networks · 2Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | JANUS: Resilient and Adaptive Data Transmission for Enabling Timely and Efficient Cross-Facility Scientific WorkflowsabstractIn modern science, the growing complexity of large-scale scientific projects has led to an increasing reliance on cross-facility scientific workflows, where resources and expertise from multiple institutions and geographic locations are leveraged to accelerate scientific discovery. These workflows often require transmitting huge amounts of scientific data through wide-area networks. Although high-speed networks like ESnet and transfer services such as Globus have improved data mobility, several challenges remain. The sheer volume of data can overwhelm network bandwidth, widely used transport protocols such as TCP suffer from inefficiencies due to retransmissions triggered by packet loss, and existing fault-tolerance mechanisms like erasure coding introduce substantial overhead. Vladislav Esaulov, Jieyang Chen, Norbert Podhorszki, Frédéric Suter, Scott Klasky, Anu G. Bourgeois, Lipeng Wan 0001 |
HPDC | 7 |
| 2025 | LegoIndex: A Scalable and Modular Indexing Framework for Efficient Analysis of Extreme-Scale Particle DataabstractParticle-in-Cell (PIC) simulations play a critical role in various scientific domains, including plasma physics, astrophysics, and fusion energy research, by enabling the modeling of complex interactions between charged particles and electromagnetic fields. As PIC simulations scale up in size and complexity, they generate massive volumes of particle data at enormous speeds (TBs/hour). This enormous amount of data presents significant challenges for post-simulation analysis, as existing analysis tools (typically designed for smaller datasets) struggle with low query performance and high resource utilization. While incorporating indexes for PIC data can alleviate some of these inefficiencies, current indexing solutions often fall short of addressing the diverse analysis needs of scientists, substantial index construction overhead during simulation runs, and inefficient small I/O operations. Ning Yan 0002, Lipeng Wan 0001, Zhichao Cao 0002 |
HPDC | 3 |
| 2025 | HPDR: High-Performance Portable Scientific Data Reduction FrameworkabstractThe rapid growth in scientific data generation is outpacing advancements in computing systems necessary for efficient storage, transfer, and analysis, particularly in the context of exascale computing. With the deployment of first-generation exascale computing systems and next-generation experimental facilities, this gap is widening and necessitates effective data reduction techniques to manage enormous data volumes. Over the past decade, various data reduction methods, including lossless compression, error-controlled lossy compression, and data refactoring, have been developed to accelerate I/O in scientific workflows. Despite significant reductions in data volume, these methods introduce considerable computational overhead, which can become the new bottleneck in data processing. To mitigate this, GPU-accelerated data reduction algorithms have been introduced. However, challenges remain in their integration into exascale workflows, including limited portability across different GPU architectures, substantial memory transfer overhead, and reduced scalability on dense multi-GPU systems. To address these challenges, we propose HPDR, a high-performance and portable data reduction framework. HPDR is designed to enable the execution of state-of-the-art reduction algorithms across diverse processor architectures while reducing memory transfer overhead to 2.3 % of the original, resulting in up to$3.5 \times$faster throughput compared to existing solutions. It also achieves up to 96% of the theoretical speedup in multi-GPU settings. In addition, evaluations on accelerating I/O operations at scale up to 1,024 nodes of the Frontier supercomputer demonstrate that HPDR can achieve up to$103\ \text{TB} / \mathrm{s}$reduction throughput, providing up to$4 \times$acceleration in parallel I/O performance compared to existing data reduction routines. This work highlights the potential of HPDR to significantly enhance data reduction efficiency in exascale computing environments. Jieyang Chen, Qian Gong, Yanliang Li, Xin Liang 0001, Lipeng Wan 0001, Qing Liu 0002, Norbert Podhorszki, Scott Klasky |
IPDPS | 5 |
| 2025 | Eureka: Enabling Fine-Grained Access and Range Queries on Compressed Scientific Data via Data-Index Co-Compression
Ning Yan 0002, Sheng Di, Lipeng Wan 0001 |
Proc. VLDB Endow. | 4 |
| 2024 | Revisiting Erasure Codes: A Configuration PerspectiveabstractErasure coding (EC) plays a crucial role in the fault tolerance of modern distributed storage systems (DSS). Inspired by recent research on storage configuration, we study the configuration sensitivity of EC in real DSS in this paper. We systematically inject faults to trigger EC recovery under various configurations, and measure the impact on recovery time and storage overhead quantitatively. Our results show that configurations may affect the EC recovery time significantly (e.g., up to 426%). More interestingly, theoretically superior codes may perform worse in DSS under certain configurations. Also, there is a system checking period before EC recovery that accounts for 41% to 58% of the overall system recovery time, which has been largely ignored in previous studies. Finally, in terms of storage overhead, EC may introduce 32.3% to 72.0% more write amplification (WA) than the theoretical expectation, and we derive a formula to help estimate WA more precisely. Our work suggests the importance of considering the context of real DSS for EC research, and we hope the methodology and findings can contribute to a firmer footing for EC optimization in practice. Runzhou Han, Tabassum Mahmud, Zeren Yang, Vladislav Esaulov, Lipeng Wan 0001, Yong Chen 0001, Jim Wayda, Matthew Wolf, Mai Zheng |
HotStorage | 6 |
| 2023 | SciLance: Mitigate Load Imbalance for Parallel Scientific Applications in Cloud EnvironmentsabstractElastic cloud computing provides new opportunities for accelerating the process of scientific discovery. However, unlike high-performance computing (HPC) systems that are built and optimized for workloads with intensive inter-node communication demands, the low-latency and high bandwidth communication capability is only enabled on a few very expensive high-end instance types in the cloud, which leads to poor cost-effectiveness. In addition, re-balancing the workload through extra data movement among compute nodes is a common way to mitigate the load imbalance issue in many scientific simulations, which further amplifies the communication pressure and makes it challenging to efficiently use cloud resources. To this end, we propose SciLance, which addresses the workload imbalance challenge by utilizing the heterogeneous and elastic resources offered by cloud platforms. Particularly, instead of moving data excessively among compute instances to balance the workload, SciLance dynamically adjusts the computer instances used for running parallel tasks based on the runtime imbalance identified through profiling. We prototype SciLance and perform extensive evaluation using adaptive mesh refinement (AMR) based scientific applications. The evaluation results demonstrate that SciLance can achieve up to 36.63% better performance with 16.91% lower cost for AMR-based simulation codes. Xinying Wang 0001, Lipeng Wan 0001, Scott Klasky, Dongfang Zhao 0001, Feng Yan 0001 |
CLUSTER | 2 |
| 2023 | Spatiotemporally Adaptive Compression for Scientific Dataset with Feature Preservation - A Case Study on Simulation Data with Extreme Climate Events AnalysisabstractScientific discoveries are increasingly constrained by limited storage space and I/O capacities. For time-series simulations and experiments, their data often need to be decimated over timesteps to accommodate storage and I/O limitations. In this paper, we propose a technique that addresses storage costs while improving post-analysis accuracy through spatiotemporal adaptive, error-controlled lossy compression. We investigate the trade-off between data precision and temporal output rates, revealing that reducing data precision and increasing timestep frequency lead to more accurate analysis outcomes. Additionally, we integrate spatiotemporal feature detection with data compression and demonstrate that performing adaptive error-bounded compression in higher dimensional space enables greater compression ratios, leveraging the error propagation theory of a transformation-based compressor. To evaluate our approach, we conduct experiments using the well-known E3SM climate simulation code and apply our method to compress variables used for cyclone tracking. Our results show a significant reduction in storage size while enhancing the quality of cyclone tracking analysis, both quantitatively and qualitatively, in comparison to the prevalent timestep decimation approach. Compared to three state-of-the-art lossy compressors lacking feature preservation capabilities, our adaptive compression framework improves perfectly matched cases in TC tracking by 26.4-51.3% at medium compression ratios and by 77.3-571.1% at large compression ratios, with a merely 5–11% computational overhead. Qian Gong, Chengzhu Zhang, Xin Liang 0001, Viktor Reshniak, Jieyang Chen, Anand Rangarajan 0001, Sanjay Ranka, Nicolas Vidal 0003, Lipeng Wan 0001, Paul Ullrich, Norbert Podhorszki, Robert Jacob, Scott Klasky |
e-Science | 9 |
| 2023 | RAPIDS: Reconciling Availability, Accuracy, and Performance in Managing Geo-Distributed Scientific DataabstractIn modern science, big data plays an increasingly important role. Many scientific applications, such as running simulations on supercomputers or conducting experiments on advanced instruments, produce huge amount of data at unprecedented speed. Analyzing and understanding such big data is the key for scientists to make scientific breakthroughs. However, data might become unavailable for scientists to access when outages or maintenance of the storage system occur, which severely hinders scientific discovery. To improve the data availability, data duplication and erasure coding (EC) are often used. But as the scientific data gets larger, using these two methods can cause considerable storage and network overhead. Lipeng Wan 0001, Jieyang Chen, Xin Liang 0001, Ana Gainaru, Qian Gong, Qing Liu 0002, Ben Whitney, Joy Arulraj, Zhengchun Liu, Ian T. Foster, Scott Klasky |
HPDC | 1 |
| 2023 | Improving Progressive Retrieval for HPC Scientific Data using Deep Neural NetworkabstractAs the disparity between compute and I/O on high-performance computing systems has continued to widen, it has become increasingly difficult to perform post-hoc data analytics on full-resolution scientific simulation data due to the high I/O cost. Error-bounded data decomposition and progressive data retrieval framework has recently been developed to address such a challenge by performing data decomposition before storage and reading only part of the decomposed data when necessary. However, the performance of the progressive retrieval framework has been suffering from the over-pessimistic error control theory, such that the achieved maximum error of recomposed data is significantly lower than the required error. Therefore, more data than required is fetched for recomposition, incurring additional I/O overhead. In order to tackle this issue, we propose a DNN-based progressive retrieval framework that can better identify the minimum amount of data to be retrieved. Our contributions are as follows: 1) We provide an in-depth investigation of the recently developed progressive retrieval framework; 2) We propose two designs of prediction models (named D-MGARD and E-MGARD) to estimate the amount of retrieved data size based on error bounds. 3) We evaluate our proposed solutions using scientific datasets generated by real-world simulations from two domains. Evaluation results demonstrate the effectiveness of our solution in accurately predicting the amount of retrieval data size, as well as the advantages of our solution over the traditional approach to reducing the I/O overhead. Based on our evaluation, our solution is shown to read significantly less data (5% - 40% with D-MGARD, 20% - 80% with E-MGARD). Jinzhen Wang, Xin Liang 0001, Ben Whitney, Jieyang Chen, Qian Gong, Xubin He, Lipeng Wan 0001, Scott Klasky, Norbert Podhorszki, Qing Liu 0002 |
ICDE | 7 |
| 2023 | Analyzing File Access Patterns on Large-Scale HPC Systems: Opportunities for File PrefetchingabstractThis paper explores the potential opportunities for implementing file prefetching techniques on large-scale high-performance computing (HPC) systems. Specifically, we investigate the file access patterns of various applications across multiple scientific domains using two years' worth of Darshan I/O traces obtained from the Summit supercomputer. We identify recurring trends and patterns which indicate that prefetching can be effectively leveraged to improve data access performance on HPC systems. This study serves as a valuable reference for system architects and developers in the HPC community, providing insights into the opportunities and challenges associated with enabling file prefetching on large-scale HPC systems. Ahmad Maroof Karimi, Arnab Kumar Paul, Jong Choi 0001, Lipeng Wan 0001, Feiyi Wang |
MASCOTS | 4 |
| 2022 | P-ckpt: Coordinated Prioritized CheckpointingabstractGood prediction accuracy and adequate lead time to failure are key to the success of failure-aware Check-point/Restart (C/R) models on current and future large-scale High-Performance Computing (HPC) systems. This paper develops a novel checkpointing technique, called p-ckpt, that aims to maintain the performance efficiency of failure-aware C/R models even when failures are predicted with a small lead time. The p-ckpt technique is developed for HPC systems with multi-level memory systems to prioritize checkpoints from vulnerable nodes (nodes with predicted failure) in the event of failure prediction. It applies coordination among the nodes within an application so that vulnerable nodes' checkpoint data is stored to the Parallel File System (PFS) first by assigning priorities based on the lead time to failure. Vulnerable nodes thus have low-latency access on the critical path to the PFS before any failure happens. Further, we create the hybrid p-ckpt model by integrating Live Migration (LM) because of its cost-effectiveness and to reduce checkpoint frequency. Our hybrid p-ckpt C/R model considers prediction lead time and checkpoint latency to the PFS to decide on a feasible proactive action such as p-ckpt and LM. Simulations of six real-world applications for the Summit supercomputer show a ≈53-65% reduction in overhead due to the hybrid p-ckpt model compared to a ≈31-61% reduction in a state-of-the-art solution. We assess our C/R models against multiple failure distributions and consider lead time variability and failure prediction accuracy. Based on this evaluation and assessment, we discuss the trade-offs of using these models and their impact on application overhead. Subhendu Behera, Lipeng Wan 0001, Frank Mueller 0001, Matthew Wolf, Scott Klasky |
IPDPS | 2 |
| 2022 | Region-adaptive, Error-controlled Scientific Data Compression using Multilevel DecompositionabstractThe increase of computer processing speed is significantly outpacing improvements in network and storage bandwidth, leading to the big data challenge in modern science, where scientific applications can quickly generate much more data than that can be transferred and stored. As a result, big scientific data must be reduced by a few orders of magnitude while the accuracy of the reduced data needs to be guaranteed for further scientific explorations. Moreover, scientists are often interested in some specific spatial/temporal regions in their data, where higher accuracy is required. The locations of the regions requiring high accuracy can sometimes be prescribed based on application knowledge, while other times they must be estimated based on general spatial/temporal variation. In this paper, we develop a novel multilevel approach which allows users to impose region-wise compression error bounds. Our method utilizes the byproduct of a multilevel compressor to detect regions where details are rich and we provide the theoretical underpinning for region-wise error control. With spatially varying precision preservation, our approach can achieve significantly higher compression ratios than single-error bounded compression approaches and control errors in the regions of interest. Qian Gong, Ben Whitney, Chengzhu Zhang, Xin Liang 0001, Anand Rangarajan 0001, Jieyang Chen, Lipeng Wan 0001, Paul Ullrich, Qing Liu 0002, Robert Jacob, Sanjay Ranka, Scott Klasky |
SSDBM | 7 |
| 2022 | MGARD+: Optimizing Multilevel Methods for Error-Bounded Scientific Data ReductionabstractNowadays, data reduction is becoming increasingly important in dealing with the large amounts of scientific data. Existing multilevel compression algorithms offer a promising way to manage scientific data at scale, but may suffer from relatively low performance and reduction quality. In this paper, we propose MGARD+, a multilevel data reduction and refactoring framework drawing on previous multilevel methods, to achieve high-performance data decomposition and high-quality error-bounded lossy compression. Our contributions are four-fold: 1) We propose to leverage a level-wise coefficient quantization method, which uses different error tolerances to quantize the multilevel coefficients. 2) We propose an adaptive decomposition method which treats the multilevel decomposition as a preconditioner and terminates the decomposition process at an appropriate level. 3) We leverage a set of algorithmic optimization strategies to significantly improve the performance of multilevel decomposition/recomposition. 4) We evaluate our proposed method using four real-world scientific datasets and compare with several state-of-the-art lossy compressors. Experiments demonstrate that our optimizations improve the decomposition/recomposition performance of the existing multilevel method by up to$70 \times$, and the proposed compression method can improve compression ratio by up to$2 \times$compared with other state-of-the-art error-bounded lossy compressors under the same level of data distortion. Xin Liang 0001, Ben Whitney, Jieyang Chen, Lipeng Wan 0001, Qing Liu 0002, Dingwen Tao, James Kress, David Pugmire, Matthew Wolf, Norbert Podhorszki, Scott Klasky |
IEEE Trans. Computers | 4 |
| 2022 | Understanding the Impact of Data Staging for Coupled Scientific WorkflowsabstractThe rate of data generated by cutting-edge experimental science facilities and large-scale simulations enabled by current high-performance computing (HPC) systems has continued to grow at a far greater pace than the development of the network and storage capabilities on which these systems rely. To cope with this challenge, scientist are moving toward the creation of autonomous experiments and HPC simulations using machine learning. However, efficiently moving, storing, and processing large amounts of data away from the point of origin presents an incredible challenge. In-memory computing, in situ analysis, data staging, and data streaming are recognized viable alternatives to traditional file-based methods for transferring data between coupled workflows. However, the performance trade-offs and limitations for these methods are not fully understood when used in HPC applications. This article presents a comprehensive performance assessment of the current solutions for data staging when applied to applications that are not necessary I/O intensive which makes them not ideal candidates for these methods. Our study is based on experiments running at scale on Oak Ridge National Laboratory's Summit supercomputer using applications and simulations that cover typical computational motifs and patterns. We investigated the usability and cost/benefit trade-offs of staging algorithms for HPC applications under different scenarios and highlight opportunities for optimizing the dataflow between coupled simulation workflows. Ana Gainaru, Lipeng Wan 0001, Eric Suchyta, Jieyang Chen, Norbert Podhorszki, James Kress, David Pugmire, Scott Klasky |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Improving I/O Performance for Exascale Applications Through Online Data Layout ReorganizationabstractThe applications being developed within the U.S. Exascale Computing Project (ECP) to run on imminent Exascale computers will generate scientific results with unprecedented fidelity and record turn-around time. Many of these codes are based on particle-mesh methods and use advanced algorithms, especially dynamic load-balancing and mesh-refinement, to achieve high performance on Exascale machines. Yet, as such algorithms improve parallel application efficiency, they raise new challenges for I/O logic due to their irregular and dynamic data distributions. Thus, while the enormous data rates of Exascale simulations already challenge existing file system write strategies, the need for efficient read and processing of generated data introduces additional constraints on the data layout strategies that can be used when writing data to secondary storage. We review these I/O challenges and introduce two online data layout reorganization approaches for achieving good tradeoffs between read and write performance. We demonstrate the benefits of using these two approaches for the ECP particle-in-cell simulation WarpX, which serves as a motif for a large class of important Exascale applications. We show that by understanding application I/O patterns and carefully designing data layouts we can increase read performance by more than 80 percent. Lipeng Wan 0001, Axel Huebl, Junmin Gu, Franz Poeschel, Ana Gainaru, Jieyang Chen, Xin Liang 0001, Dmitry Ganyushin, Todd S. Munson, Ian T. Foster, Jean-Luc Vay, Norbert Podhorszki, Kesheng Wu, Scott Klasky |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Accelerating Multigrid-based Hierarchical Scientific Data Refactoring on GPUsabstractRapid growth in scientific data and a widening gap between computational speed and I/O bandwidth make it increasingly infeasible to store and share all data produced by scientific simulations. Instead, we need methods for reducing data volumes: ideally, methods that can scale data volumes adaptively so as to enable negotiation of performance and fidelity tradeoffs in different situations. Multigrid-based hierarchical data representations hold promise as a solution to this problem, allowing for flexible conversion between different fidelities so that, for example, data can be created at high fidelity and then transferred or stored at lower fidelity via logically simple and mathematically sound operations. However, the effective use of such representations has been hindered until now by the relatively high costs of creating, accessing, reducing, and otherwise operating on such representations. We describe here highly optimized data refactoring kernels for GPU accelerators that enable efficient creation and manipulation of data in multigrid-based hierarchical forms. We demonstrate that our optimized design can achieve up to 250 TB/s aggregated data refactoring throughput-83% of theoretical peak-on 1024 nodes of the Summit supercomputer. We showcase our optimized design by applying it to a large-scale scientific visualization workflow and the MGARD lossy compression software. Jieyang Chen, Lipeng Wan 0001, Xin Liang 0001, Ben Whitney, Qing Liu 0002, David Pugmire, Nicholas Thompson, Jong Choi 0001, Matthew Wolf, Todd S. Munson, Ian T. Foster, Scott Klasky |
IPDPS | 2 |
| 2021 | Error-controlled, progressive, and adaptable retrieval of scientific data with multilevel decompositionabstractExtreme-scale simulations and high-resolution instruments have been generating an increasing amount of data, which poses significant challenges to not only data storage during the run, but also post-processing where data will be repeatedly retrieved and analyzed for a long period of time. The challenges in satisfying a wide range of post-hoc analysis needs while minimizing the I/O overhead caused by inappropriate and/or excessive data retrieval should never be left unmanaged. In this paper, we propose a data refactoring, compressing, and retrieval framework capable of 1) fine-grained data refactoring with regard to precision; 2) incrementally retrieving and recomposing the data in terms of various error bounds; and 3) adaptively retrieving data in multi-precision and multi-resolution with respect to different analysis. With the progressive data re-composition and the adaptable retrieval algorithms, our framework significantly reduces the amount of data retrieved when multiple incremental precision are requested and/or the downstream analysis time when coarse resolution is used. Experiments show that the amount of data retrieved under the same progressively requested error bound using our framework is 64% less than that using state-of-the-art single-error-bounded approaches. Parallel experiments with up to 1, 024 cores and ~ 600 GB data in total show that our approach yields 1.36× and 2.52× performance over existing approaches in writing to and reading from persistent storage systems, respectively. Xin Liang 0001, Qian Gong, Jieyang Chen, Ben Whitney, Lipeng Wan 0001, Qing Liu 0002, David Pugmire, Rick Archibald, Norbert Podhorszki, Scott Klasky |
SC | 5 |
| 2020 | Orchestrating Fault Prediction with Live Migration and CheckpointingabstractCheckpoint/Restart (C/R) is widely used to provide fault tolerance on High-Performance Computing (HPC) systems. However, Parallel File System (PFS) overhead and failure uncertainty cause significant application overhead. This paper develops an adaptive multi-level C/R model that incorporates a failure prediction and analysis model, which orchestrates failure prediction, checkpointing, checkpoint frequency, and proactive live migration along with the additional benefit of Burst Buffers (BB). It effectively reduces the overheads due to failures, checkpointing, and recovery. Simulation results for the Summit supercomputer yield a reduction of ~20%-86% in application overhead due to BBs, orchestrated failure prediction, and migration. We also observe a ~29% decrease in checkpoint writes to BBs, which can increase the longevity of the BB storage devices. Subhendu Behera, Lipeng Wan 0001, Frank Mueller 0001, Matthew Wolf, Scott Klasky |
HPDC | 2 |
| 2020 | FTRANS: energy-efficient acceleration of transformers using FPGAabstractIn natural language processing (NLP), the "Transformer" architecture was proposed as the first transduction model replying entirely on self-attention mechanisms without using sequence-aligned recurrent neural networks (RNNs) or convolution, and it achieved significant improvements for sequence to sequence tasks. The introduced intensive computation and storage of these pre-trained language representations has impeded their popularity into computation and memory constrained devices. The field-programmable gate array (FPGA) is widely used to accelerate deep learning algorithms for its high parallelism and low latency. However, the trained models are still too large to accommodate to an FPGA fabric. In this paper, we propose an efficient acceleration framework, Ftrans, for transformer-based large scale language representations. Our framework includes enhanced block-circulant matrix (BCM)-based weight representation to enable model compression on large-scale language representations at the algorithm level with few accuracy degradation, and an acceleration design at the architecture level. Experimental results show that our proposed framework significantly reduce the model size of NLP models by up to 16 times. Our FPGA design achieves 27.07× and 81 × improvement in performance and energy efficiency compared to CPU, and up to 8.80× improvement in energy efficiency compared to GPU. Santosh Pandey 0001, Haowen Fang, Yanjun Lyv, Ji Li 0006, Jieyang Chen, Mimi Xie, Lipeng Wan 0001, Hang Liu 0001, Caiwen Ding |
ISLPED | 8 |
| 2018 | A View from ORNL: Scientific Data Research Opportunities in the Big Data AgeabstractOne of the core issues across computer and computational science today is adapting to, managing, and learning from the influx of "Big Data". In the commercial space, this problem has led to a huge investment in new technologies and capabilities that are well adapted to dealing with the sorts of human-generated logs, videos, texts, and other large-data artifacts that are processed and resulted in an explosion of useful platforms and languages (Hadoop, Spark, Pandas, etc.). However, translating this work from the enterprise space to the computational science and HPC community has proven somewhat difficult, in part because of some of the fundamental differences in type and scale of data and timescales surrounding its generation and use. We describe a forward-looking research and development plan which centers around the concept of making Input/Output (I/O) intelligent for users in the scientific community, whether they are accessing scalable storage or performing in situ workflow tasks. Much of our work is based on our experience with the Adaptable I/O System (ADIOS 1.X), and our next generation version of the software ADIOS 2.X [1]. Scott Klasky, Matthew Wolf, Mark Ainsworth, Chuck Atkins, Jong Choi 0001, Greg Eisenhauer, Berk Geveci, William F. Godoy, Mark Kim, James Kress, Tahsin M. Kurç, Qing Liu 0002, Jeremy Logan, Arthur B. Maccabe, Kshitij Mehta, George Ostrouchov, Manish Parashar, Norbert Podhorszki, David Pugmire, Eric Suchyta, Lipeng Wan 0001 |
ICDCS | 21 |
| 2017 | Extending Skel to Support the Development and Optimization of Next Generation I/O SystemsabstractAs the memory and storage hierarchy get deeper and more complex, it is important to have new benchmarks and evaluation tools that allow us to explore the emerging middleware solutions to use this hierarchy. Skel is a tool aimed at automating and refining this process of studying HPC I/O performance. It works by generating application I/O kernel/benchmarks as determined by a domain-specific model. This paper provides some techniques for extending Skel to address new situations and to answer new research questions. For example, we document use cases as diverse as using Skel to troubleshoot I/O performance issues for remote users, refining an I/O system model, and facilitating the development and testing of a mechanism for runtime monitoring and performance analytics. We also discuss data oriented extensions to Skel to support the study of compression techniques for Exascale scientific data management. Jeremy Logan, Jong Choi 0001, Matthew Wolf, George Ostrouchov, Lipeng Wan 0001, Norbert Podhorszki, William F. Godoy, Scott Klasky, Erich Lohrmann, Greg Eisenhauer, Chad Wood, Kevin A. Huck |
CLUSTER | 5 |
| 2017 | Exacution: Enhancing Scientific Data Management for ExascaleabstractAs we continue toward exascale, scientific data volume is continuing to scale and becoming more burdensome to manage. In this paper, we lay out opportunities to enhance state of the art data management techniques. We emphasize well-principled data compression, and using it to achieve progressive refinement. This can both accelerate I/O and afford the user increased flexibility when she interacts with the data. The formulation naturally maps onto enabling partitioning of the progressively improving-quality representations of a data quantity into different media-type destinations, to keep the highest priority information as close as possible to the computation, and take advantage of deepening memory/storage hierarchies in ways not previously possible. Careful monitoring is requisite to our vision, not only to verify that compression has not eliminated salient features in the data, but also to better understand the performance of massively parallel scientific applications. Increased mathematical rigor would be ideal,to help bring compression on a better-understood theoretical footing, closer to the relevant scientific theory, more aware of constraints imposed by the science, and more tightly error-controlled. Throughout, we highlight pathfinding research we have begun exploring related these topics, and comment toward future work that will be needed. Scott Klasky, Eric Suchyta, Mark Ainsworth, Qing Liu 0002, Ben Whitney, Matthew Wolf, Jong Choi 0001, Ian T. Foster, Mark Kim, Jeremy Logan, Kshitij Mehta, Todd S. Munson, George Ostrouchov, Manish Parashar, Norbert Podhorszki, David Pugmire, Lipeng Wan 0001 |
ICDCS | 17 |
| 2017 | Comprehensive Measurement and Analysis of the User-Perceived I/O Performance in a Production Leadership-Class Storage SystemabstractWith the increase of the scale and intensity of the parallel I/O workloads generated by those scientific applications running on high performance computing facilities, understanding the I/O dynamics, especially the root cause of the I/O performance variability and degradation in HPC environment, have become extremely critical to the HPC community. In this paper, we run extensive I/O measuring tests on a production leadership-class storage system to capture the performance variabilities of large-scale parallel I/O. Analyzing these results and its statistic correlation revealed some valuable insights into the characteristics of the storage system and the root cause of I/O performance variability. Further, we leverage these findings and propose an I/O middleware design refactoring which can improve the performance of the parallel I/O by optimizing the data striping and placement. Our preliminary evaluation results demonstrate the proposed approach can reduce the average per-process write latency by at least 80% and the maximum per-process write latency by at least 20%. Lipeng Wan 0001, Matthew Wolf, Feiyi Wang, Jong Choi 0001, George Ostrouchov, Scott Klasky |
ICDCS | 1 |
| 2017 | Approximate Cardinality Estimation (ACE) in large-scale Internet of Things deployments
Qing Cao 0001, Yunhe Feng, Zheng Lu 0005, Hairong Qi 0001, Leon M. Tolbert, Lipeng Wan 0001, Zhibo Wang 0001, Wenjun Zhou 0001 |
Ad Hoc Networks | 6 |
| 2017 | Optimizing checkpoint data placement with guaranteed burst buffer endurance in large-scale hierarchical storage systems
Lipeng Wan 0001, Qing Cao 0001, Feiyi Wang, Sarp Oral |
J. Parallel Distributed Comput. | 1 |
| 2017 | Optimizing the performance of sensor network programs through estimation-based code profiling
Lipeng Wan 0001, Qing Cao 0001, Wenjun Zhou 0001 |
Pervasive Mob. Comput. | 1 |
| 2015 | Estimation-based profiling for code placement optimization in sensor network programsabstractIn this work, we focus on applying profiling guided code placement to programs running on resource-constrained sensor motes. Specifically, we model the execution of sensor network programs under nondeterministic inputs as discrete-time Markov processes, and propose a novel approach named Code Tomography to estimate parameters of the Markov models that reflect sensor network programs' dynamic execution behavior by only using end-to-end timing information measured at start and end points of each procedure. The parameters estimated by Code Tomography are fed back to compilers to optimize the code placement so that branch misprediction rate can be reduced. Lipeng Wan 0001, Qing Cao 0001, Wenjun Zhou 0001 |
ISPASS | 1 |
| 2015 | A practical approach to reconciling availability, performance, and capacity in provisioning extreme-scale storage systemsabstractThe increasing data demands from high-performance computing applications significantly accelerate the capacity, capability and reliability requirements of storage systems. As systems scale, component failures and repair times increase, significantly impacting data availability. A wide array of decision points must be balanced in designing such systems. Lipeng Wan 0001, Feiyi Wang, Sarp Oral, Devesh Tiwari, Sudharshan S. Vazhkudai, Qing Cao 0001 |
SC | 1 |
| 2014 | Towards approximate spatial queries for large-scale vehicle networksabstractWith advances in vehicle-to-vehicle communication, future vehicles will have access to a communication channel through which messages can be sent and received when two get close to each other. This enabling technology makes it possible for authenticated users to send queries to those vehicles of interest, such as those that are located within a geographic region, over multiple hops for various application goals. However, a naive method that requires flooding the queries to each active vehicle in a region will incur a total communication overhead that is proportional to the size of the area and the density of vehicles. In this paper, we study the problem of spatial queries for vehicle networks by investigating probabilistic methods, where we only try to obtain approximate estimates within desired confidence intervals using only sublinear overheads. We consider this to be particularly useful when spatial query results can be made approximate or not precise, as is the case with many potential applications. The proposed method has been tested on snapshots from real world vehicle network traces. Lipeng Wan 0001, Zhibo Wang 0001, Zheng Lu 0005, Hairong Qi 0001, Wenjun Zhou 0001, Qing Cao 0001 |
SIGSPATIAL/GIS | 1 |
| 2014 | SSD-optimized workload placement with adaptive learning and classification in HPC environmentsabstractIn recent years, non-volatile memory devices such as SSD drives have emerged as a viable storage solution due to their increasing capacity and decreasing cost. Due to the unique capability and capacity requirements in large scale HPC (High Performance Computing) storage environment, a hybrid configuration (SSD and HDD) may represent one of the most available and balanced solutions considering the cost and performance. Under this setting, effective data placement as well as movement with controlled overhead become a pressing challenge. In this paper, we propose an integrated object placement and movement framework and adaptive learning algorithms to address these issues. Specifically, we present a method that shuffle data objects across storage tiers to optimize the data access performance. The method also integrates an adaptive learning algorithm where realtime classification is employed to predict the popularity of data object accesses, so that they can be placed on, or migrate between SSD or HDD drives in the most efficient manner. We discuss preliminary results based on this approach using a simulator we developed to show that the proposed methods can dynamically adapt storage placements and access pattern as workloads evolve to achieve the best system level performance such as throughput. Lipeng Wan 0001, Zheng Lu 0005, Qing Cao 0001, Feiyi Wang, Sarp Oral, Bradley W. Settlemyer |
MSST | 1 |
| 2013 | Towards Instruction Level Record and Replay of Sensor Network ApplicationsabstractDebugging wireless sensor network (WSN) applications has been complicated for multiple reasons, among which the lack of visibility is one of the most challenging. To address this issue, in this paper, we present a systematic approach to record and replay WSN applications at the granularity of instructions. This approach differs from previous ones in that it is purely software based, therefore, no additional hardware component is needed. Our key idea is to combine the static, structural information of the assembly-level code with their dynamic, run-time traces as measured by timestamps and basic block counters, so that we can faithfully infer and replay the actual execution paths of applications at instruction level in a post-mortem manner. The evaluation results show that this approach is feasible despite of the resource constraints of sensor nodes. We also provide two case studies to demonstrate that our instruction level record-and-replay approach can be used to: (1) discover randomness of EEPROM writing time, (2) localize stack smashing bugs in sensor network applications. Lipeng Wan 0001, Qing Cao 0001 |
MASCOTS | 1 |
| 2012 | PhoneCon: Voice-driven SmartPhone Controllable Wireless Sensor NetworksabstractWith the widespread use of wireless sensor networks (WSN), more types of applications are emerging. However, controlling a deployed wireless sensor network after deployment still requires expertise on embedded system programming and administration. In this paper, we study how to exploit the widely available smartphones to reduce these restrictions on users' expertise. In particular, we present a system called PhoneCon, which stands for Voice-driven SmartPhone Controllable Wireless Sensor Networks. Through PhoneCon, users may simply speak to their phones to operate a deployed wireless sensor network, such as getting the status of deployed motes, without being exposed to any of system or programming level details. We have implemented PhoneCon on Android phones for a testbed of MicaZ sensor nodes. Our evaluation results show that PhoneCon can interpret users' intentions well enough through the microphone and perform corresponding actions with acceptable delays. Yanjun Yao, Lipeng Wan 0001, Qing Cao 0001, Rukun Mao |
IPCCC | 2 |