EDBT 2026 Demo / reviewers in the wild / expert
Robert Underwood
dblp:178/6201
· DBLP profile ↗
29ranked-venue papers
5as first author
26since 2021 · last 2026
0000-0002-1464-729XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 4 first-author · 23 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving
Meghana Maghyastha, Robert Underwood, Randal C. Burns, Bogdan Nicolae |
Euro-Par (2) | 2 |
| 2026 | Bridging Information Theory and Practice for Scientific Lossy CompressionabstractError-bounded lossy compressors have been developed for years to reduce the vast volumes of scientific data generated by high-performance computing (HPC) applications and advanced scientific instruments. While these compressors have been effective in mitigating the challenges posed by massive datasets, a significant gap remains in our understanding of the fundamental compressibility limits of scientific data–an issue that critically impacts the sustainable adoption and development of efficient lossy compression techniques in practice. Classical rate-distortion theory, established by Shannon, assumes stationary 1D sources with unconstrained coding–assumptions that do not hold for scientific datasets compressed under the tiling constraints imposed by modern parallel lossy compressors. This paper addresses this gap by developing a novel framework that characterizes compressibility limits for scientific datasets under realistic tiling constraints. The contribution is two-fold. First, we establish a tile-aware, finite-blocklength extension of rate–distortion theory that advances classical 1D asymptotic formulations into a rigorous framework for piecewise 2D Gaussian random fields. To our knowledge, this is the first framework to rigorously characterize lossy compressibility limits for scientific datasets and compressor, moving beyond classical asymptotic 1D source models. Second, we conduct a comprehensive validation of the proposed modeling framework using state-of-the-art error-bounded lossy compressors and diverse real-world HPC datasets, demonstrating that our theory accurately predicts rate-distortion trends and provides actionable insights for compressor design. Sujata Sinha, Sheng Di, Vishwas Rao, Robert Underwood, David Lenz 0002, Zizhe Jian, Zhuoxun Yang, Kai Zhao 0008, Lingjia Liu 0001, Franck Cappello |
HPDC | 4 |
| 2026 | TZ: Achieving High-Ratio Scientific Data Compression on GPUs with Global Data DecompositionabstractAs high-performance computing shifts toward GPU-accelerated exascale systems, the exponential growth of scientific data poses severe challenges to both storage capacity and I/O bandwidth. While current GPU-based lossy compressors attempt to address this by porting CPU algorithms to the device, they rely heavily on block-wise spatial decomposition to fit GPU parallelism. This approach suffers from a fundamental locality barrier: by partitioning data into independent blocks, these methods fail to capture global correlations and fragment the unified data patterns required for effective coding, severely limiting compression ratios. In this paper, we propose TZ, a novel GPU-native error-bounded lossy compressor that breaks this ceiling by adopting global Tucker decomposition. By prioritizing global spectral energy compaction over local approximation, TZ naturally maximizes the compression potential for scientific datasets. To render this computationally intensive approach practical for high-throughput GPU workflows, we introduce a highly optimized adaptive randomized SVD engine. This design allows TZ to achieve the superior compression ratios of global spectral decomposition while maintaining competitive execution speeds. Furthermore, the global processing nature of TZ enables a unified quantization and coding scheme that eliminates block artifacts and metadata overhead. Evaluation on production-scale scientific datasets demonstrates that TZ achieves approximately 10 × higher compression ratios than state-of-the-art GPU compressors under the same error bound, while maintaining competitive, high-throughput performance. Zhuoxun Yang, Amit N. Subrahmanya, Vishwas Rao, Sheng Di, Robert Underwood, Jinyang Liu 0003, Franck Cappello, Kai Zhao 0008 |
HPDC | 6 |
| 2026 | OPAL: On-demand Progressive Accelerated Scientific Lossy Compression
Zhuoxun Yang, Robert Underwood, Sheng Di, Daoce Wang, Jinyang Liu 0003, Jiajun Huang 0001, Franck Cappello, Kai Zhao 0008 |
HPDC | 4 |
| 2026 | pMSz: A Distributed Parallel Algorithm for Correcting Extrema and Morse-Smale Segmentations in Lossy Compression
Yuxiao Li 0002, Mingze Xia, Xin Liang 0001, Bei Wang 0001, Robert Underwood, Sheng Di, Hemant Sharma, Dishant Beniwal, Franck Cappello, Hanqi Guo 0001 |
IPDPS | 5 |
| 2026 | FFCz: Fast Fourier Correction for Spectrum-Preserving Lossy Compression of Scientific Data
Congrong Ren, Robert Underwood, Sheng Di, Emrecan Kutay, Zarija Lukic, Aylin Yener, Franck Cappello, Hanqi Guo 0001 |
IPDPS | 2 |
| 2026 | Designing Domain-Specific Compilers for Lossy Compression: A Case Study on Wafer-Scale Engine
Shihui Song, Robert Underwood, Sheng Di, Peng Jiang 0004, Franck Cappello |
IPDPS | 2 |
| 2025 | Sensitivity and Impacts on Parallel Compression of Prediction of Lossy Compression Ratios for Scientific DataabstractCompression ratio estimation is an important optimization of I/O workflows processing terabytes of data. Applications such as compression auto-tuning or lossy compressor selection require a high-throughput, accurate estimation to be fast. Prior works that utilize sampling are fast but inaccurate, while approaches that do not use sampling are accurate but slow. We present a novel, lightweight sampling technique that excels in speed and accuracy, by leveraging both statistical and spatial properties of dataset samples. Through a comprehensive sensitivity analysis we show that there is no one-size-fits-all sampling strategy because of differences in compression principles, and accounting for these differences has a direct impact on overall prediction accuracy in the methods that use these predictions. Our estimation technique demonstrates superior ($1.36 \times$to$11 \times$) prediction accuracy compared to existing sampling methods while significantly improving data throughput by$3.46 \times$compared to non-sampling approaches. Alexandra Poulos, Robert Underwood, Jon Calhoun 0001, Sheng Di, Franck Cappello |
IPDPS | 2 |
| 2025 | A Memory-Efficient and Computation-Balanced Lossy Compressor on Wafer-Scale EngineabstractCerebras system has demonstrated immense potential across various scientific domains. However, modern scientific simulations frequently generate vast volumes of data in a short time, leading to bottlenecks in runtime performance and memory footprint. While an ultra-fast error-bounded lossy compressor can mitigate such limitations with high compression ratios and guaranteed data quality, deploying it into Cerebras dataflow architecture poses significant difficulties. Specifically, Cerebras faces memory challenges, such as the absence of shared memory and limited local memory, alongside computational challenges, including specialized parallelism and sensitivity to imbalanced workloads. In this work, we propose CERESZII, an error-bounded lossy compressor that computes within Cerebras system. CereSZ-II addresses these challenges with a carefully optimized four-stage compression workflow, consisting of Pre-quantization, Lightweight Prediction, Fixed-size Huffman Encoding, and Spatial-aware Offset Computation, ensuring both memory efficiency and computational balance. Evaluation of several real-world scientific datasets shows that CERESZ-II achieves over 800 GB/s throughput, delivering high compression ratios and reliable reconstructed data quality. Shihui Song, Robert Underwood, Sheng Di, Yafan Huang, Peng Jiang 0004, Franck Cappello |
IPDPS | 2 |
| 2025 | To Compress or Not to Compress: Energy Trade-Offs and Benefits of Lossy Compressed I/OabstractModern scientific simulations generate massive volumes of data, creating significant challenges for I/O and storage systems. Error-bounded lossy compression (EBLC) offers a solution by reducing data set sizes while preserving data quality within user-specified limits. This study provides the first comprehensive energy characterization of state-of-the-art EBLC algorithms-SZ2, SZ3, ZFP, QoZ, and SZx-across various scientific data sets, CPU generations, and parallel/serial modes. We analyze the energy consumption patterns of compression and decompression operations, as well as the energy trade-offs in data I/O scenarios. Our work demonstrates the relationships between compression ratios, runtime, energy efficiency, and data quality, highlighting the importance of considering compressors and error bounds for specific use cases. We demonstrate that EBLC can significantly reduce I/O energy consumption, with savings of up to two orders of magnitude compared to uncompressed I/O for large data sets. In multi-node HPC environments, we observe energy reductions of approximately 25 % when using EBLC. We also show that EBLC can achieve compression ratios of$10-100 \times$, potentially reducing storage device requirements by nearly two orders of magnitude. This work provides a framework for system operators and computational scientists to make informed decisions about implementing EBLC for energy-efficient data management in HPC environments. Grant Wilkins, Sheng Di, Jon Calhoun 0001, Robert Underwood, Franck Cappello |
IPDPS | 4 |
| 2025 | What to Support When You're Compressing: The State of Practice Gaps and Opportunities for Scientific Data CompressionabstractOver the last nearly 20 years, lossy compression has become an essential aspect of HPC applications’ data pipelines, allowing them to overcome limitations in storage capacity and bandwidth and, in some cases, increase computational throughput and capacity. However, with the adoption of lossy compression comes the requirement to assess and control the impact lossy compression has on scientific outcomes. In this work, we take a major step forward in describing the state of practice and by characterizing workloads. We examine applications’ needs and compressors’ capabilities across 9 different supercomputing application domains. We present 24 takeaways that provide best practices for applications, operational impacts for facilities achieving compressed data, and gaps in application needs not addressed by production compressors that point towards opportunities for future compression research. Franck Cappello, Robert Underwood, Yuri Alexeev, Allison H. Baker, Ebru Bozdag, Martin Burtscher, Kyle Chard, Sheng Di, Kyle Gerard Felker, Paul Christopher O'Grady, Hanqi Guo 0001, Yafan Huang, Peng Jiang 0004, Sian Jin, Petter Johansson, Shaomeng Li, Xin Liang 0001, Erik Lindahl, Peter Lindstrom 0001, Zarija Lukic, Magnus Lundborg, Danylo Lykov, Masaru Nagaso, Kento Sato, Amarjit Singh, Seung Woo Son 0001, Shihui Song, William Tang 0002, Dingwen Tao, Jiannan Tian, Kazutomo Yoshii, Kai Zhao 0008 |
SC | 2 |
| 2025 | lsCOMP: Efficient Light Source CompressionabstractLight source facilities, which generate X-rays for probing microstructures and dynamic processes, produce intense data streams, reaching up to 250 GB/s and projected to exceed 1 TB/s by the end of this decade. Managing such massive data poses critical challenges due to limited local processing capacity and bandwidth constraints when offloading data to HPC systems. To address these challenges, we propose lsCOMP, a GPU compressor that operates within a single kernel. lsCOMP supports both lossless and configurable lossy compression, ensuring high compression ratios and preserved data quality across diverse light source applications. On a single NVIDIA A100 GPU, lsCOMP achieves compression throughputs of 380.89 to 509.21 GB/s in lossless mode, delivering up to 20 times higher performance than industry-leading GPU compressors while achieving superior compression ratios. In lossy modes, lsCOMP further improves throughput and ratios significantly. Additionally, lsCOMP demonstrates versatile performance across various integer datasets and supports TB/s-level random access throughput. Yafan Huang, Sheng Di, Robert Underwood, Peco Myint, Miaoqi Chu, Guanpeng Li, Nicholas Schwarz, Franck Cappello |
SC | 3 |
| 2025 | Multifacets of lossy compression for scientific data in the Joint-Laboratory of Extreme Scale Computing
Franck Cappello, Mario C. Acosta, Emmanuel Agullo, Hartwig Anzt, Jon Calhoun 0001, Sheng Di, Luc Giraud, Thomas Grützmacher, Sian Jin, Kentaro Sano, Kento Sato, Amarjit Singh, Dingwen Tao, Jiannan Tian, Tomohiro Ueno, Robert Underwood, Frédéric Vivien, Xavier Yepes, Kazutomo Yoshii, Boyuan Zhang 0002 |
Future Gener. Comput. Syst. | 16 |
| 2025 | LCP: Enhancing Scientific Data Management with Lossy Compression for ParticlesabstractMany scientific applications opt for particles instead of meshes as their basic primitives to model complex systems composed of billions of discrete entities. Such applications span a diverse array of scientific domains, including molecular dynamics, cosmology, computational fluid dynamics, and geology. The scale of the particles in those scientific applications increases substantially thanks to the ever-increasing computational power in high-performance computing (HPC) platforms. However, the actual gains from such increases are often undercut by obstacles in data management systems related to data storage, transfer, and processing. Lossy compression has been widely recognized as a promising solution to enhance scientific data management systems regarding such challenges, although most existing compression solutions are tailored for Cartesian grids and thus have sub-optimal results on discrete particle data. In this paper, we introduce LCP, an innovative lossy compressor designed for particle datasets, offering superior compression quality and higher speed than existing compression solutions. Specifically, our contribution is threefold. (1) We propose LCP-S, an error-bound aware block-wise spatial compressor to efficiently reduce particle data size while satisfying the pre-defined error criteria. This approach is universally applicable to particle data across various domains, eliminating the need for reliance on specific application domain characteristics. (2) We develop LCP, a hybrid compression solution for multi-frame particle data, featuring dynamic method selection and parameter optimization. It aims to maximize compression effectiveness while preserving data quality as much as possible by utilizing both spatial and temporal domains. (3) We evaluate our solution alongside eight state-of-the-art alternatives on eight real-world particle datasets from seven distinct domains. The results demonstrate that our solution achieves up to 104% improvement in compression ratios and up to 593% increase in speed compared to the second-best option, under the same error criteria. Congrong Ren, Sheng Di, Jinyang Liu 0003, Jiajun Huang 0001, Robert Underwood, Pascal Grosset, Dingwen Tao, Xin Liang 0001, Hanqi Guo 0001, Franck Cappello, Kai Zhao 0008 |
Proc. ACM Manag. Data | 7 |
| 2024 | DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language ModelsabstractLLMs have seen rapid adoption in all domains. They need to be trained on high-end high-performance computing (HPC) infrastructures and ingest massive amounts of input data. Unsurprisingly, at such a large scale, unexpected events (e.g., failures of components, instability of the software, undesirable learning patterns, etc.), are frequent and typically impact the training in a negative fashion. Thus, LLMs need to be checkpointed frequently so that they can be rolled back to a stable state and subsequently fine-tuned. However, given the large sizes of LLMs, a straightforward checkpointing solution that directly writes the model parameters and optimizer state to persistent storage (e.g., a parallel file system), incurs significant I/O overheads. To address this challenge, in this paper we study how to reduce the I/O overheads for enabling fast and scalable checkpointing for LLMs that can be applied at high frequency (up to the granularity of individual iterations) without significant impact on the training process. Specifically, we introduce a lazy asynchronous multi-level approach that takes advantage of the fact that the tensors making up the model and optimizer state shards remain immutable for extended periods of time, which makes it possible to copy their content in the background with minimal interference during the training process. We evaluate our approach at scales of up to 180 GPUs using different model sizes, parallelism settings, and checkpointing frequencies. The results show up to 48× faster checkpointing and 2.2× faster end-to-end training runtime compared with the state-of-art checkpointing approaches. Avinash Maurya, Robert Underwood, M. Mustafa Rafique, Franck Cappello, Bogdan Nicolae |
HPDC | 2 |
| 2024 | EvoStore: Towards Scalable Storage of Evolving Learning ModelsabstractDeep Learning (DL) has seen rapid adoption in all domains. Since training DL models is expensive, both in terms of time and resources, application workflows that make use of DL increasingly need to operate with a large number of derived learning models, which are obtained through transfer learning and fine-tuning. At scale, thousands of such derived DL models are accessed concurrently by a large number of processes. In this context, an important question is how to design and develop specialized DL model repositories that remain scalable under concurrent access, while addressing key challenges: how to query the DL model architectures for specific patterns? How to load/store a subset of layers/tensors from a DL model? How to efficiently share unmodified layers/tensors between DL models derived from each other through transfer learning? How to maintain provenance and answer ancestry queries? State of art leaves a gap regarding these challenges. To fill this gap, we introduce EvoStore, a distributed DL model repository with scalable data and metadata support to store and access derived DL models efficiently. Large-scale experiments on hundreds of GPUs show significant benefits over state-of-art with respect to I/O and metadata performance, as well as storage space utilization. Robert Underwood, Meghana Madhyastha, Randal C. Burns, Bogdan Nicolae |
HPDC | 1 |
| 2024 | FedSZ: Leveraging Error-Bounded Lossy Compression for Federated Learning CommunicationsabstractWith the promise of federated learning (FL) to allow for geographically-distributed and highly personalized services, the efficient exchange of model updates between clients and servers becomes crucial. FL, though decentralized, often faces communication bottlenecks, especially in resource-constrained scenarios. Existing data compression techniques like gradient sparsification, quantization, and pruning offer some solutions, but may compromise model performance or necessitate expensive retraining. In this paper, we introduce FedSZ, a specialized lossy-compression algorithm designed to minimize the size of client model updates in FL. FedSZ incorporates a comprehensive compression pipeline featuring data partitioning, lossy and lossless compression of model parameters and metadata, and serialization. We evaluate FedSZ using a suite of error-bounded lossy compressors, ultimately finding SZ2 to be the most effective across various model architectures and datasets including AlexNet, MobileNetV2, ResNet50, CIFAR-10, Caltech101, and Fashion-MNIST. Our study reveals that a relative error bound${10}^{-2}$achieves an optimal tradeoff, compressing model states between 5.55-12.61× while maintaining inference accuracy within < 0.5 % of uncompressed results. Additionally, the runtime overhead of FedSZ is < 4.7% or between of the wall-clock communication-round time, a worthwhile trade-off for reducing network transfer times by an order of magnitude for networks bandwidths < 350Mbps. Intriguingly, we also find that the error introduced by FedSZ could potentially serve as a source of differentially private noise, opening up new avenues for privacy-preserving FL. Grant Wilkins, Sheng Di, Jon Calhoun 0001, Zilinghan Li, Kibaek Kim, Robert Underwood, Richard Mortier, Franck Cappello |
ICDCS | 6 |
| 2024 | CliZ: Optimizing Lossy Compression for Climate Datasets with Adaptive Fine-tuned Data PredictionabstractBenefiting from the cutting-edge supercomputers that support extremely large-scale scientific simulations, climate research has advanced significantly over the past decades. However, new critical challenges have arisen regarding efficiently storing and transferring large-scale climate data among distributed repositories and databases for post hoc analysis. In this paper, we develop CliZ, an efficient online error-controlled lossy compression method with optimized data prediction and encoding methods for climate datasets across various climate models. On the one hand, we explored how to take advantage of particular properties of the climate datasets (such as mask-map information, dimension permutation/fusion, and data periodicity pattern) to improve the data prediction accuracy. On the other hand, CliZ features a novel multi-Huffman encoding method, which can significantly improve the encoding efficiency. Therefore significantly improving compression ratios. We evaluated CliZ versus many other state-of-the-art error-controlled lossy compressors (including SZ3, ZFP, SPERR, and QoZ) based on multiple real-world climate datasets with different models. Experiments show that CliZ outperforms the second-best compressor (SZ3, SPERR, or QoZ1.1) on climate datasets by 20%-200% in compression ratio. CliZ can significantly reduce the data transfer cost between the two remote Globus endpoints by 32%-38%. Zizhe Jian, Sheng Di, Jinyang Liu 0003, Kai Zhao 0008, Xin Liang 0001, Haiying Xu, Robert Underwood, Shixun Wu, Jiajun Huang 0001, Zizhong Chen, Franck Cappello |
IPDPS | 7 |
| 2024 | cuSZ-i: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level InterpolationabstractError-bounded lossy compression is a critical technique for significantly reducing scientific data volumes. Compared to CPU-based compressors, GPU-based compressors exhibit substantially higher throughputs, fitting better for today’s HPC applications. However, the critical limitations of existing GPU-based compressors are their low compression ratios and qualities, severely restricting their applicability. To overcome these, we introduce a new GPU-based error-bounded scientific lossy compressor named CUSZ-i, with the following contributions: (1) A novel GPU-optimized interpolation-based prediction method significantly improves the compression ratio and decompression data quality. (2) The Huffman encoding module in CUSZ-i is optimized for better efficiency. (3) CUSZ-i is the first to integrate the NVIDIA Bitcomp-lossless as an additional compression-ratio-enhancing module. Evaluations show that CUSZ-i significantly outperforms other latest GPU-based lossy compressors in compression ratio under the same error bound (hence, the desired quality), showcasing a 476% advantage over the second-best. This leads to CUSZ-i’s optimized performance in several real-world use cases. Jinyang Liu 0003, Jiannan Tian, Shixun Wu, Sheng Di, Boyuan Zhang 0002, Robert Underwood, Yafan Huang, Jiajun Huang 0001, Kai Zhao 0008, Guanpeng Li, Dingwen Tao, Zizhong Chen, Franck Cappello |
SC | 6 |
| 2023 | A Lightweight, Effective Compressibility Estimation Method for Error-bounded Lossy CompressionabstractError-bounded lossy compression turns more and more important for the data-moving intensive applications to deal with big datasets efficiently in HPC environments, which often requires knowing the compressibility of the datasets before performing the compression. However, the off-the-shelf state-of-the-art lossy compressors are often driven by error bounds, so the compression ratios cannot be forecasted until the completion of the compression operation. In this paper, we propose a lightweight, robust, easy-to-train model that estimates the compressibility of datasets for different lossy compressors accurately. Our approach combines novel predictors that measure various notions of spatial correlation and smoothness exploited by lossy compressors that are implemented efficiently on the GPU in a framework and that uses mixture model regression to improve robustness with conformal prediction to provide bounds on the estimates. We then use these models with a detailed analysis of speedup to understand the tradeoffs between high speed, consistent speed, and accuracy of the methods on real applications. We evaluate our approach in the context of 3 key applications where compression ratio estimation is highly required. Arkaprabha Ganguli, Robert Underwood, Julie Bessac, David Krasowska, Jon Calhoun 0001, Sheng Di, Franck Cappello |
CLUSTER | 2 |
| 2023 | Understanding Patterns of Deep Learning Model Evolution in Network Architecture SearchabstractNetwork Architecture Search and specifically Regularized Evolution is a common way to refine the structure of a deep learning model. However, little is known about how models empirically evolve over time which has design implications for designing caching policies, refining the search algorithm for particular applications, and other important use cases. In this work, we algorithmically analyze and quantitatively characterize the patterns of model evolution for a set of models from the Candle project and the Nasbench-201 search space. We show how the evolution of the model structure is influenced by the regularized evolution algorithm. We describe how evolutionary patterns appear in distributed settings and opportunities for caching and improved scheduling. Lastly, we describe the conditions that affect when particular model architectures rise and fall in popularity based on their frequency of acting as a donor in a sliding window. Robert Underwood, Meghana Madhyastha, Randal C. Burns, Bogdan Nicolae |
HiPC | 1 |
| 2023 | A Feature-Driven Fixed-Ratio Lossy Compression Framework for Real-World Scientific DatasetsabstractToday’s scientific applications and advanced instruments are producing extremely large volumes of data everyday, so that error-controlled lossy compression has become a critical technique to the scientific data storage and management. Existing lossy scientific data compressors, however, are designed mainly based on error-control driven mechanism, which cannot be efficiently applied in the fixed-ratio use-case, where a desired compression ratio needs to be reached because of the restricted data processing/management resources such as limited memory/storage capacity and network bandwidth. To address this gap, we propose a low-cost compressor-agnostic feature-driven fixed-ratio lossy compression framework (FXRZ). The key contributions are three-fold. (1) We perform an in-depth analysis of the correlation between diverse data features and compression ratios based on a wide range of application datasets, which is a fundamental work for our framework. (2) We propose a series of optimization strategies that can enable the framework to reach a fairly high accuracy in identifying the expected error configuration with very low computational cost. (3) We comprehensively evaluate our framework using 4 state-of-the-art error-controlled lossy compressors on 10 different snapshots and simulation configuration-based real-world scientific datasets from 4 different applications across different domains. Our experiment shows that FXRZ outperforms the state-of-the-art related work by 108×. The experiments with 4,096 cores on a supercomputer show a performance gain of 1.18∼8.71× than the related work in overall parallel data dumping. Md Hasanur Rahman 0001, Sheng Di, Kai Zhao 0008, Robert Underwood, Guanpeng Li, Franck Cappello |
ICDE | 4 |
| 2023 | DStore: A Lightweight Scalable Learning Model Repository with Fine-Grain Tensor-Level AccessabstractThe ability to share and reuse deep learning (DL) models is a key driver that facilitates the rapid adoption of artificial intelligence (AI) in both industrial and scientific applications. However, state-of-the-art approaches to store and access DL models efficiently at scale lag behind. Most often, DL models are serialized by using various formats (e.g., HDF5, SavedModel) and stored as files on POSIX file systems. While simple and portable, such an approach exhibits high serialization and I/O overheads, especially under concurrency. Additionally, the emergence of advanced AI techniques (transfer learning, sensitivity analysis, explainability, etc.) introduces the need for fine-grained access to tensors to facilitate the extraction and reuse of individual or subsets of tensors. Such patterns are underserved by state-of-the-art approaches. Requiring tensors to be read in bulk incurs suboptimal performance, scales poorly, and/or overutilizes network bandwidth. In this paper we propose a lightweight, distributed, RDMA-enabled learning model repository that addresses these challenges. Specifically we introduce several ideas: compact architecture graph representation with stable hashing and client-side metadata caching, scalable load balancing on multiple providers, RDMA-optimized data staging, and direct access to raw tensor data. We evaluate our proposal in extensive experiments that involve different access patterns using learning models of diverse shapes and sizes. Our evaluations show a significant improvement (between 2 and 30× over a variety of state-of-the-art model storage approaches while scaling to half the Cooley cluster at the Argonne Leadership Computing Facility. Meghana Madhyastha, Robert Underwood, Randal C. Burns, Bogdan Nicolae |
ICS | 2 |
| 2023 | SZ3: A Modular Framework for Composing Prediction-Based Error-Bounded Lossy CompressorsabstractToday's scientific simulations require a significant reduction of data volume because of extremely large amounts of data they produce and the limited I/O bandwidth and storage space. Error-bounded lossy compression has been considered one of the most effective solutions to the above problem. In practice, however, the best-fit compression method often needs to be customized or optimized in particular because of diverse characteristics in different datasets and various user requirements on the compression quality and performance. In this paper, we address this issue with a novel modular, composable compression framework named SZ3. Our contributions are four-folds. (1) We develop SZ3 which features an innovative modular abstraction for the prediction-based compression framework, such that compression modules can be plugged in easily to create new compressors based on characteristics of data and user requirements. (2) We create a new compression pipeline by SZ3 for GAMESS data, which significantly improves the compression ratios over state-of-the-art compressors. (3) We develop an adaptive compression pipeline by SZ3 for APS data with minimal efforts, which leads to the best rate-distortion among all existing error-bounded lossy compressors for any bit-rate. (4) We compare the sustainability of SZ3 with leading error-bounded prediction-based compressors, and then demonstrate the necessity of diverse pipelines by integrating and evaluating several compression pipelines on diverse scientific datasets from multiple disciplines. Experiments show that SZ3 incurs very limited overhead in compressor integration and our customized compression pipelines lead to up to 20% improvement in compression ratios under the same data distortion, when compared with the best existing approach. Xin Liang 0001, Kai Zhao 0008, Sheng Di, Sihuan Li, Robert Underwood, Ali Murat Gok, Jiannan Tian, Junjing Deng, Jon Calhoun 0001, Dingwen Tao, Zizhong Chen, Franck Cappello |
IEEE Trans. Big Data | 5 |
| 2022 | OptZConfig: Efficient Parallel Optimization of Lossy Compression ConfigurationabstractLossless compressors have very low compression ratios that do not meet the needs of today’s large-scale scientific applications that produce vast volumes of data. Error-bounded lossy compression (EBLC) is considered a critical technique for the success of scientific research. Although EBLC allows users to set an error bound for the compression, users have been unable to specify the requirements on the compression quality, limiting practical use. Our contributions are: (1) We formulate the problem of configuring EBLC to preserve a user-defined metric as an optimization problem. This allows many classes of new metrics to be preserved, which improves over current practices. (2) We present a framework, OptZConfig, that can adapt to improvements in the search algorithm, compressor, and metrics with minimal changes, enabling future advancements in this area. (3) We demonstrate the advantages of our approach against the leading methods to configure compressors to preserve specific metrics. Our approach improves compression ratios against a specialized compressor by up to$3\times$, has a 56× speedup over FRaZ, 1000× speedup over MGARD-QOI post tuning, and 110× speedup over systematic approaches which had not been bounded by compressors before. Robert Underwood, Jon Calhoun 0001, Sheng Di, Amy W. Apon, Franck Cappello |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | ARC: An Automated Approach to Resiliency for Lossy Compressed Data via Error Correcting CodesabstractProgress in high-performance computing (HPC) systems has led to complex applications that stress the I/O subsystem by creating vast amounts of data. Lossy compression reduces data size considerably, but a single error renders lossy compressed data unusable. This sensitivity stems from the high information content per bit in compressed data and is a critical issue as soft errors that cause bit-flips have become increasingly commonplace in HPC systems. While many works have improved lossy compressor performance, few have sought to address this critical weakness. Dakota Fulp, Alexandra Poulos, Robert Underwood, Jon Calhoun 0001 |
HPDC | 3 |
| 2020 | cuSZ: An Efficient GPU-Based Error-Bounded Lossy Compression Framework for Scientific DataabstractError-bounded lossy compression is a state-of-the-art data reduction technique for HPC applications because it not only significantly reduces storage overhead but also can retain high fidelity for postanalysis. Because supercomputers and HPC applications are becoming heterogeneous using accelerator-based architectures, in particular GPUs, several development teams have recently released GPU versions of their lossy compressors. However, existing state-of-the-art GPU-based lossy compressors suffer from either low compression and decompression throughput or low compression quality. In this paper, we present an optimized GPU version, cuSZ, for one of the best error-bounded lossy compressors-SZ. To the best of our knowledge, cuSZ is the first error-bounded lossy compressor on GPUs for scientific data. Our contributions are fourfold. (1) We propose a dual-quantization scheme to entirely remove the data dependency in the prediction step of SZ such that this step can be performed very efficiently on GPUs. (2) We develop an efficient customized Huffman coding for the SZ compressor on GPUs. (3) We implement cuSZ using CUDA and optimize its performance by improving the utilization of GPU memory bandwidth. (4) We evaluate our cuSZ on five real-world HPC application datasets from the Scientific Data Reduction Benchmarks and compare it with other state-of-the-art methods on both CPUs and GPUs. Experiments show that our cuSZ improves SZ's compression throughput by up to 370.1x and 13.1x, respectively, over the production version running on single and multiple CPU cores, respectively, while getting the same quality of reconstructed data. It also improves the compression ratio by up to 3.48x on the tested data compared with another state-of-the-art GPU supported lossy compressor. Jiannan Tian, Sheng Di, Kai Zhao 0008, Cody Rivera, Megan Hickman Fulp, Robert Underwood, Sian Jin, Xin Liang 0001, Jon Calhoun 0001, Dingwen Tao, Franck Cappello |
PACT | 6 |
| 2020 | FRaZ: A Generic High-Fidelity Fixed-Ratio Lossy Compression Framework for Scientific Floating-point DataabstractWith ever-increasing volumes of scientific floating-point data being produced by high-performance computing applications, significantly reducing scientific floating-point data size is critical, and error-controlled lossy compressors have been developed for years. None of the existing scientific floating-point lossy data compressors, however, support effective fixed-ratio lossy compression. Yet fixed-ratio lossy compression for scientific floating-point data not only compresses to the requested ratio but also respects a user-specified error bound with higher fidelity. In this paper, we present FRaZ: a generic fixed-ratio lossy compression framework respecting user-specified error constraints. The contribution is twofold. (1) We develop an efficient iterative approach to accurately determine the appropriate error settings for different lossy compressors based on target compression ratios. (2) We perform a thorough performance and accuracy evaluation for our proposed fixed-ratio compression framework with multiple state-of-the-art error-controlled lossy compressors, using several real-world scientific floating-point datasets from different domains. Experiments show that FRaZ effectively identifies the optimum error setting in the entire error setting space of any given lossy compressor. While fixed-ratio lossy compression is slower than fixed-error compression, it provides an important new lossy compression technique for users of very large scientific floating-point datasets. Robert Underwood, Sheng Di, Jon Calhoun 0001, Franck Cappello |
IPDPS | 1 |
| 2018 | Measuring Network Latency Variation Impacts to High Performance Computing Application PerformanceabstractIn this paper, we study the impacts of latency variation versus latency mean on application runtime, library performance, and packet delivery. Our contributions include the design and implementation of a network latency injector that is suitable for most QLogic and Mellanox InfiniBand cards. We fit statistical distributions of latency mean and variation to varying levels of network contention for a range of parallel application workloads. We use the statistical distributions to characterize the latency variation impacts to application degradation. The level of application degradation caused by variation in network latency depends on application characteristics, and can be significant. Observed degradation varies from no degradation for applications without communicating processes to 3.5 times slower for communication-intensive parallel applications. We support our results with statistical analysis of our experimental observations. For communication-intensive high performance computing applications, we show statistically significant evidence that changes in performance are more highly correlated with changes of variation in network latency than with changes of mean network latency alone. Robert Underwood, Amy W. Apon |
ICPE | 1 |