EDBT 2026 Demo / reviewers in the wild / expert
Huizhang Luo
dblp:159/6651
· DBLP profile ↗
28ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0003-2392-0267ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 8 first-author · 16 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Swift: High-Performance Sparse-Dense Matrix Multiplication on GPUsabstractSparse-Dense Matrix Multiplication (SpMM) on GPUs has gained significant attention because of its importance in modern applications and the increasing computing power of GPUs in the last decade. Previous SpMM studies have focused on the importance of storage format and load balance for the overall performance of SpMM on GPUs. However, very little attention has been paid to the efficacy of coalesced memory access in improving the efficiency of data loading, which incurs a notable overhead that amounts to an average of more than 32% of the overall performance, according to our experimental observation. Existing state-of-the-art (SOTA) solutions fail to adequately support coalesced memory access of both sparse and dense matrices between the global memory and threads on GPUs. In this paper, we propose an efficient algorithm called Swift.11Swift is available at https://github.com/MinttHu/Swift.git that speeds up the loading of both sparse and dense matrices of SpMM on modern GPUs. Leveraging coalesced memory access, Swift achieves high loading efficiency by sorting both the columns of the sparse matrix and elements of the dense matrix based on the number of non-zero elements and balancing the load by handling the regular and irregular parts differently and judiciously. Swift takes the Compressed Sparse Column format as an implementation case study to prove the concept and gain insights. We conduct a comprehensive comparison of Swift with four SOTA solutions: ASpT, cuSPARSE, RoDe, and Sputnik, using the full SuiteSparse Matrix Collection as the workload. The experimental results on RTX 4080s, RTX 3090Ti, A100, and V100 demonstrate that our method outperforms the baselines significantly. Jinyu Hu, Huizhang Luo, Hong Jiang 0001, Marc Casas, Kenli Li 0001, Chubo Liu |
HPCA | 2 |
| 2026 | AEIS: A New Energy Efficiency Improvement Scheme for MLC STT-MRAMabstractSpin Transfer Torque-Magnetic Random Access Memory (STT-MRAM), as a new non-volatile memory technology with lower leakage power and higher density, is widely considered to be a new generation of memory technology that may replace SRAM in the cache. STT-MRAM is divided into Single-Level Cell (SLC) STT-MRAM and Multi-Level Cell (MLC) STT-MRAM. Compared with SLC STT-MRAM, MLC STT-MRAM has further improved its storage density. However, MLC STT-MRAM has a high energy consumption and write latency due to its unique two-step state transitions (TTs) issue. State-of-the-art approaches mitigate this issue by eliminating TTs with expansion coding methods. Unfortunately, they focus more on eliminating TTs and have limited improvement in reducing energy consumption. To this end, we propose a new scheme, AEIS, which further reduces energy consumption while eliminating TTs. Our work begins with exploring the general rules of (M,N)-based expansion coding methods that eliminate TTs. Based on the discovered rules, the minimum energy coding method is found. To further improve energy efficiency, we segment the cache lines according to the data pattern. We only apply the expansion coding to those flipping segments to reduce the expansion coding overhead. The evaluation results show that AEIS can eliminate TTs in MLC STT-MRAM, reduce energy consumption by 28.5%, and increase the lifetime by 24.8%, while the total number of bits used for the cache only increases by 5.7%. Huizhang Luo, Yan Ding 0004, Chubo Liu, Wenchao Zhao, Kenli Li 0001 |
IEEE Trans. Computers | 2 |
| 2026 | Computational Burst Buffers: Accelerating HPC I/O via In-Storage Compression OffloadingabstractBurst buffers (BBs) act as an intermediate storage layer between compute nodes and parallel file systems (PFS), effectively alleviating the I/O performance gap in high-performance computing (HPC). As scientific simulations and AI workloads generate larger checkpoints and analysis outputs, BB capacity shortages and PFS bandwidth bottlenecks are emerging, and CPU-based compression is not an effective solution due to its high overhead. We introduceComputational Burst Buffers(CBBs), a storage paradigm that embeds hardware compression engines such as application-specific integrated circuit (ASIC) inside computational storage drives (CSDs) at the BB tier. CBB transparently offloads both lossless and error-bounded lossy compression from CPUs to CSDs, thereby (i) expanding effective SSD-backed BB capacity, (ii) reducing BB–PFS traffic, and (iii) eliminating contention and energy overheads of CPU-based compression. Unlike prior CSD-based compression designs targeting databases or flash caching, CBB co-designs the burst-buffer layer and CSD hardware for HPC and quantitatively evaluates compression offload in BB–PFS hierarchies. We prototype CBB using a PCIe 5.0 CSD with an ASIC Zstd-like compressor and an FPGA prototype of an SZ entropy encoder, and evaluate CBB on a 16-node cluster. Experiments with four representative HPC applications and a large-scale workflow simulator show up to 61% lower application runtime, 8–12× higher cache hit ratios, and substantially reduced compute-node CPU utilization compared to software compression and conventional BBs. These results demonstrate that compression-aware BBs with CSDs provide a practical, scalable path to next-generation HPC storage. Xiang Chen 0028, Bing Lu 0001, Haoquan Long, Huizhang Luo, Yili Ma, Guangming Tan, Dingwen Tao, Fei Wu 0005, Tao Lu 0014 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2026 | cuFastTuckerPlusTC: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor CoresabstractSparse tensors are prevalent in real-world applications, often characterized by their large-scale, high-order, and high dimensional nature. Directly handling raw tensors is impractical due to the significant memory and computational overhead involved. The current mainstream approach involves compressing or decomposing the original tensor. One popular tensor decomposition algorithm is the Tucker decomposition. However, existing state-of-the-art algorithms for large-scale Tucker decomposition typically relax the original optimization problem into multiple convex optimization problems to ensure polynomial convergence. Unfortunately, these algorithms tend to converge slowly. In contrast, tensor decomposition exhibits a simple optimization landscape, making local search algorithms capable of converging to a global (approximate) optimum much faster. In this paper, we propose the FastTuckerPlus algorithm, which decomposes the original optimization problem into two non-convex optimization problems and solves them alternately using the Stochastic Gradient Descent method. Furthermore, we introduce cuFastTuckerPlusTC, a fine-grained parallel algorithm designed for GPU platforms, leveraging the performance of tensor cores. This algorithm minimizes memory access overhead and computational costs, surpassing the state-of-the-art algorithms. Our experimental results demonstrate that the proposed method achieves a 2× to 8× improvement in convergence speed and a 3× to 5× improvement in per-iteration execution speed compared with state-of-the-art algorithms. Mingxing Duan, Huizhang Luo, Wangdong Yang, Kenli Li 0001, Keqin Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | STREAM: Spatiotemporal Similarity-based Efficient Approximate Median with Tunable GranularityabstractThe median (MED) is a crucial statistic for measuring the central tendency. However, exact MED computation remains costly, with even state-of-the-art (SOTA) algorithms failing to meet (near) real-time processing demands. While approximate MED algorithm has arisen as a promising candidate, existing approaches ignore the potential opportunity of spatiotemporal similarity within the application and fail to provide applicationspecific trade-offs between execution time and accuracy. Our goal is to design an enhanced approximate MED algorithm STREAM, which is capable of exploiting the spatiotemporal similarity to achieve bucket reuse and establish a tunable-grained bucket mechanism to meet the accuracy of application-specific requirements. Experimental results show that while maintaining nearly identical accuracy, STREAM outperforms the SOTA approximate methods DDSketch (up to $10 \times 4.7 \times$ on average) and KLL (up to $71.2 \times 10.1 \times$ on average). Fenfang Li, Huizhang Luo, Weichen Liu 0001, Anthony T. Chronopoulos, Kenli Li 0001, Chubo Liu |
DAC | 2 |
| 2025 | HSMU-SpGEMM: Achieving High Shared Memory Utilization for Parallel Sparse General Matrix-Matrix Multiplication on Modern GPUsabstractSparse general matrix-matrix multiplication (SpGEMM) is a core primitive for numerous scientific applications. Traditional hash-based approaches fail to strike a balance between reducing hash collisions and efficiently utilizing fast shared memory, which significantly undermines the performance of executing SpGEMM on GPUs. To address this issue, this paper introduces a novel accumulator design that achieves high shared memory utilization on modern GPUs. For the proposed high shared memory utilization algorithm, i.e., HSMU-SpGEMM1, we further optimize different symbolic stages. Our evaluations with four state-of-the-art hash-based SpGEMM libraries (Nsparse, spECK, OpSparse, and NVIDIA’s cuSPARSE) on three NVIDIA GPUs (Ampere, Ada Lovelace, Turing) demonstrate significant performance benefits from HSMU-SpGEMM.1HSMU-SpGEMM is available at https://github.com/wuminqaq/HSMUSpGEMM Huizhang Luo, Fenfang Li, Zhuo Tang, Kenli Li 0001, Jeff Zhang 0001, Chubo Liu |
HPCA | 2 |
| 2025 | CSubBT: A modular execution framework with self-adjusting capability for mobile manipulation system
Huihui Guo, Huizhang Luo, Huilong Pi, Mingxing Duan, Kenli Li 0001, Chubo Liu |
Neurocomputing | 2 |
| 2025 | Optimizing both performance and tail latency for B+tree on persistent memory
Xianyu He, Chaoshu Yang, Runyu Zhang 0002, Huizhang Luo, Zhichao Cao 0002, Jeff Zhang 0001 |
J. Syst. Archit. | 4 |
| 2025 | An efficient lossy compression framework for density partitioning in AMR applications
Yida Li 0001, Huizhang Luo, Yufeng Zhang 0001, Keqin Li 0001, Kenli Li 0001 |
J. Supercomput. | 2 |
| 2024 | zeroTT: A Two-Step State Transition Avoidance Scheme for MLC STT-RAMabstractCompared with conventional SRAM, Spin-Transfer Torque Random Access Memory(STT-RAM) is expected to play a crucial role in future memory technologies with the increasing demands for higher storage density and lower power consumption for modern embedded systems. Moreover, Multi-Level Cell (MLC) STT-RAM outperforms Single-Level Cell (SLC) STT-RAM since it has higher bit density. However, MLC STT-RAM suffers from write performance due to the two-step state transitions (TTs) in memory cells' soft domain. State-of-the-art approaches mitigate this issue by reducing TTs with efficient data coding. Unfortunately, none of the existing works can fully eliminate the TTs. In this work, zeroTT, an optimal (3, 4)-based expansion coding method that eliminates TTs for MLC STT-RAM. The design of ZeroTT considers space overhead and coding complexity, and our experimental results demonstrate that zeroTT can completely avoid TTs, leading to a more efficient MLC STT-RAM memory in terms of access latency, energy consumption, and device lifetime. Huizhang Luo, Jeff Zhang 0001, Mingxing Duan, Wangdong Yang, Zhuo Tang, Kenli Li 0001 |
DAC | 2 |
| 2024 | Hierarchical Explanations for Text Classification Models: Fast and EffectiveabstractGenerating explanations for deep neural networks (DNNs) can make them more trustworthy in real-world applications. For a text classification task, existing methods visualize the contributions of words or word interactions layer by layer in a traversal manner, to assist users in understanding the decision-making of models. However, all these methods only focus on the explanation performance while ignoring inefficiencies in explaining due to the traversal manner. This means that the explanation is not available to users in a timely manner, so users may no longer use it due to the big time cost. To overcome this problem, we propose HETSG, an interaction-based method for explaining text classification models quickly and faithfully, by a simple and effective two-step building strategy. Such a strategy captures the important interaction by first determining the important position and then confirming the direction of interaction, without iterating over all word interactions. We also provide a novel metric to accurately evaluate the performance of each interpretation method. The proposed method is compared with baseline methods (baselines) on six text classification datasets to explain three natural language processing (NLP) models. Experimental results show that our method outperforms all baselines with higher efficiency and it is also competitive in performance. Zhenyu Nie, Huizhang Luo, Anthony T. Chronopoulos |
ICDM | 3 |
| 2024 | AMP: Total Variation Reduction for Lossless Compression via Approximate Median-based PreconditioningabstractWith the increasing scale of cloud computing applications of next-generation embedded systems, a major challenge that domain scientists are facing is how to efficiently store and analyze the vast volume of output data. Compression can reduce the amount of data that needs to be transferred and stored. However, most of the large datasets are in floating-point format, which exhibits high entropy. As a result, existing lossless compressors cannot provide enough performance for such applications. To address this problem, we propose a total variation reduction method for improving the compression ratio of lossless compressors (namely, FPC + and FPZIP + ), which employs a median-based hyperplane to precondition the data. In particular, we first try to exploit the space-filling curve (SFC), a well-known technique to preserve data locality for a multi-dimensional dataset. We show and explain why a raw SFC, such as Hilbert and Z-order curves, cannot improve the compression ratio. Then, we explore the opportunity and theoretical feasibility of the proposed total variation reduction-based algorithm. The experiment results show the effectiveness of the proposed method. The compression ratios are improved up to 48.2% (20.6% on average) for FPZIP and 42.4% (18.4% on average) for FPC. Moreover, through observing the time composition of the proposed method, it is found that the median finding holds a high percentage of the execution time. Hence, we further introduce an approximate median finding algorithm, providing a linear-time overhead reduction scheme. The experiment results clearly demonstrate that this algorithm reduces execution time by an average of 56.7% and 40.7% compared to FPC + and FPZIP + , respectively. Fenfang Li, Huizhang Luo, Junqi Wang 0002, Yida Li 0001, Zhuo Tang, Kenli Li 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2024 | FastLoad: Speeding Up Data Loading of Both Sparse Matrix and Vector for SpMV on GPUsabstractSparse Matrix-Vector Multiplication (SpMV) on GPUs has gained significant attention because of SpMV's importance in modern applications and the increasing computing power of GPUs in the last decade. Previous studies have emphasized the importance of data loading for the overall performance of SpMV and demonstrated the efficacy of coalesced memory access in enhancing data loading efficiency. However, existing approaches fall far short of reaching the full potential of data loading on modern GPUs. In this paper, we propose an efficient algorithm called FastLoad, that speeds up the loading of both sparse matrices and input vectors of SpMV on modern GPUs. Leveraging coalesced memory access, FastLoad achieves high loading efficiency and load balance by sorting both the columns of the sparse matrix and elements of the input vector based on the number of non-zero elements while organizing non-zero elements in blocks to avoid thread divergence. FastLoad takes the Compressed Sparse Column (CSC) format as an implementation case to prove the concept and gain insights. We conduct a comprehensive comparison of FastLoad with the CSC-based SpMV, cuSPARSE, CSR5, and TileSpMV, using the full SuiteSparse Matrix Collection as workload. The experimental results on RTX 3090 Ti demonstrate that our method outperforms the others in most matrices, with geometric speedup means over CSC-based, cuSPARSE, CSR5, and TileSpMV being 2.12×, 2.98×, 2.88×, and 1.22×, respectively. Jinyu Hu, Huizhang Luo, Hong Jiang 0001, Guoqing Xiao 0001, Kenli Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | ZFP-X: Efficient Embedded Coding for Accelerating Lossy Floating Point CompressionabstractToday’s scientific simulations are confronting seriously limited I/O bandwidth, network bandwidth, and storage capacity because of immense volumes of data generated in high-performance computing systems. Data compression has emerged as one of the most effective approaches to resolve the issue of the exponential increase of scientific data. However, existing state-of-the-art compressors also are confronting the issue of low throughput, especially under the trend of growing disparities between the compute and I/O rates. Among them, embedded coding is widely applied, which contributes to the dominant running time for the corresponding compressors. In this work, we propose a new kind of embedded coding algorithm, and apply it as the backend embedded coding of ZFP, one of the most successful lossy compressors. Our embedded coding algorithm uses bit groups instead of bit planes to store the compressed data, avoiding the time overhead of generating bit planes and group tests of bit planes, which significantly reduces the running time of ZFP. Our embedded coding algorithm can also accelerate the decompression of ZFP, because the costly procedures of the reverse of group tests and reconstructing bit planes are also avoided. Moreover, we provide theoretical proof that the proposed coding algorithm has the same compression ratio as the baseline ZFP. Experiments with four representative real-world scientific simulation datasets show that the compression and decompression throughput of our solution is up to 2.5× (2.1× on average), and up to 2.1× (1.5× on average) as those of ZFP, respectively. Bing Lu 0001, Yida Li 0001, Junqi Wang 0002, Huizhang Luo, Kenli Li 0001 |
IPDPS | 4 |
| 2023 | A Data-driven Approach to Harvesting Latent Reduced Models to Precondition Lossy Compression for Scientific DataabstractIn this paper, we propose and evaluate the idea that data need to be preconditioned prior to compression, such that they can better match the design philosophies of lossy compressors for HPC scientific data. In particular, we aim to identify a reduced model that can be utilized to transform the original data into a more compressible form. We begin with two PDE applications as a proof of concept, in which we demonstrate that a reduced model can indeed reside in the full model output, and can be utilized to improve compression ratios. A mathematical proof is also presented to show how the compression ratio is improved by the reduced model. We further explore more general dimension reduction techniques to extract the reduced model, including principal component analysis, singular value decomposition, and discrete wavelet transform. After preconditioning, the reduced model in conjunction with difference between the reduced model and full model is stored, which results in higher compression ratios. We evaluate the reduced models on ten scientific datasets, and the results show the effectiveness of our approaches. Given that there is no single method that consistently achieves the best performance, we further propose a selection strategy that guides users to select the best reduced model prior to data reduction. Huizhang Luo, Junqi Wang 0002, Zhenlu Qin, Dan Huang 0001, Qing Liu 0002, MengChu Zhou, Hong Jiang 0001 |
IEEE Trans. Big Data | 1 |
| 2023 | LAMP: Improving Compression Ratio for AMR Applications via Level Associated Mapping-Based PreconditioningabstractData compression can efficiently reduce the memory and persistence storage cost, which is highly desirable in modern computing systems, such as enterprise, cloud, and High-Performance Computing (HPC) environments. However, the main challenges of existing data compressors are the insufficient compression ratio and low throughput. This paper focuses on improving the compression ratio of state-of-the-art lossy compression algorithms from the view of applications. Besides, we also use the characteristics of the applications to reduce the runtime overhead. To this end, we explore the idea with Adaptive Mesh Refinement (AMR), which is widely adopted as a computational technique to reduce the amount of computation and memory required in scientific simulations. We propose Level Associated Mapping-based Preconditioning (LAMP) to improve the storage efficiency of AMR applications. The main idea is twofold. First, we utilize the high similarities among the adjacent AMR levels to precondition the data prior to compression. Second, AMR has a unique characteristic of grid structures. We utilize grid structures to rebuild a level associated mapping table, which significantly reduces the runtime overhead of LAMP. Thanks to the optimization techniques of General Matrix Multiplication (GEMM), we further accelerate the process of rebuilding AMR hierarchy for LAMP. Besides, we also block multiple adjacent coordinates within a box and further improve cache locality. The experimental results show that the compression ratios of LAMP are improved up to 63.8% compared to directly compressing the data. Yida Li 0001, Huizhang Luo, Fenfang Li, Junqi Wang 0002, Kenli Li 0001 |
IEEE Trans. Computers | 2 |
| 2023 | COFFEE: Cross-Layer Optimization for Fast and Efficient Executions of Sinkhorn-Knopp Algorithm on HPC SystemsabstractIn this paper, we present COFFEE, cross-layer optimization for fast and efficient executions of the Sinkhorn-Knopp (SK) algorithm on HPC systems with clusters of compute nodes by exploring some architectural features of the system. By analyzing the performance of a typical implementation of the SK algorithm on such a system, a huge performance gap is observed between the row rescaling and column rescaling of the algorithm, where the latter requires much more time than the former. We also found that the costly MPI communication of the column rescaling seriously hinders the exploitation of parallelism. By observing and leveraging unique architectural characteristics across different system optimizations, such as column rescaling redesign, data blocking, micro-kernel design, enhanced intra-node and inter-node communication in MPI, etc., COFFEE is able to explore cross-layer optimization opportunities that enable fast and efficient execution of the SK algorithm. Our experimental results show that COFFEE provides up to 7.5X with an average of 2.0X performance improvement over the typical implementation on a single node, and up to 2.9X with an average of 1.6X performance improvement over the state-of-the-art MPI Allreduce algorithms on Tianhe-1 supercomputer. Chengyu Sun 0001, Huizhang Luo, Hong Jiang 0001, Jeff Zhang 0001, Kenli Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | zMesh: Theories and Methods to Exploring Application Characteristics to Improve Lossy Compression Ratio for Adaptive Mesh RefinementabstractScientific simulations on high-performance computing systems produce vast amounts of data that need to be stored and analyzed efficiently. Lossy compression significantly reduces the data volume by trading accuracy for performance. Despite the recent success of lossy compressions, such as ZFP and SZ, the compression performance is still far from being able to keep up with the exponential growth of data. This article aims to further take advantage of application characteristics, an area that is often under-explored, to improve the compression ratios of adaptive mesh refinement (AMR) - a widely used numerical solver that allows for an improved resolution in limited regions. We propose a level reordering techniquezMeshto reduce the storage footprint of AMR applications. In particular, we group the data points that are mapped to the same or adjacent geometric coordinates such that the dataset is smoother and more compressible. Unlike the prior work where the compression performance is affected by the overhead of metadata, this work re-generates the restore recipe using a chained tree structure, thus involving no extra storage overhead for compressed data, which substantially improves the compression ratios. We further derive a mathematical proof that lays the foundation for our method. The results demonstrate that zMesh can improve the smoothness of data by 67.9% and 71.3% for Z-ordering and Hilbert, respectively. Overall, zMesh improves the compression ratios by up to 16.5% and 133.7% for ZFP and SZ, respectively. Despite that zMesh involves additional compute overhead for tree and restore recipe construction, we show that the cost can be amortized as the number of quantities to be compressed increases. Huizhang Luo, Junqi Wang 0002, Qing Liu 0002, Jieyang Chen, Scott Klasky, Norbert Podhorszki |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | zMesh: Exploring Application Characteristics to Improve Lossy Compression Ratio for Adaptive Mesh RefinementabstractScientific simulations on high-performance computing systems produce vast amounts of data that need to be stored and analyzed efficiently. Lossy compression significantly reduces the data volume by trading accuracy for performance. Despite the recent success of lossy compression, such as ZFP and SZ, the compression performance is still far from being able to keep up with the exponential growth of data. This paper aims to further take advantage of application characteristics, an area that is often under-explored, to improve the compression ratios of adaptive mesh refinement (AMR) - a widely used numerical solver that allows for an improved resolution in limited regions. We propose a level reordering technique zMesh to reduce the storage footprint of AMR applications. In particular, we group the data points that are mapped to the same or adjacent geometric coordinates such that the dataset is smoother and more compressible. Unlike the prior work where the compression performance is affected by the overhead of metadata, this work re-generates restore recipe using a chained tree structure, thus involving no extra storage overhead for compressed data, which substantially improves the compression ratios. The results demonstrate that zMesh can improve the smoothness of data by 67.9% and 71.3% for Z-ordering and Hilbert, respectively. Overall, zMesh improves the compression ratios by up to 16.5% and 133.7% for ZFP and SZ, respectively. Despite that zMesh involves additional compute overhead for tree and restore recipe construction, we show that the cost can be amortized as the number of quantities to be compressed increases. Huizhang Luo, Junqi Wang 0002, Qing Liu 0002, Jieyang Chen, Scott Klasky, Norbert Podhorszki |
IPDPS | 1 |
| 2020 | Compression Ratio Modeling and Estimation across Error Bounds for Lossy CompressionabstractScientific simulations on high-performance computing (HPC) systems generate vast amounts of floating-point data that need to be reduced in order to lower the storage and I/O cost. Lossy compressors trade data accuracy for reduction performance and have been demonstrated to be effective in reducing data volume. However, a key hurdle to wide adoption of lossy compressors is that the trade-off between data accuracy and compression performance, particularly the compression ratio, is not well understood. Consequently, domain scientists often need to exhaust many possible error bounds before they can figure out an appropriate setup. The current practice of using lossy compressors to reduce data volume is, therefore, through trial and error, which is not efficient for large datasets which take a tremendous amount of computational resources to compress. This paper aims to analyze and estimate the compression performance of lossy compressors on HPC datasets. In particular, we predict the compression ratios of two modern lossy compressors that achieve superior performance, SZ and ZFP, on HPC scientific datasets at various error bounds, based upon the compressors' intrinsic metrics collected under a given base error bound. We evaluate the estimation scheme using twenty real HPC datasets and the results confirm the effectiveness of our approach. Jinzhen Wang, Tong Liu 0030, Qing Liu 0002, Xubin He, Huizhang Luo, Weiming He |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2019 | Identifying Latent Reduced Models to Precondition Lossy CompressionabstractWith the high volume and velocity of scientific data produced on high-performance computing systems, it has become increasingly critical to improve the compression performance. Leveraging the general tolerance of reduced accuracy in applications, lossy compressors can achieve much higher compression ratios with a user-prescribed error bound. However, they are still far from satisfying the reduction requirements from applications. In this paper, we propose and evaluate the idea that data need to be preconditioned prior to compression, such that they can better match the design philosophies of a compressor. In particular, we aim to identify a reduced model that can be utilized to transform the original data to a more compressible form. We begin with a case study of Heat3d as a proof of concept, in which we demonstrate that a reduced model can indeed reside in the full model output, and can be utilized to improve compression ratios. We further explore more general dimension reduction techniques to extract the reduced model, including principal component analysis, singular value decomposition, and discrete wavelet transform. After preconditioning, the reduced model in conjunction with difference between the reduced model and full model is stored, which results in higher compression ratios. We evaluate the reduced models on nine scientific datasets, and the results show the effectiveness of our approaches. Huizhang Luo, Dan Huang 0001, Qing Liu 0002, Zhenbo Qiao, Hong Jiang 0001, Jing Bi 0001, Haitao Yuan 0001, MengChu Zhou, Jinzhen Wang, Zhenlu Qin |
IPDPS | 1 |
| 2019 | Load-aware Elastic Data Reduction and Re-computation for Adaptive Mesh RefinementabstractThe increasing performance gap between computation and I/O creates huge data management challenges for simulation-based scientific discovery. Data reduction, among others, is deemed to be a promising technique to bridge the gap through reducing the amount of data migrated to persistent storage. However, the reduction performance is still far from what is being demanded from production applications. To this end, we propose a new methodology that aggressively reduces data despite the substantial loss of information, and re-computes the original accuracy on-demand. As a result, our scheme creates an illusion of a fast and large storage medium with the availability of high-accuracy data. We further design a load-aware data reduction strategy that monitors the I/O overhead at runtime, and dynamically adjusts the reduction ratio. We verify the efficacy of our methodology through adaptive mesh refinement, a popular numerical technique for solving partial differential equations. We evaluate data reduction and selective data re-computation on Titan, using a real application in FLASH and mini-applications in Chombo. To clearly demonstrate the benefits of re-computation, we compare it with other state-of-the-art data reduction methods including SZ, ZFP, FPC and deduplication, and it is shown to be superior in both write and read speeds, particularly when a small amount of data (e.g., 1%) need to be retrieved, as well as reduction ratio. Our results confirm that data reduction and selective data re-computation can 1) reduce the performance gap between I/O and compute via aggressively reducing AMR levels, and more importantly 2) can recover the target accuracy efficiently for AMR through re-computation. Mengxiao Wang, Huizhang Luo, Qing Liu 0002, Hong Jiang 0001 |
NAS | 2 |
| 2018 | Energy, latency, and lifetime improvements in MLC NVM with enhanced WOM codeabstractNon-volatile memories (NVMs), such as phase change memory (PCM) and resistive random access memory (ReRAM), have emerged as promising memory technologies for replacements of DRAM due to their advantages, such as better scalability, zero cell leakage, and DRAM-comparable read latency. Furthermore, multiple level cell (MLC) NVMs offer high data density and memory capacity over single level cell (SLC) NVM-s. However, the adoption of MLC NVMs is limited by their high programming energy and latency as well as the low endurance. In this paper, we propose an enhanced (23}2/4 WOM code for ML-C NVMs, which exploits the asymmetric characteristic in MLC NVM cell state transitions. Unlike the conventional WOM codes that focus on eliminating the worst-case latency writes, we propose to enlarge the best-case latency writes in MLC NVM cell state transitions. After data shaping with the enhanced WOM code, proportion of the best-case latency writes is maximized. In this way, the enhanced WOM code simultaneously reduces energy and latency, and improves lifetime with no memory and logic overheads. Evaluations show exciting improvement from the proposed approach. Huizhang Luo, Liang Shi 0001, Qiao Li 0001, Chun Jason Xue, Edwin H.-M. Sha |
ASP-DAC | 1 |
| 2018 | Understanding and Modeling Lossy Compression Schemes on HPC Scientific DataabstractScientific simulations generate large amounts of floating-point data, which are often not very compressible using the traditional reduction schemes, such as deduplication or lossless compression. The emergence of lossy floating-point compression holds promise to satisfy the data reduction demand from HPC applications; however, lossy compression has not been widely adopted in science production. We believe a fundamental reason is that there is a lack of understanding of the benefits, pitfalls, and performance of lossy compression on scientific data. In this paper, we conduct a comprehensive study on state-of-the-art lossy compression, including ZFP, SZ, and ISABELA, using real and representative HPC datasets. Our evaluation reveals the complex interplay between compressor design, data features and compression performance. The impact of reduced accuracy on data analytics is also examined through a case study of fusion blob detection, offering domain scientists with the insights of what to expect from fidelity loss. Furthermore, the trial and error approach to understanding compression performance involves substantial compute and storage overhead. To this end, we propose a sampling based estimation method that extrapolates the reduction ratio from data samples, to guide domain scientists to make more informed data reduction decisions. Tao Lu 0014, Qing Liu 0002, Xubin He, Huizhang Luo, Eric Suchyta, Jong Choi 0001, Norbert Podhorszki, Scott Klasky, Matthew Wolf, Tong Liu 0030, Zhenbo Qiao |
IPDPS | 4 |
| 2018 | Write Energy Reduction for PCM via Pumping Efficiency ImprovementabstractThe emerging Phase Change Memory (PCM) is considered to be a promising candidate to replace DRAM as the next generation main memory due to its higher scalability and lower leakage power. However, the high write power consumption has become a major challenge in adopting PCM as main memory. In addition to the fact that writing to PCM cells requires high write current and voltage, current loss in the charge pumps also contributes a large percentage of high power consumption. The pumping efficiency of a PCM chip is a concave function of the write current. Leveraging the characteristics of the concave function, the overall pumping efficiency can be improved if the write current is uniform. In this article, we propose a peak-to-average (PTA) write scheme, which smooths the write current fluctuation by regrouping write units. In particular, we calculate the current requirements for each write unit by their values when they are evicted from the last level cache (LLC). When the write units are waiting in the memory controller, we regroup the write units by LLC-assisted PTA to reach the current-uniform goal. Experimental results show that LLC-assisted PTA achieved 13.4% of overall energy saving compared to the baseline. Huizhang Luo, Qing Liu 0002, Jingtong Hu, Qiao Li 0001, Liang Shi 0001, Qingfeng Zhuge, Edwin H.-M. Sha |
ACM Trans. Storage | 1 |
| 2016 | Peak-to-average pumping efficiency improvement for charge pump in Phase Change MemoriesabstractThe emerging Phase Change Memory (PCM) is considered as a promising candidate to replace DRAM as the next generation main memory since it has better scalability and lower leakage power. However, the high write power consumption has become a main challenge in adopting PCM as main memory. In addition to the fact that writing to PCM cells requires high write current and voltage, current loss in the charge pumps (CPs) also contributes a large percentage of the high power consumption. The pumping efficiency of a PCM chip is a concave function of the write current. Based on the characteristics of the concave function, the overall pumping efficiency can be improved if the write current is uniform. In this paper, we propose the peak-to-average (PTA) write scheme, which smooths the write current fluctuation by regrouping write units. An off-line optimal Integer Programming (IP) formulation and an efficient online algorithm are proposed to achieve this goal. Experimental results show that PTA can improve the charge pump efficiency to ∼40% with little overhead. Meanwhile, PTA can achieve 17.0% energy reduction on average. Huizhang Luo, Jingtong Hu, Liang Shi 0001, Chun Jason Xue, Qingfeng Zhuge |
ASP-DAC | 1 |
| 2016 | Two-step state transition minimization for lifetime and performance improvement on MLC STT-RAMabstractSpin-transfer torque random access memory (STT-RAM) is considered as a promising candidate to replace SRAM as the next generation cache memory since it has better scalability and lower leakage power. Recently, 2-bit multi-level cell (MLC) STT-RAM has been proposed to further increase data density. However, a key drawback for MLC STT-RAM is that the magnetization directions of its hard and soft domains cannot be flipped to two opposite directions simultaneously, which leads to the two-step problem in state transitions. Two-step state transitions would significantly impact the lifetime of MLC STT-RAM due to the wasted flips in the soft domains. To solve the problem, this paper proposes a novel two-step state transition minimization (TSTM) scheme, to improve the lifetime of MLC STT-RAM when it is employed in cache design. The basic idea is by sacrificing certain cells as auxiliary flags, the two-step state transitions in STT-RAM can be well eliminated. Experimental results show that the proposed scheme can improve the lifetime of MLC STT-RAM to 318.5%. Huizhang Luo, Jingtong Hu, Liang Shi 0001, Chun Jason Xue, Qingfeng Zhuge |
DAC | 1 |
| 2016 | Write reconstruction for write throughput improvement on MLC PCM based main memory
Huizhang Luo, Penglin Dai, Liang Shi 0001, Chun Jason Xue, Qingfeng Zhuge, Edwin H.-M. Sha |
J. Syst. Archit. | 1 |