Jinzhen Wang

dblp:156/9939 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0001-6317-2940ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Rethinking the Potential of Layer Freezing for DNN Training Efficiency
abstract
With the growing scale of deep neural networks and datasets, training has become increasingly expensive. Layer freezing reduces this cost by stopping updates to selected layers, but frozen layers still require forward propagation to generate activations for later layers. Caching these activations as a surrogate dataset can eliminate this redundant computation, but it faces two key challenges: effectively augmenting cached features and reducing the storage overhead of high-dimensional activations. This paper provides the first systematic study of these challenges and proposes practical solutions. We introduce Similarity-Aware Channel Augmentation to preserve accuracy by caching transformation-sensitive channels with limited overhead. We further incorporate lossy compression and design a progressive compression strategy that exploits the higher compressibility of deeper-layer activations. Our method reduces computation cost, memory usage, and training time while maintaining accuracy. Experiments on NVIDIA Orin Edge GPU further demonstrate training acceleration and significant power savings, highlighting its practicality for resource-constrained training.
Chence Yang, Ningxi Cheng, Ci Zhang, Qitao Tan, Sheng Li 0019, Ao Li 0004, Xulong Tang, Shaoyi Huang, Jinzhen Wang, Jundong Li, Xiaoming Zhai, Jin Lu 0001, Geng Yuan
ACM Great Lakes Symposium on VLSI10
2025 NeurLZ: An Online Neural Learning-based Method to Enhance Scientific Lossy Compression
abstract
SZ3 (0.07MB, PSNR = 29.2) (c)NeurLZ (0.07MB, PSNR = 39.1)
Wenqi Jia 0003, Zhewen Hu, Youyuan Liu, Boyuan Zhang 0002, Jinzhen Wang, Jinyang Liu 0003, Wei Niu 0002, Stavros Kalafatis, Junzhou Huang, Sian Jin, Daoce Wang, Jiannan Tian, Miao Yin
ICS5
2024 QualityNet: Error-bounded Lossy Compression Quality Prediction via Deep Surrogate
abstract
As scientific simulations generate increasingly large datasets, efficient data compression becomes essential to mitigate storage and transmission challenges. Error-bounded Lossy compression algorithms reduce data volume at the expense of data fidelity. However, the evaluation of compressed data quality can be resource-intensive and time-consuming. This paper proposes a surrogate-based framework to predict key compression quality assessment metrics, allowing users to evaluate the quality of compressed data quickly without extensive trial and error. We evaluate our proposed framework on real-world scientific application datasets. Our surrogate-based framework shows superior prediction accuracy, with the RMSE generally lower than 5%. On the other hand, our framework demonstrates robust generalization for prediction across data fields within the same application. Our results show that our approach significantly reduces the error bound selection time and accelerates the downstream evaluation process for decompressed data, making it easier for domain scientists to select optimal error bounds of lossy compression while maintaining data quality that meets their needs. This work bridges the gap between compression ratio optimization and data fidelity, offering a scalable solution for scientific applications that rely on large-scale lossy compression.
Khondoker Mirazul Mumenin, Dong Dai 0001, Jinzhen Wang, Sheng Di
IEEE Big Data3
2024 A Generic Performance Metric for Scientific Data Compression
abstract
The performance outcomes of scientific data compression are complex in nature, often with one performance metric being more prioritized/optimized over others in the algorithm. To balance between different performance metrics, this paper formulates a new generic metric and uses it for searching an optimal configuration that leads to a balanced compression performance.
Zhenlu Qin, Jessie Gu, Maggie Liu, Jinzhen Wang, Hongjian Zhu
e-Science4
2024 Tango: A Cross-layer Approach to Managing I/O Interference over Local Ephemeral Storage
abstract
As simulation-based scientific discovery advances to exascale, a major question that the community is striving to answer is how to co-design data storage and complex physicsrich analytics in a way that the time to knowledge can be minimized for post-processing. A particular challenge is how to accommodate a broad spectrum of data analytics needsparticularly those that become clear only until very late during the post-processing, a scenario where existing methods, such as in situ processing, are unable or less effective in supporting data analytics. As HPC storage systems have become deeper and more complex with the recent addition of NVMe, die-stacked memory, and burst buffer, it requires fundamentally rethinking new paradigms and methods for data storage and analysis. This paper aims to address the issue of I/O interference for data analytics over local ephemeral storage, which is shared by multiple applications in a non-exclusive node usage scenario-often configured for small- to medium-sized clusters. At the core of this work is a coordinated cross-layer approach that reacts to storage interference from both storage and application layers. By decomposing and distributing analysis data across the storage hierarchy, data analytics can adapt to the interference by reducing or completely avoiding access to lower tiers whenever there is a high interference, while maintaining a prescribed error bound to limit the information loss. Meanwhile, proper actions are also taken at the storage layer to ensure sufficient bandwidth is allocated for retrieving an augmentation, which is based upon the cardinality and accuracy of the augmentation as well as the nature of an application. We evaluate three realworld data analytics, XGC, GenASiS, and CFD, on Chameleon, and quantitatively demonstrate that the I/O performance can be vastly improved, e.g., by 52% versus no adaptivity and 36% versus single-layer adaptivity, while maintaining acceptable outcomes of data analysis.
Zhenbo Qiao, Qirui Tian, Zhenlu Qin, Jinzhen Wang, Qing Liu 0002, Norbert Podhorszki, Scott Klasky, Hongjian Zhu
SC4
2023 Improving Progressive Retrieval for HPC Scientific Data using Deep Neural Network
abstract
As the disparity between compute and I/O on high-performance computing systems has continued to widen, it has become increasingly difficult to perform post-hoc data analytics on full-resolution scientific simulation data due to the high I/O cost. Error-bounded data decomposition and progressive data retrieval framework has recently been developed to address such a challenge by performing data decomposition before storage and reading only part of the decomposed data when necessary. However, the performance of the progressive retrieval framework has been suffering from the over-pessimistic error control theory, such that the achieved maximum error of recomposed data is significantly lower than the required error. Therefore, more data than required is fetched for recomposition, incurring additional I/O overhead. In order to tackle this issue, we propose a DNN-based progressive retrieval framework that can better identify the minimum amount of data to be retrieved. Our contributions are as follows: 1) We provide an in-depth investigation of the recently developed progressive retrieval framework; 2) We propose two designs of prediction models (named D-MGARD and E-MGARD) to estimate the amount of retrieved data size based on error bounds. 3) We evaluate our proposed solutions using scientific datasets generated by real-world simulations from two domains. Evaluation results demonstrate the effectiveness of our solution in accurately predicting the amount of retrieval data size, as well as the advantages of our solution over the traditional approach to reducing the I/O overhead. Based on our evaluation, our solution is shown to read significantly less data (5% - 40% with D-MGARD, 20% - 80% with E-MGARD).
Jinzhen Wang, Xin Liang 0001, Ben Whitney, Jieyang Chen, Qian Gong, Xubin He, Lipeng Wan 0001, Scott Klasky, Norbert Podhorszki, Qing Liu 0002
ICDE1
2023 High-Ratio Lossy Compression: Exploring the Autoencoder to Compress Scientific Data
abstract
Scientific simulations on high-performance computing (HPC) systems can generate large amounts of floating-point data per run. To mitigate the data storage bottleneck and lower the data volume, it is common for floating-point compressors to be employed. As compared to lossless compressors, lossy compressors, such as SZ and ZFP, can reduce data volume more aggressively while maintaining the usefulness of the data. However, a reduction ratio of more than two orders of magnitude is almost impossible without seriously distorting the data. In deep learning, the autoencoder technique has shown great potential for data compression, in particular with images. Whether the autoencoder can deliver similar performance on scientific data, however, is unknown. In this article, we for the first time conduct a comprehensive study on the use of autoencoders to compress real-world scientific data and illustrate several key findings on using autoencoders for scientific data reduction. We implement an autoencoder-based compression prototype to reduce floating-point data. Our study shows that the out-of-the-box implementation needs to be further tuned in order to achieve high compression ratios and satisfactory error bounds. Our evaluation results show that, for most of the test datasets, the tuned autoencoder outperforms SZ by up to 4X, and ZFP by up to 50X in compression ratios, respectively. Our practices and lessons learned in this work can direct future optimizations for using autoencoders to compress scientific data.
Tong Liu 0030, Jinzhen Wang, Qing Liu 0002, Shakeel Alibhai, Tao Lu 0014, Xubin He
IEEE Trans. Big Data2
2023 zPerf: A Statistical Gray-Box Approach to Performance Modeling and Extrapolation for Scientific Lossy Compression
abstract
With the scaling up of simulation-based scientific discovery on high-performance computing systems, the disparity between compute and I/O has increased, forcing domain scientists to save only a small amount of simulation data to persistent storage. This can result in the loss of essential physics fields that are needed for data analysis. While error-bounded lossy compression has made tremendous progress in bridging the gap between compute and I/O, the lack of understanding of compression performance remains a key hurdle to its wide adoption. In this work, we present zPerf, a statistical gray-box performance modeling approach for scientific lossy compression. Our contributions are threefold: 1) We develop zPerf to estimate the performance of lossy compression techniques, based on in-depth understanding and statistical modeling for data features and core compression metrics; 2) We demonstrate the in-detailed implementation of zPerf using two case studies, where we derive the performance modeling for SZ and ZFP, two leading lossy compressors; 3) We evaluate the effectiveness of zPerf on real-world datasets across various domains. Based on the evaluation, we demonstrate the efficacy of the zPerf performance model; 4) We further discuss three case studies where zPerf is applied to extrapolate the compression ratio of SZ and ZFP with alternative encoding schemes as well as ZFP with an alternative transform scheme. Through the case studies, we demonstrate the potential of zPerf for exploring the design space of lossy compression, which has hardly been studied in the literature.
Jinzhen Wang, Tong Liu 0030, Qing Liu 0002, Xubin He
IEEE Trans. Computers1
2022 Locality-based transfer learning on compression autoencoder for efficient scientific data lossy compression
Tong Liu 0030, Jinzhen Wang, Qing Liu 0002, Shakeel Alibhai, Xubin He
J. Netw. Comput. Appl.3
2021 Reducing the Training Overhead of the HPC Compression Autoencoder via Dataset Proportioning
abstract
As the storage overhead of high-performance computing (HPC) data reaches into the petabyte or even exabyte scale, it could be useful to find new methods of compressing such data. The compression autoencoder (CAE) has recently been proposed to compress HPC data with a very high compression ratio. However, this machine learning-based method suffers from the major drawback of lengthy training time. In this paper, we attempt to mitigate this problem by proposing a proportioning scheme to reduce the amount of data that is used for training relative to the amount of data to be compressed. We show that this method drastically reduces the training time without, in most cases, significantly increasing the error. We further explain how this scheme can even improve the accuracy of the CAE on certain datasets. Finally, we provide some guidance on how to determine a suitable proportion of the training dataset to use in order to train the CAE for a given dataset.
Tong Liu 0030, Shakeel Alibhai, Jinzhen Wang, Qing Liu 0002, Xubin He
NAS3
2020 Compression Ratio Modeling and Estimation across Error Bounds for Lossy Compression
abstract
Scientific simulations on high-performance computing (HPC) systems generate vast amounts of floating-point data that need to be reduced in order to lower the storage and I/O cost. Lossy compressors trade data accuracy for reduction performance and have been demonstrated to be effective in reducing data volume. However, a key hurdle to wide adoption of lossy compressors is that the trade-off between data accuracy and compression performance, particularly the compression ratio, is not well understood. Consequently, domain scientists often need to exhaust many possible error bounds before they can figure out an appropriate setup. The current practice of using lossy compressors to reduce data volume is, therefore, through trial and error, which is not efficient for large datasets which take a tremendous amount of computational resources to compress. This paper aims to analyze and estimate the compression performance of lossy compressors on HPC datasets. In particular, we predict the compression ratios of two modern lossy compressors that achieve superior performance, SZ and ZFP, on HPC scientific datasets at various error bounds, based upon the compressors' intrinsic metrics collected under a given base error bound. We evaluate the estimation scheme using twenty real HPC datasets and the results confirm the effectiveness of our approach.
Jinzhen Wang, Tong Liu 0030, Qing Liu 0002, Xubin He, Huizhang Luo, Weiming He
IEEE Trans. Parallel Distributed Syst.1
2019 Identifying Latent Reduced Models to Precondition Lossy Compression
abstract
With the high volume and velocity of scientific data produced on high-performance computing systems, it has become increasingly critical to improve the compression performance. Leveraging the general tolerance of reduced accuracy in applications, lossy compressors can achieve much higher compression ratios with a user-prescribed error bound. However, they are still far from satisfying the reduction requirements from applications. In this paper, we propose and evaluate the idea that data need to be preconditioned prior to compression, such that they can better match the design philosophies of a compressor. In particular, we aim to identify a reduced model that can be utilized to transform the original data to a more compressible form. We begin with a case study of Heat3d as a proof of concept, in which we demonstrate that a reduced model can indeed reside in the full model output, and can be utilized to improve compression ratios. We further explore more general dimension reduction techniques to extract the reduced model, including principal component analysis, singular value decomposition, and discrete wavelet transform. After preconditioning, the reduced model in conjunction with difference between the reduced model and full model is stored, which results in higher compression ratios. We evaluate the reduced models on nine scientific datasets, and the results show the effectiveness of our approaches.
Huizhang Luo, Dan Huang 0001, Qing Liu 0002, Zhenbo Qiao, Hong Jiang 0001, Jing Bi 0001, Haitao Yuan 0001, MengChu Zhou, Jinzhen Wang, Zhenlu Qin
IPDPS9
2019 Exploring Transfer Learning to Reduce Training Overhead of HPC Data in Machine Learning
abstract
Nowadays, scientific simulations on high-performance computing (HPC) systems can generate large amounts of data (in the scale of terabytes or petabytes) per run. When this huge amount of HPC data is processed by machine learning applications, the training overhead will be significant. Typically, the training process for a neural network can take several hours to complete, if not longer. When machine learning is applied to HPC scientific data, the training time can take several days or even weeks. Transfer learning, an optimization usually used to save training time or achieve better performance, has potential for reducing this large training overhead. In this paper, we apply transfer learning to a machine learning HPC application. We find that transfer learning can reduce training time without, in most cases, significantly increasing the error. This indicates transfer learning can be very useful for working with HPC datasets in machine learning applications.
Tong Liu 0030, Shakeel Alibhai, Jinzhen Wang, Qing Liu 0002, Xubin He, Chentao Wu
NAS3
2015 Parameter estimation of chirp signal under low SNR
Jinzhen Wang, Shaoying Su, Zengping Chen
Sci. China Inf. Sci.1