Boyuan Zhang 0002

dblp:122/4939-2 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
15since 2021 · last 2026
0009-0003-8937-4067ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 7 first-author · 15 since 2021
YearPublicationVenuePosition
2026 GPZ: GPU-Accelerated Lossy Compressor for Particle Data
abstract
Particle-based simulations and point-cloud applications generate massive, irregular datasets that challenge storage, I/O, and real-time analytics. Traditional compression techniques struggle with irregular particle distributions and GPU architectural constraints, often resulting in limited throughput and suboptimal compression ratios. In this paper, we present GPZ, a high-performance, error-bounded lossy compressor designed specifically for large-scale particle data on modern GPUs. GPZ employs a novel four-stage parallel pipeline that synergistically balances high compression efficiency with the architectural demands of massively parallel hardware. We introduce a suite of targeted optimizations for computation, memory access, and GPU occupancy that enable GPZ to achieve near-hardware-limit throughput. We conduct an extensive evaluation on three distinct GPU architectures (workstation, data center, and edge) using six large-scale, real-world scientific datasets from four distinct domains. The results demonstrate that GPZ consistently and significantly outperforms four state-of-the-art GPU compressors, delivering up to 8x higher end-to-end throughput while achieving superior compression ratios and data quality.
Yafan Huang, Zhuoxun Yang, Sheng Di, Boyuan Zhang 0002, Jiajun Huang 0001, Jinyang Liu 0003, Jiannan Tian, Guanpeng Li, Fengguang Song, Hanqi Guo 0001, Franck Cappello, Kai Zhao 0008
ICS6
2026 Accelerating AI Compression through Lightweight Lossless Encoding and Pipelined Workflows
Boyuan Zhang 0002, Luanzheng Guo, Jiannan Tian, Jinyang Liu 0003, Daoce Wang, Chengming Zhang 0006, Bo Fang 0002, Fengguang Song, Jan Strube 0001, Nathan R. Tallent, Dingwen Tao
IPDPS1
2026 Near-Zero Cost KV Cache Compression for Large Language Model Inference
Boyuan Zhang 0002, Yafan Huang, Shihui Song, Jinda Jia, Chengming Zhang 0006, Zhi Zhang 0005
IPDPS1
2025 BMQSim: Overcoming Memory Constraints in Quantum Circuit Simulation with a High-Fidelity Compression Framework
Boyuan Zhang 0002, Bo Fang 0002, Fanjiang Ye, Luanzheng Guo, Fengguang Song, Nathan R. Tallent, Dingwen Tao
ICS1
2025 Pushing the Limits of GPU Lossy Compression: A Hierarchical Delta Approach
Boyuan Zhang 0002, Yafan Huang, Sheng Di, Fengguang Song, Guanpeng Li, Franck Cappello
ICS1
2025 NeurLZ: An Online Neural Learning-based Method to Enhance Scientific Lossy Compression
abstract
SZ3 (0.07MB, PSNR = 29.2) (c)NeurLZ (0.07MB, PSNR = 39.1)
Wenqi Jia 0003, Zhewen Hu, Youyuan Liu, Boyuan Zhang 0002, Jinzhen Wang, Jinyang Liu 0003, Wei Niu 0002, Stavros Kalafatis, Junzhou Huang, Sian Jin, Daoce Wang, Jiannan Tian, Miao Yin
ICS4
2025 High-performance Visual Semantics Compression for AI-Driven Science
abstract
Scientific images play a crucial role in many experimental sciences; however, the large volumes of data generated present significant challenges. Effective image compression must be fast, achieve high compression ratios, and preserve critical domain-specific features. Existing compressors, such as JPEG and SZ, often distort important textures when operating at high compression ratios. Conversely, AI-based compressors offer superior image quality and higher compression ratios but are significantly slower than traditional methods. To address this trade-off, we developed ViSemZ, a high-performance AI-based compressor specifically designed to preserve visual semantics. Our approach enhances AI compression by integrating sparse encoding with variable-length integer truncation, optimized lossless encoding using bitshuffle and a decoupled lookback prefix-sum, and pipelining techniques to enable efficient data streaming and asynchronous processing. Evaluations on scientific datasets demonstrate that, at comparable compression ratios, ViSemZ achieves performance almost on par with existing AI-based compressors while delivering a 9.6× overall compression speedup. These results effectively bridge the performance gap between traditional and AI-based compression methods.
Boyuan Zhang 0002, Luanzheng Guo, Jiannan Tian, Jinyang Liu 0003, Daoce Wang, Fanjiang Ye, Chengming Zhang 0006, Jan Strube 0001, Nathan R. Tallent, Dingwen Tao
PPoPP1
2025 COMPSO: Optimizing Gradient Compression for Distributed Training with Second-Order Optimizers
abstract
Second-order optimization methods have been developed to enhance convergence and generalization in deep neural network (DNN) training compared to first-order methods like Stochastic Gradient Descent (SGD). However, these methods face challenges in distributed settings due to high communication overhead. Gradient compression, a technique commonly used to accelerate communication for first-order approaches, often results in low communication reduction ratios, decreased model accuracy, and/or high compression overhead when applied to second-order methods. To address these limitations, we introduce a novel gradient compression method for second-order optimizers called COMPSO. This method effectively reduces communication costs while preserving the advantages of second-order optimization. COMPSO employs stochastic rounding to maintain accuracy and filters out minor gradients to improve compression ratios. Additionally, we develop GPU optimizations to minimize compression overhead and performance modeling to ensure end-to-end performance gains across various systems. Evaluation of COMPSO on different DNN models shows that it achieves a compression ratio of 22.1×, reduces communication time by 14.2×, and improves overall performance by 1.9×, all without any drop in model accuracy.
Baixi Sun, Weijin Liu, J. Gregory Pauloski, Jiannan Tian, Jinda Jia, Daoce Wang, Boyuan Zhang 0002, Mingkai Zheng, Sheng Di, Sian Jin, Zhao Zhang 0007, Xiaodong Yu 0001, Kamil Iskra, Pete Beckman, Guangming Tan, Dingwen Tao
PPoPP7
2025 STZ: A High Quality and High Speed Streaming Lossy Compression Framework for Scientific Data
abstract
Error-bounded lossy compression is one of the most efficient solutions to reduce the volume of scientific data. For lossy compression, progressive decompression and random-access decompression are critical features that enable on-demand data access and flexible analysis workflows. However, these features can severely degrade compression quality and speed. To address these limitations, we propose a novel streaming compression framework that supports both progressive decompression and random-access decompression while maintaining high compression quality and speed. Our contributions are three-fold: (1) we design the first compression framework that simultaneously enables both progressive decompression and random-access decompression; (2) we introduce a hierarchical partitioning strategy to enable both streaming features, along with a hierarchical prediction mechanism that mitigates the impact of partitioning and achieves high compression quality—even comparable to state-of-the-art (SOTA) non-streaming compressor SZ3; and (3) our framework delivers high compression and decompression speed, up to 6.7 × faster than SZ3.
Daoce Wang, Pascal Grosset, Jesus Pulido, Jiannan Tian, Tushar M. Athawale, Jinda Jia, Baixi Sun, Boyuan Zhang 0002, Sian Jin, Kai Zhao 0008, James P. Ahrens, Fengguang Song
SC8
2025 Multifacets of lossy compression for scientific data in the Joint-Laboratory of Extreme Scale Computing
Franck Cappello, Mario C. Acosta, Emmanuel Agullo, Hartwig Anzt, Jon Calhoun 0001, Sheng Di, Luc Giraud, Thomas Grützmacher, Sian Jin, Kentaro Sano, Kento Sato, Amarjit Singh, Dingwen Tao, Jiannan Tian, Tomohiro Ueno, Robert Underwood, Frédéric Vivien, Xavier Yepes, Kazutomo Yoshii, Boyuan Zhang 0002
Future Gener. Comput. Syst.20
2024 MASC: A Memory-Efficient Adjoint Sensitivity Analysis through Compression Using Novel Spatiotemporal Prediction
abstract
Adjoint sensitivity analysis is critical in modern integrated circuit design and verification, but its computational intensity grows significantly with the circuit size, the number of objective functions, and the accumulation of time points. This growth can impede its wider application. The intimate link between the forward integration in transient analysis and the reverse integration in adjoint sensitivity analysis allows for the retention of Jacobian matrices from transient analysis, thereby speeding up sensitivity analysis. However, Jacobian matrices across multiple timesteps are often so large that they cannot be stored in memory during the forward integration process, necessitating disk storage and incurring significant I/O overhead. To address this, we develop a memory-efficient sensitivity analysis method that utilizes data compression to minimize memory overhead during simulation and enhance analysis efficiency. Our compression method can efficiently compress the sparse tensor that contains the Jacobian matrices over time by exploiting the spatiotemporal characteristics of the data and circuit attributes. It also introduces a shared-indices technique, a cutting-edge spatiotemporal prediction model, and robust residual encoding. We evaluate our compression method on 7 datasets from real-world simulations and demonstrate that it can reduce memory requirements by more than 16x on average, which is significantly more efficient than other state-of-the-art compression techniques.
Boyuan Zhang 0002, Yongqiang Duan, Zuochang Ye, Weifeng Liu 0002, Dingwen Tao, Zhou Jin 0001
DAC2
2024 cuSZ-i: High-Ratio Scientific Lossy Compression on GPUs with Optimized Multi-Level Interpolation
abstract
Error-bounded lossy compression is a critical technique for significantly reducing scientific data volumes. Compared to CPU-based compressors, GPU-based compressors exhibit substantially higher throughputs, fitting better for today’s HPC applications. However, the critical limitations of existing GPU-based compressors are their low compression ratios and qualities, severely restricting their applicability. To overcome these, we introduce a new GPU-based error-bounded scientific lossy compressor named CUSZ-i, with the following contributions: (1) A novel GPU-optimized interpolation-based prediction method significantly improves the compression ratio and decompression data quality. (2) The Huffman encoding module in CUSZ-i is optimized for better efficiency. (3) CUSZ-i is the first to integrate the NVIDIA Bitcomp-lossless as an additional compression-ratio-enhancing module. Evaluations show that CUSZ-i significantly outperforms other latest GPU-based lossy compressors in compression ratio under the same error bound (hence, the desired quality), showcasing a 476% advantage over the second-best. This leads to CUSZ-i’s optimized performance in several real-world use cases.
Jinyang Liu 0003, Jiannan Tian, Shixun Wu, Sheng Di, Boyuan Zhang 0002, Robert Underwood, Yafan Huang, Jiajun Huang 0001, Kai Zhao 0008, Guanpeng Li, Dingwen Tao, Zizhong Chen, Franck Cappello
SC5
2024 Accelerating Communication in Deep Learning Recommendation Model Training with Dual-Level Adaptive Lossy Compression
abstract
DLRM is a state-of-the-art recommendation system model that has gained widespread adoption across various industry applications. The large size of DLRM models, however, necessitates the use of multiple devices/GPUs for efficient training. A significant bottleneck in this process is the time-consuming all-to-all communication required to collect embedding data from all devices. To mitigate this, we introduce a method that employs error-bounded lossy compression to reduce the communication data size and accelerate DLRM training. We develop a novel error-bounded lossy compression algorithm, informed by an in-depth analysis of embedding data features, to achieve high compression ratios. Moreover, we introduce a dual-level adaptive strategy for error-bound adjustment, spanning both table-wise and iteration-wise aspects, to balance the compression benefits with the potential impacts on accuracy. We further optimize our compressor for PyTorch tensors on GPUs, minimizing compression overhead. Evaluation shows that our method achieves a 1.38 × training speedup with a minimal accuracy impact.
Boyuan Zhang 0002, Fanjiang Ye, Min Si, Ching-Hsiang Chu, Jiannan Tian, Chunxing Yin, Summer Deng, Yuchen Hao, Pavan Balaji, Tong Geng, Dingwen Tao
SC2
2023 FZ-GPU: A Fast and High-Ratio Lossy Compressor for Scientific Computing Applications on GPUs
abstract
Today's large-scale scientific applications running on high-performance computing (HPC) systems generate vast data volumes. Thus, data compression is becoming a critical technique to mitigate the storage burden and data-movement cost. However, existing lossy compressors for scientific data cannot achieve a high compression ratio and throughput simultaneously, hindering their adoption in many applications requiring fast compression, such as in-memory compression. To this end, in this work, we develop a fast and high- ratio error-bounded lossy compressor on GPUs for scientific data (called FZ-GPU). Specifically, we first design a new compression pipeline that consists of fully parallelized quantization, bitshuffle, and our newly designed fast encoding. Then, we propose a series of deep architectural optimizations for each kernel in the pipeline to take full advantage of CUDA architectures. We propose a warp-level optimization to avoid data conflicts for bit-wise operations in bitshuffle, maximize shared memory utilization, and eliminate unnecessary data movements by fusing different compression kernels. Finally, we evaluate FZ-GPU on two NVIDIA GPUs (i.e., A100 and RTX A4000) using six representative scientific datasets from SDRBench. Results on the A100 GPU show that FZ-GPU achieves an average speedup of 4.2× over cuSZ and an average speedup of 37.0× over a multi-threaded CPU implementation of our algorithm under the same error bound. FZ-GPU also achieves an average speedup of 2.3× and an average compression ratio improvement of 2.0× over cuZFP under the same data distortion.
Boyuan Zhang 0002, Jiannan Tian, Sheng Di, Xiaodong Yu 0001, Yunhe Feng, Xin Liang 0001, Dingwen Tao, Franck Cappello
HPDC1
2023 GPULZ: Optimizing LZSS Lossless Compression for Multi-byte Data on Modern GPUs
abstract
Today's graphics processing unit (GPU) applications produce vast volumes of data, which are challenging to store and transfer efficiently. Thus, data compression is becoming a critical technique to mitigate the storage burden and communication cost. LZSS is the core algorithm in many widely used compressors, such as Deflate. However, existing GPU-based LZSS compressors suffer from low throughput due to the sequential nature of the LZSS algorithm. Moreover, many GPU applications produce multi-byte data (e.g., int16/int32 index, floating-point numbers), while the current LZSS compression only takes single-byte data as input. To this end, in this work, we propose gpuLZ, a highly efficient LZSS compression on modern GPUs for multi-byte data. The contribution of our work is fourfold: First, we perform an in-depth analysis of existing LZ compressors for GPUs and investigate their main issues. Then, we propose two main algorithm-level optimizations. Specifically, we (1) change prefix sum from one pass to two passes and fuse multiple kernels to reduce data movement between shared memory and global memory, and (2) optimize existing pattern-matching approach for multi-byte symbols to reduce computation complexity and explore longer repeated patterns. Third, we perform architectural performance optimizations, such as maximizing shared memory utilization by adapting data partitions to different GPU architectures. Finally, we evaluate gpuLZ on six datasets of various types with NVIDIA A100 and A4000 GPUs. Results show that gpuLZ achieves up to 272.1× speedup on A4000 and up to 1.4× higher compression ratio compared to state-of-the-art solutions.
Boyuan Zhang 0002, Jiannan Tian, Sheng Di, Xiaodong Yu 0001, D. Martin Swany, Dingwen Tao, Franck Cappello
ICS1