EDBT 2026 Demo / reviewers in the wild / expert
Zhuoxun Yang
dblp:399/6586
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0004-3232-4963ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging Information Theory and Practice for Scientific Lossy CompressionabstractError-bounded lossy compressors have been developed for years to reduce the vast volumes of scientific data generated by high-performance computing (HPC) applications and advanced scientific instruments. While these compressors have been effective in mitigating the challenges posed by massive datasets, a significant gap remains in our understanding of the fundamental compressibility limits of scientific data–an issue that critically impacts the sustainable adoption and development of efficient lossy compression techniques in practice. Classical rate-distortion theory, established by Shannon, assumes stationary 1D sources with unconstrained coding–assumptions that do not hold for scientific datasets compressed under the tiling constraints imposed by modern parallel lossy compressors. This paper addresses this gap by developing a novel framework that characterizes compressibility limits for scientific datasets under realistic tiling constraints. The contribution is two-fold. First, we establish a tile-aware, finite-blocklength extension of rate–distortion theory that advances classical 1D asymptotic formulations into a rigorous framework for piecewise 2D Gaussian random fields. To our knowledge, this is the first framework to rigorously characterize lossy compressibility limits for scientific datasets and compressor, moving beyond classical asymptotic 1D source models. Second, we conduct a comprehensive validation of the proposed modeling framework using state-of-the-art error-bounded lossy compressors and diverse real-world HPC datasets, demonstrating that our theory accurately predicts rate-distortion trends and provides actionable insights for compressor design. Sujata Sinha, Sheng Di, Vishwas Rao, Robert Underwood, David Lenz 0002, Zizhe Jian, Zhuoxun Yang, Kai Zhao 0008, Lingjia Liu 0001, Franck Cappello |
HPDC | 7 |
| 2026 | TZ: Achieving High-Ratio Scientific Data Compression on GPUs with Global Data DecompositionabstractAs high-performance computing shifts toward GPU-accelerated exascale systems, the exponential growth of scientific data poses severe challenges to both storage capacity and I/O bandwidth. While current GPU-based lossy compressors attempt to address this by porting CPU algorithms to the device, they rely heavily on block-wise spatial decomposition to fit GPU parallelism. This approach suffers from a fundamental locality barrier: by partitioning data into independent blocks, these methods fail to capture global correlations and fragment the unified data patterns required for effective coding, severely limiting compression ratios. In this paper, we propose TZ, a novel GPU-native error-bounded lossy compressor that breaks this ceiling by adopting global Tucker decomposition. By prioritizing global spectral energy compaction over local approximation, TZ naturally maximizes the compression potential for scientific datasets. To render this computationally intensive approach practical for high-throughput GPU workflows, we introduce a highly optimized adaptive randomized SVD engine. This design allows TZ to achieve the superior compression ratios of global spectral decomposition while maintaining competitive execution speeds. Furthermore, the global processing nature of TZ enables a unified quantization and coding scheme that eliminates block artifacts and metadata overhead. Evaluation on production-scale scientific datasets demonstrates that TZ achieves approximately 10 × higher compression ratios than state-of-the-art GPU compressors under the same error bound, while maintaining competitive, high-throughput performance. Zhuoxun Yang, Amit N. Subrahmanya, Vishwas Rao, Sheng Di, Robert Underwood, Jinyang Liu 0003, Franck Cappello, Kai Zhao 0008 |
HPDC | 1 |
| 2026 | OPAL: On-demand Progressive Accelerated Scientific Lossy Compression
Zhuoxun Yang, Robert Underwood, Sheng Di, Daoce Wang, Jinyang Liu 0003, Jiajun Huang 0001, Franck Cappello, Kai Zhao 0008 |
HPDC | 3 |
| 2026 | GPZ: GPU-Accelerated Lossy Compressor for Particle DataabstractParticle-based simulations and point-cloud applications generate massive, irregular datasets that challenge storage, I/O, and real-time analytics. Traditional compression techniques struggle with irregular particle distributions and GPU architectural constraints, often resulting in limited throughput and suboptimal compression ratios. In this paper, we present GPZ, a high-performance, error-bounded lossy compressor designed specifically for large-scale particle data on modern GPUs. GPZ employs a novel four-stage parallel pipeline that synergistically balances high compression efficiency with the architectural demands of massively parallel hardware. We introduce a suite of targeted optimizations for computation, memory access, and GPU occupancy that enable GPZ to achieve near-hardware-limit throughput. We conduct an extensive evaluation on three distinct GPU architectures (workstation, data center, and edge) using six large-scale, real-world scientific datasets from four distinct domains. The results demonstrate that GPZ consistently and significantly outperforms four state-of-the-art GPU compressors, delivering up to 8x higher end-to-end throughput while achieving superior compression ratios and data quality. Yafan Huang, Zhuoxun Yang, Sheng Di, Boyuan Zhang 0002, Jiajun Huang 0001, Jinyang Liu 0003, Jiannan Tian, Guanpeng Li, Fengguang Song, Hanqi Guo 0001, Franck Cappello, Kai Zhao 0008 |
ICS | 4 |
| 2025 | IPComp: Interpolation Based Progressive Lossy Compression for Scientific ApplicationsabstractCompression is a crucial solution for data reduction in modern scientific applications due to the exponential growth of data from simulations, experiments, and observations. Compression with progressive retrieval capability allows users to quickly access coarse approximations of data and then incrementally refine these approximations to higher fidelity. Existing progressive compression solutions suffer from low reduction ratios or high operation costs, effectively undermining the approach's benefits. In this paper, we propose our interpolation-based progressive lossy compression solution that has both high reduction ratios and low operation costs. The interpolation-based algorithm has been verified as one of the best for scientific data reduction, but previously, no effort exists to make it support progressive retrieval. Our contributions are threefold: (1) We thoroughly analyze the error characteristics of the interpolation algorithm and propose our solution, IPComp, with multi-level bitplane and predictive coding. (2) We derive optimized strategies toward minimum data retrieval under different fidelity levels indicated by users through error bounds and bitrates. (3) We evaluate the proposed solution using six real-world datasets from four diverse domains. Experimental results demonstrate our solution archives up to 487% higher compression ratios and 698% faster speed than other state-of-the-art progressive compressors, and reduces the data volume for retrieval by up to 83% compared to baselines under the same error bound, and reduces the error by up to 99% under the same bitrate. Zhuoxun Yang, Sheng Di, Ximiao Li, Jiajun Huang 0001, Jinyang Liu 0003, Franck Cappello, Kai Zhao 0008 |
HPDC | 1 |