VLDB 2026 Research / reviewers in the wild / expert
Pu Jiao
dblp:337/3134
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0001-6829-1463ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Time-Varying Vector Field Compression with Preserved Critical Point Trajectories
Mingze Xia, Yuxiao Li 0002, Pu Jiao, Bei Wang 0001, Xin Liang 0001, Hanqi Guo 0001 |
ICDE | 3 |
| 2026 | Mitigating Artifacts in Pre-quantization Based Scientific Data Compressors with Quantization-aware Interpolation
Pu Jiao, Sheng Di, Jiannan Tian, Mingze Xia, Yang Zhang 0031, Xin Liang 0001, Franck Cappello |
IPDPS | 1 |
| 2025 | TspSZ: An Efficient Parallel Error-Bounded Lossy Compressor for Topological Skeleton PreservationabstractData compression is a powerful solution for addressing big data challenges in database and data management. In scientific data compression for vector fields, preserving topological information is essential for accurate analysis and visualization. The topological skeleton, a fundamental component of vector field topology, consists of critical points and their connectivity (i.e., separatrices). While previous work has focused on preserving critical points in error-controlled lossy compression, little attention has been given to preserving separatrices, which are equally important. In this work, we introduce TspSZ, an efficient error-bounded lossy compression framework designed to preserve both critical points and separatrices. Our key contributions are threefold. First, we propose TspSZ, a topological-skeleton-preserving lossy compression framework that integrates two algorithms, enabling existing critical-point-preserving compressors to also retain separatrices, significantly enhancing their topology preservation capabilities. Second, we optimize TspSZ for efficiency through tailored improvements and parallelization. Specifically, we introduce a new error control mechanism to achieve high compression ratios and implement a shared-memory parallelization strategy to boost compression throughput. Third, we evaluate TspSZ against state-of-the-art lossy and lossless compressors using four real-world scientific datasets. Experimental results show that TspSZ achieves compression ratios of up to 7.7× while effectively preserving the topological skeleton, ensuring efficient storage and transmission of scientific data without compromising topological integrity. Mingze Xia, Bei Wang 0001, Yuxiao Li 0002, Pu Jiao, Xin Liang 0001, Hanqi Guo 0001 |
ICDE | 4 |
| 2025 | Improving the Efficiency of Interpolation-based Scientific Data Compressors with Adaptive Quantization Index PredictionabstractLarge-scale scientific simulations produce unprecedented amounts of data using high-performance computing systems, leading to severe problems in data storage, I/O, and communication. To address the data movement challenge, errorcontrolled lossy compression has been proposed to significantly reduce the data size while retaining the data quality. Recently, interpolation-based compressors, including MGARD, SZ3, QoZ, and HPEZ, have stood out due to their efficiency in obtaining relatively high compression ratios with decent compression and decompression throughput. Nevertheless, these methods focus on data decorrelation in the compression pipeline yet overlook the correlation of the quantization indices generated after decorrelation. In this paper, we develop a generic framework that can use the correlation of quantization indices to significantly improve the compression ratios for state-of-the-art interpolation-based error-bounded lossy compressors. Our contributions are threefold: (1) We carefully characterized the quantization index array produced by the interpolation-based compressors and identified the unused correlation; (2) We designed a generic quantization index prediction method to exploit such correlation, which leads to improved compression ratio with only minor degradation in throughput; (3) We integrate our method into 4 state-of-theart interpolation-based compressors and evaluate them using 5 real-world datasets. Experimental results demonstrate that the proposed method improves the compression ratios of the base compressors by up to 95% while keeping the same quality. It also leads to 16% improvement in end-to-end data transfer performance under a parallel setting. Pu Jiao, Sheng Di, Mingze Xia, Jinyang Liu 0003, Xin Liang 0001, Franck Cappello |
IPDPS | 1 |
| 2025 | Enabling Efficient Error-Controlled Lossy Compression for Unstructured Scientific DataabstractToday's scientific applications are producing vast amounts of data with cutting-edge high-performance computing systems and high-resolution instruments, causing severe problems in data transmission and storage. While error-controlled data compression is regarded as a direct way to solve the problem, most existing compressors are designed for data from structured meshes. In this work, we propose a generic framework to enable efficient error-controlled compression for scientific data from unstructured meshes. The contributions are four-fold: (1) We design a prediction-based framework with additional preprocessing stages to better incorporate mesh information. (2) We propose three families of prediction methods for unstructured meshes and integrate them into the framework, which yields high prediction accuracy and thus significantly improves the compression ratios and quality. (3) We enhance our framework by enabling invalid node processing and feature preservation. (4) We evaluate our approaches using five datasets from real-world applications and compare them with state-of-the-art error-controlled lossy compressors. Experiments demonstrate that the proposed compression methods deliver up to$26.36 \times, 8.24 \times$, and$2.56 \times$compression ratios over existing compressors under the same error bound, Peak Signal-to-Noise Ratios, and critical point preservation levels, respectively. This leads to$1.63 \times$performance speedup in the end-to-end data transfer on Globus. Sheng Di, Congrong Ren, Pu Jiao, Mingze Xia, Hanqi Guo 0001, Xin Liang 0001, Franck Cappello |
IPDPS | 4 |
| 2025 | QPET: A Versatile and Portable Quantity-of-Interest-preservation Framework for Error-Bounded Lossy CompressionabstractError-bounded lossy compression has been widely adopted in many scientific domains because it can address the challenges in storing, transferring, and analyzing unprecedented amounts of scientific data. However, general error-bounded lossy compressors may fail to meet additional quality requirements for downstream analysis, a.k.a. Quantities of Interest (QoIs). This may lead to uncertainties and even misinterpretations in scientific discoveries, significantly limiting the use of lossy compression in practice. In this paper, we propose QPET, a novel, versatile, and portable framework for QoI-preserving error-bounded lossy compression, which overcomes the challenges of modeling diverse QoIs by leveraging numerical strategies. QPET features (1) high portability to multiple existing lossy compressors, (2) versatile preservation to most differentiable univariate and multivariate QoIs, and (3) significant compression improvements in QoI-preservation tasks. Experiments with six real-world datasets demonstrate that integrating QPET into state-of-the-art error-bounded lossy compressors can gain 2x to 10x compression speedups of existing QoI-preserving error-bounded lossy compression solutions, up to 1000% compression ratio improvements to general-purpose compressors, and up to 133% compression ratio improvements to existing QoI-integrated scientific compressors. Jinyang Liu 0003, Pu Jiao, Kai Zhao 0008, Xin Liang 0001, Sheng Di, Franck Cappello |
Proc. VLDB Endow. | 2 |
| 2024 | Preserving Topological Feature with Sign-of-Determinant Predicates in Lossy Compression: A Case Study of Vector Field Critical PointsabstractLossy compression has been employed to reduce the unprecedented amount of data produced by today's large-scale scientific simulations and high-resolution instruments. To avoid loss of critical information, state-of-the-art scientific lossy compressors provide error controls on relatively simple metrics such as absolute error bound. However, preserving these metrics does not translate to the preservation of topological features, such as critical points in vector fields. To address this problem, we investigate how to effectively preserve the sign of determinant in error-controlled lossy compression, as it is an important quantity of interest used for the robust detection of many topological features. Our contribution is three-fold. (1) We develop a generic theory to derive the allowable perturbation for one row of a matrix while preserving its sign of the determinant. As a practical use-case, we apply this theory to preserve critical points in vector fields because critical point detection can be reduced to the result of the point-in-simplex test that purely relies on the sign of determinants. (2) We optimize this algorithm with a speculative compression scheme to allow for high compression ratios and efficiently parallelize it in distributed environments. (3) We perform solid experiments with real-world datasets, demonstrating that our method achieves up to 440% improvements in compression ratios over state-of-the-art lossy compressors when all critical points need to be preserved. Using the parallelization strategies, our method delivers up to 1.25 x and 4.38 x performance speedup in data writing and reading compared with the vanilla approach without compression. Mingze Xia, Sheng Di, Franck Cappello, Pu Jiao, Kai Zhao 0008, Jinyang Liu 0003, Xin Liang 0001, Hanqi Guo 0001 |
ICDE | 4 |
| 2023 | Characterization and Detection of Artifacts for Error-Controlled Lossy CompressorsabstractToday's scientific high-performance computing (HPC) applications are often running on large-scale environments, producing extremely large volumes of data that need to be compressed effectively for efficient storage or data transfer. Error-bounded lossy compression is arguably the most efficient way to this end, because it can get very high compression ratios while controlling the data distortion strictly based on user requirements for compression errors. However, error-bounded lossy compressors may have serious artifact issues in situations with relatively large error bound or high compression ratios, which is highly undesirable to users. In this paper, we compre-hensively characterize the artifacts for multiple state-of-the-art error-bounded lossy compressors (including SZ-1.4, SZ-2.1, SZ-3.0, FPZIP, ZFP, MGARD) and provide an in-depth analysis for the root cause of these artifacts. We summarize the artifact issue into three types and also develop an efficient artifact detection algorithm for each type of artifact. We finally evaluate our artifact detection methods using four scientific datasets, which demonstrates that the proposed methods are able to detect artifact issues under linear time complexity. Pu Jiao, Sheng Di, Jinyang Liu 0003, Xin Liang 0001, Franck Cappello |
HiPC | 1 |
| 2022 | Toward Quantity-of-Interest Preserving Lossy Compression for Scientific DataabstractToday's scientific simulations and instruments are producing a large amount of data, leading to difficulties in storing, transmitting, and analyzing these data. While error-controlled lossy compressors are effective in significantly reducing data volumes and efficiently developing databases for multiple scientific applications, they mainly support error controls on raw data, which leaves a significant gap between the data and user's downstream analysis. This may cause unqualified uncertainties in the outcomes of the analysis, a.k.a quantities of interest (QoIs), which are the major concerns of users in adopting lossy compression in practice. In this paper, we propose rigorous mathematical theories to preserve four families of QoIs that are widely used in scientific analysis during lossy compression along with practical implementations. Specifically, we first develop the error control theory for univariate QoIs which are essential for computing physical properties such as kinetic energy, followed by multivariate QoIs that are more commonly used in real-world applications. The proposed method is integrated into a state-of-the-art compression framework in a modular fashion, which could easily adapt to new QoIs and new compression algorithms. Experiments on real-world datasets demonstrate that the proposed method provides faithful error control on important QoIs including kinetic energy, regional average, and isosurface without trials and errors, while offering compression ratios that are up to 4X of the compression ratios provided by state-of-the-art compressors. Pu Jiao, Sheng Di, Hanqi Guo 0001, Kai Zhao 0008, Jiannan Tian, Dingwen Tao, Xin Liang 0001, Franck Cappello |
Proc. VLDB Endow. | 1 |