EDBT 2026 Demo / reviewers in the wild / expert
Jinzhao Xiao
dblp:327/7238
· DBLP profile ↗
4ranked-venue papers
3as first author
4since 2021 · last 2025
0009-0005-6255-7773ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BOS: Bit-Packing with Outlier SeparationabstractBit-packing serves as the fundamental operator in various data encoding and compression methods. The idea is to use a fixed bit-width to represent all the (processed) values in a sequence. Some extremely large values, known as outliers, obviously amplify the bit-width, and thus lead to wasted bits for most other small values. We notice that not only the large values (upper outliers) but also the small ones (lower outliers) could incur wasted bit-width. In this paper, we propose to store both the upper and lower outliers separately, namely Bit-packing with Outlier Separation (BOS). While the remaining center values have a narrow spread, i.e., condensed bit-width, the separated outliers need some extra cost to denote their positions. The problem is thus how to determine better thresholds for separating the upper and lower outliers, yielding smaller storage cost. Rather than enumerating all the possible values as upper and lower outlier separators, in$O(n^{2})$time, we consider bit-width as the separators, with$O(n\log n)$search time. Theoretical analysis illustrates all the possible cases such that the bit-width separation still returns the optimal solution as the value separation, and further leads to an approximate separation strategy with both median and bit-width, in$O(n)$time. Remarkably, our BOS is compatible to any existing compression methods using Bit-packing, and has replaced Bit-packing in Apache IoTDB and Apache TsFile. The extensive experiments on many real-world datasets demonstrate that by replacing Bit-packing with the proposed BOS in various compression methods, the compression ratio is significantly improved from about 2.75 to 3.25. Jinzhao Xiao, Shaoxu Song |
ICDE | 1 |
| 2024 | REGER: Reordering Time Series Data for Regression EncodingabstractRegression models are employed in lossless compression of time series data, by storing the residual of each point, known as regression encoding. Owing to value fluctuation, the regression residuals could be large and thus occupy huge space. It is worth noting that compared to the fluctuating values, time intervals are often regular and easy to compress, especially in the IoT scenarios where sensor data are collected in a preset frequency. In this sense, there is a trade-off between storing the regular timestamps and fluctuating values. Intuitively, rather than in time order, we may exchange the data points in the series such that the nearby ones have both smoother timestamps and values, leading to lower residuals. In this paper, we propose to reorder the time series data for better regression encoding. Rather than recomputing from scratch, efficient updates of residuals after moving some points are devised. The experimental comparison over various real-world datasets, either public or collected by our industrial partners, illustrates the superiority of the proposal in compression ratio. The method, REGression Encoding with Reordering (REGER), has now become an encoding method in an open-source time series database, Apache IoTDB. Jinzhao Xiao, Wendi He, Shaoxu Song, Xiangdong Huang 0001, Chen Wang 0018, Jianmin Wang 0001 |
ICDE | 1 |
| 2024 | Time series data encoding in Apache IoTDB: comparative analysis and recommendation
Tianrui Xia, Jinzhao Xiao, Yuxiang Huang 0001, Shaoxu Song, Xiangdong Huang 0001, Jianmin Wang 0001 |
VLDB J. | 2 |
| 2022 | Time Series Data Encoding for Efficient Storage: A Comparative Analysis in Apache IoTDBabstractNot only the vast applications but also the distinct features of time series data stimulate the booming growth of time series database management systems, such as Apache IoTDB, InfluxDB, OpenTSDB and so on. Almost all these systems employ columnar storage, with effective encoding of time series data. Given the distinct features of various time series data, it is not surprising that different encoding strategies may perform variously. In this study, we first summarize the features of time series data that may affect encoding performance, including scale, delta, repeat and increase. Then, we introduce the storage scheme of a typical time series database, Apache IoTDB, prescribing the limits to implementing encoding algorithms in the system. A qualitative analysis of encoding effectiveness regarding to various data features is then presented for the studied algorithms. To this end, we develop a benchmark for evaluating encoding algorithms, including a data generator regarding the aforesaid data features and several real-world datasets from our industrial partners. Finally, we present an extensive experimental evaluation using the benchmark. Remarkably, a quantitative analysis of encoding effectiveness regarding to various data features is conducted in Apache IoTDB. Jinzhao Xiao, Yuxiang Huang 0001, Shaoxu Song, Xiangdong Huang 0001, Jianmin Wang 0001 |
Proc. VLDB Endow. | 1 |