VLDB 2026 Research / reviewers in the wild / expert
Zijie Chen 0009
dblp:135/0704-9
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
0009-0005-8535-379XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OneRoundSTL: In-Database Seasonal-Trend DecompositionabstractSeasonal-trend decomposition has been widely used in time series analysis, e.g., time series forecasting and anomaly detection. Existing seasonal-trend decomposition methods, such as STL and its variations, assume that the time series is complete and sorted by timestamp. However, popular time series databases usually adopt LSM-Tree based storage, which stores data in pages not necessarily in time order. Moreover, time series stored in databases often suffer from missing values due to sensor failures, further compromising their integrity. A straightforward idea is to first merge and sort the data of different pages, and then decompose them. It obviously leads to heavy online computation, repeated calculations for multiple queries, and still cannot deal with the remaining missing data. In this paper, we propose OneRoundSTL, which pre-calculates offline some results in each individual page and concatenates the pre-calculated results online at query time to obtain the decomposition outcome. OneRoundSTL has been deployed and included as a function in an open source time series database, Apache IoTDB. Experiments on synthetic and real-world datasets in the system show that our OneRoundSTL exhibits high efficiency, far exceeding the state-of-the-art methods, while keeping decomposition effect. Zijie Chen 0009, Shaoxu Song, Jianmin Wang 0001 |
ICDE | 1 |
| 2025 | Cleaning Time Series under Seasonal and Trend ConstraintsabstractTime series data are often found to be dirty, e.g., with anomalies or sensor failures. Such dirty data obviously hinder the downstream analysis tasks such as forecasting, clustering or classification. Simply discarding the potentially dirty data points is not an option, making the time series incomplete and incompatible to machine learning models. While many time series data cleaning techniques have been developed in the last decade, e.g., with the help of constraints on value fluctuation, the seasonal features are surprisingly ignored. In this paper, we propose to clean time series by first capturing seasonal and trend constraints, and then enforcing them for cleaning. Unfortunately, directly applying existing seasonal-trend decomposition methods is found imprecise (itself affected by errors) and incomplete (not computed at the beginning or end of the series). Moreover, unlike efficient cleaning with simple value fluctuation constraints, the time series cleaning problem with seasonal and trend constraints is proved to be NP-complete. In this sense, we first improve seasonal and trend filter with tolerance to errors and extension on two directions. Then, an efficient heuristic is designed to iteratively repair the time series and refine the seasonal and trend constraints. The approach has now become a built-in function in a product system Apache IoTDB. Experiments on real-world datasets demonstrate the superiority of our proposal in cleaning seasonal time series and improving downstream applications. Zijie Chen 0009, Aoqian Zhang, Shaoxu Song |
Proc. ACM Manag. Data | 1 |
| 2024 | On Reducing Space Amplification with Multi-Column Compaction in Apache IoTDBabstractLog-structured merge trees (LSM-trees) are commonly employed as the storage engines for write-intensive workloads in modern time series databases including Apache IoTDB. Following append-only principle, LSM-trees can handle intensive writes and updates, but consequently suffer high space amplification (SA). To reduce SA in LSM-tree, compaction is triggered periodically to reorganize a large number of immutable files on disk to eliminate redundancy. This issue is further complicated in the Internet of Things (IoT) scenarios, where frequent out-of-order data insertions and data updates introduce duplicated keys, obsolete values and overlapping bitmaps in multi-column data, thereby exacerbating SA concerns. To mitigate SA in such contexts, this paper presents a Multi-Column Compaction (MCC) strategy in Apache IoTDB, an open-source time series database utilizing LSM-tree architecture and supporting multi-column storage. We take into consideration both the separate insertions (out-of-order data) and updates of multi-column data, and analyze the hardness of selecting proper files with the maximum space reduction in compaction. We then propose a heuristic method designed to improve the file selection, thus reducing SA. To enhance the efficiency of this approach, we further devise File Prefetcher and Compaction Cache. The proposed MCC has been implemented in Apache IoTDB. Experimental results demonstrate that our proposed MCC achieves better performance in reducing space amplification. Chenguang Fang, Zijie Chen 0009, Shaoxu Song, Xiangdong Huang 0001, Chen Wang 0018, Jianmin Wang 0001 |
Proc. VLDB Endow. | 2 |