Abduvoris Abduvakhobov

dblp:381/6104 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0005-4160-418XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 50% High-performance computing · 50%
Databases, data mining, and information retrieval
2 papers
Spatial and temporal data management · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Energy systems and smart grids · 100%
Computer networks
1 paper
Edge and fog computing · 100%

Topics — the 5 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Spatial and temporal data management
time series compression
1.522024
Scalable Model-Based Management of Massive High Frequency Wind Turbine Data with ModelarDB · Proc. VLDB Endow. 2024
Why Model-Based Lossy Compression is Great for Wind Turbine Analytics · ICDE 2024
Storage systems
storage reliability
1.522024
Scalable Model-Based Management of Massive High Frequency Wind Turbine Data with ModelarDB · Proc. VLDB Endow. 2024
Why Model-Based Lossy Compression is Great for Wind Turbine Analytics · ICDE 2024
Spatial and temporal data management
time series data management
0.812024
Scalable Model-Based Management of Massive High Frequency Wind Turbine Data with ModelarDB · Proc. VLDB Endow. 2024
High-performance computing › lossy compression
error-bounded lossy compression
0.812024
Scalable Model-Based Management of Massive High Frequency Wind Turbine Data with ModelarDB · Proc. VLDB Endow. 2024
High-performance computing
lossy compression
0.812024
Why Model-Based Lossy Compression is Great for Wind Turbine Analytics · ICDE 2024

Methods — techniques the papers use, named apart from their topics

per-value error bounds · 2.3model-based lossy compression · 2.3model-based compression · 2.3
YearPublicationVenuePosition
2026 Compressing High-Frequency Time Series Through Multiple Models and Stealing From Residuals
abstract
Wind turbines are equipped with high-quality sensors that generate vast volumes of high-frequency time series. The time series are ingested on the edge and transferred to the cloud for later analytics. This process is complicated by challenges like low network bandwidth and high cloud storage costs. ModelarDB was proposed as a solution to efficiently manage time series across the entire pipeline by using so-called models for lossless or error-bounded lossy compression of time series. However, ModelarDB’s compression can be further improved through: 1) avoiding models that only represent few values by storing residuals (i.e., values that models fail to compress) explicitly with them; 2) exploiting error bounds even more through preprocessing; and 3) timestamp compression specialized for regular and irregular time series. We propose the multi-model compression method Fauna which uses 1) the novel model fitting method Platypus; 2) PMC and Swing for compressing values and; 3) the novel Macaque for compressing residuals and timestamps. Platypus is a model fitting method that uses different models for specialized compression of values and residuals. We then evaluate state-of-the-art lossless compression methods for 32-bit floats and propose preprocessing methods to add support for error-bounded compression. We present Macaque that includes MacaqueV and MacaqueTS. MacaqueV modifies Facebook Gorilla’s lossless compression method for 32-bit floats (GorillaV) and combines it with our novel preprocessing methods to now also enable error-bounded lossy compression. MacaqueTS is a lossless compression method for timestamps. Using only Platypus reduces ModelarDB’s storage use by up to 1.8x and significantly simplifies using the system. While also up to 7x better for lossless compression, ModelarDB with Fauna uses up to 2.5x less storage than ModelarDB and up to 14.5x, 7.2x, 17.5x and 14.2x less storage than ClickHouse, Apache IoTDB, Apache Parquet and TimescaleDB, respectively, with a realistic 1% error bound.
Abduvoris Abduvakhobov, Søren Kejser Jensen, Christian Thomsen 0001, Torben Bach Pedersen
ICDE1
2024 Why Model-Based Lossy Compression is Great for Wind Turbine Analytics
abstract
Modern wind turbines are equipped with wired high-quality sensors that produce high-frequency sensor data in the form of time series as shown in Figure 1 a. From working with multiple different practitioners, we have learned that relatively few but very long high-quality time series are produced. The time series are either univariate, i.e., have one value per timestamp, or multivariate, i.e., have multiple values per timestamp. Further, they are either regular, i.e., have a fixed time interval between consecutive data points, or irregular. Despite these differences, the volume and velocity of the time series that are being produced are generally major challenges. For example, if the sensors are sampled at 100Hz, a single park of 100 wind turbines generates more than 11 PiB of data each year [1]. The sensor data is collected by weak edge devices and then transferred to powerful cloud servers over a relatively slow connection as shown in Figure 2. However, it is infeasible to transfer and store the raw time series due to their volume and velocity. Renewable energy system installations use low-end commodity PCs on the edge, e.g., 4 CPU cores, 4 GiB RAM, and an HDD [1]. In addition, the bandwidth between the edge and the cloud can be as low as 0.5-5 Mbit/s [1]. Thus, practitioners use simple aggregates, e.g., 10-minute averages, which remove valuable outliers and fluctuations as shown in Figure 1b. To remedy this, practitioners want to use lossy compression with a per-value error bound (E) to collect more high-frequency time series and thus improve their analytics.
Søren Kejser Jensen, Christian Thomsen 0001, Torben Bach Pedersen, Carlos Muñiz Cuza, Abduvoris Abduvakhobov
ICDE5
2024 Scalable Model-Based Management of Massive High Frequency Wind Turbine Data with ModelarDB
abstract
Modern wind turbines are monitored by sensors that generate massive amounts of high frequency time series that are ingested on the edge and then transferred to the cloud where they are stored and analyzed. This results in at least four challenges: (1) Limited hardware makes efficient ingestion necessary to keep up; (2) Limited bandwidth makes data compression necessary; (3) High storage costs as all data must be stored; and (4) Low data quality due to lossy compression methods without error bounds. Practitioners currently use solutions that only solve some of these. In this paper, we evaluate the Time Series Management System ModelarDB, a solution that meets all four challenges by efficiently managing time series across the entire pipeline. We compare it to three commonly used alternatives and evaluate different aspects of them in a realistic edge-to-cloud scenario with real-life datasets. For lossless compression, ModelarDB achieves up to 2x better compression and 1.2x better transfer efficiency. For lossy compression, ModelarDB achieves up to 4.6x better compression and 10x better transfer efficiency, or similar compression with orders of magnitude less error.
Abduvoris Abduvakhobov, Søren Kejser Jensen, Torben Bach Pedersen, Christian Thomsen 0001
Proc. VLDB Endow.1