Rui Kang 0004

dblp:89/1167-4 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0003-4910-8410ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Exploring SIMD Vectorization in Aggregation Pipelines for Encoded IoT Data
abstract
Time-series databases have been critical for collecting and analyzing data in industries where sensors send large amounts of IoT data by network devices. Both data received from networks and data collected in database storage are sufficiently encoded to reduce I/O occupation and latency. The IoT encoders successively combine the Delta, Repeat, and Packing operators, yielding a higher compression ratio than simply adopting each. However, efficient compression makes query execution even harder, requiring serial decoding before processing queries. Among them, selective aggregations, such as down-sampling, are the core of time series analytical queries. This paper identifies operators to process and accelerate IoT aggregation queries based on encoded data arrays, extensible to integrate thread-level and instruction-level designs. In addition, encoded data could aggregate directly in parallel without decoding, and encoding statistics can help to reduce unnecessary computation. Identified operators construct a pipeline query engine to integrate into an existing open source database, the Apache IoTDB. Remarkably, our systemic evaluations show vast improvements in the efficiency of selective aggregation over existing works.
Rui Kang 0004, Shaoxu Song, Jianmin Wang 0001
ICDE1
2024 Optimizing Time Series Queries with Versions
abstract
We show that the time-series database for industrial IoT data management exhibits intrinsic demands for integrating an automatic version control system, which introduces advanced data semantics and query optimization. In deployed IoT database instances, IoT data managed by an LSM tree is multi-leveled and multi-versioned due to network issues and erroneous IoT readings. For data semantics, each query merges versioned data according to query expressions or data block levels. For query optimization, we find that existing time-series databases relying on write-ahead-logs suboptimally execute data queries, due to their performance bottlenecks in merging numerous versioned data. In this paper, an algebra consisting of version operators addresses the semantics for time-series applications to evaluate and optimize physical query plans. We propose version reducibility as a key feature of executing consistent plans and evaluate the benefits of putting off data merges. We also show the integration of version queries to existing relational databases by translating them to standard SQL based on relational reducibility. Finally, our extended experiments show the effectiveness of optimizing execution plans over versioned data.
Rui Kang 0004, Shaoxu Song
Proc. ACM Manag. Data1
2023 Dynamic Relation Repairing for Knowledge Enhancement
abstract
As the prosperity of unstructured data in networks, knowledge extraction tools have been designed for new knowledges from unstructured data streams. The generated RDF streams by knowledge extraction are always containing much errorous tuples causing inconsistency to knowledge graph engine.To enable the completeness of information from unstructured streams, dynamically repairing the violated RDF tuples is the best way to process. Observed this, we propose dynamic relation repair process to find and eliminate violations in errorous RDF stream. RDF data, arranged as graphs, leads to computation hardness when trying to find constraints and repairing metrics. In this paper, we consider graph repairing process with implicit graph constraints enabling RDF candidates validation and repairing through subgraph matching with the sample of localized subgraphs from graph engine with the same relation labels. We also propose approximated graph matching process through dynamic graph embedding for time efficiency. Cold start problem is also well analyzed to avoid inefficient repairing. Experimental results on real datasets demonstrate that our work can capture and repair violation in RDF streams dynamically and effectively.
Rui Kang 0004, Hongzhi Wang 0001
IEEE Trans. Knowl. Data Eng.1
2022 Conditional Regression Rules
abstract
Mixed data distribution is widely observed, for example, the bird migration data consist of the observed locations of various birds in different years, varying in data distribution. Learning a single regression model over such a mixed data distribution is often ineffective, while manually segmenting the data, e.g., by bird, date or region, for learning individual models is truly labor-intensive. In this paper, we propose to automatically discover the regression models that apply conditionally to only a part of the data, namely conditional regression rules (CRRs), enlightened by the conditional functional dependencies (CFDs) that are FDs hold only in some data. Remarkably, a regression model may apply in different parts of data, e.g., the seasonal migration of birds is similar in different years. To capture the shared regression models, we investigate the inference of CRRs. An algorithm is devised to learn and discover CRRs from data, with the help of CRR inference. Extensive experiments on real-world datasets demonstrate that the discovered conditional regression rules are more effective than the regression models without conditions. In particular, with the inference of CRRs, the number of learned CRRs is significantly reduced without sacrificing rule semantics.
Rui Kang 0004, Shaoxu Song, Chaokun Wang
ICDE1