Yuyuan Kang

dblp:274/6467 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0003-4734-4882ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 95% Performance modeling and evaluation · 5%
Databases, data mining, and information retrieval
1 paper
Indexing and storage engines · 70% Data stream processing · 30%
Computer networks
1 paper
Network measurement and analytics · 100%

Topics — the 9 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › networked storage
storage networking
0.912025
Understanding and Profiling NVMe-over-TCP Using ntprof · NSDI 2025
Indexing and storage engines
LSM-tree
0.812024
Time Series Representation for Visualization in Apache IoTDB · Proc. ACM Manag. Data 2024
Indexing and storage engines
time series storage
0.812024
Time Series Representation for Visualization in Apache IoTDB · Proc. ACM Manag. Data 2024
Storage systems › key-value storage
compaction
0.612022
Separation or Not: On Handing Out-of-Order Time-Series Data in Leveled LSM-Tree · ICDE 2022
Storage systems
key-value storage
0.612022
Separation or Not: On Handing Out-of-Order Time-Series Data in Leveled LSM-Tree · ICDE 2022
Storage systems › key-value storage
LSM-tree
0.612022
Separation or Not: On Handing Out-of-Order Time-Series Data in Leveled LSM-Tree · ICDE 2022
Storage systems › flash and SSD › flash memory management › garbage collection
write amplification
0.612022
Separation or Not: On Handing Out-of-Order Time-Series Data in Leveled LSM-Tree · ICDE 2022
Performance modeling and evaluation
analytical modeling
0.212022
Separation or Not: On Handing Out-of-Order Time-Series Data in Leveled LSM-Tree · ICDE 2022
Storage systems › data management › database storage
time series storage
0.212022
Separation or Not: On Handing Out-of-Order Time-Series Data in Leveled LSM-Tree · ICDE 2022

Methods — techniques the papers use, named apart from their topics

intra-chunk indexing · 0.8chunk merging · 0.8LSM-tree · 0.8workload analysis · 0.6analytical modeling · 0.6
YearPublicationVenuePosition
2025 Understanding and Profiling NVMe-over-TCP Using ntprof
Yuyuan Kang, Ming Liu 0027
NSDI1
2024 Time Series Representation for Visualization in Apache IoTDB
abstract
When analyzing time series, often interactively, the analysts frequently demand to visualize instantly large-scale data stored in databases. M4 visualization selects the first, last, bottom and top data points in each pixel column to ensure pixel-perfectness of the two-color line chart visualization. While M4 already shows its preciseness of encasing time series in different scales into a fixed size of pixels, how to efficiently support M4 representation in a time series native database is still absent. It is worth noting that, to enable fast writes, the commodity time series database systems, such as Apache IoTDB or InfluxDB, employ LSM-Tree based storage. That is, a time series is segmented and stored in a number of chunks, with possibly out-of-order arrivals, i.e., disordered on timestamps. To implement M4, a natural idea is to merge online the chunks as a whole series, with costly merge sort on timestamps, and then perform M4 representation as in relational databases. In this study, we propose a novel chunk merge free approach called M4-LSM to accelerate M4 representation and visualization. In particular, we utilize the metadata of chunks to prune and avoid the costly merging of any chunk. Moreover, intra-chunk indexing and pruning are enabled for efficiently accessing the representation points, referring to the special properties of time series. Remarkably, the time series database native operator M4-LSM has been implemented in Apache IoTDB, an open-source time series database, and deployed in companies across various industries. In the experiments over real-world datasets, the proposed M4-LSM operator demonstrates high efficiency without sacrificing preciseness.
Lei Rui, Xiangdong Huang 0001, Shaoxu Song, Yuyuan Kang, Chen Wang 0018, Jianmin Wang 0001
Proc. ACM Manag. Data4
2022 Separation or Not: On Handing Out-of-Order Time-Series Data in Leveled LSM-Tree
abstract
LSM-Tree is widely adopted for storing time-series data in Internet of Things. According to conventional policy (denoted by$\pi_{c}$), when writing, the data will first be buffered in MemTable in memory. When it is full, the data will be written to the disk to form SSTables. Compaction is triggered to sort the data in each layer of the LSM-Tree on the disk. However, the arrival of data can be unordered due to reasons such as transition delay. Apache IoTDB uses in-order and out-of-order MemTables to separately buffer the in-order and out-of-order data to accelerate queries, namely the separation policy (denoted by$\pi_{s}$). However, given a specific space of memory budget to buffer the data, write amplification (WA) of the leveled LSM-Tree will be influenced by$\pi_{s}$. Whether the influence by separation is positive or negative, and how intense WA is influenced, depend on the properties of workloads and the capacity of the in-order and out-of-order MemTables. It is highly demanded to build robust models for estimating the expected amount of data rewritten in each compaction, and predicting the WA under$\pi_{c}$and$\pi_{s}$. Note that as an industrial paper, rather than proposing novel techniques for research problems, we focus on the practice of whether separating or not for lower write amplification. Experiments on synthetic and real-world datasets show that the models for estimating WA are accurate under various delay distributions. In addition, based on the estimation models, we implement an analyzer module in the open-source Apache IoTDB, for choosing the policy with lower WA. We apply the method in the use case of our industrial partner, a service provider of engineering machinery. The use case verifies the effectiveness of deciding whether separation or not by WA estimation.
Yuyuan Kang, Xiangdong Huang 0001, Shaoxu Song, Lingzhe Zhang, Jialin Qiao, Chen Wang 0018, Jianmin Wang 0001, Julian Feinauer
ICDE1
2020 Heterogeneous Replicas for Multi-dimensional Data Management
Jialin Qiao, Yuyuan Kang, Xiangdong Huang 0001, Lei Rui, Jianmin Wang 0001, Philip S. Yu
DASFAA (1)2