Kaitian Hu

dblp:415/9952 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 87% Storage systems · 13%
Databases, data mining, and information retrieval
1 paper
Data stream processing · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data stream processing
state management
0.912025
Disaggregated State Management in Apache Flink 2.0 · Proc. VLDB Endow. 2025
Data stream processing
stream processing systems
0.912025
Disaggregated State Management in Apache Flink 2.0 · Proc. VLDB Endow. 2025
Distributed systems
fault tolerance
0.912025
Disaggregated State Management in Apache Flink 2.0 · Proc. VLDB Endow. 2025
Distributed systems › fault tolerance
rollback recovery
0.912025
Disaggregated State Management in Apache Flink 2.0 · Proc. VLDB Endow. 2025
Storage systems › distributed storage
disaggregated storage
0.312025
Disaggregated State Management in Apache Flink 2.0 · Proc. VLDB Endow. 2025

Methods — techniques the papers use, named apart from their topics

remote distributed file system · 1.7asynchronous runtime · 1.7
YearPublicationVenuePosition
2025 Disaggregated State Management in Apache Flink 2.0
abstract
We present Apache Flink 2.0, an evolution of the popular stream processing system's architecture that decouples computation from state management. Flink 2.0 relies on a remote distributed file system (DFS) for primary state storage and uses local disks as a secondary cache, with state updates streamed continuously and directly to the DFS. To address the latency implications of remote storage, Flink 2.0 incorporates an asynchronous runtime execution model. Furthermore, Flink 2.0 introduces ForSt, a novel state store featuring a unified file system that enables faster and lightweight checkpointing, recovery, and reconfiguration with minimal intrusion to the existing Flink runtime architecture. Using a comprehensive set of Nexmark benchmarks and a large-scale stateful production workload, we evaluate Flink 2.0's large-state processing, checkpointing, and recovery mechanisms. Our results show significant performance improvements and reduced resource utilization compared to the baseline Flink 1.20 implementation. Specifically, we observe up to 94% reduction in checkpoint duration, up to 49× faster recovery after failures or a rescaling operation, and up to 50% cost savings.
Zhaoqian Lan, Yanfei Lei, Han Yin, Kaitian Hu, Paris Carbone, Vasiliki Kalavri
Proc. VLDB Endow.7