VLDB 2026 Research / reviewers in the wild / expert
Kaitian Hu
dblp:415/9952
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 87% Storage systems · 13% | |
| Databases, data mining, and information retrieval
1 paper |
Data stream processing · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data stream processing
state management |
0.9 | 1 | 2025 | Disaggregated State Management in Apache Flink 2.0 · Proc. VLDB Endow. 2025 |
Data stream processing
stream processing systems |
0.9 | 1 | 2025 | Disaggregated State Management in Apache Flink 2.0 · Proc. VLDB Endow. 2025 |
Distributed systems
fault tolerance |
0.9 | 1 | 2025 | Disaggregated State Management in Apache Flink 2.0 · Proc. VLDB Endow. 2025 |
Distributed systems › fault tolerance
rollback recovery |
0.9 | 1 | 2025 | Disaggregated State Management in Apache Flink 2.0 · Proc. VLDB Endow. 2025 |
Storage systems › distributed storage
disaggregated storage |
0.3 | 1 | 2025 | Disaggregated State Management in Apache Flink 2.0 · Proc. VLDB Endow. 2025 |
Methods — techniques the papers use, named apart from their topics
remote distributed file system · 1.7asynchronous runtime · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Disaggregated State Management in Apache Flink 2.0abstractWe present Apache Flink 2.0, an evolution of the popular stream processing system's architecture that decouples computation from state management. Flink 2.0 relies on a remote distributed file system (DFS) for primary state storage and uses local disks as a secondary cache, with state updates streamed continuously and directly to the DFS. To address the latency implications of remote storage, Flink 2.0 incorporates an asynchronous runtime execution model. Furthermore, Flink 2.0 introduces ForSt, a novel state store featuring a unified file system that enables faster and lightweight checkpointing, recovery, and reconfiguration with minimal intrusion to the existing Flink runtime architecture. Using a comprehensive set of Nexmark benchmarks and a large-scale stateful production workload, we evaluate Flink 2.0's large-state processing, checkpointing, and recovery mechanisms. Our results show significant performance improvements and reduced resource utilization compared to the baseline Flink 1.20 implementation. Specifically, we observe up to 94% reduction in checkpoint duration, up to 49× faster recovery after failures or a rescaling operation, and up to 50% cost savings. Zhaoqian Lan, Yanfei Lei, Han Yin, Kaitian Hu, Paris Carbone, Vasiliki Kalavri |
Proc. VLDB Endow. | 7 |