VLDB 2026 Research / reviewers in the wild / expert
Jiahong Shen
dblp:281/5585
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2023
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Storage systems · 62% Performance modeling and evaluation · 38% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
benchmarking |
0.7 | 1 | 2023 | An Empirical Evaluation of Columnar Storage Formats · Proc. VLDB Endow. 2023 |
Storage systems › data management › database storage
columnar storage |
0.7 | 1 | 2023 | An Empirical Evaluation of Columnar Storage Formats · Proc. VLDB Endow. 2023 |
Storage systems
data compression |
0.2 | 1 | 2023 | An Empirical Evaluation of Columnar Storage Formats · Proc. VLDB Endow. 2023 |
Storage systems › data compression
dictionary encoding |
0.2 | 1 | 2023 | An Empirical Evaluation of Columnar Storage Formats · Proc. VLDB Endow. 2023 |
Methods — techniques the papers use, named apart from their topics
benchmarking · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | An Empirical Evaluation of Columnar Storage FormatsabstractColumnar storage is a core component of a modern data analytics system. Although many database management systems (DBMSs) have proprietary storage formats, most provide extensive support to open-source storage formats such as Parquet and ORC to facilitate cross-platform data sharing. But these formats were developed over a decade ago, in the early 2010s, for the Hadoop ecosystem. Since then, both the hardware and workload landscapes have changed. In this paper, we revisit the most widely adopted open-source columnar storage formats (Parquet and ORC) with a deep dive into their internals. We designed a benchmark to stress-test the formats' performance and space efficiency under different workload configurations. From our comprehensive evaluation of Parquet and ORC, we identify design decisions advantageous with modern hardware and real-world data distributions. These include using dictionary encoding by default, favoring decoding speed over compression ratio for integer encoding algorithms, making block compression optional, and embedding finer-grained auxiliary data structures. We also point out the inefficiencies in the format designs when handling common machine learning workloads and using GPUs for decoding. Our analysis identified important considerations that may guide future formats to better fit modern technology trends. Xinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo, Wes McKinney, Huanchen Zhang |
Proc. VLDB Endow. | 3 |