Global Hash Tables Strike Back! An Analysis of Parallel GROUP BY Aggregation

vldb26-1080 · Experiment, Analysis & Benchmark (EA&B) · Daniel Xue, Ryan Marcus
Abstract

Efficiently computing group aggregations (i.e., GROUP BY) on modern architectures is critical for analytic database systems. Today, hash-based methods predominantly use a partitioned approach, in which incoming data is partitioned by key so that every row for a particular key is sent to the same partition. In this paper, we revisit a simpler strategy: a fully concurrent aggregation technique using a shared hash table. While approaches using general-purpose concurrent hash tables have generally been found to perform worse than partitioning-based approaches, we argue that the key ingredient is customizing the concurrent hash table for the specific task of group aggregation. Through experiments on synthetic workloads (varying key cardinality, skew, and thread count), we demonstrate that in morsel-driven systems, a purpose-built concurrent hash table can match or surpass partitioning-based techniques. We also analyze the operational characteristics of both techniques, including resizing costs and memory pressure. In the process, we derive practical guidelines for database implementers. Overall, our analysis indicates that fully concurrent group aggregation is a viable alternative to partitioning.

Assigned reviewers

No reviewers assigned yet.

Candidates from the panel ranked by taxonomy affinity

#ReviewerMatchLoadWhy