Marius Dumitru

dblp:82/5967 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 82% Parallel and multicore computing · 18%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › query execution › hardware-accelerated query processing
GPU-accelerated query processing
0.912025
GPU Acceleration of SQL Analytics on Compressed Data · Proc. VLDB Endow. 2025
GPUs and heterogeneous computing
GPU computing
0.912025
GPU Acceleration of SQL Analytics on Compressed Data · Proc. VLDB Endow. 2025
Query processing and optimization
aggregate query processing
0.712023
Cache-Efficient Top-k Aggregation over High Cardinality Large Datasets · Proc. VLDB Endow. 2023
Query processing and optimization › query execution
cache-conscious query processing
0.712023
Cache-Efficient Top-k Aggregation over High Cardinality Large Datasets · Proc. VLDB Endow. 2023
Query processing and optimization › top-k query processing
top-k aggregation
0.712023
Cache-Efficient Top-k Aggregation over High Cardinality Large Datasets · Proc. VLDB Endow. 2023
Parallel and multicore computing › parallel query processing
multicore query processing
0.212023
Cache-Efficient Top-k Aggregation over High Cardinality Large Datasets · Proc. VLDB Endow. 2023

Methods — techniques the papers use, named apart from their topics

compression · 1.7partitioning · 1.3hashing · 1.3adaptive multi-pass algorithm · 1.3
YearPublicationVenuePosition
2025 GPU Acceleration of SQL Analytics on Compressed Data
Zezhou Huang, Krystian Sakowski, Hans Lehnert, Carlo Curino, Matteo Interlandi, Marius Dumitru, Rathijit Sen
Proc. VLDB Endow.7
2023 Cache-Efficient Top-k Aggregation over High Cardinality Large Datasets
abstract
Top-k aggregation queries are widely used in data analytics for summarizing and identifying important groups from large amounts of data. These queries are usually processed by first computing exact aggregates for all groups and then selecting the groups with the top-k aggregate values. However, such an approach can be inefficient for high-cardinality large datasets where intermediate results may not fit within the local cache of multi-core processors leading to excessive data movement. To address this problem, we have developed Zippy, a new cache-conscious aggregation framework that leverages the skew in the data distribution to minimize data movements. This is achieved by designing cache-resident data structures and an adaptive multi-pass algorithm that quickly identifies candidate groups during processing, and performs exact aggregations for these groups. The non-candidate groups are pruned cheaply using efficient hashing and partitioning techniques without performing exact aggregations. We develop techniques to improve robustness over adversarial data distributions and have optimized the framework to reuse computations incrementally for rolling (or paginated) top-k aggregate queries. Our extensive evaluation using both real-world and synthetic datasets demonstrate that Zippy can achieve a median speed-up of more than 3× for monotonic aggregation functions across typical ranges of k values (e.g., 1 to 100) and 1.4× for non-monotonic functions when compared with state-of-the-art cache-conscious aggregation techniques.
Tarique Siddiqui, Vivek R. Narasayya, Marius Dumitru, Surajit Chaudhuri
Proc. VLDB Endow.3