Marko Kabic

dblp:247/9619 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0006-5567-0580ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
GPUs and heterogeneous computing · 33% Performance modeling and evaluation · 23% Distributed systems · 15%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems › communication optimization
communication-optimal algorithms
0.922021
On the parallel I/O optimality of linear algebra kernels: near-optimal matrix factorizations · SC 2021
Red-blue pebbling revisited: near optimal parallel matrix-matrix multiplication · SC 2019
Data integration and cleaning
heterogeneous query processing
0.912025
Maximus: A Modular Accelerated Query Engine for Data Analytics on Heterogeneous Systems · Proc. ACM Manag. Data 2025
GPUs and heterogeneous computing › GPU-accelerated data processing
GPU-accelerated data analytics
0.912025
Powerful GPUs or Fast Interconnects: Analyzing Relational Workloads on Modern GPUs · Proc. VLDB Endow. 2025
GPUs and heterogeneous computing › GPU query processing
GPU query engine
0.912025
Maximus: A Modular Accelerated Query Engine for Data Analytics on Heterogeneous Systems · Proc. ACM Manag. Data 2025
Hardware accelerators and domain-specific architectures › database accelerator
query accelerator
0.912025
Maximus: A Modular Accelerated Query Engine for Data Analytics on Heterogeneous Systems · Proc. ACM Manag. Data 2025
High-performance computing › numerical linear algebra
matrix factorization
0.512021
On the parallel I/O optimality of linear algebra kernels: near-optimal matrix factorizations · SC 2021
Parallel and multicore computing › parallel algorithms › parallel matrix algorithms
parallel matrix multiplication
0.412019
Red-blue pebbling revisited: near optimal parallel matrix-matrix multiplication · SC 2019
Performance modeling and evaluation
benchmarking
0.312025
Powerful GPUs or Fast Interconnects: Analyzing Relational Workloads on Modern GPUs · Proc. VLDB Endow. 2025

Methods — techniques the papers use, named apart from their topics

substrait · 1.7kernel fusion · 1.7cost modeling · 0.9i/o lower bound analysis · 0.5red-blue pebble game · 0.4i/o optimality analysis · 0.4
YearPublicationVenuePosition
2026 End-to-End Declarative Data Analytics: Co-designing Engines, Interfaces, and Cloud Infrastructure
Pinghe Li, Tom Kuchler, Marko Kabic, Tobias Stocker, Gustavo Alonso, Ana Klimovic
CIDR3
2025 Maximus: A Modular Accelerated Query Engine for Data Analytics on Heterogeneous Systems
abstract
Several trends are changing the underlying fabric for data processing in fundamental ways. On the hardware side, machines are becoming heterogeneous with smart NICs, TPUs, DPUs, etc., but specially with GPUs taking a more dominant role. On the software side, the diversity in workloads, data sources, and data formats has given rise to the notion of composable data processing where the data is processed across a variety of engines and platforms. Finally, on the infrastructure side, different storage types, disaggregated storage, disaggregated memory, networking, and interconnects are all rapidly evolving, which demands a degree of customization to optimize data movement well beyond established techniques. To tackle these challenges, in this paper, we present Maximus, a modular data processing engine that embraces heterogeneity from the ground up. Maximus can run queries on CPUs and GPUs, can split execution between CPUs and GPUs, import and export data in a variety of formats, interact with a wide range of query engines through Substrait, and efficiently manage the execution of complex data processing pipelines. Through the concept of operator-level integration, Maximus can use operators from third-party engines and achieve even better performance with these operators than when they are used with their native engines. The current version of Maximus supports all TPC-H queries on both the GPU and the CPU and optimizes the data movement and kernel execution between them, enabling the overlap of communication and computation to achieve performance comparable to that of the best systems available, but with a far higher degree of completeness and flexibility.
Marko Kabic, Shriram Chandran, Gustavo Alonso
Proc. ACM Manag. Data1
2025 Powerful GPUs or Fast Interconnects: Analyzing Relational Workloads on Modern GPUs
abstract
In this study we explore the impact of different combinations of GPU models (RTX3090, A100, H100, GraceHoppers - GH200) and interconnects (PCIe 3.0, PCIe 4.0, PCIe 5.0, and NVLink 4.0) on various relational data analytics workloads (TPC-H, H2O-G, ClickBench). We present MaxBench, a comprehensive framework designed for benchmarking, profiling, and modeling these workloads on GPUs. Beyond delivering detailed performance metrics, MaxBench estimates query execution performance using a novel cost model. With this model, we move beyond traditional metrics such as arithmetic intensity and GFlop/s and suggest using instead the notions of characteristic query complexity and characteristic GPU efficiency , as more suitable metrics for data analytics workloads. We conduct an extensive experimental analysis with MaxBench across different combinations of GPU models and interconnects on various data analytics workloads. The insights from this analysis reveal the trade-offs between GPU computing capacity and interconnect bandwidth on query processing. Using this cost model, we also examine future trends by investigating how enhancements in interconnect bandwidth or GPU efficiency would affect performance in the future.
Marko Kabic, Bowen Wu 0003, Jonas Dann, Gustavo Alonso
Proc. VLDB Endow.1
2021 On the parallel I/O optimality of linear algebra kernels: near-optimal matrix factorizations
Grzegorz Kwasniewski, Marko Kabic, Tal Ben-Nun, Alexandros Nikolaos Ziogas, Jens Eirik Saethre, André Gaillard, Timo Schneider, Maciej Besta, Anton Kozhevnikov, Joost VandeVondele, Torsten Hoefler
SC2
2019 Red-blue pebbling revisited: near optimal parallel matrix-matrix multiplication
abstract
We propose COSMA: a parallel matrix-matrix multiplication algorithm that is near communication-optimal for all combinations of matrix dimensions, processor counts, and memory sizes. The key idea behind COSMA is to derive an optimal (up to a factor of 0.03% for 10MB of fast memory) sequential schedule and then parallelize it, preserving I/O optimality. To achieve this, we use the red-blue pebble game to precisely model MMM dependencies and derive a constructive and tight sequential and parallel I/O lower bound proofs. Compared to 2D or 3D algorithms, which fix processor decomposition upfront and then map it to the matrix dimensions, it reduces communication volume by up to √ times. COSMA outperforms the established ScaLAPACK, CARMA, and CTF algorithms in all scenarios up to 12.8x (2.2x on average), achieving up to 88% of Piz Daint's peak performance. Our work does not require any hand tuning and is maintained as an open source implementation.
Grzegorz Kwasniewski, Marko Kabic, Maciej Besta, Joost VandeVondele, Raffaele Solcà, Torsten Hoefler
SC2