Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ivan Ilic

dblp:251/0103 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 33% Interconnection networks and networks-on-chip · 33% Performance modeling and evaluation · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
benchmarking
0.612022
Evaluating Multi-GPU Sorting with Modern Interconnects · SIGMOD Conference 2022
Interconnection networks and networks-on-chip › interconnect architecture
GPU interconnect
0.612022
Evaluating Multi-GPU Sorting with Modern Interconnects · SIGMOD Conference 2022
GPUs and heterogeneous computing
multi-GPU computing
0.612022
Evaluating Multi-GPU Sorting with Modern Interconnects · SIGMOD Conference 2022

Methods — techniques the papers use, named apart from their topics

radix sort · 0.6p2p sort · 0.6HET sort · 0.6
YearPublicationVenuePosition
2022 Evaluating Multi-GPU Sorting with Modern Interconnects
abstract
GPUs have become a mainstream accelerator for database operations such as sorting. Most GPU sorting algorithms are single-GPU approaches. They neither harness the full computational power nor exploit the high-bandwidth P2P interconnects of modern multi-GPU platforms. The latest NVLink 2.0 and NVLink 3.0-based NVSwitch interconnects promise unparalleled multi-GPU acceleration. So far, multi-GPU sorting has only been evaluated on systems with PCIe 3.0. In this paper, we analyze serial, parallel, and bidirectional data transfer rates to, from, and between multiple GPUs on systems with PCIe 3.0/4.0, NVLink 2.0/3.0, and NVSwitch. We measure up to 35x higher parallel P2P throughput with NVLink 3.0-based NVSwitch over PCIe 3.0. To study GPU-accelerated sorting on today's hardware, we implement a P2P-based GPU-only (P2P sort) and a heterogeneous (HET sort) multi-GPU sorting algorithm and evaluate them on three modern platforms. We observe speedups over state-of-the-art parallel CPU radix sort of up to 14x for P2P sort and 9x for HET sort. On systems with fast P2P interconnects, P2P sort outperforms HET sort up to 1.65x. Finally, we show that overlapping GPU copy/compute operations does not mitigate the transfer bottleneck when sorting large out-of-core data.
Tobias Maltenberger, Ivan Ilic, Ilin Tolovski, Tilmann Rabl
SIGMOD Conference2