Ward Vermeulen

dblp:428/1093 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0001-9184-8087ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 61% GPUs and heterogeneous computing · 39%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU programming
1.012026
Flexible Performant Tensor Contractions on GPUs · IEEE Trans. Parallel Distributed Syst. 2026
High-performance computing
performance optimization at scale
1.012026
Flexible Performant Tensor Contractions on GPUs · IEEE Trans. Parallel Distributed Syst. 2026
High-performance computing › tensor computation
tensor contractions
1.012026
Flexible Performant Tensor Contractions on GPUs · IEEE Trans. Parallel Distributed Syst. 2026
GPUs and heterogeneous computing › GPU computing
tensor cores
0.312026
Flexible Performant Tensor Contractions on GPUs · IEEE Trans. Parallel Distributed Syst. 2026

Methods — techniques the papers use, named apart from their topics

julia programming · 1.0GEMM-like kernel adaptation · 1.0
YearPublicationVenuePosition
2026 Flexible Performant Tensor Contractions on GPUs
abstract
Tensor contractions extend the concept of the General Matrix Multiplication (GEMM) to high-dimensional spaces. They enable sophisticated computations in various scientific disciplines. Graphics Processing Units (GPUs) are commonly used to accelerate tensor contraction algorithms due to their inherent parallelisability. NVIDIA's cuTENSOR stands as a state-of-the-art library for GPU-based tensor contractions. However, its lack of flexibility limits researchers in tailoring contraction kernels to their specific research needs. This paper presents a novel and flexible implementation of the GEMM-like Tensor Tensor (GETT) multiplication algorithm for tensor contractions in Julia. By repurposing and adapting components of GemmKernels.jl, a versatile library offering customisable and high-performance GEMM kernels for CUDA-enabled GPUs, we construct GEMM-like kernels that cater to the unique requirements of tensor contractions. Despite being entirely written in high-level Julia code and not yet exploiting a range of modern CUDA hardware features, the average performance of our library on standard tensor contractions compares favourably to cuTENSOR's hand-optimised implementations, with outliers in both directions (faster and slower). When flexibility is needed, e.g. to fuse arbitrary elementwise operations into kernels, our library performs up to an order of magnitude faster than cuTENSOR, even on recent, data centre-grade devices such as the RTX 6000 Ada.
Thomas Faingnaert, Ward Vermeulen, Tim Besard, Bjorn De Sutter
IEEE Trans. Parallel Distributed Syst.2