Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Thomas Faingnaert

dblp:275/3481 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2026
0000-0002-6420-6476ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 61% GPUs and heterogeneous computing · 39%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU programming
1.622026
Flexible Performant Tensor Contractions on GPUs · IEEE Trans. Parallel Distributed Syst. 2026
Flexible Performant GEMM Kernels on GPUs · IEEE Trans. Parallel Distributed Syst. 2022
High-performance computing
performance optimization at scale
1.622026
Flexible Performant Tensor Contractions on GPUs · IEEE Trans. Parallel Distributed Syst. 2026
Flexible Performant GEMM Kernels on GPUs · IEEE Trans. Parallel Distributed Syst. 2022
High-performance computing › tensor computation
tensor contractions
1.012026
Flexible Performant Tensor Contractions on GPUs · IEEE Trans. Parallel Distributed Syst. 2026
High-performance computing › numerical linear algebra
matrix multiplication
0.612022
Flexible Performant GEMM Kernels on GPUs · IEEE Trans. Parallel Distributed Syst. 2022
GPUs and heterogeneous computing › GPU computing
tensor cores
0.522026
Flexible Performant Tensor Contractions on GPUs · IEEE Trans. Parallel Distributed Syst. 2026
Flexible Performant GEMM Kernels on GPUs · IEEE Trans. Parallel Distributed Syst. 2022

Methods — techniques the papers use, named apart from their topics

julia programming · 1.6GEMM-like kernel adaptation · 1.0abstraction and interface design · 0.6
YearPublicationVenuePosition
2026 Flexible Performant Tensor Contractions on GPUs
abstract
Tensor contractions extend the concept of the General Matrix Multiplication (GEMM) to high-dimensional spaces. They enable sophisticated computations in various scientific disciplines. Graphics Processing Units (GPUs) are commonly used to accelerate tensor contraction algorithms due to their inherent parallelisability. NVIDIA's cuTENSOR stands as a state-of-the-art library for GPU-based tensor contractions. However, its lack of flexibility limits researchers in tailoring contraction kernels to their specific research needs. This paper presents a novel and flexible implementation of the GEMM-like Tensor Tensor (GETT) multiplication algorithm for tensor contractions in Julia. By repurposing and adapting components of GemmKernels.jl, a versatile library offering customisable and high-performance GEMM kernels for CUDA-enabled GPUs, we construct GEMM-like kernels that cater to the unique requirements of tensor contractions. Despite being entirely written in high-level Julia code and not yet exploiting a range of modern CUDA hardware features, the average performance of our library on standard tensor contractions compares favourably to cuTENSOR's hand-optimised implementations, with outliers in both directions (faster and slower). When flexibility is needed, e.g. to fuse arbitrary elementwise operations into kernels, our library performs up to an order of magnitude faster than cuTENSOR, even on recent, data centre-grade devices such as the RTX 6000 Ada.
Thomas Faingnaert, Ward Vermeulen, Tim Besard, Bjorn De Sutter
IEEE Trans. Parallel Distributed Syst.1
2022 Flexible Performant GEMM Kernels on GPUs
abstract
General Matrix Multiplication or GEMM kernels take centre place in high performance computing and machine learning. Recent NVIDIA GPUs include GEMM accelerators, such as NVIDIA’s Tensor Cores. Their exploitation is hampered by the two-language problem: it requires either low-level programming which implies low programmer productivity or using libraries that only offer a limited set of components. Because rephrasing algorithms in terms of established components often introduces overhead, the libraries’ lack of flexibility limits the freedom to explore new algorithms. Researchers using GEMMs can hence not enjoy programming productivity, high performance, and research flexibility at once. In this paper we solve this problem. We present three sets of abstractions and interfaces to program GEMMs within the scientific Julia programming language. The interfaces and abstractions are co-designed for researchers’ needs and Julia’s features to achieve sufficient separation of concerns and flexibility to easily extend basic GEMMs in many different ways without paying a performance price. Comparing our GEMMs to state-of-the-art libraries cuBLAS and CUTLASS, we demonstrate that our performance is in the same ballpark of the libraries, and in some cases even exceeds it, without having to write a single line of code in CUDA C++ or assembly, and without facing flexibility limitations.
Thomas Faingnaert, Tim Besard, Bjorn De Sutter
IEEE Trans. Parallel Distributed Syst.1