EDBT 2026 Demo / reviewers in the wild / expert
Tim Besard
dblp:170/6783
· DBLP profile ↗
3ranked-venue papers
1as first author
2since 2021 · last 2026
0000-0001-7826-8021ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 53% GPUs and heterogeneous computing · 47% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 77% Programming languages and type systems · 23% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU programming |
2.0 | 3 | 2026 | Flexible Performant Tensor Contractions on GPUs · IEEE Trans. Parallel Distributed Syst. 2026 Flexible Performant GEMM Kernels on GPUs · IEEE Trans. Parallel Distributed Syst. 2022 Effective Extensible Programming: Unleashing Julia on GPUs · IEEE Trans. Parallel Distributed Syst. 2019 |
High-performance computing
performance optimization at scale |
1.6 | 2 | 2026 | Flexible Performant Tensor Contractions on GPUs · IEEE Trans. Parallel Distributed Syst. 2026 Flexible Performant GEMM Kernels on GPUs · IEEE Trans. Parallel Distributed Syst. 2022 |
High-performance computing › tensor computation
tensor contractions |
1.0 | 1 | 2026 | Flexible Performant Tensor Contractions on GPUs · IEEE Trans. Parallel Distributed Syst. 2026 |
High-performance computing › numerical linear algebra
matrix multiplication |
0.6 | 1 | 2022 | Flexible Performant GEMM Kernels on GPUs · IEEE Trans. Parallel Distributed Syst. 2022 |
GPUs and heterogeneous computing › GPU computing
tensor cores |
0.5 | 2 | 2026 | Flexible Performant Tensor Contractions on GPUs · IEEE Trans. Parallel Distributed Syst. 2026 Flexible Performant GEMM Kernels on GPUs · IEEE Trans. Parallel Distributed Syst. 2022 |
Compilers and program optimization › accelerator compilation
GPU compiler |
0.4 | 1 | 2019 | Effective Extensible Programming: Unleashing Julia on GPUs · IEEE Trans. Parallel Distributed Syst. 2019 |
GPUs and heterogeneous computing › GPU programming
high-level language compilation |
0.4 | 1 | 2019 | Effective Extensible Programming: Unleashing Julia on GPUs · IEEE Trans. Parallel Distributed Syst. 2019 |
Methods — techniques the papers use, named apart from their topics
julia programming · 1.6GEMM-like kernel adaptation · 1.0compiler infrastructure · 0.8abstraction and interface design · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Flexible Performant Tensor Contractions on GPUsabstractTensor contractions extend the concept of the General Matrix Multiplication (GEMM) to high-dimensional spaces. They enable sophisticated computations in various scientific disciplines. Graphics Processing Units (GPUs) are commonly used to accelerate tensor contraction algorithms due to their inherent parallelisability. NVIDIA's cuTENSOR stands as a state-of-the-art library for GPU-based tensor contractions. However, its lack of flexibility limits researchers in tailoring contraction kernels to their specific research needs. This paper presents a novel and flexible implementation of the GEMM-like Tensor Tensor (GETT) multiplication algorithm for tensor contractions in Julia. By repurposing and adapting components of GemmKernels.jl, a versatile library offering customisable and high-performance GEMM kernels for CUDA-enabled GPUs, we construct GEMM-like kernels that cater to the unique requirements of tensor contractions. Despite being entirely written in high-level Julia code and not yet exploiting a range of modern CUDA hardware features, the average performance of our library on standard tensor contractions compares favourably to cuTENSOR's hand-optimised implementations, with outliers in both directions (faster and slower). When flexibility is needed, e.g. to fuse arbitrary elementwise operations into kernels, our library performs up to an order of magnitude faster than cuTENSOR, even on recent, data centre-grade devices such as the RTX 6000 Ada. Thomas Faingnaert, Ward Vermeulen, Tim Besard, Bjorn De Sutter |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | Flexible Performant GEMM Kernels on GPUsabstractGeneral Matrix Multiplication or GEMM kernels take centre place in high performance computing and machine learning. Recent NVIDIA GPUs include GEMM accelerators, such as NVIDIA’s Tensor Cores. Their exploitation is hampered by the two-language problem: it requires either low-level programming which implies low programmer productivity or using libraries that only offer a limited set of components. Because rephrasing algorithms in terms of established components often introduces overhead, the libraries’ lack of flexibility limits the freedom to explore new algorithms. Researchers using GEMMs can hence not enjoy programming productivity, high performance, and research flexibility at once. In this paper we solve this problem. We present three sets of abstractions and interfaces to program GEMMs within the scientific Julia programming language. The interfaces and abstractions are co-designed for researchers’ needs and Julia’s features to achieve sufficient separation of concerns and flexibility to easily extend basic GEMMs in many different ways without paying a performance price. Comparing our GEMMs to state-of-the-art libraries cuBLAS and CUTLASS, we demonstrate that our performance is in the same ballpark of the libraries, and in some cases even exceeds it, without having to write a single line of code in CUDA C++ or assembly, and without facing flexibility limitations. Thomas Faingnaert, Tim Besard, Bjorn De Sutter |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | Effective Extensible Programming: Unleashing Julia on GPUsabstractGPUs and other accelerators are popular devices for accelerating compute-intensive, parallelizable applications. However, programming these devices is a difficult task. Writing efficient device code is challenging, and is typically done in a low-level programming language. High-level languages are rarely supported, or do not integrate with the rest of the high-level language ecosystem. To overcome this, we propose compiler infrastructure to efficiently add support for new hardware or environments to an existing programming language. We evaluate our approach by adding support for NVIDIA GPUs to the Julia programming language. By integrating with the existing compiler, we significantly lower the cost to implement and maintain the new compiler, and facilitate reuse of existing application code. Moreover, use of the high-level Julia programming language enables new and dynamic approaches for GPU programming. This greatly improves programmer productivity, while maintaining application performance similar to that of the official NVIDIA CUDA toolkit. Tim Besard, Christophe Foket, Bjorn De Sutter |
IEEE Trans. Parallel Distributed Syst. | 1 |