VLDB 2026 Research / reviewers in the wild / expert
Bruno Turcksin
dblp:139/0707
· DBLP profile ↗
6ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0001-5954-6313ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 since 2021Theory of computation · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The ArborX Library: Version 2.0abstractThis article provides an overview of the 2.0 release of the ArborX library, a performance portable geometric search library based on Kokkos. We describe the major changes in ArborX 2.0 including a new interface for the library to support a wider range of user problems, new search data structures (brute force and distributed), support for user functions to be executed on the results (callbacks), and an expanded set of the supported algorithms (ray tracing and clustering). Andrey Prokopenko, Daniel Arndt 0003, Damien Lebrun-Grandié, Bruno Turcksin |
ACM Trans. Math. Softw. | 4 |
| 2022 | Kokkos 3: Programming Model Extensions for the Exascale EraabstractAs the push towards exascale hardware has increased the diversity of system architectures, performance portability has become a critical aspect for scientific software. We describe the Kokkos Performance Portable Programming Model that allows developers to write single source applications for diverse high-performance computing architectures. Kokkos provides key abstractions for both the compute and memory hierarchy of modern hardware. We describe the novel abstractions that have been added to Kokkos version 3 such as hierarchical parallelism, containers, task graphs, and arbitrary-sized atomic operations to prepare for exascale era architectures. We demonstrate the performance of these new features with reproducible benchmarks on CPUs and GPUs. Christian Trott, Damien Lebrun-Grandié, Daniel Arndt 0003, Jan Ciesko, Vinh Q. Dang, Nathan D. Ellingwood, Rahulkumar Gayatri, Evan Harvey, Daisy S. Hollman, Daniel Ibanez, Nevin Liber, Jonathan R. Madsen, Jeff Miles, David Poliakoff, Amy Powell, Sivasankaran Rajamanickam, Mikael Simberg, Daniel Sunderland, Bruno Turcksin, Jeremiah J. Wilke |
IEEE Trans. Parallel Distributed Syst. | 19 |
| 2021 | ArborX: A Performance Portable Geometric Search LibraryabstractSearching for geometric objects that are close in space is a fundamental component of many applications. The performance of search algorithms comes to the forefront as the size of a problem increases both in terms of total object count as well as in the total number of search queries performed. Scientific applications requiring modern leadership-class supercomputers also pose an additional requirement of performance portability, i.e., being able to efficiently utilize a variety of hardware architectures. In this article, we introduce a new open-source C++ search library, ArborX, which we have designed for modern supercomputing architectures. We examine scalable search algorithms with a focus on performance, including a highly efficient parallel bounding volume hierarchy implementation, and propose a flexible interface making it easy to integrate with existing applications. We demonstrate the performance portability of ArborX on multi-core CPUs and GPUs and compare it to the state-of-the-art libraries such as Boost.Geometry.Index and nanoflann. Damien Lebrun-Grandié, Andrey Prokopenko, Bruno Turcksin, Stuart R. Slattery |
ACM Trans. Math. Softw. | 3 |
| 2020 | A parallel strategy for density functional theory computations on accelerated nodes
Massimiliano Lupo Pasini, Bruno Turcksin, Wenjun Ge, Jean-Luc Fattebert |
Parallel Comput. | 2 |
| 2019 | Fast, scalable and accurate finite-element based ab initio calculations using mixed precision computing: 46 PFLOPS simulation of a metallic dislocation systemabstractAccurate large-scale first principles calculations based on density functional theory (DFT) in metallic systems are prohibitively expensive due to the asymptotic cubic scaling computational complexity with number of electrons. Using algorithmic advances in employing finite-element discretization for DFT (DFT-FE) in conjunction with efficient computational methodologies and mixed precision strategies, we delay the onset of this cubic scaling by significantly reducing the computational prefactor while increasing the arithmetic intensity and lowering the data movement costs. This has enabled fast, accurate and massively parallel DFT calculations on large-scale metallic systems on both many-core and heterogeneous architectures, with time-to-solution being an order of magnitude faster than state-of-the-art plane-wave DFT codes. We demonstrate an unprecedented sustained performance of 46 PFLOPS (27.8% peak FP64 performance) on a dislocation system in Magnesium containing 105,080 electrons using 3,800 GPU nodes of Summit supercomputer, which is the highest performance to-date among DFT codes. Sambit Das, Phani Motamarri, Vikram Gavini, Bruno Turcksin, Ying Wai Li, Brent Leback |
SC | 4 |
| 2016 | WorkStream - A Design Pattern for Multicore-Enabled Finite Element ComputationsabstractMany operations that need to be performed in modern finite element codes can be described as an operation that needs to be done independently on every cell, followed by a reduction of these local results into a global data structure. For example, matrix assembly, estimating discretization errors, or converting nodal values into data structures that can be output in visualization file formats all fall into this class of operations. Using this realization, we identify a software design pattern that we call WorkStream and that can be used to model such operations and enables the use of multicore shared memory parallel processing. We also describe in detail how this design pattern can be efficiently implemented, and we provide numerical scalability results from its use in the deal .II software library. Bruno Turcksin, Martin Kronbichler 0002, Wolfgang Bangerth |
ACM Trans. Math. Softw. | 1 |