Daniel Arndt 0003

dblp:85/11357-3 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0001-8773-4901ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 75% High-performance computing · 25%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel programming models › structured parallelism
hierarchical parallelism
0.612022
Kokkos 3: Programming Model Extensions for the Exascale Era · IEEE Trans. Parallel Distributed Syst. 2022
Parallel and multicore computing
parallel programming models
0.612022
Kokkos 3: Programming Model Extensions for the Exascale Era · IEEE Trans. Parallel Distributed Syst. 2022
High-performance computing › performance engineering
performance portability
0.612022
Kokkos 3: Programming Model Extensions for the Exascale Era · IEEE Trans. Parallel Distributed Syst. 2022
Parallel and multicore computing › parallel programming models
portable programming models
0.612022
Kokkos 3: Programming Model Extensions for the Exascale Era · IEEE Trans. Parallel Distributed Syst. 2022

Methods — techniques the papers use, named apart from their topics

benchmarking · 0.6
YearPublicationVenuePosition
2025 The ArborX Library: Version 2.0
abstract
This article provides an overview of the 2.0 release of the ArborX library, a performance portable geometric search library based on Kokkos. We describe the major changes in ArborX 2.0 including a new interface for the library to support a wider range of user problems, new search data structures (brute force and distributed), support for user functions to be executed on the results (callbacks), and an expanded set of the supported algorithms (ray tracing and clustering).
Andrey Prokopenko, Daniel Arndt 0003, Damien Lebrun-Grandié, Bruno Turcksin
ACM Trans. Math. Softw.2
2023 Fast tree-based algorithms for DBSCAN for low-dimensional data on GPUs
abstract
DBSCAN is a well-known density-based clustering algorithm to discover arbitrary shape clusters. While conceptually simple in serial, the algorithm is challenging to efficiently parallelize on manycore GPU architectures. Common pitfalls, such as asynchronous range query calls, result in high thread execution divergence in many implementations. In this paper, we propose a new framework for GPU-accelerated DBSCAN, and describe two tree-based algorithms within that framework. Both algorithms fuse the search for neighbors with updating cluster information, but differ in their treatment of dense regions of the data. We show that the time taken to compute clusters is at most twice that of determination of the neighbors. We compare the proposed algorithms with existing CPU and GPU implementations, and demonstrate their competitiveness and performance using a fast traversal structure (bounding volume hierarchy) for low dimensional data. We also show that the memory usage can be reduced by processing object neighbors dynamically without storing them.
Andrey Prokopenko, Damien Lebrun-Grandié, Daniel Arndt 0003
ICPP3
2022 Kokkos 3: Programming Model Extensions for the Exascale Era
abstract
As the push towards exascale hardware has increased the diversity of system architectures, performance portability has become a critical aspect for scientific software. We describe the Kokkos Performance Portable Programming Model that allows developers to write single source applications for diverse high-performance computing architectures. Kokkos provides key abstractions for both the compute and memory hierarchy of modern hardware. We describe the novel abstractions that have been added to Kokkos version 3 such as hierarchical parallelism, containers, task graphs, and arbitrary-sized atomic operations to prepare for exascale era architectures. We demonstrate the performance of these new features with reproducible benchmarks on CPUs and GPUs.
Christian Trott, Damien Lebrun-Grandié, Daniel Arndt 0003, Jan Ciesko, Vinh Q. Dang, Nathan D. Ellingwood, Rahulkumar Gayatri, Evan Harvey, Daisy S. Hollman, Daniel Ibanez, Nevin Liber, Jonathan R. Madsen, Jeff Miles, David Poliakoff, Amy Powell, Sivasankaran Rajamanickam, Mikael Simberg, Daniel Sunderland, Bruno Turcksin, Jeremiah J. Wilke
IEEE Trans. Parallel Distributed Syst.3