Fabian Knorr

dblp:277/9750 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0003-4193-374XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Bridging usability and performance: High-level abstractions for advanced accelerator cluster programming
Philip Salzmann, Fabian Knorr, Peter Thoman, Philipp Gschwandtner, Thomas Fahringer
Future Gener. Comput. Syst.2
2023 An Asynchronous Dataflow-Driven Execution Model For Distributed Accelerator Computing
abstract
While domain-specific HPC software packages continue to thrive and are vital to many scientific communities, a general purpose high-productivity GPU cluster programming model that facilitates experimentation for non-experts remains elusive. We demonstrate how Celerity, a high-level C++ programming model for distributed accelerator computing based on the open SYCL standard, allows for the quick development of - and experimentation with - distributed applications. To achieve scalability on large machines, we replace Celerity's existing master/worker scheduling model with a fully distributed scheme that reduces the worst-case scheduling complexity from quadratic to linear while maintaining the existing programming interface. We then show how this declarative, data-flow based API paired with a point-to-point communication model with eager data pushing can effectively expose and leverage opportunities for latency hiding and computation/communication overlapping with minimal or no manual guidance. We demonstrate how Celerity exhibits very good scalability on multiple benchmarks from several scientific domains and up to 128 GPUs.
Philip Salzmann, Fabian Knorr, Peter Thoman, Philipp Gschwandtner, Biagio Cosenza, Thomas Fahringer
CCGrid2
2021 ndzip: A High-Throughput Parallel Lossless Compressor for Scientific Data
abstract
Publikationen von Forschenden. Knorr, Fabian; Thoman, Peter; Fahringer, Thomas: ndzip: a high-throughput parallel lossless compressor for scientific data. In: Proceedings 2021 Data Compression Conference (DCC) / Bilgin, Ali; Marcellin, Michael W.; Serra-Sagrista, Joan; Storer, James A. IEEE, 2021
Fabian Knorr, Peter Thoman, Thomas Fahringer
DCC1
2021 Porting Real-World Applications to GPU Clusters: A Celerity and Cronos Case Study
abstract
Accelerator clusters are an ongoing trend in high performance computing, continuously gaining traction and forming a ubiquitous hardware resource for domain scientists to run large-scale simulations on. However, there is often a gap between new hardware technologies and adoption by legacy code bases. Porting real-world applications to new programming models is a difficult undertaking, aggravated by the need for support for both distributed-memory and accelerator parallelism. In this work, we present a case study of porting Cronos, a real-world code from the field of magnetohydrodynamics, to Celerity, a high-level programming model for distributed-memory accelerator clusters. We discuss the numerical, algorithmic and implementation properties of the application and motivate our decisions for adapting them where necessary. Preliminary results show a parallel efficiency of up to 87% for 16 GPUs.
Philipp Gschwandtner, Ralf Kissmann, David Huber 0003, Philip Salzmann, Fabian Knorr, Peter Thoman, Thomas Fahringer
e-Science5
2021 ndzip-gpu: efficient lossless compression of scientific floating-point data on GPUs
abstract
Lossless data compression is a promising software approach for reducing the bandwidth requirements of scientific applications on accelerator clusters without introducing approximation errors. Suitable compressors must be able to effectively compact floating-point data while saturating the system interconnect to avoid introducing unnecessary latencies.
Fabian Knorr, Peter Thoman, Thomas Fahringer
SC1