EDBT 2026 Demo / reviewers in the wild / expert
Fabian Knorr
dblp:277/9750
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0003-4193-374XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging usability and performance: High-level abstractions for advanced accelerator cluster programming
Philip Salzmann, Fabian Knorr, Peter Thoman, Philipp Gschwandtner, Thomas Fahringer |
Future Gener. Comput. Syst. | 2 |
| 2023 | An Asynchronous Dataflow-Driven Execution Model For Distributed Accelerator ComputingabstractWhile domain-specific HPC software packages continue to thrive and are vital to many scientific communities, a general purpose high-productivity GPU cluster programming model that facilitates experimentation for non-experts remains elusive. We demonstrate how Celerity, a high-level C++ programming model for distributed accelerator computing based on the open SYCL standard, allows for the quick development of - and experimentation with - distributed applications. To achieve scalability on large machines, we replace Celerity's existing master/worker scheduling model with a fully distributed scheme that reduces the worst-case scheduling complexity from quadratic to linear while maintaining the existing programming interface. We then show how this declarative, data-flow based API paired with a point-to-point communication model with eager data pushing can effectively expose and leverage opportunities for latency hiding and computation/communication overlapping with minimal or no manual guidance. We demonstrate how Celerity exhibits very good scalability on multiple benchmarks from several scientific domains and up to 128 GPUs. Philip Salzmann, Fabian Knorr, Peter Thoman, Philipp Gschwandtner, Biagio Cosenza, Thomas Fahringer |
CCGrid | 2 |
| 2021 | ndzip: A High-Throughput Parallel Lossless Compressor for Scientific DataabstractPublikationen von Forschenden. Knorr, Fabian; Thoman, Peter; Fahringer, Thomas: ndzip: a high-throughput parallel lossless compressor for scientific data. In: Proceedings 2021 Data Compression Conference (DCC) / Bilgin, Ali; Marcellin, Michael W.; Serra-Sagrista, Joan; Storer, James A. IEEE, 2021 Fabian Knorr, Peter Thoman, Thomas Fahringer |
DCC | 1 |
| 2021 | Porting Real-World Applications to GPU Clusters: A Celerity and Cronos Case StudyabstractAccelerator clusters are an ongoing trend in high performance computing, continuously gaining traction and forming a ubiquitous hardware resource for domain scientists to run large-scale simulations on. However, there is often a gap between new hardware technologies and adoption by legacy code bases. Porting real-world applications to new programming models is a difficult undertaking, aggravated by the need for support for both distributed-memory and accelerator parallelism. In this work, we present a case study of porting Cronos, a real-world code from the field of magnetohydrodynamics, to Celerity, a high-level programming model for distributed-memory accelerator clusters. We discuss the numerical, algorithmic and implementation properties of the application and motivate our decisions for adapting them where necessary. Preliminary results show a parallel efficiency of up to 87% for 16 GPUs. Philipp Gschwandtner, Ralf Kissmann, David Huber 0003, Philip Salzmann, Fabian Knorr, Peter Thoman, Thomas Fahringer |
e-Science | 5 |
| 2021 | ndzip-gpu: efficient lossless compression of scientific floating-point data on GPUsabstractLossless data compression is a promising software approach for reducing the bandwidth requirements of scientific applications on accelerator clusters without introducing approximation errors. Suitable compressors must be able to effectively compact floating-point data while saturating the system interconnect to avoid introducing unnecessary latencies. Fabian Knorr, Peter Thoman, Thomas Fahringer |
SC | 1 |