EDBT 2026 Demo / reviewers in the wild / expert
Nicolai Stawinoga
dblp:224/6707
· DBLP profile ↗
4ranked-venue papers
3as first author
2since 2021 · last 2026
0000-0002-3806-2691ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
High-performance computing · 45% GPUs and heterogeneous computing · 25% Hardware accelerators and domain-specific architectures · 23% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
compiler infrastructure |
1.0 | 1 | 2026 | An MLIR Lowering Pipeline for Stencils at Wafer-Scale · ASPLOS (2) 2026 |
Compilers and program optimization › compiler infrastructure
MLIR |
1.0 | 1 | 2026 | An MLIR Lowering Pipeline for Stencils at Wafer-Scale · ASPLOS (2) 2026 |
GPUs and heterogeneous computing
GPU compilation |
0.3 | 1 | 2018 | Predictable Thread Coarsening · ACM Trans. Archit. Code Optim. 2018 |
High-performance computing
scientific computing systems |
0.3 | 1 | 2026 | An MLIR Lowering Pipeline for Stencils at Wafer-Scale · ASPLOS (2) 2026 |
High-performance computing
stencil computation |
0.3 | 1 | 2026 | An MLIR Lowering Pipeline for Stencils at Wafer-Scale · ASPLOS (2) 2026 |
Hardware accelerators and domain-specific architectures › many-core accelerator
wafer-scale engine |
0.3 | 1 | 2026 | An MLIR Lowering Pipeline for Stencils at Wafer-Scale · ASPLOS (2) 2026 |
Methods — techniques the papers use, named apart from their topics
MLIR · 2.0static analysis · 0.7occupancy modeling · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An MLIR Lowering Pipeline for Stencils at Wafer-ScaleabstractThe Cerebras Wafer-Scale Engine (WSE) delivers performance at an unprecedented scale of over 900,000 compute units, all connected via a single-wafer on-chip interconnect. Initially designed for AI, the WSE architecture is also well-suited for High Performance Computing (HPC). However, its distributed asynchronous programming model diverges significantly from the simple sequential or bulk-synchronous programs that one would typically derive for a given mathematical program description. Targeting the WSE requires a bespoke re-implementation when porting existing code. The absence of WSE support in compilers such as MLIR, meant that there was little hope for automating this process. Nicolai Stawinoga, David Katz, Anton Lydike, Justs Zarins, Nick Brown 0002, George Bisbas, Tobias Grosser |
ASPLOS (2) | 1 |
| 2026 | A Portable Compiler-Runtime Approach for Scalability PredictionabstractHighly scalable parallel applications can efficiently solve expensive computational problems when run on a large number of compute nodes. However, selecting the optimal number of nodes for a compute job of a given size is non-trivial, and allocating too few or too many nodes may not yield the expected performance. Knowing the scaling behavior of an application in advance enables us, for example, to make optimal use of the available hardware resources. We introduce a novel, portable approach to predict the scalability of parallel applications written in modern high-level programming models. We propose a predictive compiler-runtime framework based on Celerity, a task-based distributed runtime system that enables executing SYCL codes on clusters. The framework targets a broad range of computing systems, from CPU to GPU clusters, and proposes a model that combines machine learning, communication modeling and DAG heuristics. Experimental results on two large-scale clusters, JUWELS and Marconi-100, show accurate scalability prediction of unseen single and multi-task applications. Nicolai Stawinoga, Sohan Lal, Biagio Cosenza, Philip Salzmann, Peter Thoman, Thomas Fahringer |
Future Gener. Comput. Syst. | 1 |
| 2020 | SYCL-Bench: A Versatile Cross-Platform Benchmark Suite for Heterogeneous Computing
Sohan Lal, Aksel Alpay, Philip Salzmann, Biagio Cosenza, Alexander Hirsch, Nicolai Stawinoga, Peter Thoman, Thomas Fahringer, Vincent Heuveline |
Euro-Par | 6 |
| 2018 | Predictable Thread CoarseningabstractThread coarsening on GPUs combines the work of several threads into one. We show how thread coarsening can be implemented as a fully automated compile-time optimisation that estimates the optimal coarsening factor based on a low-cost, approximate static analysis of cache line re-use and an occupancy prediction model. We evaluate two coarsening strategies on three different NVidia GPU architectures. For NVidia reduction kernels we achieve a maximum speedup of 5.08x, and for the Rodinia benchmarks we achieve a mean speedup of 1.30x over 8 of 19 kernels that were determined safe to coarsen. Nicolai Stawinoga, Tony Field |
ACM Trans. Archit. Code Optim. | 1 |