Nicolai Stawinoga

dblp:224/6707 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
2since 2021 · last 2026
0000-0002-3806-2691ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 45% GPUs and heterogeneous computing · 25% Hardware accelerators and domain-specific architectures · 23%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
compiler infrastructure
1.012026
An MLIR Lowering Pipeline for Stencils at Wafer-Scale · ASPLOS (2) 2026
Compilers and program optimization › compiler infrastructure
MLIR
1.012026
An MLIR Lowering Pipeline for Stencils at Wafer-Scale · ASPLOS (2) 2026
GPUs and heterogeneous computing
GPU compilation
0.312018
Predictable Thread Coarsening · ACM Trans. Archit. Code Optim. 2018
High-performance computing
scientific computing systems
0.312026
An MLIR Lowering Pipeline for Stencils at Wafer-Scale · ASPLOS (2) 2026
High-performance computing
stencil computation
0.312026
An MLIR Lowering Pipeline for Stencils at Wafer-Scale · ASPLOS (2) 2026
Hardware accelerators and domain-specific architectures › many-core accelerator
wafer-scale engine
0.312026
An MLIR Lowering Pipeline for Stencils at Wafer-Scale · ASPLOS (2) 2026

Methods — techniques the papers use, named apart from their topics

MLIR · 2.0static analysis · 0.7occupancy modeling · 0.7
YearPublicationVenuePosition
2026 An MLIR Lowering Pipeline for Stencils at Wafer-Scale
abstract
The Cerebras Wafer-Scale Engine (WSE) delivers performance at an unprecedented scale of over 900,000 compute units, all connected via a single-wafer on-chip interconnect. Initially designed for AI, the WSE architecture is also well-suited for High Performance Computing (HPC). However, its distributed asynchronous programming model diverges significantly from the simple sequential or bulk-synchronous programs that one would typically derive for a given mathematical program description. Targeting the WSE requires a bespoke re-implementation when porting existing code. The absence of WSE support in compilers such as MLIR, meant that there was little hope for automating this process.
Nicolai Stawinoga, David Katz, Anton Lydike, Justs Zarins, Nick Brown 0002, George Bisbas, Tobias Grosser
ASPLOS (2)1
2026 A Portable Compiler-Runtime Approach for Scalability Prediction
abstract
Highly scalable parallel applications can efficiently solve expensive computational problems when run on a large number of compute nodes. However, selecting the optimal number of nodes for a compute job of a given size is non-trivial, and allocating too few or too many nodes may not yield the expected performance. Knowing the scaling behavior of an application in advance enables us, for example, to make optimal use of the available hardware resources. We introduce a novel, portable approach to predict the scalability of parallel applications written in modern high-level programming models. We propose a predictive compiler-runtime framework based on Celerity, a task-based distributed runtime system that enables executing SYCL codes on clusters. The framework targets a broad range of computing systems, from CPU to GPU clusters, and proposes a model that combines machine learning, communication modeling and DAG heuristics. Experimental results on two large-scale clusters, JUWELS and Marconi-100, show accurate scalability prediction of unseen single and multi-task applications.
Nicolai Stawinoga, Sohan Lal, Biagio Cosenza, Philip Salzmann, Peter Thoman, Thomas Fahringer
Future Gener. Comput. Syst.1
2020 SYCL-Bench: A Versatile Cross-Platform Benchmark Suite for Heterogeneous Computing
Sohan Lal, Aksel Alpay, Philip Salzmann, Biagio Cosenza, Alexander Hirsch, Nicolai Stawinoga, Peter Thoman, Thomas Fahringer, Vincent Heuveline
Euro-Par6
2018 Predictable Thread Coarsening
abstract
Thread coarsening on GPUs combines the work of several threads into one. We show how thread coarsening can be implemented as a fully automated compile-time optimisation that estimates the optimal coarsening factor based on a low-cost, approximate static analysis of cache line re-use and an occupancy prediction model. We evaluate two coarsening strategies on three different NVidia GPU architectures. For NVidia reduction kernels we achieve a maximum speedup of 5.08x, and for the Rodinia benchmarks we achieve a mean speedup of 1.30x over 8 of 19 kernels that were determined safe to coarsen.
Nicolai Stawinoga, Tony Field
ACM Trans. Archit. Code Optim.1