EDBT 2026 Demo / reviewers in the wild / expert
Dominik Grewe
dblp:66/9298
· DBLP profile ↗
6ranked-venue papers
3as first author
1since 2021 · last 2025
0009-0008-6483-3841ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 69% Speech recognition and synthesis · 31% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 84% GPUs and heterogeneous computing · 16% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
0.9 | 1 | 2025 | PartIR: Composing SPMD Partitioning Strategies for Machine Learning · ASPLOS (1) 2025 |
Machine learning › Efficient and distributed learning › distributed training › parallelization
model partitioning |
0.9 | 1 | 2025 | PartIR: Composing SPMD Partitioning Strategies for Machine Learning · ASPLOS (1) 2025 |
Parallel and multicore computing
parallelization strategies |
0.9 | 1 | 2025 | PartIR: Composing SPMD Partitioning Strategies for Machine Learning · ASPLOS (1) 2025 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.3 | 1 | 2018 | Parallel WaveNet: Fast High-Fidelity Speech Synthesis · ICML 2018 |
Natural language and speech › Speech recognition and synthesis › speech synthesis
neural speech synthesis |
0.3 | 1 | 2018 | Parallel WaveNet: Fast High-Fidelity Speech Synthesis · ICML 2018 |
Natural language and speech › Speech recognition and synthesis › speech synthesis › neural speech synthesis
parallel speech synthesis |
0.3 | 1 | 2018 | Parallel WaveNet: Fast High-Fidelity Speech Synthesis · ICML 2018 |
Natural language and speech › Speech recognition and synthesis
speech synthesis |
0.3 | 1 | 2018 | Parallel WaveNet: Fast High-Fidelity Speech Synthesis · ICML 2018 |
Compilers and program optimization
code generation |
0.2 | 1 | 2014 | Automatic and Portable Mapping of Data Parallel Programs to OpenCL for GPU-Based Heterogeneous Systems · ACM Trans. Archit. Code Optim. 2014 |
GPUs and heterogeneous computing
GPU programming |
0.2 | 1 | 2014 | Automatic and Portable Mapping of Data Parallel Programs to OpenCL for GPU-Based Heterogeneous Systems · ACM Trans. Archit. Code Optim. 2014 |
Machine learning › Efficient and distributed learning
inference acceleration |
0.1 | 1 | 2018 | Parallel WaveNet: Fast High-Fidelity Speech Synthesis · ICML 2018 |
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP |
0.1 | 1 | 2014 | Automatic and Portable Mapping of Data Parallel Programs to OpenCL for GPU-Based Heterogeneous Systems · ACM Trans. Archit. Code Optim. 2014 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2014 | Automatic and Portable Mapping of Data Parallel Programs to OpenCL for GPU-Based Heterogeneous Systems · ACM Trans. Archit. Code Optim. 2014 |
Methods — techniques the papers use, named apart from their topics
schedule-like API · 1.7intermediate representation rewriting · 1.7machine learning-based predictive model · 0.4data transformation · 0.4probability density distillation · 0.3autoregressive modeling · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PartIR: Composing SPMD Partitioning Strategies for Machine LearningabstractTraining modern large neural networks (NNs) requires a combination of parallelization strategies, including data, model, or optimizer sharding. To address the growing complexity of these strategies, we introduce PartIR, a hardware-and-runtime agnostic NN partitioning system. PartIR is: 1) Expressive: It allows for the composition of multiple sharding strategies, whether user-defined or automatically derived; 2) Decoupled: the strategies are separate from the ML implementation; and 3) Predictable: It follows a set of well-defined general rules to partition the NN. PartIR utilizes a schedule-like API that incrementally rewrites the ML program intermediate representation (IR) after each strategy, allowing simulators and users to verify the strategy's performance. PartIR has been successfully used both for training large models and across diverse model architectures, demonstrating its predictability, expressiveness, and performance. Sami Alabed, Daniel Belov, Bart Chrzaszcz, Juliana Franco, Dominik Grewe, Dougal Maclaurin, James Molloy, Tom Natan, Tamara Norman, Xiaoyue Pan, Adam Paszke, Norman A. Rink, Michael Schaarschmidt, Timur Sitdikov, Agnieszka Swietlik, Dimitrios Vytiniotis, Joel Wee |
ASPLOS (1) | 5 |
| 2018 | Parallel WaveNet: Fast High-Fidelity Speech SynthesisabstractThe recently-developed WaveNet architecture is the current state of the art in realistic speech synthesis, consistently rated as more natural sounding for many different languages than any previous system. However, because WaveNet relies on sequential generation of one audio sample at a time, it is poorly suited to today’s massively parallel computers, and therefore hard to deploy in a real-time production setting. This paper introduces Probability Density Distillation, a new method for training a parallel feed-forward network from a trained WaveNet with no significant difference in quality. The resulting system is capable of generating high-fidelity speech samples at more than 20 times faster than real-time, a 1000x speed up relative to the original WaveNet, and capable of serving multiple English and Japanese voices in a production setting. Aäron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche 0002, Edward Lockhart, Luis C. Cobo, Florian Stimberg, Norman Casagrande, Dominik Grewe, Seb Noury, Sander Dieleman, Erich Elsen, Nal Kalchbrenner, Heiga Zen, Alex Graves, Helen King, Tom Walters, Daniel Belov, Demis Hassabis |
ICML | 12 |
| 2014 | Automatic and Portable Mapping of Data Parallel Programs to OpenCL for GPU-Based Heterogeneous SystemsabstractGeneral-purpose GPU-based systems are highly attractive, as they give potentially massive performance at little cost. Realizing such potential is challenging due to the complexity of programming. This article presents a compiler-based approach to automatically generate optimized OpenCL code from data parallel OpenMP programs for GPUs. A key feature of our scheme is that it leverages existing transformations, especially data transformations, to improve performance on GPU architectures and uses automatic machine learning to build a predictive model to determine if it is worthwhile running the OpenCL code on the GPU or OpenMP code on the multicore host. We applied our approach to the entire NAS parallel benchmark suite and evaluated it on distinct GPU-based systems. We achieved average (up to) speedups of 4.51× and 4.20× (143× and 67×) on Core i7/NVIDIA GeForce GTX580 and Core i7/AMD Radeon 7970 platforms, respectively, over a sequential baseline. Our approach achieves, on average, greater than 10× speedups over two state-of-the-art automatic GPU code generators. Zheng Wang 0001, Dominik Grewe, Michael F. P. O'Boyle |
ACM Trans. Archit. Code Optim. | 2 |
| 2013 | Portable mapping of data parallel programs to OpenCL for heterogeneous systemsabstractGeneral purpose GPU based systems are highly attractive as they give potentially massive performance at little cost. Re-alizing such potential is challenging due to the complexity of programming. This paper presents a compiler based approach to automatically generate optimized OpenCL code from data-parallel OpenMP programs for GPUs. Such an approach brings together the benefits of a clear high levellanguage (OpenMP) and an emerging standard (OpenCL) for heterogeneous multi-cores. A key feature of our scheme is that it leverages existing transformations, especially data transformations, to improve performance on GPU architectures and uses predictive modeling to automatically determine if it is worthwhile running the OpenCL code on the GPU or OpenMP code on the multi-core host. We applied our approach to the entire NAS parallel benchmark suite and evaluated it on two distinct GPU based systems: Core i7/NVIDIA GeForce GTX 580 and Core 17/AMD Radeon 7970. We achieved average (up to) speedups of 4.51x and 4.20x (143x and 67x) respectively over a sequential baseline. This is, on average, a factor 1.63 and 1.56 times faster than a hand-coded, GPU-specific OpenCL implementation developed by independent expert programmers. Dominik Grewe, Zheng Wang 0001, Michael F. P. O'Boyle |
CGO | 1 |
| 2011 | A Static Task Partitioning Approach for Heterogeneous Systems Using OpenCL
Dominik Grewe, Michael F. P. O'Boyle |
CC | 1 |
| 2011 | A workload-aware mapping approach for data-parallel programsabstractMuch compiler-orientated work in the area of mapping parallel programs to parallel architectures has ignored the issue of external workload. Given that the majority of platforms will not be dedicated to just one task at a time, the impact of other jobs needs to be addressed. As mapping is highly dependent on the underlying machine, a technique that is easily portable across platforms is also desirable. Dominik Grewe, Zheng Wang 0001, Michael F. P. O'Boyle |
HiPEAC | 1 |