Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Dominik Grewe

dblp:66/9298 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
1since 2021 · last 2025
0009-0008-6483-3841ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 69% Speech recognition and synthesis · 31%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 84% GPUs and heterogeneous computing · 16%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
0.912025
PartIR: Composing SPMD Partitioning Strategies for Machine Learning · ASPLOS (1) 2025
Machine learning › Efficient and distributed learning › distributed training › parallelization
model partitioning
0.912025
PartIR: Composing SPMD Partitioning Strategies for Machine Learning · ASPLOS (1) 2025
Parallel and multicore computing
parallelization strategies
0.912025
PartIR: Composing SPMD Partitioning Strategies for Machine Learning · ASPLOS (1) 2025
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.312018
Parallel WaveNet: Fast High-Fidelity Speech Synthesis · ICML 2018
Natural language and speech › Speech recognition and synthesis › speech synthesis
neural speech synthesis
0.312018
Parallel WaveNet: Fast High-Fidelity Speech Synthesis · ICML 2018
Natural language and speech › Speech recognition and synthesis › speech synthesis › neural speech synthesis
parallel speech synthesis
0.312018
Parallel WaveNet: Fast High-Fidelity Speech Synthesis · ICML 2018
Natural language and speech › Speech recognition and synthesis
speech synthesis
0.312018
Parallel WaveNet: Fast High-Fidelity Speech Synthesis · ICML 2018
Compilers and program optimization
code generation
0.212014
Automatic and Portable Mapping of Data Parallel Programs to OpenCL for GPU-Based Heterogeneous Systems · ACM Trans. Archit. Code Optim. 2014
GPUs and heterogeneous computing
GPU programming
0.212014
Automatic and Portable Mapping of Data Parallel Programs to OpenCL for GPU-Based Heterogeneous Systems · ACM Trans. Archit. Code Optim. 2014
Machine learning › Efficient and distributed learning
inference acceleration
0.112018
Parallel WaveNet: Fast High-Fidelity Speech Synthesis · ICML 2018
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP
0.112014
Automatic and Portable Mapping of Data Parallel Programs to OpenCL for GPU-Based Heterogeneous Systems · ACM Trans. Archit. Code Optim. 2014
Parallel and multicore computing
parallel programming models
0.112014
Automatic and Portable Mapping of Data Parallel Programs to OpenCL for GPU-Based Heterogeneous Systems · ACM Trans. Archit. Code Optim. 2014

Methods — techniques the papers use, named apart from their topics

schedule-like API · 1.7intermediate representation rewriting · 1.7machine learning-based predictive model · 0.4data transformation · 0.4probability density distillation · 0.3autoregressive modeling · 0.3
YearPublicationVenuePosition
2025 PartIR: Composing SPMD Partitioning Strategies for Machine Learning
abstract
Training modern large neural networks (NNs) requires a combination of parallelization strategies, including data, model, or optimizer sharding. To address the growing complexity of these strategies, we introduce PartIR, a hardware-and-runtime agnostic NN partitioning system. PartIR is: 1) Expressive: It allows for the composition of multiple sharding strategies, whether user-defined or automatically derived; 2) Decoupled: the strategies are separate from the ML implementation; and 3) Predictable: It follows a set of well-defined general rules to partition the NN. PartIR utilizes a schedule-like API that incrementally rewrites the ML program intermediate representation (IR) after each strategy, allowing simulators and users to verify the strategy's performance. PartIR has been successfully used both for training large models and across diverse model architectures, demonstrating its predictability, expressiveness, and performance.
Sami Alabed, Daniel Belov, Bart Chrzaszcz, Juliana Franco, Dominik Grewe, Dougal Maclaurin, James Molloy, Tom Natan, Tamara Norman, Xiaoyue Pan, Adam Paszke, Norman A. Rink, Michael Schaarschmidt, Timur Sitdikov, Agnieszka Swietlik, Dimitrios Vytiniotis, Joel Wee
ASPLOS (1)5
2018 Parallel WaveNet: Fast High-Fidelity Speech Synthesis
abstract
The recently-developed WaveNet architecture is the current state of the art in realistic speech synthesis, consistently rated as more natural sounding for many different languages than any previous system. However, because WaveNet relies on sequential generation of one audio sample at a time, it is poorly suited to today’s massively parallel computers, and therefore hard to deploy in a real-time production setting. This paper introduces Probability Density Distillation, a new method for training a parallel feed-forward network from a trained WaveNet with no significant difference in quality. The resulting system is capable of generating high-fidelity speech samples at more than 20 times faster than real-time, a 1000x speed up relative to the original WaveNet, and capable of serving multiple English and Japanese voices in a production setting.
Aäron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche 0002, Edward Lockhart, Luis C. Cobo, Florian Stimberg, Norman Casagrande, Dominik Grewe, Seb Noury, Sander Dieleman, Erich Elsen, Nal Kalchbrenner, Heiga Zen, Alex Graves, Helen King, Tom Walters, Daniel Belov, Demis Hassabis
ICML12
2014 Automatic and Portable Mapping of Data Parallel Programs to OpenCL for GPU-Based Heterogeneous Systems
abstract
General-purpose GPU-based systems are highly attractive, as they give potentially massive performance at little cost. Realizing such potential is challenging due to the complexity of programming. This article presents a compiler-based approach to automatically generate optimized OpenCL code from data parallel OpenMP programs for GPUs. A key feature of our scheme is that it leverages existing transformations, especially data transformations, to improve performance on GPU architectures and uses automatic machine learning to build a predictive model to determine if it is worthwhile running the OpenCL code on the GPU or OpenMP code on the multicore host. We applied our approach to the entire NAS parallel benchmark suite and evaluated it on distinct GPU-based systems. We achieved average (up to) speedups of 4.51× and 4.20× (143× and 67×) on Core i7/NVIDIA GeForce GTX580 and Core i7/AMD Radeon 7970 platforms, respectively, over a sequential baseline. Our approach achieves, on average, greater than 10× speedups over two state-of-the-art automatic GPU code generators.
Zheng Wang 0001, Dominik Grewe, Michael F. P. O'Boyle
ACM Trans. Archit. Code Optim.2
2013 Portable mapping of data parallel programs to OpenCL for heterogeneous systems
abstract
General purpose GPU based systems are highly attractive as they give potentially massive performance at little cost. Re-alizing such potential is challenging due to the complexity of programming. This paper presents a compiler based approach to automatically generate optimized OpenCL code from data-parallel OpenMP programs for GPUs. Such an approach brings together the benefits of a clear high levellanguage (OpenMP) and an emerging standard (OpenCL) for heterogeneous multi-cores. A key feature of our scheme is that it leverages existing transformations, especially data transformations, to improve performance on GPU architectures and uses predictive modeling to automatically determine if it is worthwhile running the OpenCL code on the GPU or OpenMP code on the multi-core host. We applied our approach to the entire NAS parallel benchmark suite and evaluated it on two distinct GPU based systems: Core i7/NVIDIA GeForce GTX 580 and Core 17/AMD Radeon 7970. We achieved average (up to) speedups of 4.51x and 4.20x (143x and 67x) respectively over a sequential baseline. This is, on average, a factor 1.63 and 1.56 times faster than a hand-coded, GPU-specific OpenCL implementation developed by independent expert programmers.
Dominik Grewe, Zheng Wang 0001, Michael F. P. O'Boyle
CGO1
2011 A Static Task Partitioning Approach for Heterogeneous Systems Using OpenCL
Dominik Grewe, Michael F. P. O'Boyle
CC1
2011 A workload-aware mapping approach for data-parallel programs
abstract
Much compiler-orientated work in the area of mapping parallel programs to parallel architectures has ignored the issue of external workload. Given that the majority of platforms will not be dedicated to just one task at a time, the impact of other jobs needs to be addressed. As mapping is highly dependent on the underlying machine, a technique that is easily portable across platforms is also desirable.
Dominik Grewe, Zheng Wang 0001, Michael F. P. O'Boyle
HiPEAC1