Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Gregory Frederick Diamos

dblp:06/1767 · also Greg Diamos, Gregory F. Diamos · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
1since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 5 first-authorArtificial intelligence and machine learning · 7 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 4Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Speech recognition and synthesis · 45% Efficient and distributed learning · 23% Trustworthy machine learning · 17%
Computer architecture, parallel and distributed computing, and storage systems
9 papers
GPUs and heterogeneous computing · 39% Performance modeling and evaluation · 24% Hardware accelerators and domain-specific architectures · 15%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 70% Programming languages and type systems · 30%

Topics — the 30 heaviest of 42, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
Data-centric AI
0.712023
DataPerf: Benchmarks for Data-Centric AI Development · NeurIPS 2023
Natural language and speech › Speech recognition and synthesis
text-to-speech synthesis
0.622017
Deep Voice 2: Multi-Speaker Neural Text-to-Speech · NIPS 2017
Deep Voice: Real-time Neural Text-to-Speech · ICML 2017
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.522020
Beyond human-level accuracy: computational challenges in deep learning · PPoPP 2019
MLPerf Inference Benchmark · ISCA 2020
Performance modeling and evaluation
benchmarking
0.412020
MLPerf Inference Benchmark · ISCA 2020
GPUs and heterogeneous computing
GPU computing
0.422016
Persistent RNNs: Stashing Recurrent Weights On-Chip · ICML 2016
Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU Computation · MICRO 2012
Performance modeling and evaluation
workload characterization
0.412019
Beyond human-level accuracy: computational challenges in deep learning · PPoPP 2019
Machine learning › Efficient and distributed learning
low-precision training
0.312018
Mixed Precision Training · ICLR (Poster) 2018
Natural language and speech › Speech recognition and synthesis › pronunciation modeling
grapheme-to-phoneme conversion
0.312017
Deep Voice: Real-time Neural Text-to-Speech · ICML 2017
Machine learning › Efficient and distributed learning
model compression
0.312017
Exploring Sparsity in Recurrent Neural Networks · ICLR (Poster) 2017
Natural language and speech › Speech recognition and synthesis › speech synthesis
neural speech synthesis
0.312017
Deep Voice: Real-time Neural Text-to-Speech · ICML 2017
Natural language and speech › Speech recognition and synthesis › speech synthesis
neural vocoder
0.312017
Deep Voice 2: Multi-Speaker Neural Text-to-Speech · NIPS 2017
Natural language and speech › Speech recognition and synthesis
prosody prediction
0.312017
Deep Voice: Real-time Neural Text-to-Speech · ICML 2017
Machine learning › Efficient and distributed learning › model compression
sparsity
0.312017
Exploring Sparsity in Recurrent Neural Networks · ICLR (Poster) 2017
Natural language and speech › Speech recognition and synthesis › speaker recognition
speaker embedding
0.312017
Deep Voice 2: Multi-Speaker Neural Text-to-Speech · NIPS 2017
Machine learning › Deep learning architectures and training › neural network training
end-to-end deep learning
0.212016
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
end-to-end speech recognition
0.212016
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016
GPUs and heterogeneous computing
deep learning on GPUs
0.212016
Persistent RNNs: Stashing Recurrent Weights On-Chip · ICML 2016
Compilers and program optimization › loop transformation
loop fusion
0.112012
Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU Computation · MICRO 2012
Compilers and program optimization › deep learning compiler
operator fusion
0.112012
Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU Computation · MICRO 2012
GPUs and heterogeneous computing
GPU microarchitecture
0.112012
Simultaneous branch and warp interweaving for sustained GPU performance · ISCA 2012
GPUs and heterogeneous computing › GPU kernel optimization
kernel fusion
0.112012
Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU Computation · MICRO 2012
GPUs and heterogeneous computing › GPU microarchitecture
SIMT execution
0.112012
Simultaneous branch and warp interweaving for sustained GPU performance · ISCA 2012
Programming languages and type systems
control flow
0.112011
SIMD re-convergence at thread frontiers · MICRO 2011
GPUs and heterogeneous computing › control flow divergence
branch divergence
0.112011
SIMD re-convergence at thread frontiers · MICRO 2011
Processor architecture and microarchitecture
SIMD
0.112011
SIMD re-convergence at thread frontiers · MICRO 2011
Machine learning › Deep learning architectures and training
scaling laws
0.112019
Beyond human-level accuracy: computational challenges in deep learning · PPoPP 2019
Natural language and speech › Language models and text generation › language modeling › language model architecture
sequence-to-sequence model
0.112017
Deep Voice 2: Multi-Speaker Neural Text-to-Speech · NIPS 2017
Parallel and multicore computing › parallel scheduling
runtime scheduling
0.112008
Harmony: an execution model and runtime for heterogeneous many core systems · HPDC 2008
High-performance computing
performance optimization at scale
0.112016
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016
Query processing and optimization › query execution › hardware-accelerated query processing
GPU-accelerated query processing
0.012012
Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU Computation · MICRO 2012

Methods — techniques the papers use, named apart from their topics

scaling projection · 0.8benchmarking · 0.7batch dispatch · 0.5mixed-precision training · 0.3SIMD · 0.3producer-consumer dependence classification · 0.3wavenet · 0.3speaker embedding · 0.3post-processing · 0.3neural vocoder · 0.3deep neural network · 0.3connectionist temporal classification · 0.3recurrent neural network · 0.2persistent kernel · 0.2GPU-based inference · 0.2multithreading · 0.2multi-threading · 0.2thread reconvergence · 0.1
YearPublicationVenuePosition
2023 DataPerf: Benchmarks for Data-Centric AI Development
abstract
Machine learning research has long focused on models rather than datasets, and prominent datasets are used for common ML tasks without regard to the breadth, difficulty, and faithfulness of the underlying problems. Neglecting the fundamental importance of data has given rise to inaccuracy, bias, and fragility in real-world applications, and research is hindered by saturation across existing dataset benchmarks. In response, we present DataPerf, a community-led benchmark suite for evaluating ML datasets and data-centric algorithms. We aim to foster innovation in data-centric AI through competition, comparability, and reproducibility. We enable the ML community to iterate on datasets, instead of just architectures, and we provide an open, online platform with multiple rounds of challenges to support this iterative development. The first iteration of DataPerf contains five benchmarks covering a wide spectrum of data-centric techniques, tasks, and modalities in vision, speech, acquisition, debugging, and diffusion prompting, and we support hosting new contributed benchmarks from the community. The benchmarks, online evaluation platform, and baseline implementations are open source, and the MLCommons Association will maintain DataPerf to ensure long-term benefits to academia and industry.
Mark Mazumder, Colby R. Banbury, Xiaozhe Yao, Bojan Karlas, William Gaviria Rojas, Sudnya Frederick Diamos, Gregory Frederick Diamos, Lynn He, Alicia Parrish, Hannah Kirk, Jessica Quaye, Charvi Rastogi, Douwe Kiela, David Jurado, David Kanter, Rafael Mosquera, Will Cukierski, Juan Ciro, Lora Aroyo, Bilge Acun, Lingjiao Chen, Mehul Raje, Max Bartolo, Sabri Eyuboglu, Amirata Ghorbani, Emmett D. Goodman, Addison Howard, Oana Inel, Tariq Kane, Christine R. Kirkpatrick, D. Sculley, Tzu-Sheng Kuo, Jonas Mueller 0001, Tristan Thrush, Joaquin Vanschoren, Margaret Warren, Adina Williams, Serena Yeung-Levy, Newsha Ardalani, Praveen K. Paritosh, Ce Zhang 0001, James Zou 0001, Carole-Jean Wu, Cody Coleman, Andrew Y. Ng, Peter Mattson, Vijay Janapa Reddi
NeurIPS7
2020 MLPerf Inference Benchmark
abstract
Machine-learning (ML) hardware and software system demand is burgeoning. Driven by ML applications, the number of different ML inference systems has exploded. Over 100 organizations are building ML inference chips, and the systems that incorporate existing models span at least three orders of magnitude in power consumption and five orders of magnitude in performance; they range from embedded devices to data-center solutions. Fueling the hardware are a dozen or more software frameworks and libraries. The myriad combinations of ML hardware and ML software make assessing ML-system performance in an architecture-neutral, representative, and reproducible manner challenging. There is a clear need for industry-wide standard ML benchmarking and evaluation criteria. MLPerf Inference answers that call. In this paper, we present our benchmarking method for evaluating ML inference systems. Driven by more than 30 organizations as well as more than 200 ML engineers and practitioners, MLPerf prescribes a set of rules and best practices to ensure comparability across systems with wildly differing architectures. The first call for submissions garnered more than 600 reproducible inference-performance measurements from 14 organizations, representing over 30 systems that showcase a wide range of capabilities. The submissions attest to the benchmark’s flexibility and adaptability.
Vijay Janapa Reddi, David Kanter, Peter Mattson, Guenther Schmuelling, Carole-Jean Wu, Maximilien Breughe, Mark Charlebois, William Chou, Ramesh Chukka, Cody Coleman, Sam Davis, Gregory Frederick Diamos, Jared Duke, David Fick, J. Scott Gardner, Itay Hubara, Sachin Idgunji, Thomas B. Jablin, Jeff Jiao, Tom St. John, Pankaj Kanwar, Jeffery Liao, Anton Lokhmotov, Francisco Massa, Peng Meng, Paulius Micikevicius, Colin Osborne, Gennady Pekhimenko, Arun Tejusve Raghunath Rajan, Dilip Sequeira, Ashish Sirasao, Fei Sun 0002, Michael Thomson, Frank Wei, Ephrem Wu, Lingjie Xu, Koichi Yamada, George Yuan, Aaron Zhong, Peizhao Zhang
ISCA15
2019 Language Modeling at Scale
abstract
We show how Zipf's Law can be used to scale up language modeling (LM) to take advantage of more training data and more GPUs. LM plays a key role in many important natural language applications such as speech recognition and machine translation. Scaling up LM is important since it is widely accepted by the community that there is no data like more data. Eventually, we would like to train on terabytes (TBs) of text (trillions of words). Modern training methods are far from this goal, because of various bottlenecks, especially memory (within GPUs) and communication (across GPUs). This paper shows how Zipf's Law can address these bottlenecks by grouping parameters for common words and character sequences, because U ≪ N, where U is the number of unique words (types) and N is the size of the training set (tokens). For a local batch size K with G GPUs and a D-dimension embedding matrix, we reduce the original per-GPU memory and communication asymptotic complexity from Θ(GKD) to Θ(GK + UD). Empirically, we find U ∝ (GK)^0.64 on four publicly available large datasets. When we scale up the number of GPUs to 64, a factor of 8, training time speeds up by factors up to 6.7× (for character LMs) and 6.3× (for word LMs) with negligible loss of accuracy. Our weak scaling on 192 GPUs on the Tieba dataset shows a 35% improvement in LM prediction accuracy by training on 93 GB of data (2.5× larger than publicly available SOTA dataset), but taking only 1.25× increase in training time, compared to 3 GB of the same dataset running on 6 GPUs.
Md. Mostofa Ali Patwary, Milind Chabbi, Heewoo Jun, Jiaji Huang, Gregory Frederick Diamos, Kenneth Church 0001
IPDPS5
2019 Beyond human-level accuracy: computational challenges in deep learning
abstract
Deep learning (DL) research yields accuracy and product improvements from both model architecture changes and scale: larger data sets and models, and more computation. For hardware design, it is difficult to predict DL model changes. However, recent prior work shows that as dataset sizes grow, DL model accuracy and model size grow predictably. This paper leverages the prior work to project the dataset and model size growth required to advance DL accuracy beyond human-level, to frontier targets defined by machine learning experts. Datasets will need to grow 33--971×, while models will need to grow 6.6--456× to achieve target accuracies.
Joel Hestness, Newsha Ardalani, Gregory Frederick Diamos
PPoPP3
2019 Fast Spectrogram Inversion Using Multi-Head Convolutional Neural Networks
abstract
We propose the multi-head convolutional neural network (MCNN) for waveform synthesis from spectrograms. Nonlinear interpolation in MCNN is employed with transposed convolution layers in parallel heads. MCNN enables significantly better utilization of modern multi-core processors than commonly used iterative algorithms like Griffin-Lim, and yields very fast (more than 300 × real time) runtime. For training of MCNN, we use a large-scale speech recognition dataset and losses defined on waveforms that are related to perceptual audio quality. We demonstrate that MCNN constitutes a very promising approach for high-quality speech synthesis, without any iterative algorithms or autoregression in computations.
Sercan Ö. Arik, Heewoo Jun, Gregory Frederick Diamos
IEEE Signal Process. Lett.3
2018 Mixed Precision Training
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Frederick Diamos, Erich Elsen, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh
ICLR (Poster)4
2017 Exploring Sparsity in Recurrent Neural Networks
Sharan Narang, Gregory Frederick Diamos, Shubho Sengupta, Erich Elsen
ICLR (Poster)2
2017 Deep Voice: Real-time Neural Text-to-Speech
abstract
We present Deep Voice, a production-quality text-to-speech system constructed entirely from deep neural networks. Deep Voice lays the groundwork for truly end-to-end neural speech synthesis. The system comprises five major building blocks: a segmentation model for locating phoneme boundaries, a grapheme-to-phoneme conversion model, a phoneme duration prediction model, a fundamental frequency prediction model, and an audio synthesis model. For the segmentation model, we propose a novel way of performing phoneme boundary detection with deep neural networks using connectionist temporal classification (CTC) loss. For the audio synthesis model, we implement a variant of WaveNet that requires fewer parameters and trains faster than the original. By using a neural network for each component, our system is simpler and more flexible than traditional text-to-speech systems, where each component requires laborious feature engineering and extensive domain expertise. Finally, we show that inference with our system can be performed faster than real time and describe optimized WaveNet inference kernels on both CPU and GPU that achieve up to 400x speedups over existing implementations.
Sercan Ö. Arik, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Andrew Gibiansky, Yongguo Kang, John Miller 0001, Andrew Y. Ng, Jonathan Raiman, Shubho Sengupta, Mohammad Shoeybi
ICML4
2017 Deep Voice 2: Multi-Speaker Neural Text-to-Speech
abstract
We introduce a technique for augmenting neural text-to-speech (TTS) with low-dimensional trainable speaker embeddings to generate different voices from a single model. As a starting point, we show improvements over the two state-of-the-art approaches for single-speaker neural TTS: Deep Voice 1 and Tacotron. We introduce Deep Voice 2, which is based on a similar pipeline with Deep Voice 1, but constructed with higher performance building blocks and demonstrates a significant audio quality improvement over Deep Voice 1. We improve Tacotron by introducing a post-processing neural vocoder, and demonstrate a significant audio quality improvement. We then demonstrate our technique for multi-speaker speech synthesis for both Deep Voice 2 and Tacotron on two multi-speaker TTS datasets. We show that a single neural TTS system can learn hundreds of unique voices from less than half an hour of data per speaker, while achieving high audio quality synthesis and preserving the speaker identities almost perfectly.
Andrew Gibiansky, Sercan Ö. Arik, Gregory Frederick Diamos, John Miller 0001, Kainan Peng, Wei Ping, Jonathan Raiman, Yanqi Zhou
NIPS3
2016 Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin
abstract
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, enabling experiments that previously took weeks to now run in days. This allows us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale.
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Erich Elsen, Jesse H. Engel, Linxi Fan, Christopher Fougner, Awni Y. Hannun, Billy Jun, Tony Han, Patrick LeGresley, Xiangang Li, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Sheng Qian, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Chong Wang 0002, Zhiqian Wang, Dani Yogatama, Zhenyao Zhu
ICML12
2016 Persistent RNNs: Stashing Recurrent Weights On-Chip
abstract
This paper introduces a new technique for mapping Deep Recurrent Neural Networks (RNN) efficiently onto GPUs. We show how it is possi- ble to achieve substantially higher computational throughput at low mini-batch sizes than direct implementations of RNNs based on matrix multiplications. The key to our approach is the use of persistent computational kernels that exploit the GPU’s inverted memory hierarchy to reuse network weights over multiple timesteps. Our initial implementation sustains 2.8 TFLOP/s at a mini-batch size of 4 on an NVIDIA TitanX GPU. This provides a 16x reduction in activation memory footprint, enables model training with 12x more parameters on the same hardware, allows us to strongly scale RNN training to 128 GPUs, and allows us to efficiently explore end-to-end speech recognition models with over 100 layers.
Gregory Frederick Diamos, Shubho Sengupta, Bryan Catanzaro, Mike Chrzanowski, Adam Coates 0002, Erich Elsen, Jesse H. Engel, Awni Y. Hannun, Sanjeev Satheesh
ICML1
2014 Red Fox: An Execution Environment for Relational Query Processing on GPUs
Haicheng Wu, Gregory Frederick Diamos, Tim Sheard, Molham Aref, Sean Baxter, Michael Garland, Sudhakar Yalamanchili
CGO2
2013 Relational algorithms for multi-bulk-synchronous processors
abstract
Relational databases remain an important application infrastructure for organizing and analyzing massive volumes of data. At the same time, processor architectures are increasingly gravitating towards Multi-Bulk-Synchronous processor (Multi-BSP) architectures employing throughput-optimized memory systems, lightweight multi-threading, and Single-Instruction Multiple-Data (SIMD) core organizations. This paper explores the mapping of primitive relational algebra operations onto such architectures to improve the throughput of data warehousing applications built on relational databases.
Gregory Frederick Diamos, Haicheng Wu, Jin Wang 0010, Ashwin Sanjay Lele, Sudhakar Yalamanchili
PPoPP1
2012 Dynamic compilation of data-parallel kernels for vector processors
abstract
Modern processors enjoy augmented throughput and power efficiency through specialized functional units leveraged via instruction set extensions. These functional units accelerate performance for specific types of operations but must be programmed explicitly. Moreover, applications targeting these specialized units will not take advantage of future ISA extensions and tend not to be portable across multiple ISAs. As architecture designers increasingly rely on heterogeneity for performance improvements, the challenges of leveraging specialized functional units will only become more critical. In particular, exploiting software parallelism without sacrificing portability across the spectrum of commodity and multi-core SIMD processors remains elusive.
Andrew Kerr, Gregory Frederick Diamos, Sudhakar Yalamanchili
CGO2
2012 Simultaneous branch and warp interweaving for sustained GPU performance
abstract
Instruction Multiple-Thread (SIMT) micro-architectures implemented in Graphics Processing Units (GPUs) run fine-grained threads in lockstep by grouping them into units, referred to as warps, to amortize the cost of instruction fetch, decode and control logic over multiple execution units. As individual threads take divergent execution paths, their processing takes place sequentially, defeating part of the efficiency advantage of SIMD execution. We present two complementary techniques that mitigate the impact of thread divergence on SIMT micro-architectures. Both techniques relax the SIMD execution model by allowing two distinct instructions to be scheduled to disjoint subsets of the the same row of execution units, instead of one single instruction. They increase flexibility by providing more thread grouping opportunities than SIMD, while preserving the affinity between threads to avoid introducing extra memory divergence. We consider (1) co-issuing instructions from different divergent paths of the same warp and (2) co-issuing instructions from different warps. To support (1), we introduce a novel thread reconvergence technique that ensures threads are run back in lockstep at control-flow reconvergence points without hindering their ability to run branches in parallel. We propose a lane shuffling technique to allow solution (2) to benefit from inter-warp correlations in divergence patterns. The combination of all these techniques improves performance by 23% on a set of regular GPGPU applications and by 40% on irregular applications, while maintaining the same instruction-fetch and processing-unit resource requirements as the contemporary Fermi GPU architecture.
Nicolas Brunie, Caroline Collange, Gregory Frederick Diamos
ISCA3
2012 Kernel Weaver: Automatically Fusing Database Primitives for Efficient GPU Computation
abstract
Data warehousing applications represent an emerging application arena that requires the processing of relational queries and computations over massive amounts of data. Modern general purpose GPUs are high bandwidth architectures that potentially offer substantial improvements in throughput for these applications. However, there are significant challenges that arise due to the overheads of data movement through the memory hierarchy and between the GPU and host CPU. This paper proposes data movement optimizations to address these challenges. Inspired in part by loop fusion optimizations in the scientific computing community, we propose kernel fusion as a basis for data movement optimizations. Kernel fusion fuses the code bodies of two GPU kernels to i) reduce data footprint to cut down data movement throughout GPU and CPU memory hierarchy, and ii) enlarge compiler optimization scope. We classify producer consumer dependences between compute kernels into three types, i) fine-grained thread-to-thread dependences, ii) medium-grained thread block dependences, and iii) coarse-grained kernel dependences. Based on this classification, we propose a compiler framework, Kernel Weaver, that can automatically fuse relational algebra operators thereby eliminating redundant data movement. The experiments on NVIDIA Fermi platforms demonstrate that kernel fusion achieves 2.89x speedup in GPU computation and a 2.35x speedup in PCIe transfer time on average across the micro-benchmarks tested. We present key insights, lessons learned, measurements from our compiler implementation, and opportunities for further improvements.
Haicheng Wu, Gregory Frederick Diamos, Srihari Cadambi, Sudhakar Yalamanchili
MICRO2
2011 SIMD re-convergence at thread frontiers
abstract
Hardware and compiler techniques for mapping data-parallel programs with divergent control flow to SIMD architectures have recently enabled the emergence of new GPGPU programming models such as CUDA, OpenCL, and DirectX Compute. The impact of branch divergence can be quite different depending upon whether the program's control flow is structured or unstructured. In this paper, we show that unstructured control flow occurs frequently in applications and can lead to significant code expansion when executed using existing approaches for handling branch divergence.
Gregory Frederick Diamos, Benjamin Ashbaugh, Subramaniam Maiyuran, Andrew Kerr, Haicheng Wu, Sudhakar Yalamanchili
MICRO1
2010 Ocelot: a dynamic optimization framework for bulk-synchronous applications in heterogeneous systems
abstract
Ocelot is a dynamic compilation framework designed to map the explicitly data parallel execution model used by NVIDIA CUDA applications onto diverse multithreaded platforms. Ocelot includes a dynamic binary translator from Parallel Thread eXecution ISA (PTX) to many-core processors that leverages the Low Level Virtual Machine (LLVM) code generator to target x86 and other ISAs. The dynamic compiler is able to execute existing CUDA binaries without recompilation from source and supports switching between execution on an NVIDIA GPU and a many-core CPU at runtime. It has been validated against over 130 applications taken from the CUDA SDK, the UIUC Parboil benchmarks [1], the Virginia Rodinia benchmarks [2], the GPU-VSIPL signal and image processing library [3], the Thrust library [4], and several domain specific applications.
Gregory Frederick Diamos, Andrew Kerr, Sudhakar Yalamanchili, Nathan Clark
PACT1
2010 Speculative execution on multi-GPU systems
abstract
The lag of parallel programming models and languages behind the advance of heterogeneous many-core processors has left a gap between the computational capability of modern systems and the ability of applications to exploit them. Emerging programming models, such as CUDA and OpenCL, force developers to explicitly partition applications into components (kernels) and assign them to accelerators in order to utilize them effectively. An accelerator is a processor with a different ISA and micro-architecture than the main CPU. These static partitioning schemes are effective when targeting a system with only a single accelerator. However, they are not robust to changes in the number of accelerators or the performance characteristics of future generations of accelerators. In previous work, we presented the Harmony execution model for computing on heterogeneous systems with several CPUs and accelerators. In this paper, we extend Harmony to target systems with multiple accelerators using control speculation to expose parallelism. We refer to this technique as Kernel Level Speculation (KLS). We argue that dynamic parallelization techniques such as KLS are sufficient to scale applications across several accelerators based on the intuition that there will be fewer distinct accelerators than cores within each accelerator. In this paper, we use a complete prototype of the Harmony runtime that we developed to explore the design decisions and trade-offs in the implementation of KLS. We show that KLS improves parallelism to a sufficient degree while retaining a sequential programming model. We accomplish this by demonstrating good scaling of KLS on a highly heterogeneous system with three distinct accelerator types and ten processors.
Gregory Frederick Diamos, Sudhakar Yalamanchili
IPDPS1
2008 Harmony: an execution model and runtime for heterogeneous many core systems
abstract
The emergence of heterogeneous many core architectures presents a unique opportunity for delivering order of magnitude performance increases to high performance applications by matching certain classes of algorithms to specifically tailored architectures. Their ubiquitous adoption, however, has been limited by a lack of programming models and management frameworks designed to reduce the high degree of complexity of software development intrinsic to heterogeneous architectures. This paper proposes Harmony, a runtime supported programming and execution model that provides: (1) semantics for simplifying parallelism management, (2) dynamic scheduling of compute intensive kernels to heterogeneous processor resources, and (3) online monitoring driven performance optimization for heterogeneous many core systems. We are particulably concerned with simplifying development and ensuring binary portability and scalability across system configurations and sizes. Initial results from ongoing development demonstrate the binary compatibility with variable number of cores, as well as dynamic adaptation of schedules to data sets. We present preliminary results of key features for some benchmark applications.
Gregory Frederick Diamos, Sudhakar Yalamanchili
HPDC1