Erich Elsen

dblp:94/3606 · DBLP profile ↗
← Back
23ranked-venue papers
2as first author
7since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 1 first-author · 7 since 2021Systems, architecture and hardware · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
21 papers
Efficient and distributed learning · 46% Deep learning architectures and training · 16% Speech recognition and synthesis · 11%
Computer architecture, parallel and distributed computing, and storage systems
6 papers
Hardware accelerators and domain-specific architectures · 34% GPUs and heterogeneous computing · 34% High-performance computing · 23%

Topics — the 30 heaviest of 52, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
2.882022
Sparse GPU kernels for deep learning · SC 2020
Top-KAST: Top-K Always Sparse Training · NeurIPS 2020
Rigging the Lottery: Making All Tickets Winners · ICML 2020
Machine learning › Efficient and distributed learning › model compression
sparse training
2.252022
The State of Sparse Training in Deep Reinforcement Learning · ICML 2022
Practical Real Time Recurrent Learning with a Sparse Approximation · ICLR 2021
Top-KAST: Top-K Always Sparse Training · NeurIPS 2020
Machine learning › Efficient and distributed learning › model compression
sparse neural network
1.542022
Sparse GPU kernels for deep learning · SC 2020
Rigging the Lottery: Making All Tickets Winners · ICML 2020
Fast Sparse ConvNets · CVPR 2020
Machine learning › Deep learning architectures and training
scaling laws
1.122022
An empirical analysis of compute-optimal large language model training · NeurIPS 2022
Unified Scaling Laws for Routed Language Models · ICML 2022
Natural language and speech › Speech recognition and synthesis
text-to-speech synthesis
0.822021
End-to-end Adversarial Text-to-Speech · ICLR 2021
Efficient Neural Audio Synthesis · ICML 2018
Natural language and speech › Speech recognition and synthesis
speech synthesis
0.822020
High Fidelity Speech Synthesis with Adversarial Networks · ICLR 2020
Parallel WaveNet: Fast High-Fidelity Speech Synthesis · ICML 2018
Machine learning › Generative modeling
generative adversarial network
0.622021
High Fidelity Speech Synthesis with Adversarial Networks · ICLR 2020
End-to-end Adversarial Text-to-Speech · ICLR 2021
Machine learning › Efficient and distributed learning › efficient training
compute-optimal training
0.612022
An empirical analysis of compute-optimal large language model training · NeurIPS 2022
Machine learning › Reinforcement learning
deep reinforcement learning
0.612022
The State of Sparse Training in Deep Reinforcement Learning · ICML 2022
Machine learning › Generative modeling
diffusion model
0.612022
Step-unrolled Denoising Autoencoders for Text Generation · ICLR 2022
Natural language and speech › Question answering and dialogue systems
knowledge-intensive tasks
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Machine learning › Deep learning architectures and training
mixture of experts
0.612022
Unified Scaling Laws for Routed Language Models · ICML 2022
Natural language and speech › Language models and text generation
retrieval-augmented language models
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Natural language and speech › Language models and text generation
text generation
0.612022
Step-unrolled Denoising Autoencoders for Text Generation · ICLR 2022
Information retrieval
document retrieval
0.612022
Improving Language Models by Retrieving from Trillions of Tokens · ICML 2022
Machine learning › Deep learning architectures and training › recurrent neural network
real-time recurrent learning
0.512021
Practical Real Time Recurrent Learning with a Sparse Approximation · ICLR 2021
Machine learning › Deep learning architectures and training
recurrent neural network
0.512021
Practical Real Time Recurrent Learning with a Sparse Approximation · ICLR 2021
Machine learning › Efficient and distributed learning › model compression › sparsity
dynamic sparsity
0.412020
Top-KAST: Top-K Always Sparse Training · NeurIPS 2020
Machine learning › Learning theory
generalization
0.412020
On the Generalization Benefit of Noise in Stochastic Gradient Descent · ICML 2020
Machine learning › Learning theory
generalization bounds
0.412020
On the Generalization Benefit of Noise in Stochastic Gradient Descent · ICML 2020
Computer vision › 3D vision › point cloud processing
sparse convolution
0.412020
Fast Sparse ConvNets · CVPR 2020
Machine learning › Optimization for machine learning
stochastic gradient descent
0.412020
On the Generalization Benefit of Noise in Stochastic Gradient Descent · ICML 2020
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.412020
Sparse GPU kernels for deep learning · SC 2020
Machine learning › Generative modeling
autoregressive model
0.412019
Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset · ICLR 2019
Machine learning › Generative modeling
music generation
0.412019
Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset · ICLR 2019
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.312018
Parallel WaveNet: Fast High-Fidelity Speech Synthesis · ICML 2018
Machine learning › Efficient and distributed learning
low-precision training
0.312018
Mixed Precision Training · ICLR (Poster) 2018
Natural language and speech › Speech recognition and synthesis › speech synthesis
neural speech synthesis
0.312018
Parallel WaveNet: Fast High-Fidelity Speech Synthesis · ICML 2018
Natural language and speech › Speech recognition and synthesis › speech synthesis › neural speech synthesis
parallel speech synthesis
0.312018
Parallel WaveNet: Fast High-Fidelity Speech Synthesis · ICML 2018
Machine learning › Efficient and distributed learning › model compression › pruning
weight pruning
0.312018
Efficient Neural Audio Synthesis · ICML 2018

Methods — techniques the papers use, named apart from their topics

differentiable encoder · 1.1chunked cross-attention · 1.1pruning · 0.9recurrent neural network · 0.8step unrolling · 0.6sparse training · 0.6scaling law estimation · 0.6power-law scaling · 0.6effective parameter count · 0.6denoising autoencoder · 0.6sparse matrix-dense matrix multiplication · 0.4sparse kernels · 0.4sampled dense-dense matrix multiplication · 0.4depthwise separable convolution · 0.4factorized modeling · 0.4autoregressive generation · 0.4persistent kernel · 0.2batch dispatch · 0.2
YearPublicationVenuePosition
2022 Step-unrolled Denoising Autoencoders for Text Generation
Nikolay Savinov, Junyoung Chung, Mikolaj Binkowski, Erich Elsen, Aäron van den Oord
ICLR4
2022 Improving Language Models by Retrieving from Trillions of Tokens
abstract
We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a 2 trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance to GPT-3 and Jurassic-1 on the Pile, despite using 25{\texttimes} fewer parameters. After fine-tuning, RETRO performance translates to downstream knowledge-intensive tasks such as question answering. RETRO combines a frozen Bert retriever, a differentiable encoder and a chunked cross-attention mechanism to predict tokens based on an order of magnitude more data than what is typically consumed during training. We typically train RETRO from scratch, yet can also rapidly RETROfit pre-trained transformers with retrieval and still achieve good performance. Our work opens up new avenues for improving language models through explicit memory at unprecedented scale.
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche 0002, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Albin Cassirer, Andrew Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, Laurent Sifre
ICML27
2022 Unified Scaling Laws for Routed Language Models
abstract
The performance of a language model has been shown to be effectively modeled as a power-law in its parameter count. Here we study the scaling behaviors of Routing Networks: architectures that conditionally use only a subset of their parameters while processing an input. For these models, parameter count and computational requirement form two independent axes along which an increase leads to better performance. In this work we derive and justify scaling laws defined on these two variables which generalize those known for standard language models and describe the performance of a wide range of routing architectures trained via three different techniques. Afterwards we provide two applications of these laws: first deriving an Effective Parameter Count along which all models scale at the same rate, and then using the scaling coefficients to give a quantitative comparison of the three routing techniques considered. Our analysis derives from an extensive evaluation of Routing Networks across five orders of magnitude of size, including models with hundreds of experts and hundreds of billions of parameters.
Aidan Clark, Diego de Las Casas, Aurelia Guy, Arthur Mensch, Michela Paganini, Jordan Hoffmann, Bogdan Damoc, Blake A. Hechtman, Trevor Cai, Sebastian Borgeaud, George van den Driessche 0002, Eliza Rutherford, Tom Hennigan, Matthew J. Johnson 0002, Albin Cassirer, Elena Buchatskaya, David Budden, Laurent Sifre, Simon Osindero, Oriol Vinyals, Marc'Aurelio Ranzato, Jack W. Rae, Erich Elsen, Koray Kavukcuoglu, Karen Simonyan
ICML24
2022 The State of Sparse Training in Deep Reinforcement Learning
abstract
The use of sparse neural networks has seen rapid growth in recent years, particularly in computer vision. Their appeal stems largely from the reduced number of parameters required to train and store, as well as in an increase in learning efficiency. Somewhat surprisingly, there have been very few efforts exploring their use in Deep Reinforcement Learning (DRL). In this work we perform a systematic investigation into applying a number of existing sparse training techniques on a variety of DRL agents and environments. Our results corroborate the findings from sparse training in the computer vision domain {–}sparse networks perform better than dense networks for the same parameter count{–} in the DRL domain. We provide detailed analyses on how the various components in DRL are affected by the use of sparse networks and conclude by suggesting promising avenues for improving the effectiveness of sparse training methods, as well as for advancing their use in DRL.
Laura Graesser, Utku Evci, Erich Elsen, Pablo Samuel Castro
ICML3
2022 An empirical analysis of compute-optimal large language model training
abstract
We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant. By training over 400 language models ranging from 70 million to over 16 billion parameters on 5 to 500 billion tokens, we find that for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled. We test this hypothesis by training a predicted compute-optimal model, Chinchilla, that uses the same compute budget as Gopher but with 70B parameters and 4$\times$ more data. Chinchilla uniformly and significantly outperformsGopher (280B), GPT-3 (175B), Jurassic-1 (178B), and Megatron-Turing NLG (530B) on a large range of downstream evaluation tasks. This also means that Chinchilla uses substantially less compute for fine-tuning and inference, greatly facilitating downstream usage. As a highlight, Chinchilla reaches a state-of-the-art average accuracy of 67.5% on the MMLU benchmark, a 7% improvement over Gopher.
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche 0002, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack W. Rae, Laurent Sifre
NeurIPS19
2021 End-to-end Adversarial Text-to-Speech
Jeff Donahue, Sander Dieleman, Mikolaj Binkowski, Erich Elsen, Karen Simonyan
ICLR4
2021 Practical Real Time Recurrent Learning with a Sparse Approximation
Jacob Menick, Erich Elsen, Utku Evci, Simon Osindero, Karen Simonyan, Alex Graves
ICLR2
2020 Fast Sparse ConvNets
abstract
Historically, the pursuit of efficient inference has been one of the driving forces behind the research into new deep learning architectures and building blocks. Some of the recent examples include: the squeeze-and-excitation module, depthwise separable convolutions in Xception, and the inverted bottleneck in MobileNet v2. Notably, in all of these cases, the resulting building blocks enabled not only higher efficiency, but also higher accuracy, and found wide adoption in the field. In this work, we further expand the arsenal of efficient building blocks for neural network architectures; but instead of combining standard primitives (such as convolution), we advocate for the replacement of these dense primitives with their sparse counterparts. While the idea of using sparsity to decrease the parameter count is not new, the conventional wisdom is that this reduction in theoretical FLOPs does not translate into real-world efficiency gains. We aim to correct this misconception by introducing a family of efficient sparse kernels for several hardware platforms, which we plan to open source for the benefit of the community. Equipped with our efficient implementation of sparse primitives, we show that sparse versions of MobileNet v1 and MobileNet v2 architectures substantially outperform strong dense baselines on the efficiency-accuracy curve. On Snapdragon 835 our sparse networks outperform their dense equivalents by 1.3 - 2.4× - equivalent to approximately one entire generation of improvement. We hope that our findings will facilitate wider adoption of sparsity as a tool for creating efficient and accurate deep learning architectures.
Erich Elsen, Marat Dukhan, Trevor Gale, Karen Simonyan
CVPR1
2020 High Fidelity Speech Synthesis with Adversarial Networks
Mikolaj Binkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C. Cobo, Karen Simonyan
ICLR5
2020 Rigging the Lottery: Making All Tickets Winners
abstract
Many applications require sparse neural networks due to space or inference time restrictions. There is a large body of work on training dense networks to yield sparse networks for inference, but this limits the size of the largest trainable sparse model to that of the largest trainable dense model. In this paper we introduce a method to train sparse neural networks with a fixed parameter count and a fixed computational cost throughout training, without sacrificing accuracy relative to existing dense-to-sparse training methods. Our method updates the topology of the sparse network during training by using parameter magnitudes and infrequent gradient calculations. We show that this approach requires fewer floating-point operations (FLOPs) to achieve a given level of accuracy compared to prior techniques. We demonstrate state-of-the-art sparse training results on a variety of networks and datasets, including ResNet-50, MobileNets on Imagenet-2012, and RNNs on WikiText-103. Finally, we provide some insights into why allowing the topology to change during the optimization can overcome local minima encountered when the topology remains static.
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, Erich Elsen
ICML5
2020 On the Generalization Benefit of Noise in Stochastic Gradient Descent
abstract
It has long been argued that minibatch stochastic gradient descent can generalize better than large batch gradient descent in deep neural networks. However recent papers have questioned this claim, arguing that this effect is simply a consequence of suboptimal hyperparameter tuning or insufficient compute budgets when the batch size is large. In this paper, we perform carefully designed experiments and rigorous hyperparameter sweeps on a range of popular models, which verify that small or moderately large batch sizes can substantially outperform very large batches on the test set. This occurs even when both models are trained for the same number of iterations and large batches achieve smaller training losses. Our results confirm that the noise in stochastic gradients can enhance generalization. We study how the optimal learning rate schedule changes as the epoch budget grows, and we provide a theoretical account of our observations based on the stochastic differential equation perspective of SGD dynamics.
Samuel L. Smith, Erich Elsen, Soham De
ICML2
2020 Top-KAST: Top-K Always Sparse Training
abstract
Sparse neural networks are becoming increasingly important as the field seeks to improve the performance of existing models by scaling them up, while simultaneously trying to reduce power consumption and computational footprint. Unfortunately, most existing methods for inducing performant sparse models still entail the instantiation of dense parameters, or dense gradients in the backward-pass, during training. For very large models this requirement can be prohibitive. In this work we propose Top-KAST, a method that preserves constant sparsity throughout training (in both the forward and backward-passes). We demonstrate the efficacy of our approach by showing that it performs comparably to or better than previous works when training models on the established ImageNet benchmark, whilst fully maintaining sparsity. In addition to our ImageNet results, we also demonstrate our approach in the domain of language modeling where the current best performing architectures tend to have tens of billions of parameters and scaling up does not yet seem to have saturated performance. Sparse versions of these architectures can be run with significantly fewer resources, making them more widely accessible and applicable. Furthermore, in addition to being effective, our approach is straightforward and can easily be implemented in a wide range of existing machine learning frameworks with only a few additional lines of code. We therefore hope that our contribution will help enable the broader community to explore the potential held by massive models, without incurring massive computational cost.
Siddhant M. Jayakumar, Razvan Pascanu, Jack W. Rae, Simon Osindero, Erich Elsen
NeurIPS5
2020 Sparse GPU kernels for deep learning
abstract
Scientific workloads have traditionally exploited high levels of sparsity to accelerate computation and reduce memory requirements. While deep neural networks can be made sparse, achieving practical speedups on GPUs is difficult because these applications have relatively moderate levels of sparsity that are not sufficient for existing sparse kernels to outperform their dense counterparts. In this work, we study sparse matrices from deep learning applications and identify favorable properties that can be exploited to accelerate computation. Based on these insights, we develop high-performance GPU kernels for two sparse matrix operations widely applicable in neural networks: sparse matrix-dense matrix multiplication and sampled dense- dense matrix multiplication. Our kernels reach 27% of singleprecision peak on Nvidia V100 GPUs. Using our kernels, we demonstrate sparse Transformer and MobileNet models that achieve 1.2-2.1× speedups and up to 12.8× memory savings without sacrificing accuracy.
Trevor Gale, Matei Zaharia, Cliff Young, Erich Elsen
SC4
2019 Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset
Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse H. Engel, Douglas Eck
ICLR7
2018 Mixed Precision Training
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Frederick Diamos, Erich Elsen, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh
ICLR (Poster)5
2018 Efficient Neural Audio Synthesis
abstract
Sequential models achieve state-of-the-art results in audio, visual and textual domains with respect to both estimating the data distribution and generating desired samples. Efficient sampling for this class of models at the cost of little to no loss in quality has however remained an elusive problem. With a focus on text-to-speech synthesis, we describe a set of general techniques for reducing sampling time while maintaining high output quality. We first describe a single-layer recurrent neural network, the WaveRNN, with a dual softmax layer that matches the quality of the state-of-the-art WaveNet model. The compact form of the network makes it possible to generate 24 kHz 16-bit audio 4 times faster than real time on a GPU. Secondly, we apply a weight pruning technique to reduce the number of weights in the WaveRNN. We find that, for a constant number of parameters, large sparse networks perform better than small dense networks and this relationship holds past sparsity levels of more than 96%. The small number of weights in a Sparse WaveRNN makes it possible to sample high-fidelity audio on a mobile phone CPU in real time. Finally, we describe a new dependency scheme for sampling that lets us trade a constant number of non-local, distant dependencies for the ability to generate samples in batches. The Batch WaveRNN produces 8 samples per step without loss of quality and offers orthogonal ways of further increasing sampling efficiency.
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aäron van den Oord, Sander Dieleman, Koray Kavukcuoglu
ICML2
2018 Parallel WaveNet: Fast High-Fidelity Speech Synthesis
abstract
The recently-developed WaveNet architecture is the current state of the art in realistic speech synthesis, consistently rated as more natural sounding for many different languages than any previous system. However, because WaveNet relies on sequential generation of one audio sample at a time, it is poorly suited to today’s massively parallel computers, and therefore hard to deploy in a real-time production setting. This paper introduces Probability Density Distillation, a new method for training a parallel feed-forward network from a trained WaveNet with no significant difference in quality. The resulting system is capable of generating high-fidelity speech samples at more than 20 times faster than real-time, a 1000x speed up relative to the original WaveNet, and capable of serving multiple English and Japanese voices in a production setting.
Aäron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche 0002, Edward Lockhart, Luis C. Cobo, Florian Stimberg, Norman Casagrande, Dominik Grewe, Seb Noury, Sander Dieleman, Erich Elsen, Nal Kalchbrenner, Heiga Zen, Alex Graves, Helen King, Tom Walters, Daniel Belov, Demis Hassabis
ICML15
2017 DSD: Dense-Sparse-Dense Training for Deep Neural Networks
Song Han 0003, Jeff Pool, Sharan Narang, Huizi Mao, Enhao Gong, Shijian Tang, Erich Elsen, Peter Vajda, Manohar Paluri, Bryan Catanzaro, William J. Dally
ICLR (Poster)7
2017 Exploring Sparsity in Recurrent Neural Networks
Sharan Narang, Gregory Frederick Diamos, Shubho Sengupta, Erich Elsen
ICLR (Poster)4
2016 Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin
abstract
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, enabling experiments that previously took weeks to now run in days. This allows us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale.
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Erich Elsen, Jesse H. Engel, Linxi Fan, Christopher Fougner, Awni Y. Hannun, Billy Jun, Tony Han, Patrick LeGresley, Xiangang Li, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Sheng Qian, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Chong Wang 0002, Zhiqian Wang, Dani Yogatama, Zhenyao Zhu
ICML13
2016 Persistent RNNs: Stashing Recurrent Weights On-Chip
abstract
This paper introduces a new technique for mapping Deep Recurrent Neural Networks (RNN) efficiently onto GPUs. We show how it is possi- ble to achieve substantially higher computational throughput at low mini-batch sizes than direct implementations of RNNs based on matrix multiplications. The key to our approach is the use of persistent computational kernels that exploit the GPU’s inverted memory hierarchy to reuse network weights over multiple timesteps. Our initial implementation sustains 2.8 TFLOP/s at a mini-batch size of 4 on an NVIDIA TitanX GPU. This provides a 16x reduction in activation memory footprint, enables model training with 12x more parameters on the same hardware, allows us to strongly scale RNN training to 128 GPUs, and allows us to efficiently explore end-to-end speech recognition models with over 100 layers.
Gregory Frederick Diamos, Shubho Sengupta, Bryan Catanzaro, Mike Chrzanowski, Adam Coates 0002, Erich Elsen, Jesse H. Engel, Awni Y. Hannun, Sanjeev Satheesh
ICML6
2011 Liszt: a domain specific language for building portable mesh-based PDE solvers
abstract
Heterogeneous computers with processors and accelerators are becoming widespread in scientific computing. However, it is difficult to program hybrid architectures and there is no commonly accepted programming model. Ideally, applications should be written in a way that is portable to many platforms, but providing this portability for general programs is a hard problem.
Zach DeVito, Niels Joubert, Francisco Palacios Ortega, Stephen Oakley, Montserrat Medina, Mike Barrientos, Erich Elsen, Frank Ham, Alex Aiken, Karthik Duraisamy, Eric Darve, Juan J. Alonso, Pat Hanrahan
SC7
2006 Poster reception - N-Body simulation on GPUs
abstract
Commercial graphics processors (GPUs) have high compute capacity at very low cost, which makes them attractive for general purpose scientic computing. In this poster we show how graphics processors can be used for N-body simulations to obtain large improvements in performance over current generation CPUs. We have developed a highly optimized algorithm for performing the O(N^2) force calculations that constitute the major part of stellar and molecular dynamics simulations. In the calculations, we achieve sustained performance of nearly 100 GFlops on an ATI X1900XTX. The performance on GPUs 25x an Intel Pentium4, and 2x specialized hardware such as GRAPE-6A, but at a fraction of the cost. Furthermore, the wide availability of GPUs has signicant implications for cluster computing and distributed computing efforts like [email protected]
Erich Elsen, Mike Houston, Vaidyanathan Vishal, Eric Darve, Pat Hanrahan, Vijay S. Pande
SC1