Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Mike Chrzanowski

dblp:173/5380 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Deep learning architectures and training · 46% Speech recognition and synthesis · 27% Trustworthy machine learning · 16%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 77% High-performance computing · 12% Cloud and datacenter computing · 12%

Topics — the 18 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
attention mechanism
1.132020
Towards Robust Image Classification Using Sequential Attention Models · CVPR 2020
Towards Interpretable Reinforcement Learning Using Attention Augmented Agents · NeurIPS 2019
Relational recurrent neural networks · NeurIPS 2018
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.412020
Towards Robust Image Classification Using Sequential Attention Models · CVPR 2020
Machine learning › Trustworthy machine learning › interpretability
explainable reinforcement learning
0.412019
Towards Interpretable Reinforcement Learning Using Attention Augmented Agents · NeurIPS 2019
Machine learning › Deep learning architectures and training
memory-augmented neural networks
0.312018
Relational recurrent neural networks · NeurIPS 2018
Machine learning › Deep learning architectures and training › attention mechanism
multi-head attention
0.312018
Relational recurrent neural networks · NeurIPS 2018
Machine learning › Deep learning architectures and training
recurrent neural network
0.312018
Relational recurrent neural networks · NeurIPS 2018
Natural language and speech › Speech recognition and synthesis › pronunciation modeling
grapheme-to-phoneme conversion
0.312017
Deep Voice: Real-time Neural Text-to-Speech · ICML 2017
Natural language and speech › Speech recognition and synthesis › speech synthesis
neural speech synthesis
0.312017
Deep Voice: Real-time Neural Text-to-Speech · ICML 2017
Natural language and speech › Speech recognition and synthesis
prosody prediction
0.312017
Deep Voice: Real-time Neural Text-to-Speech · ICML 2017
Natural language and speech › Speech recognition and synthesis
text-to-speech synthesis
0.312017
Deep Voice: Real-time Neural Text-to-Speech · ICML 2017
Machine learning › Deep learning architectures and training › neural network training
end-to-end deep learning
0.212016
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
end-to-end speech recognition
0.212016
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016
GPUs and heterogeneous computing
deep learning on GPUs
0.212016
Persistent RNNs: Stashing Recurrent Weights On-Chip · ICML 2016
GPUs and heterogeneous computing
GPU computing
0.212016
Persistent RNNs: Stashing Recurrent Weights On-Chip · ICML 2016
Machine learning › Reinforcement learning
deep reinforcement learning
0.112019
Towards Interpretable Reinforcement Learning Using Attention Augmented Agents · NeurIPS 2019
Natural language and speech › Language models and text generation
language modeling
0.112018
Relational recurrent neural networks · NeurIPS 2018
Machine learning › Reinforcement learning
partially observable reinforcement learning
0.112018
Relational recurrent neural networks · NeurIPS 2018
High-performance computing
performance optimization at scale
0.112016
Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016

Methods — techniques the papers use, named apart from their topics

top-down sequential process · 0.4recurrent attention · 0.4adversarial training · 0.4soft attention · 0.4deep q-network · 0.4relational reasoning · 0.3dot product attention · 0.3wavenet · 0.3deep neural network · 0.3connectionist temporal classification · 0.3recurrent neural network · 0.2persistent kernel · 0.2batch dispatch · 0.2GPU-based inference · 0.2
YearPublicationVenuePosition
2020 Towards Robust Image Classification Using Sequential Attention Models
abstract
In this paper we propose to augment a modern neuralnetwork architecture with an attention model inspired by human perception. Specifically, we adversarially train and analyze a neural model incorporating a human inspired, visual attention component that is guided by a recurrent top-down sequential process. Our experimental evaluation uncovers several notable findings about the robustness and behavior of this new model. First, introducing attention to the model significantly improves adversarial robustness resulting in state-of-the-art ImageNet accuracies under a wide range of random targeted attack strengths. Second, we show that by varying the number of attention steps (glances/fixations) for which the model is unrolled, we are able to make its defense capabilities stronger, even in light of stronger attacks - resulting in a “computational race” between the attacker and the defender. Finally, we show that some of the adversarial examples generated by attacking our model are quite different from conventional adversarial examples - they contain global, salient and spatially coherent structures coming from the target class that would be recognizable even to a human, and work by distracting the attention of the model away from the main object in the original image.
Daniel Zoran, Mike Chrzanowski, Po-Sen Huang, Sven Gowal, Alex Mott, Pushmeet Kohli
CVPR2
2019 Towards Interpretable Reinforcement Learning Using Attention Augmented Agents
abstract
Inspired by recent work in attention models for image captioning and question answering, we present a soft attention model for the reinforcement learning domain. This model bottlenecks the view of an agent by a soft, top-down attention mechanism, forcing the agent to focus on task-relevant information by sequentially querying its view of the environment. The output of the attention mechanism allows direct observation of the information used by the agent to select its actions, enabling easier interpretation of this model than of traditional models. We analyze the different strategies the agents learn and show that a handful of strategies arise repeatedly across different games. We also show that the model learns to query separately about space and content (where'' vs.what''). We demonstrate that an agent using this mechanism can achieve performance competitive with state-of-the-art models on ATARI tasks while still being interpretable.
Alex Mott, Daniel Zoran, Mike Chrzanowski, Daan Wierstra, Danilo Jimenez Rezende
NeurIPS3
2018 Relational recurrent neural networks
abstract
Memory-based neural networks model temporal data by leveraging an ability to remember information for long periods. It is unclear, however, whether they also have an ability to perform complex relational reasoning with the information they remember. Here, we first confirm our intuitions that standard memory architectures may struggle at tasks that heavily involve an understanding of the ways in which entities are connected -- i.e., tasks involving relational reasoning. We then improve upon these deficits by using a new memory module -- a Relational Memory Core (RMC) -- which employs multi-head dot product attention to allow memories to interact. Finally, we test the RMC on a suite of tasks that may profit from more capable relational reasoning across sequential information, and show large gains in RL domains (BoxWorld & Mini PacMan), program evaluation, and language modeling, achieving state-of-the-art results on the WikiText-103, Project Gutenberg, and GigaWord datasets.
Adam Santoro, Ryan Faulkner 0001, David Raposo, Jack W. Rae, Mike Chrzanowski, Theophane Weber, Daan Wierstra, Oriol Vinyals, Razvan Pascanu, Timothy P. Lillicrap
NeurIPS5
2017 Deep Voice: Real-time Neural Text-to-Speech
abstract
We present Deep Voice, a production-quality text-to-speech system constructed entirely from deep neural networks. Deep Voice lays the groundwork for truly end-to-end neural speech synthesis. The system comprises five major building blocks: a segmentation model for locating phoneme boundaries, a grapheme-to-phoneme conversion model, a phoneme duration prediction model, a fundamental frequency prediction model, and an audio synthesis model. For the segmentation model, we propose a novel way of performing phoneme boundary detection with deep neural networks using connectionist temporal classification (CTC) loss. For the audio synthesis model, we implement a variant of WaveNet that requires fewer parameters and trains faster than the original. By using a neural network for each component, our system is simpler and more flexible than traditional text-to-speech systems, where each component requires laborious feature engineering and extensive domain expertise. Finally, we show that inference with our system can be performed faster than real time and describe optimized WaveNet inference kernels on both CPU and GPU that achieve up to 400x speedups over existing implementations.
Sercan Ö. Arik, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Andrew Gibiansky, Yongguo Kang, John Miller 0001, Andrew Y. Ng, Jonathan Raiman, Shubho Sengupta, Mohammad Shoeybi
ICML2
2016 Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin
abstract
We show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, enabling experiments that previously took weeks to now run in days. This allows us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale.
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Erich Elsen, Jesse H. Engel, Linxi Fan, Christopher Fougner, Awni Y. Hannun, Billy Jun, Tony Han, Patrick LeGresley, Xiangang Li, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Sheng Qian, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Chong Wang 0002, Zhiqian Wang, Dani Yogatama, Zhenyao Zhu
ICML10
2016 Persistent RNNs: Stashing Recurrent Weights On-Chip
abstract
This paper introduces a new technique for mapping Deep Recurrent Neural Networks (RNN) efficiently onto GPUs. We show how it is possi- ble to achieve substantially higher computational throughput at low mini-batch sizes than direct implementations of RNNs based on matrix multiplications. The key to our approach is the use of persistent computational kernels that exploit the GPU’s inverted memory hierarchy to reuse network weights over multiple timesteps. Our initial implementation sustains 2.8 TFLOP/s at a mini-batch size of 4 on an NVIDIA TitanX GPU. This provides a 16x reduction in activation memory footprint, enables model training with 12x more parameters on the same hardware, allows us to strongly scale RNN training to 128 GPUs, and allows us to efficiently explore end-to-end speech recognition models with over 100 layers.
Gregory Frederick Diamos, Shubho Sengupta, Bryan Catanzaro, Mike Chrzanowski, Adam Coates 0002, Erich Elsen, Jesse H. Engel, Awni Y. Hannun, Sanjeev Satheesh
ICML4