EDBT 2026 Demo / reviewers in the wild / expert
Mike Chrzanowski
dblp:173/5380
· DBLP profile ↗
6ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Deep learning architectures and training · 46% Speech recognition and synthesis · 27% Trustworthy machine learning · 16% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
GPUs and heterogeneous computing · 77% High-performance computing · 12% Cloud and datacenter computing · 12% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
attention mechanism |
1.1 | 3 | 2020 | Towards Robust Image Classification Using Sequential Attention Models · CVPR 2020 Towards Interpretable Reinforcement Learning Using Attention Augmented Agents · NeurIPS 2019 Relational recurrent neural networks · NeurIPS 2018 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.4 | 1 | 2020 | Towards Robust Image Classification Using Sequential Attention Models · CVPR 2020 |
Machine learning › Trustworthy machine learning › interpretability
explainable reinforcement learning |
0.4 | 1 | 2019 | Towards Interpretable Reinforcement Learning Using Attention Augmented Agents · NeurIPS 2019 |
Machine learning › Deep learning architectures and training
memory-augmented neural networks |
0.3 | 1 | 2018 | Relational recurrent neural networks · NeurIPS 2018 |
Machine learning › Deep learning architectures and training › attention mechanism
multi-head attention |
0.3 | 1 | 2018 | Relational recurrent neural networks · NeurIPS 2018 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 1 | 2018 | Relational recurrent neural networks · NeurIPS 2018 |
Natural language and speech › Speech recognition and synthesis › pronunciation modeling
grapheme-to-phoneme conversion |
0.3 | 1 | 2017 | Deep Voice: Real-time Neural Text-to-Speech · ICML 2017 |
Natural language and speech › Speech recognition and synthesis › speech synthesis
neural speech synthesis |
0.3 | 1 | 2017 | Deep Voice: Real-time Neural Text-to-Speech · ICML 2017 |
Natural language and speech › Speech recognition and synthesis
prosody prediction |
0.3 | 1 | 2017 | Deep Voice: Real-time Neural Text-to-Speech · ICML 2017 |
Natural language and speech › Speech recognition and synthesis
text-to-speech synthesis |
0.3 | 1 | 2017 | Deep Voice: Real-time Neural Text-to-Speech · ICML 2017 |
Machine learning › Deep learning architectures and training › neural network training
end-to-end deep learning |
0.2 | 1 | 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
end-to-end speech recognition |
0.2 | 1 | 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016 |
GPUs and heterogeneous computing
deep learning on GPUs |
0.2 | 1 | 2016 | Persistent RNNs: Stashing Recurrent Weights On-Chip · ICML 2016 |
GPUs and heterogeneous computing
GPU computing |
0.2 | 1 | 2016 | Persistent RNNs: Stashing Recurrent Weights On-Chip · ICML 2016 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.1 | 1 | 2019 | Towards Interpretable Reinforcement Learning Using Attention Augmented Agents · NeurIPS 2019 |
Natural language and speech › Language models and text generation
language modeling |
0.1 | 1 | 2018 | Relational recurrent neural networks · NeurIPS 2018 |
Machine learning › Reinforcement learning
partially observable reinforcement learning |
0.1 | 1 | 2018 | Relational recurrent neural networks · NeurIPS 2018 |
High-performance computing
performance optimization at scale |
0.1 | 1 | 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and Mandarin · ICML 2016 |
Methods — techniques the papers use, named apart from their topics
top-down sequential process · 0.4recurrent attention · 0.4adversarial training · 0.4soft attention · 0.4deep q-network · 0.4relational reasoning · 0.3dot product attention · 0.3wavenet · 0.3deep neural network · 0.3connectionist temporal classification · 0.3recurrent neural network · 0.2persistent kernel · 0.2batch dispatch · 0.2GPU-based inference · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Towards Robust Image Classification Using Sequential Attention ModelsabstractIn this paper we propose to augment a modern neuralnetwork architecture with an attention model inspired by human perception. Specifically, we adversarially train and analyze a neural model incorporating a human inspired, visual attention component that is guided by a recurrent top-down sequential process. Our experimental evaluation uncovers several notable findings about the robustness and behavior of this new model. First, introducing attention to the model significantly improves adversarial robustness resulting in state-of-the-art ImageNet accuracies under a wide range of random targeted attack strengths. Second, we show that by varying the number of attention steps (glances/fixations) for which the model is unrolled, we are able to make its defense capabilities stronger, even in light of stronger attacks - resulting in a “computational race” between the attacker and the defender. Finally, we show that some of the adversarial examples generated by attacking our model are quite different from conventional adversarial examples - they contain global, salient and spatially coherent structures coming from the target class that would be recognizable even to a human, and work by distracting the attention of the model away from the main object in the original image. Daniel Zoran, Mike Chrzanowski, Po-Sen Huang, Sven Gowal, Alex Mott, Pushmeet Kohli |
CVPR | 2 |
| 2019 | Towards Interpretable Reinforcement Learning Using Attention Augmented AgentsabstractInspired by recent work in attention models for image captioning and question answering, we present a soft attention model for the reinforcement learning domain. This model bottlenecks the view of an agent by a soft, top-down attention mechanism, forcing the agent to focus on task-relevant information by sequentially querying its view of the environment. The output of the attention mechanism allows direct observation of the information used by the agent to select its actions, enabling easier interpretation of this model than of traditional models. We analyze the different strategies the agents learn and show that a handful of strategies arise repeatedly across different games. We also show that the model learns to query separately about space and content (where'' vs.what''). We demonstrate that an agent using this mechanism can achieve performance competitive with state-of-the-art models on ATARI tasks while still being interpretable. Alex Mott, Daniel Zoran, Mike Chrzanowski, Daan Wierstra, Danilo Jimenez Rezende |
NeurIPS | 3 |
| 2018 | Relational recurrent neural networksabstractMemory-based neural networks model temporal data by leveraging an ability to remember information for long periods. It is unclear, however, whether they also have an ability to perform complex relational reasoning with the information they remember. Here, we first confirm our intuitions that standard memory architectures may struggle at tasks that heavily involve an understanding of the ways in which entities are connected -- i.e., tasks involving relational reasoning. We then improve upon these deficits by using a new memory module -- a Relational Memory Core (RMC) -- which employs multi-head dot product attention to allow memories to interact. Finally, we test the RMC on a suite of tasks that may profit from more capable relational reasoning across sequential information, and show large gains in RL domains (BoxWorld & Mini PacMan), program evaluation, and language modeling, achieving state-of-the-art results on the WikiText-103, Project Gutenberg, and GigaWord datasets. Adam Santoro, Ryan Faulkner 0001, David Raposo, Jack W. Rae, Mike Chrzanowski, Theophane Weber, Daan Wierstra, Oriol Vinyals, Razvan Pascanu, Timothy P. Lillicrap |
NeurIPS | 5 |
| 2017 | Deep Voice: Real-time Neural Text-to-SpeechabstractWe present Deep Voice, a production-quality text-to-speech system constructed entirely from deep neural networks. Deep Voice lays the groundwork for truly end-to-end neural speech synthesis. The system comprises five major building blocks: a segmentation model for locating phoneme boundaries, a grapheme-to-phoneme conversion model, a phoneme duration prediction model, a fundamental frequency prediction model, and an audio synthesis model. For the segmentation model, we propose a novel way of performing phoneme boundary detection with deep neural networks using connectionist temporal classification (CTC) loss. For the audio synthesis model, we implement a variant of WaveNet that requires fewer parameters and trains faster than the original. By using a neural network for each component, our system is simpler and more flexible than traditional text-to-speech systems, where each component requires laborious feature engineering and extensive domain expertise. Finally, we show that inference with our system can be performed faster than real time and describe optimized WaveNet inference kernels on both CPU and GPU that achieve up to 400x speedups over existing implementations. Sercan Ö. Arik, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Andrew Gibiansky, Yongguo Kang, John Miller 0001, Andrew Y. Ng, Jonathan Raiman, Shubho Sengupta, Mohammad Shoeybi |
ICML | 2 |
| 2016 | Deep Speech 2 : End-to-End Speech Recognition in English and MandarinabstractWe show that an end-to-end deep learning approach can be used to recognize either English or Mandarin Chinese speech–two vastly different languages. Because it replaces entire pipelines of hand-engineered components with neural networks, end-to-end learning allows us to handle a diverse variety of speech including noisy environments, accents and different languages. Key to our approach is our application of HPC techniques, enabling experiments that previously took weeks to now run in days. This allows us to iterate more quickly to identify superior architectures and algorithms. As a result, in several cases, our system is competitive with the transcription of human workers when benchmarked on standard datasets. Finally, using a technique called Batch Dispatch with GPUs in the data center, we show that our system can be inexpensively deployed in an online setting, delivering low latency when serving users at scale. Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Jingdong Chen, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Erich Elsen, Jesse H. Engel, Linxi Fan, Christopher Fougner, Awni Y. Hannun, Billy Jun, Tony Han, Patrick LeGresley, Xiangang Li, Libby Lin, Sharan Narang, Andrew Y. Ng, Sherjil Ozair, Ryan Prenger, Sheng Qian, Jonathan Raiman, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Chong Wang 0002, Zhiqian Wang, Dani Yogatama, Zhenyao Zhu |
ICML | 10 |
| 2016 | Persistent RNNs: Stashing Recurrent Weights On-ChipabstractThis paper introduces a new technique for mapping Deep Recurrent Neural Networks (RNN) efficiently onto GPUs. We show how it is possi- ble to achieve substantially higher computational throughput at low mini-batch sizes than direct implementations of RNNs based on matrix multiplications. The key to our approach is the use of persistent computational kernels that exploit the GPU’s inverted memory hierarchy to reuse network weights over multiple timesteps. Our initial implementation sustains 2.8 TFLOP/s at a mini-batch size of 4 on an NVIDIA TitanX GPU. This provides a 16x reduction in activation memory footprint, enables model training with 12x more parameters on the same hardware, allows us to strongly scale RNN training to 128 GPUs, and allows us to efficiently explore end-to-end speech recognition models with over 100 layers. Gregory Frederick Diamos, Shubho Sengupta, Bryan Catanzaro, Mike Chrzanowski, Adam Coates 0002, Erich Elsen, Jesse H. Engel, Awni Y. Hannun, Sanjeev Satheesh |
ICML | 4 |