Rajat Monga

dblp:99/10669 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4Graphics, computer vision, multimedia, augmented reality and games · 3Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Deep learning architectures and training · 37% Efficient and distributed learning · 34% Video understanding and tracking · 20%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 100%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
0.522018
Dynamic control flow in large-scale machine learning · EuroSys 2018
Large Scale Distributed Deep Networks · NIPS 2012
Distributed systems
distributed machine learning
0.522018
Dynamic control flow in large-scale machine learning · EuroSys 2018
Large Scale Distributed Deep Networks · NIPS 2012
Machine learning › Efficient and distributed learning
large-scale learning
0.212016
TensorFlow: A System for Large-Scale Machine Learning · OSDI 2016
Distributed systems
large-scale machine learning systems
0.212016
TensorFlow: A System for Large-Scale Machine Learning · OSDI 2016
Machine learning › Deep learning architectures and training
convolutional neural network
0.212015
Beyond short snippets: Deep networks for video classification · CVPR 2015
Computer vision › Video understanding and tracking › temporal modeling
long-range temporal modeling
0.212015
Beyond short snippets: Deep networks for video classification · CVPR 2015
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM
0.212015
Beyond short snippets: Deep networks for video classification · CVPR 2015
Machine learning › Deep learning architectures and training
recurrent neural network
0.212015
Beyond short snippets: Deep networks for video classification · CVPR 2015
Computer vision › Video understanding and tracking
video classification
0.212015
Beyond short snippets: Deep networks for video classification · CVPR 2015
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.112012
Building high-level features using large scale unsupervised learning · ICML 2012

Methods — techniques the papers use, named apart from their topics

data flow graphs · 0.6data flow graph · 0.6asynchronous SGD · 0.3Sandblaster L-BFGS · 0.3Downpour SGD · 0.3optical flow · 0.2LSTM · 0.2CNN · 0.2
YearPublicationVenuePosition
2018 Dynamic control flow in large-scale machine learning
abstract
Many recent machine learning models rely on fine-grained dynamic control flow for training and inference. In particular, models based on recurrent neural networks and on reinforcement learning depend on recurrence relations, data-dependent conditional execution, and other features that call for dynamic control flow. These applications benefit from the ability to make rapid control-flow decisions across a set of computing devices in a distributed system. For performance, scalability, and expressiveness, a machine learning system must support dynamic control flow in distributed and heterogeneous environments.
Martín Abadi, Paul Barham 0001, Eugene Brevdo, Michael Burrows, Andy Davis, Jeffrey Dean, Sanjay Ghemawat, Tim Harley, Peter Hawkins, Michael Isard, Manjunath Kudlur, Rajat Monga, Derek Gordon Murray, Xiaoqiang Zheng
EuroSys13
2016 TensorFlow: A System for Large-Scale Machine Learning
Martín Abadi, Paul Barham 0001, Jianmin Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Xiaoqiang Zheng
OSDI13
2015 Beyond short snippets: Deep networks for video classification
abstract
Convolutional neural networks (CNNs) have been extensively applied for image recognition problems giving state-of-the-art results on recognition, detection, segmentation and retrieval. In this work we propose and evaluate several deep neural network architectures to combine image information across a video over longer time periods than previously attempted. We propose two methods capable of handling full length videos. The first method explores various convolutional temporal feature pooling architectures, examining the various design choices which need to be made when adapting a CNN for this task. The second proposed method explicitly models the video as an ordered sequence of frames. For this purpose we employ a recurrent neural network that uses Long Short-Term Memory (LSTM) cells which are connected to the output of the underlying CNN. Our best networks exhibit significant performance improvements over previously published results on the Sports 1 million dataset (73.1% vs. 60.9%) and the UCF-101 datasets with (88.6% vs. 88.0%) and without additional optical flow information (82.6% vs. 73.0%).
Joe Yue-Hei Ng, Matthew J. Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, George Toderici
CVPR5
2014 Sequence discriminative distributed training of long short-term memory recurrent neural networks
Hasim Sak, Oriol Vinyals, Georg Heigold, Andrew W. Senior, Erik McDermott, Rajat Monga, Mark Z. Mao
INTERSPEECH6
2013 On rectified linear units for speech processing
abstract
Deep neural networks have recently become the gold standard for acoustic modeling in speech recognition systems. The key computational unit of a deep network is a linear projection followed by a point-wise non-linearity, which is typically a logistic function. In this work, we show that we can improve generalization and make training of deep networks faster and simpler by substituting the logistic units with rectified linear units. These units are linear when their input is positive and zero otherwise. In a supervised setting, we can successfully train very deep nets from random initialization on a large vocabulary speech recognition task achieving lower word error rates than using a logistic network with the same topology. Similarly in an unsupervised setting, we show how we can learn sparse features that can be useful for discriminative tasks. All our experiments are executed in a distributed environment using several hundred machines and several hundred hours of speech data.
Matthew D. Zeiler, Marc'Aurelio Ranzato, Rajat Monga, Mark Z. Mao, Quoc V. Le, Patrick Nguyen, Andrew W. Senior, Vincent Vanhoucke, Jeffrey Dean, Geoffrey E. Hinton
ICASSP3
2012 Building high-level features using large scale unsupervised learning
Quoc V. Le, Marc'Aurelio Ranzato, Rajat Monga, Matthieu Devin, Gregory S. Corrado, Kai Chen 0010, Jeffrey Dean, Andrew Y. Ng
ICML3
2012 Large Scale Distributed Deep Networks
abstract
Recent work in unsupervised feature learning and deep learning has shown that being able to train large models can dramatically improve performance. In this paper, we consider the problem of training a deep network with billions of parameters using tens of thousands of CPU cores. We have developed a software framework called DistBelief that can utilize computing clusters with thousands of machines to train large models. Within this framework, we have developed two algorithms for large-scale distributed training: (i) Downpour SGD, an asynchronous stochastic gradient descent procedure supporting a large number of model replicas, and (ii) Sandblaster, a framework that supports for a variety of distributed batch optimization procedures, including a distributed implementation of L-BFGS. Downpour SGD and Sandblaster L-BFGS both increase the scale and speed of deep network training. We have successfully used our system to train a deep network 100x larger than previously reported in the literature, and achieves state-of-the-art performance on ImageNet, a visual object recognition task with 16 million images and 21k categories. We show that these same techniques dramatically accelerate the training of a more modestly sized deep network for a commercial speech recognition service. Although we focus on and report performance of these methods as applied to training large neural networks, the underlying algorithms are applicable to any gradient-based machine learning algorithm.
Jeffrey Dean, Gregory S. Corrado, Rajat Monga, Kai Chen 0010, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc'Aurelio Ranzato, Andrew W. Senior, Paul A. Tucker, Andrew Y. Ng
NIPS3