VLDB 2026 Research / reviewers in the wild / expert
Rajat Monga
dblp:99/10669
· DBLP profile ↗
7ranked-venue papers
0as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4Graphics, computer vision, multimedia, augmented reality and games · 3Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Deep learning architectures and training · 37% Efficient and distributed learning · 34% Video understanding and tracking · 20% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Distributed systems · 100% |
Topics — the 10 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
0.5 | 2 | 2018 | Dynamic control flow in large-scale machine learning · EuroSys 2018 Large Scale Distributed Deep Networks · NIPS 2012 |
Distributed systems
distributed machine learning |
0.5 | 2 | 2018 | Dynamic control flow in large-scale machine learning · EuroSys 2018 Large Scale Distributed Deep Networks · NIPS 2012 |
Machine learning › Efficient and distributed learning
large-scale learning |
0.2 | 1 | 2016 | TensorFlow: A System for Large-Scale Machine Learning · OSDI 2016 |
Distributed systems
large-scale machine learning systems |
0.2 | 1 | 2016 | TensorFlow: A System for Large-Scale Machine Learning · OSDI 2016 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.2 | 1 | 2015 | Beyond short snippets: Deep networks for video classification · CVPR 2015 |
Computer vision › Video understanding and tracking › temporal modeling
long-range temporal modeling |
0.2 | 1 | 2015 | Beyond short snippets: Deep networks for video classification · CVPR 2015 |
Machine learning › Deep learning architectures and training › recurrent neural network
LSTM |
0.2 | 1 | 2015 | Beyond short snippets: Deep networks for video classification · CVPR 2015 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.2 | 1 | 2015 | Beyond short snippets: Deep networks for video classification · CVPR 2015 |
Computer vision › Video understanding and tracking
video classification |
0.2 | 1 | 2015 | Beyond short snippets: Deep networks for video classification · CVPR 2015 |
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning |
0.1 | 1 | 2012 | Building high-level features using large scale unsupervised learning · ICML 2012 |
Methods — techniques the papers use, named apart from their topics
data flow graphs · 0.6data flow graph · 0.6asynchronous SGD · 0.3Sandblaster L-BFGS · 0.3Downpour SGD · 0.3optical flow · 0.2LSTM · 0.2CNN · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Dynamic control flow in large-scale machine learningabstractMany recent machine learning models rely on fine-grained dynamic control flow for training and inference. In particular, models based on recurrent neural networks and on reinforcement learning depend on recurrence relations, data-dependent conditional execution, and other features that call for dynamic control flow. These applications benefit from the ability to make rapid control-flow decisions across a set of computing devices in a distributed system. For performance, scalability, and expressiveness, a machine learning system must support dynamic control flow in distributed and heterogeneous environments. Martín Abadi, Paul Barham 0001, Eugene Brevdo, Michael Burrows, Andy Davis, Jeffrey Dean, Sanjay Ghemawat, Tim Harley, Peter Hawkins, Michael Isard, Manjunath Kudlur, Rajat Monga, Derek Gordon Murray, Xiaoqiang Zheng |
EuroSys | 13 |
| 2016 | TensorFlow: A System for Large-Scale Machine Learning
Martín Abadi, Paul Barham 0001, Jianmin Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Xiaoqiang Zheng |
OSDI | 13 |
| 2015 | Beyond short snippets: Deep networks for video classificationabstractConvolutional neural networks (CNNs) have been extensively applied for image recognition problems giving state-of-the-art results on recognition, detection, segmentation and retrieval. In this work we propose and evaluate several deep neural network architectures to combine image information across a video over longer time periods than previously attempted. We propose two methods capable of handling full length videos. The first method explores various convolutional temporal feature pooling architectures, examining the various design choices which need to be made when adapting a CNN for this task. The second proposed method explicitly models the video as an ordered sequence of frames. For this purpose we employ a recurrent neural network that uses Long Short-Term Memory (LSTM) cells which are connected to the output of the underlying CNN. Our best networks exhibit significant performance improvements over previously published results on the Sports 1 million dataset (73.1% vs. 60.9%) and the UCF-101 datasets with (88.6% vs. 88.0%) and without additional optical flow information (82.6% vs. 73.0%). Joe Yue-Hei Ng, Matthew J. Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, George Toderici |
CVPR | 5 |
| 2014 | Sequence discriminative distributed training of long short-term memory recurrent neural networks
Hasim Sak, Oriol Vinyals, Georg Heigold, Andrew W. Senior, Erik McDermott, Rajat Monga, Mark Z. Mao |
INTERSPEECH | 6 |
| 2013 | On rectified linear units for speech processingabstractDeep neural networks have recently become the gold standard for acoustic modeling in speech recognition systems. The key computational unit of a deep network is a linear projection followed by a point-wise non-linearity, which is typically a logistic function. In this work, we show that we can improve generalization and make training of deep networks faster and simpler by substituting the logistic units with rectified linear units. These units are linear when their input is positive and zero otherwise. In a supervised setting, we can successfully train very deep nets from random initialization on a large vocabulary speech recognition task achieving lower word error rates than using a logistic network with the same topology. Similarly in an unsupervised setting, we show how we can learn sparse features that can be useful for discriminative tasks. All our experiments are executed in a distributed environment using several hundred machines and several hundred hours of speech data. Matthew D. Zeiler, Marc'Aurelio Ranzato, Rajat Monga, Mark Z. Mao, Quoc V. Le, Patrick Nguyen, Andrew W. Senior, Vincent Vanhoucke, Jeffrey Dean, Geoffrey E. Hinton |
ICASSP | 3 |
| 2012 | Building high-level features using large scale unsupervised learning
Quoc V. Le, Marc'Aurelio Ranzato, Rajat Monga, Matthieu Devin, Gregory S. Corrado, Kai Chen 0010, Jeffrey Dean, Andrew Y. Ng |
ICML | 3 |
| 2012 | Large Scale Distributed Deep NetworksabstractRecent work in unsupervised feature learning and deep learning has shown that being able to train large models can dramatically improve performance. In this paper, we consider the problem of training a deep network with billions of parameters using tens of thousands of CPU cores. We have developed a software framework called DistBelief that can utilize computing clusters with thousands of machines to train large models. Within this framework, we have developed two algorithms for large-scale distributed training: (i) Downpour SGD, an asynchronous stochastic gradient descent procedure supporting a large number of model replicas, and (ii) Sandblaster, a framework that supports for a variety of distributed batch optimization procedures, including a distributed implementation of L-BFGS. Downpour SGD and Sandblaster L-BFGS both increase the scale and speed of deep network training. We have successfully used our system to train a deep network 100x larger than previously reported in the literature, and achieves state-of-the-art performance on ImageNet, a visual object recognition task with 16 million images and 21k categories. We show that these same techniques dramatically accelerate the training of a more modestly sized deep network for a commercial speech recognition service. Although we focus on and report performance of these methods as applied to training large neural networks, the underlying algorithms are applicable to any gradient-based machine learning algorithm. Jeffrey Dean, Gregory S. Corrado, Rajat Monga, Kai Chen 0010, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc'Aurelio Ranzato, Andrew W. Senior, Paul A. Tucker, Andrew Y. Ng |
NIPS | 3 |