Paul A. Tucker

dblp:39/814 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1Computer networks · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 90% Optimization for machine learning · 10%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
large-scale learning
0.212016
TensorFlow: A System for Large-Scale Machine Learning · OSDI 2016
Distributed systems
large-scale machine learning systems
0.212016
TensorFlow: A System for Large-Scale Machine Learning · OSDI 2016
Machine learning › Efficient and distributed learning
distributed training
0.112012
Large Scale Distributed Deep Networks · NIPS 2012
Distributed systems
distributed machine learning
0.112012
Large Scale Distributed Deep Networks · NIPS 2012

Methods — techniques the papers use, named apart from their topics

asynchronous SGD · 0.3Sandblaster L-BFGS · 0.3Downpour SGD · 0.3data flow graphs · 0.2data flow graph · 0.2
YearPublicationVenuePosition
2016 TensorFlow: A System for Large-Scale Machine Learning
Martín Abadi, Paul Barham 0001, Jianmin Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Xiaoqiang Zheng
OSDI17
2012 Large Scale Distributed Deep Networks
abstract
Recent work in unsupervised feature learning and deep learning has shown that being able to train large models can dramatically improve performance. In this paper, we consider the problem of training a deep network with billions of parameters using tens of thousands of CPU cores. We have developed a software framework called DistBelief that can utilize computing clusters with thousands of machines to train large models. Within this framework, we have developed two algorithms for large-scale distributed training: (i) Downpour SGD, an asynchronous stochastic gradient descent procedure supporting a large number of model replicas, and (ii) Sandblaster, a framework that supports for a variety of distributed batch optimization procedures, including a distributed implementation of L-BFGS. Downpour SGD and Sandblaster L-BFGS both increase the scale and speed of deep network training. We have successfully used our system to train a deep network 100x larger than previously reported in the literature, and achieves state-of-the-art performance on ImageNet, a visual object recognition task with 16 million images and 21k categories. We show that these same techniques dramatically accelerate the training of a more modestly sized deep network for a commercial speech recognition service. Although we focus on and report performance of these methods as applied to training large neural networks, the underlying algorithms are applicable to any gradient-based machine learning algorithm.
Jeffrey Dean, Gregory S. Corrado, Rajat Monga, Kai Chen 0010, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc'Aurelio Ranzato, Andrew W. Senior, Paul A. Tucker, Andrew Y. Ng
NIPS10
2002 Minimax problems with bitonic matrices
abstract
Abstract The minimax problem is a new optimization problem which substitutes maximum for addition in the constraint inequalities of a linear program. We show how a problem with n variables and m constraints can be reduced to a set cover problem with nm variables and m constraints. We define the mountain property on the coefficients in each column of the constraint matrix of a minimax problem and its generalization—the bitonic property. It is shown that these are equivalent, respectively, to set cover problems with consecutive 1's in each column or circular 1's in each column thus solving the periodic scheduling problem. We present a shortest path algorithm to solve a minimax problem with the mountain property in time O(mn + log m). For bitonic matrix problems, we present an algorithm of complexity O(m2(n + log n)). The same algorithms are used to solve a set cover problem on v sets and r elements to be covered in O(v + r log r) for a problem with consecutive 1's in each column and in O(v(v + r log r)) for a problem with circular 1's in each column. We further establish that r is at least \documentclass{article}\pagestyle{empty} \begin{document}$\Omega(\sqrt{v})$\end{document} and at most O(v). We also provide an efficient algorithm for recognizing bitonic matrices in O(mn log m) time. © 2002 Wiley Periodicals, Inc.
Dorit S. Hochbaum, Paul A. Tucker
Networks2