Samuel S. Schoenholz

dblp:190/7108 · also Samuel Stern Schoenholz · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
5since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
20 papers
Deep learning architectures and training · 44% Learning theory · 21% Optimization for machine learning · 10%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Computational science and engineering · 86% Bioinformatics and computational biology · 14%

Topics — the 30 heaviest of 47, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
2.352022
Fast Finite Width Neural Tangent Kernel · ICML 2022
Finite Versus Infinite Neural Networks: an Empirical Study · NeurIPS 2020
Disentangling Trainability and Generalization in Deep Neural Networks · ICML 2020
Machine learning › Deep learning architectures and training
weight initialization
1.542022
Deep equilibrium networks are sensitive to initialization statistics · ICML 2022
MetaInit: Initializing learning by learning to initialize · NeurIPS 2019
Mean Field Residual Networks: On the Edge of Chaos · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
mean-field approximation
1.342019
A Mean Field Theory of Batch Normalization · ICLR (Poster) 2019
Dynamical Isometry and a Mean Field Theory of CNNs: How to Train 10, 000-Layer Vanilla Convolutional Neural Networks · ICML 2018
Dynamical Isometry and a Mean Field Theory of RNNs: Gating Enables Signal Propagation in Recurrent Neural Networks · ICML 2018
Machine learning › Learning theory
generalization
1.022021
Whitening and Second Order Optimization Both Make Information in the Dataset Unusable During Training, and Can Reduce or Prevent Generalization · ICML 2021
Tilting the playing field: Dynamical loss functions for machine learning · ICML 2021
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.822020
Finite Versus Infinite Neural Networks: an Empirical Study · NeurIPS 2020
Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent · NeurIPS 2019
Machine learning › Deep learning architectures and training
convolutional neural network
0.822020
Disentangling Trainability and Generalization in Deep Neural Networks · ICML 2020
Dynamical Isometry and a Mean Field Theory of CNNs: How to Train 10, 000-Layer Vanilla Convolutional Neural Networks · ICML 2018
Machine learning › Deep learning architectures and training › training dynamics
signal propagation
0.722018
Dynamical Isometry and a Mean Field Theory of CNNs: How to Train 10, 000-Layer Vanilla Convolutional Neural Networks · ICML 2018
Dynamical Isometry and a Mean Field Theory of RNNs: Gating Enables Signal Propagation in Recurrent Neural Networks · ICML 2018
Machine learning › Deep learning architectures and training › weight initialization
dynamical isometry
0.622018
Dynamical Isometry and a Mean Field Theory of RNNs: Gating Enables Signal Propagation in Recurrent Neural Networks · ICML 2018
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice · NIPS 2017
Machine learning › Deep learning architectures and training › equilibrium models
deep equilibrium model
0.612022
Deep equilibrium networks are sensitive to initialization statistics · ICML 2022
Computer vision › 3D vision
higher-order statistics
0.612022
Deep equilibrium networks are sensitive to initialization statistics · ICML 2022
Machine learning › Deep learning architectures and training › weight initialization
initialization sensitivity
0.612022
Deep equilibrium networks are sensitive to initialization statistics · ICML 2022
Machine learning › Learning theory › information-theoretic analysis
information-theoretic bounds
0.512021
Whitening and Second Order Optimization Both Make Information in the Dataset Unusable During Training, and Can Reduce or Prevent Generalization · ICML 2021
Machine learning › Optimization for machine learning
learned optimizer
0.512021
Learn2Hop: Learned Optimization on Rough Landscapes · ICML 2021
Machine learning › Deep learning architectures and training
loss landscape
0.512021
Tilting the playing field: Dynamical loss functions for machine learning · ICML 2021
Machine learning › Transfer learning and domain adaptation
meta-learning
0.512021
Learn2Hop: Learned Optimization on Rough Landscapes · ICML 2021
Machine learning › Optimization for machine learning
non-convex optimization
0.512021
Learn2Hop: Learned Optimization on Rough Landscapes · ICML 2021
Machine learning › Optimization for machine learning
second-order optimization
0.512021
Whitening and Second Order Optimization Both Make Information in the Dataset Unusable During Training, and Can Reduce or Prevent Generalization · ICML 2021
Machine learning › Deep learning architectures and training › overparameterized neural network
wide neural networks
0.412020
Finite Versus Infinite Neural Networks: an Empirical Study · NeurIPS 2020
Computational science and engineering › scientific machine learning
differentiable simulation
0.412020
JAX MD: A Framework for Differentiable Physics · NeurIPS 2020
Computational science and engineering › computational chemistry › molecular simulation
molecular dynamics
0.412020
JAX MD: A Framework for Differentiable Physics · NeurIPS 2020
Computational science and engineering › statistical physics
statistical physics simulation
0.412020
JAX MD: A Framework for Differentiable Physics · NeurIPS 2020
Programming languages and type systems
library design
0.412020
Neural Tangents: Fast and Easy Infinite Neural Networks in Python · ICLR 2020
Machine learning › Graph learning
graph neural network
0.422020
Neural Message Passing for Quantum Chemistry · ICML 2017
JAX MD: A Framework for Differentiable Physics · NeurIPS 2020
Machine learning › Learning theory
random matrix theory
0.422018
Dynamical Isometry and a Mean Field Theory of RNNs: Gating Enables Signal Propagation in Recurrent Neural Networks · ICML 2018
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice · NIPS 2017
Machine learning › Deep learning architectures and training › normalization
batch normalization
0.412019
A Mean Field Theory of Batch Normalization · ICLR (Poster) 2019
Machine learning › Deep learning architectures and training
normalization
0.412019
A Mean Field Theory of Batch Normalization · ICLR (Poster) 2019
Machine learning › Deep learning architectures and training
training dynamics
0.412019
Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent · NeurIPS 2019
Machine learning › Deep learning architectures and training › weight initialization
critical initialization
0.312018
Dynamical Isometry and a Mean Field Theory of RNNs: Gating Enables Signal Propagation in Recurrent Neural Networks · ICML 2018
Machine learning › Deep learning architectures and training › recurrent neural network
gated recurrent network
0.312018
Dynamical Isometry and a Mean Field Theory of RNNs: Gating Enables Signal Propagation in Recurrent Neural Networks · ICML 2018
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.312018
Deep Neural Networks as Gaussian Processes · ICLR (Poster) 2018

Methods — techniques the papers use, named apart from their topics

neural tangent kernel · 1.9meta-learning · 1.4gradient-based optimization · 1.0mean-field analysis · 1.0mean-field theory · 0.9graph neural network · 0.9automatic differentiation · 0.9random matrix theory · 0.6orthogonal initialization · 0.6jacobian computation · 0.6regularized second-order optimization · 0.5hessian analysis · 0.5supervised learning · 0.3message passing · 0.3aggregation · 0.3macroscopic analytical modeling · 0.1
YearPublicationVenuePosition
2022 Deep equilibrium networks are sensitive to initialization statistics
abstract
Deep equilibrium networks (DEQs) are a promising way to construct models which trade off memory for compute. However, theoretical understanding of these models is still lacking compared to traditional networks, in part because of the repeated application of a single set of weights. We show that DEQs are sensitive to the higher order statistics of the matrix families from which they are initialized. In particular, initializing with orthogonal or symmetric matrices allows for greater stability in training. This gives us a practical prescription for initializations which allow for training with a broader range of initial weight scales.
Atish Agarwala, Samuel S. Schoenholz
ICML2
2022 Fast Finite Width Neural Tangent Kernel
abstract
The Neural Tangent Kernel (NTK), defined as the outer product of the neural network (NN) Jacobians, has emerged as a central object of study in deep learning. In the infinite width limit, the NTK can sometimes be computed analytically and is useful for understanding training and generalization of NN architectures. At finite widths, the NTK is also used to better initialize NNs, compare the conditioning across models, perform architecture search, and do meta-learning. Unfortunately, the finite width NTK is notoriously expensive to compute, which severely limits its practical utility. We perform the first in-depth analysis of the compute and memory requirements for NTK computation in finite width networks. Leveraging the structure of neural networks, we further propose two novel algorithms that change the exponent of the compute and memory requirements of the finite width NTK, dramatically improving efficiency. Our algorithms can be applied in a black box fashion to any differentiable function, including those implementing neural networks. We open-source our implementations within the Neural Tangents package at https://github.com/google/neural-tangents.
Roman Novak, Jascha Sohl-Dickstein, Samuel S. Schoenholz
ICML3
2021 Learn2Hop: Learned Optimization on Rough Landscapes
abstract
Optimization of non-convex loss surfaces containing many local minima remains a critical problem in a variety of domains, including operations research, informatics, and material design. Yet, current techniques either require extremely high iteration counts or a large number of random restarts for good performance. In this work, we propose adapting recent developments in meta-learning to these many-minima problems by learning the optimization algorithm for various loss landscapes. We focus on problems from atomic structural optimization—finding low energy configurations of many-atom systems—including widely studied models such as bimetallic clusters and disordered silicon. We find that our optimizer learns a hopping behavior which enables efficient exploration and improves the rate of low energy minima discovery. Finally, our learned optimizers show promising generalization with efficiency gains on never before seen tasks (e.g. new elements or compositions). Code is available at https://learn2hop.page.link/github.
Amil Merchant, Luke Metz, Samuel S. Schoenholz, Ekin Dogus Cubuk
ICML3
2021 Tilting the playing field: Dynamical loss functions for machine learning
abstract
We show that learning can be improved by using loss functions that evolve cyclically during training to emphasize one class at a time. In underparameterized networks, such dynamical loss functions can lead to successful training for networks that fail to find deep minima of the standard cross-entropy loss. In overparameterized networks, dynamical loss functions can lead to better generalization. Improvement arises from the interplay of the changing loss landscape with the dynamics of the system as it evolves to minimize the loss. In particular, as the loss function oscillates, instabilities develop in the form of bifurcation cascades, which we study using the Hessian and Neural Tangent Kernel. Valleys in the landscape widen and deepen, and then narrow and rise as the loss landscape changes during a cycle. As the landscape narrows, the learning rate becomes too large and the network becomes unstable and bounces around the valley. This process ultimately pushes the system into deeper and wider regions of the loss landscape and is characterized by decreasing eigenvalues of the Hessian. This results in better regularized models with improved generalization performance.
Miguel Ruiz-Garcia, Samuel S. Schoenholz, Andrea J. Liu
ICML3
2021 Whitening and Second Order Optimization Both Make Information in the Dataset Unusable During Training, and Can Reduce or Prevent Generalization
abstract
Machine learning is predicated on the concept of generalization: a model achieving low error on a sufficiently large training set should also perform well on novel samples from the same distribution. We show that both data whitening and second order optimization can harm or entirely prevent generalization. In general, model training harnesses information contained in the sample-sample second moment matrix of a dataset. For a general class of models, namely models with a fully connected first layer, we prove that the information contained in this matrix is the only information which can be used to generalize. Models trained using whitened data, or with certain second order optimization schemes, have less access to this information, resulting in reduced or nonexistent generalization ability. We experimentally verify these predictions for several architectures, and further demonstrate that generalization continues to be harmed even when theoretical requirements are relaxed. However, we also show experimentally that regularized second order optimization can provide a practical tradeoff, where training is accelerated but less information is lost, and generalization can in some circumstances even improve.
Neha S. Wadia, Daniel Duckworth, Samuel S. Schoenholz, Ethan Dyer, Jascha Sohl-Dickstein
ICML3
2020 Neural Tangents: Fast and Easy Infinite Neural Networks in Python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee 0001, Alexander A. Alemi, Jascha Sohl-Dickstein, Samuel S. Schoenholz
ICLR7
2020 Disentangling Trainability and Generalization in Deep Neural Networks
abstract
A longstanding goal in the theory of deep learning is to characterize the conditions under which a given neural network architecture will be trainable, and if so, how well it might generalize to unseen data. In this work, we provide such a characterization in the limit of very wide and very deep networks, for which the analysis simplifies considerably. For wide networks, the trajectory under gradient descent is governed by the Neural Tangent Kernel (NTK), and for deep networks the NTK itself maintains only weak data dependence. By analyzing the spectrum of the NTK, we formulate necessary conditions for trainability and generalization across a range of architectures, including Fully Connected Networks (FCNs) and Convolutional Neural Networks (CNNs). We identify large regions of hyperparameter space for which networks can memorize the training set but completely fail to generalize. We find that CNNs without global average pooling behave almost identically to FCNs, but that CNNs with pooling have markedly different and often better generalization performance. These theoretical results are corroborated experimentally on CIFAR10 for a variety of network architectures. We include a \href{https://colab.research.google.com/github/google/neural-tangents/blob/master/notebooks/disentangling_trainability_and_generalization.ipynb}{colab} notebook that reproduces the essential results of the paper.
Lechao Xiao, Jeffrey Pennington, Samuel S. Schoenholz
ICML3
2020 Finite Versus Infinite Neural Networks: an Empirical Study
abstract
We perform a careful, thorough, and large scale empirical study of the correspondence between wide neural networks and kernel methods. By doing so, we resolve a variety of open questions related to the study of infinitely wide neural networks. Our experimental results include: kernel methods outperform fully-connected finite-width networks, but underperform convolutional finite width networks; neural network Gaussian process (NNGP) kernels frequently outperform neural tangent (NT) kernels; centered and ensembled finite networks have reduced posterior variance and behave more similarly to infinite networks; weight decay and the use of a large learning rate break the correspondence between finite and infinite networks; the NTK parameterization outperforms the standard parameterization for finite width networks; diagonal regularization of kernels acts similarly to early stopping; floating point precision limits kernel performance beyond a critical dataset size; regularized ZCA whitening improves accuracy; finite network performance depends non-monotonically on width in ways not captured by double descent phenomena; equivariance of CNNs is only beneficial for narrow networks far from the kernel regime. Our experiments additionally motivate an improved layer-wise scaling for weight decay which improves generalization in finite-width networks. Finally, we develop improved best practices for using NNGP and NT kernels for prediction, including a novel ensembling technique. Using these best practices we achieve state-of-the-art results on CIFAR-10 classification for kernels corresponding to each architecture class we consider.
Jaehoon Lee 0001, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, Jascha Sohl-Dickstein
NeurIPS2
2020 JAX MD: A Framework for Differentiable Physics
abstract
We introduce JAX MD, a software package for performing differentiable physics simulations with a focus on molecular dynamics. JAX MD includes a number of statistical physics simulation environments as well as interaction potentials and neural networks that can be integrated into these environments without writing any additional code. Since the simulations themselves are differentiable functions, entire trajectories can be differentiated to perform meta-optimization. These features are built on primitive operations, such as spatial partitioning, that allow simulations to scale to hundreds-of-thousands of particles on a single GPU. These primitives are flexible enough that they can be used to scale up workloads outside of molecular dynamics. We present several examples that highlight the features of JAX MD including: integration of graph neural networks into traditional simulations, meta-optimization through minimization of particle packings, and a multi-agent flocking simulation. JAX MD is available at www.github.com/google/jax-md.
Samuel S. Schoenholz, Ekin Dogus Cubuk
NeurIPS1
2019 A Mean Field Theory of Batch Normalization
Greg Yang, Jeffrey Pennington, Vinay Rao, Jascha Sohl-Dickstein, Samuel S. Schoenholz
ICLR (Poster)5
2019 MetaInit: Initializing learning by learning to initialize
abstract
Deep learning models frequently trade handcrafted features for deep features learned with much less human intervention using gradient descent. While this paradigm has been enormously successful, deep networks are often difficult to train and performance can depend crucially on the initial choice of parameters. In this work, we introduce an algorithm called MetaInit as a step towards automating the search for good initializations using meta-learning. Our approach is based on a hypothesis that good initializations make gradient descent easier by starting in regions that look locally linear with minimal second order effects. We formalize this notion via a quantity that we call the gradient quotient, which can be computed with any architecture or dataset. MetaInit minimizes this quantity efficiently by using gradient descent to tune the norms of the initial weight matrices. We conduct experiments on plain and residual networks and show that the algorithm can automatically recover from a class of bad initializations. MetaInit allows us to train networks and achieve performance competitive with the state-of-the-art without batch normalization or residual connections. In particular, we find that this approach outperforms normalization for networks without skip connections on CIFAR-10 and can scale to Resnet-50 models on Imagenet.
Yann N. Dauphin, Samuel S. Schoenholz
NeurIPS2
2019 Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent
abstract
A longstanding goal in deep learning research has been to precisely characterize training and generalization. However, the often complex loss landscapes of neural networks have made a theory of learning dynamics elusive. In this work, we show that for wide neural networks the learning dynamics simplify considerably and that, in the infinite width limit, they are governed by a linear model obtained from the first-order Taylor expansion of the network around its initial parameters. Furthermore, mirroring the correspondence between wide Bayesian neural networks and Gaussian processes, gradient-based training of wide neural networks with a squared loss produces test set predictions drawn from a Gaussian process with a particular compositional kernel. While these theoretical results are only exact in the infinite width limit, we nevertheless find excellent empirical agreement between the predictions of the original network and those of the linearized version even for finite practically-sized networks. This agreement is robust across different architectures, optimization methods, and loss functions.
Jaehoon Lee 0001, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, Jeffrey Pennington
NeurIPS3
2018 The emergence of spectral universality in deep networks
abstract
Recent work has shown that tight concentration of the entire spectrum of singular values of a deep network’s input-output Jacobian around one at initialization can speed up learning by orders of magnitude. Therefore, to guide important design choices, it is important to build a full theoretical understanding of the spectra of Jacobians at initialization. To this end, we leverage powerful tools from free probability theory to provide a detailed analytic understanding of how a deep network’s Jacobian spectrum depends on various hyperparameters including the nonlinearity, the weight and bias distributions, and the depth. For a variety of nonlinearities, our work reveals the emergence of new universal limiting spectral distributions that remain concentrated around one even as the depth goes to infinity.
Jeffrey Pennington, Samuel S. Schoenholz, Surya Ganguli
AISTATS2
2018 Deep Neural Networks as Gaussian Processes
Jaehoon Lee 0001, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, Jascha Sohl-Dickstein
ICLR (Poster)4
2018 Dynamical Isometry and a Mean Field Theory of RNNs: Gating Enables Signal Propagation in Recurrent Neural Networks
abstract
Recurrent neural networks have gained widespread use in modeling sequence data across various domains. While many successful recurrent architectures employ a notion of gating, the exact mechanism that enables such remarkable performance is not well understood. We develop a theory for signal propagation in recurrent networks after random initialization using a combination of mean field theory and random matrix theory. To simplify our discussion, we introduce a new RNN cell with a simple gating mechanism that we call the minimalRNN and compare it with vanilla RNNs. Our theory allows us to define a maximum timescale over which RNNs can remember an input. We show that this theory predicts trainability for both recurrent architectures. We show that gated recurrent networks feature a much broader, more robust, trainable region than vanilla RNNs, which corroborates recent experimental findings. Finally, we develop a closed-form critical initialization scheme that achieves dynamical isometry in both vanilla RNNs and minimalRNNs. We show that this results in significantly improved training dynamics. Finally, we demonstrate that the minimalRNN achieves comparable performance to its more complex counterparts, such as LSTMs or GRUs, on a language modeling task.
Minmin Chen, Jeffrey Pennington, Samuel S. Schoenholz
ICML3
2018 Dynamical Isometry and a Mean Field Theory of CNNs: How to Train 10, 000-Layer Vanilla Convolutional Neural Networks
abstract
In recent years, state-of-the-art methods in computer vision have utilized increasingly deep convolutional neural network architectures (CNNs), with some of the most successful models employing hundreds or even thousands of layers. A variety of pathologies such as vanishing/exploding gradients make training such deep networks challenging. While residual connections and batch normalization do enable training at these depths, it has remained unclear whether such specialized architecture designs are truly necessary to train deep CNNs. In this work, we demonstrate that it is possible to train vanilla CNNs with ten thousand layers or more simply by using an appropriate initialization scheme. We derive this initialization scheme theoretically by developing a mean field theory for signal propagation and by characterizing the conditions for dynamical isometry, the equilibration of singular values of the input-output Jacobian matrix. These conditions require that the convolution operator be an orthogonal transformation in the sense that it is norm-preserving. We present an algorithm for generating such random initial orthogonal convolution kernels and demonstrate empirically that they enable efficient training of extremely deep architectures.
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel S. Schoenholz, Jeffrey Pennington
ICML4
2017 Deep Information Propagation
Samuel S. Schoenholz, Justin Gilmer, Surya Ganguli, Jascha Sohl-Dickstein
ICLR (Poster)1
2017 Neural Message Passing for Quantum Chemistry
abstract
Supervised learning on molecules has incredible potential to be useful in chemistry, drug discovery, and materials science. Luckily, several promising and closely related neural network models invariant to molecular symmetries have already been described in the literature. These models learn a message passing algorithm and aggregation procedure to compute a function of their entire input graph. At this point, the next step is to find a particularly effective variant of this general approach and apply it to chemical prediction benchmarks until we either solve them or reach the limits of the approach. In this paper, we reformulate existing models into a single common framework we call Message Passing Neural Networks (MPNNs) and explore additional novel variations within this framework. Using MPNNs we demonstrate state of the art results on an important molecular property prediction benchmark; these results are strong enough that we believe future work should focus on datasets with larger molecules or more accurate ground truth labels.
Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, George E. Dahl
ICML2
2017 Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
abstract
It is well known that weight initialization in deep networks can have a dramatic impact on learning speed. For example, ensuring the mean squared singular value of a network's input-output Jacobian is O(1) is essential for avoiding exponentially vanishing or exploding gradients. Moreover, in deep linear networks, ensuring that all singular values of the Jacobian are concentrated near 1 can yield a dramatic additional speed-up in learning; this is a property known as dynamical isometry. However, it is unclear how to achieve dynamical isometry in nonlinear deep networks. We address this question by employing powerful tools from free probability theory to analytically compute the {\it entire} singular value distribution of a deep network's input-output Jacobian. We explore the dependence of the singular value distribution on the depth of the network, the weight initialization, and the choice of nonlinearity. Intriguingly, we find that ReLU networks are incapable of dynamical isometry. On the other hand, sigmoidal networks can achieve isometry, but only with orthogonal weight initialization. Moreover, we demonstrate empirically that deep nonlinear networks achieving dynamical isometry learn orders of magnitude faster than networks that do not. Indeed, we show that properly-initialized deep sigmoidal networks consistently outperform deep ReLU networks. Overall, our analysis reveals that controlling the entire distribution of Jacobian singular values is an important design consideration in deep learning.
Jeffrey Pennington, Samuel S. Schoenholz, Surya Ganguli
NIPS2
2017 Mean Field Residual Networks: On the Edge of Chaos
abstract
We study randomly initialized residual networks using mean field theory and the theory of difference equations. Classical feedforward neural networks, such as those with tanh activations, exhibit exponential behavior on the average when propagating inputs forward or gradients backward. The exponential forward dynamics causes rapid collapsing of the input space geometry, while the exponential backward dynamics causes drastic vanishing or exploding gradients. We show, in contrast, that by adding skip connections, the network will, depending on the nonlinearity, adopt subexponential forward and backward dynamics, and in many cases in fact polynomial. The exponents of these polynomials are obtained through analytic methods and proved and verified empirically to be correct. In terms of the "edge of chaos" hypothesis, these subexponential and polynomial laws allow residual networks to "hover over the boundary between stability and chaos," thus preserving the geometry of the input space and the gradient information flow. In our experiments, for each activation function we study here, we initialize residual networks with different hyperparameters and train them on MNIST. Remarkably, our initialization time theory can accurately predict test time performance of these networks, by tracking either the expected amount of gradient explosion or the expected squared distance between the images of two input vectors. Importantly, we show, theoretically as well as empirically, that common initializations such as the Xavier or the He schemes are not optimal for residual networks, because the optimal initialization variances depend on the depth. Finally, we have made mathematical contributions by deriving several new identities for the kernels of powers of ReLU functions by relating them to the zeroth Bessel function of the second kind.
Greg Yang, Samuel S. Schoenholz
NIPS2
2009 Specialization as an optimal strategy under varying external conditions
abstract
We present an investigation of specialization when considering the execution of collaborative tasks by a robot swarm. Specifically, we consider the stick-pulling problem first proposed by Martinoli et al. [1], [2] and develop a macroscopic analytical model for the swarm executing a set of tasks that require the collaboration of two robots. We show, for constant external conditions, maximum productivity can be achieved by a single species swarm with carefully chosen operational parameters. While the same applies for a two species swarm, we show how specialization is a strategy best employed for changing external conditions.
M. Ani Hsieh, Ádám M. Halász, Ekin Dogus Cubuk, Samuel S. Schoenholz, Alcherio Martinoli
ICRA4