Florian Rossmannek

dblp:255/5083 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2022
0000-0001-5772-5086ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2022 A proof of convergence for gradient descent in the training of artificial neural networks for constant target functions
Patrick Cheridito, Arnulf Jentzen, Adrian Riekert, Florian Rossmannek
J. Complex.4
2022 Efficient Approximation of High-Dimensional Functions With Neural Networks
abstract
In this article, we develop a framework for showing that neural networks can overcome the curse of dimensionality in different high-dimensional approximation problems. Our approach is based on the notion of a catalog network, which is a generalization of a standard neural network in which the nonlinear activation functions can vary from layer to layer as long as they are chosen from a predefined catalog of functions. As such, catalog networks constitute a rich family of continuous functions. We show that under appropriate conditions on the catalog, catalog networks can efficiently be approximated with rectified linear unit-type networks and provide precise estimates on the number of parameters needed for a given approximation accuracy. As special cases of the general results, we obtain different classes of functions that can be approximated with recitifed linear unit networks without the curse of dimensionality.
Patrick Cheridito, Arnulf Jentzen, Florian Rossmannek
IEEE Trans. Neural Networks Learn. Syst.3
2021 Non-convergence of stochastic gradient descent in the training of deep neural networks
abstract
Deep neural networks have successfully been trained in various application areas with stochastic gradient descent. However, there exists no rigorous mathematical explanation why this works so well. The training of neural networks with stochastic gradient descent has four different discretization parameters: (i) the network architecture; (ii) the amount of training data; (iii) the number of gradient steps; and (iv) the number of randomly initialized gradient trajectories. While it can be shown that the approximation error converges to zero if all four parameters are sent to infinity in the right order, we demonstrate in this paper that stochastic gradient descent fails to converge for ReLU networks if their depth is much larger than their width and the number of random initializations does not increase to infinity fast enough.
Patrick Cheridito, Arnulf Jentzen, Florian Rossmannek
J. Complex.3