Michael Kohler

dblp:97/3780 · DBLP profile ↗
← Back
30ranked-venue papers
17as first author
7since 2021 · last 2026
0000-0002-8372-9811ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 18 · 10 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 1 first-author
YearPublicationVenuePosition
2026 On the Rate of Convergence of an Over-Parametrized Deep Neural Network Regression Estimate Learned by Gradient Descent
Michael Kohler
IEEE Trans. Inf. Theory1
2025 Analysis of the Rate of Convergence of an Over-Parametrized Deep Neural Network Estimate Learned by Gradient Descent
abstract
Estimation of a regression function from independent and identically distributed random variables is considered. The$L_{2}$error with integration with respect to the design measure is used as an error criterion. Over-parametrized deep neural network estimates are defined which are based on a special network topology, which use a special random initialization and where all the weights are learned by the gradient descent. It is shown that the expected$L_{2}$error of these estimates converges to zero with the rate close to$n^{-1/(1+d)}$in case that the regression function is Hölder smooth with Hölder exponent$p \in [{1/2,1}]$. In case of an interaction model where the regression function is assumed to be a sum of Hölder smooth functions where each of the functions depends only on$d^{*}$of ofdcomponents of the design variable, it is shown that these estimates achieve the corresponding$d^{*}$-dimensional rate of convergence.
Michael Kohler, Adam Krzyzak
IEEE Trans. Inf. Theory1
2025 Regularized Over-Parametrized Neural Networks Learned by Gradient Descent Can Generalize Well
abstract
Estimation of univariate regression function by a neural network with one hidden layer is considered, where the weight vector is determined by applying gradient descent to a regularized empirical$L_{2}$risk. Here the number of hidden neurons is allowed to be much larger than the sample size. It is shown that the estimate nevertheless generalizes well in case that the Fourier transform of the regression function decays suitably fast, and that in this case over-parametrization leads to a particular good rate of convergence.
Michael Kohler, Adam Krzyzak
IEEE Trans. Inf. Theory1
2024 Rate of Convergence of an Over-Parametrized Convolutional Neural Network Image Classifier Learned by Gradient Descent
abstract
Image classifiers based on over-parametrized deep convolutional neural networks with an average-pooling are proposed. The weights of the network are learned by gradient descent. We present the bound on the rate of convergence of the difference between the expected misclassification risk of the plug-in classifier and the Bayes risk. The obtained rate of convergence is independent of image dimension under appropriate constraints on the image distribution.
Michael Kohler, Benjamin Walter 0001, Adam Krzyzak
ISIT1
2023 Analysis of Convolutional Neural Network Image Classifiers in a Rotationally Symmetric Model
abstract
Convolutional neural network image classifiers are defined and the rate of convergence of the misclassification risk of the estimates towards the optimal misclassification risk is analyzed. Here we consider images as random variables with values in some functional space, where we only observe discrete samples as function values on some finite grid. Under suitable structural and smoothness assumptions on the functional a posteriori probability, which includes some kind of symmetry against rotation of subparts of the input image, it is shown that least squares plug-in classifiers based on convolutional neural networks are able to circumvent the curse of dimensionality in binary image classification if we neglect a resolution-dependent error term. The finite sample size behavior of the classifier is analyzed by applying it to simulated and real data.
Michael Kohler, Benjamin Kohler
IEEE Trans. Inf. Theory1
2022 On the Rate of Convergence of a Classifier Based on a Transformer Encoder
abstract
Pattern recognition based on a high-dimensional predictor is considered. A classifier is defined which is based on a Transformer encoder. The rate of convergence of the misclassification probability of the classifier towards the optimal misclassification probability is analyzed. It is shown that this classifier is able to circumvent the curse of dimensionality provided the a posteriori probability satisfies a suitable hierarchical composition model. Furthermore, the difference between the Transformer classifiers theoretically analyzed in this paper and the ones used in practice today is illustrated by means of classification problems in natural language processing.
Iryna Gurevych, Michael Kohler, Gözde Gül Sahin
IEEE Trans. Inf. Theory2
2022 Estimation of a Function of Low Local Dimensionality by Deep Neural Networks
abstract
Deep neural networks (DNNs) achieve impressive results for complicated tasks like object detection on images and speech recognition. Motivated by this practical success, there is now a strong interest in showing good theoretical properties of DNNs. To describe for which tasks DNNs perform well and when they fail, it is a key challenge to understand their performance. The aim of this paper is to contribute to the current statistical theory of DNNs. We apply DNNs on high dimensional data and we show that the least squares regression estimates using DNNs are able to achieve dimensionality reduction in case that the regression function has locally low dimensionality. Consequently, the rate of convergence of the estimate does not depend on its input dimension$d$, but on its local dimension$d^{*}$and the DNNs are able to circumvent the curse of dimensionality in case that$d^{*}$is much smaller than$d$. In our simulation study we provide numerical experiments to support our theoretical result and we compare our estimate with other conventional nonparametric regression estimates. The performance of our estimates is also validated in experiments with real data.
Michael Kohler, Adam Krzyzak, Sophie Langer
IEEE Trans. Inf. Theory1
2020 Common Voice: A Massively-Multilingual Speech Corpus
abstract
The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other domains (e.g. language identification). To achieve scale and sustainability, the Common Voice project employs crowdsourcing for both data collection and data validation. The most recent release includes 29 languages, and as of November 2019 there are a total of 38 languages collecting data. Over 50,000 individuals have participated so far, resulting in 2,500 hours of collected audio. To our knowledge this is the largest audio corpus in the public domain for speech recognition, both in terms of number of hours and number of languages. As an example use case for Common Voice, we present speech recognition experiments using Mozilla’s DeepSpeech Speech-to-Text toolkit. By applying transfer learning from a source English model, we find an average Character Error Rate improvement of 5.99 ± 5.48 for twelve target languages (German, French, Italian, Turkish, Catalan, Slovenian, Welsh, Irish, Breton, Tatar, Chuvash, and Kabyle). For most of these languages, these are the first ever published results on end-to-end Automatic Speech Recognition.
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis M. Tyers, Gregor Weber
LREC4
2019 Estimation of a Density From an Imperfect Simulation Model
abstract
Uncertainty quantification of a technical system can be done using density estimation. We usually start with a stochastic model, which is fitted to the technical system, and the density estimation is done using data from this stochastic model. However, in any application, such a stochastic model will not be perfect, and the estimation of the density should take into account the inadequacy of the stochastic model. In this paper, we show how observed data of the real system together with an imperfect simulation model can be used to derive confidence bands for the density of the technical system. Our main result is that the newly introduced confidence bands allow to derive lower and upper bounds on the probability of intervals in the technical system. Furthermore, we present an upper bound on the area of the confidence band in case of a smooth density. The results are illustrated by applying the estimates to simulated and real data.
Michael Kohler, Adam Krzyzak
IEEE Trans. Inf. Theory1
2018 Adaptive Estimation of Quantiles in a Simulation Model
abstract
Let X be an Rdvalued random variable, let m : Rd→ R be a measurable function and set Y = m(X). Given a sample of (X, Y) of size n, we consider the problem of estimating the quantile of Y of a given level α ∈ (0, 1). A method for choosing the parameter of a surrogate model of m is introduced, and it is shown that the corresponding surrogate quantile estimate achieves the rate of convergence bounded by the sum of the minimal rate of convergence of the quantile estimates corresponding to the given surrogate estimates and a term of order log(n)/n. The finite sample size behavior of this quantile estimate is illustrated by applying it to simulated data and to a quantile estimation problem in mechanical engineering.
Michael Kohler, Adam Krzyzak
IEEE Trans. Inf. Theory1
2017 Nonparametric Regression Based on Hierarchical Interaction Models
abstract
In this paper, we introduce the so-called hierarchical interaction models, where we assume that the computation of the value of a function m : ℝd→ ℝ is done in several layers, where in each layer a function of at most d* inputs computed by the previous layer is evaluated. We investigate two different regression estimates based on polynomial splines and on neural networks, and show that if the regression function satisfies a hierarchical interaction model and all occurring functions in the model are smooth, the rate of convergence of these estimates depends on d* (and not on d). Hence, in this case, the estimates can achieve good rate of convergence even for large d, and are in this sense able to circumvent the so-called curse of dimensionality.
Michael Kohler, Adam Krzyzak
IEEE Trans. Inf. Theory1
2016 The rates of convergence of neural network estimates of hierarchical interaction regression models
abstract
Regression estimation often suffers from the curse of dimensionality. In the present paper we circumvent this problem by introducing a class of models called hierarchical interaction models where the values of a function m ∶ ℝd→ ℝ are computed in a feed-forward manner in several layers, where in each layer a function of at most d* inputs produced by the previous layer is computed. We introduce regression estimates based on neural networks with two hidden layers and apply them to estimation of regression functions from a class of hierarchical interaction models. Under smoothness condition imposed on all functions occurring in the model we show that the rate of convergence of these estimates depends on d*, which is typically much smaller than d.
Michael Kohler, Adam Krzyzak
ISIT1
2016 Efficient automotive grid maps using a sensor ray based refinement process
abstract
The occupancy grid mapping technique is widely used for environmental mapping of moving vehicles. Occupancy grid maps with fixed cell size have been extended using the quadtree implementation with adaptive cell size. Adaptive grid maps have proven to be more resource efficient than fixed cell size grid maps. Dynamic cell sizes introduce the necessity of a split and merge process to trigger the refinement of grid cells. This paper presents a novel ray-based refinement process in order to choose the appropriate resolution for the sensor observation. Based on measurement conflicts some approaches use an iterative refinement process until all conflicts are solved. In contrast this paper presents an non-iterative approach based on the sensor resolution. Using the measurement data efficiently we propose an algorithm, which solves the problem of partially free cells in an adaptive grid map. The proposed algorithm is compared against other widely used algorithms and methodologies.
Ruben Jungnickel, Michael Kohler, Franz Korf
Intelligent Vehicles Symposium2
2016 Nonparametric Quantile Estimation Based on Surrogate Models
abstract
Nonparametric estimation of a quantile qm(X),αof a random variable m(X) is considered, where m : ℝd→ ℝ is a function, which is costly to compute and X is an ℝd-valued random variable with known distribution. Monte Carlo surrogate quantile estimates are considered, where in a first step, the function m is estimated by some estimate (surrogate) mn, and then, the quantile qm(X),αis estimated by a Monte Carlo estimate of the quantile qmn(X),α. A general error bound on the error of this quantile estimate is derived, which depends on the local error of the function estimate mn, and the rates of convergence of the corresponding Monte Carlo surrogate quantile estimates are analyzed for two different function estimates. The finite sample size behavior of the estimates is investigated in simulations.
Georg C. Enss, Michael Kohler, Adam Krzyzak, Roland Platz
IEEE Trans. Inf. Theory2
2015 Adaptive Density Estimation From Data With Small Measurement Errors
abstract
In this paper, we study the problem of density estimation from data that contains small measurement errors. The only assumption on these errors is that the maximal measurement error is bounded by some real number converging to zero for sample size tending to infinity. In particular, we do not assume that the measurement errors are independent with expectation zero. We estimate the density by a standard kernel density estimate applied to data with measurement errors and derive a data-driven method to choose its bandwidth. We derive an adaptation result for this estimate and analyze the expected L1error of our density estimate depending on the smoothness of the density and the size of the maximal measurement error. The results are applied in a density estimation problem in a simulation model, where we show under suitable assumptions that the L1error of our newly proposed estimate converges to zero much faster than the L1error of the standard kernel density estimate if both are based on the same number of observations in the simulation model. The performance of the method in case of finite sample size is analyzed using simulated data.
Tina Felber, Michael Kohler, Adam Krzyzak
IEEE Trans. Inf. Theory2
2014 Density estimation using real and artificial data
abstract
Let X, X1, X2, ... be independent and identically distributed ℝd-valued random variables and let m : ℝd→ ℝ be an unknown measurable function such that a density f of Y = m(X) exists. In this paper we consider estimating f based on i.i.d. sample (X1, Y1);...; (Xn, Yn) of (X, Y) and on additional independent observations of X. We compare the standard kernel density estimate based on the y-values of the sample of (X, Y) and a kernel density estimate based on artificially generated y-values corresponding to the additional observations of X. It is shown that under suitable smoothness assumptions on f and m the rate of convergence of the L1error of the latter estimate is better than that of the standard kernel density estimate. Furthermore, a density estimate defined as convex combination of these two estimates is considered and a data-driven choice of the bandwidths and the weight of the convex combination is proposed and investigated.
Tina Felber, Michael Kohler, Adam Krzyzak
ISIT2
2013 Estimation of a Density Using Real and Artificial Data
abstract
LetX,X1,X2, ... be independent and identically distributed Rd-valued random variables and letm: Rd→ R be a measurable function such that a densityfofY=m(X) exists. Given a sample of the distribution of (X,Y) and additional independent observations ofX, we are interested in estimatingf. We apply a regression estimate to the sample of (X,Y) and use this estimate to generate additional artificial observations ofY. Using these artificial observations together with the real observations ofY, we construct a density estimate offby using a convex combination of two kernel density estimates. It is shown that if the bandwidths satisfy the usual conditions and if in addition the supremum norm error of the regression estimate converges almost surely faster toward zero than the bandwidth of the kernel density estimate applied to the artificial data, then the convex combination of the two density estimates isL1-consistent. The performance of the estimate for finite sample size is illustrated by simulated data, and the usefulness of the procedure is demonstrated by applying it to a density estimation problem in a simulation model.
Luc Devroye, Tina Felber, Michael Kohler
IEEE Trans. Inf. Theory3
2012 An EHR Prototype Using Structured ISO/EN 13606 Documents to Respond to Identified Clinical Information Needs of Diabetes Specialists: A Controlled Study on Feasibility and Impact
Gudrun Hübner-Bloder, Georg Duftschmid, Michael Kohler, Christoph Rinner, Samrend Saboor, Elske Ammenwerth
AMIA3
2012 Weakly Universally Consistent Forecasting of Stationary and Ergodic Time Series
abstract
Static forecasting of stationary and ergodic time series is considered, i.e., inference of the conditional expectation of the response variable at time zero given the infinite past. It is shown that the mean squared error of a combination of suitably defined localized least squares estimates converges to zero for all distributions where the response variable is square integrable.
Michael Kohler, Harro Walk
IEEE Trans. Inf. Theory2
2011 Analysis of the rate of convergence of least squares neural network regression estimates in case of measurement errors
Michael Kohler, Jens Mehnert
Neural Networks1
2010 An L2-boosting algorithm for estimation of a regression function
abstract
AnL2-boosting algorithm for estimation of a regression function from random design is presented, which consists of fitting repeatedly a function from a fixed nonlinear function space to the residuals of the data by least squares and by defining the estimate as a linear combination of the resulting least squares estimates. Splitting of the sample is used to decide after how many iterations of smoothing of the residuals the algorithm terminates. The rate of convergence of the algorithm is analyzed in case of an unbounded response variable. The method is used to fit a sum of maxima of minima of linear functions to a given data set, and is compared with other nonparametric regression estimates using simulated data.
Adil M. Bagirov, Conny Clausen, Michael Kohler
IEEE Trans. Inf. Theory3
2009 On application of nonparametric regression estimation to options pricing
abstract
We consider American options also called Bermudan options in discrete time.We use the dual approach to derive upper bounds on the price of such options using only a reduced number of nested Monte Carlo steps. The key idea is to use nonparametric regression to estimate continuation values and all other required conditional expectations and to combine the resulting estimate with another estimate computed by using only a reduced number of nested Monte Carlo steps. The mean value of the resulting estimate is an upper bound on the option price. One can show that the estimates of the option prices are universally consistent, i.e., they converge to the true price regardless of the structure of the continuation values. The finite sample behavior is validated by experiments on simulated data.
Michael Kohler, Adam Krzyzak, Harro Walk
ISIT1
2009 Estimation of a Regression Function by Maxima of Minima of Linear Functions
abstract
In this paper, estimation of a regression function from independent and identically distributed random variables is considered. Estimates are defined by minimization of the empirical L2risk over a class of functions, which are defined as maxima of minima of linear functions. Results concerning the rate of convergence of the estimates are derived. In particular, it is shown that for smooth regression functions satisfying the assumption of single index models, the estimate is able to achieve (up to some logarithmic factor) the corresponding optimal one-dimensional rate of convergence. Hence, under these assumptions, the estimate is able to circumvent the so-called curse of dimensionality. The small sample behavior of the estimates is illustrated by applying them to simulated data.
Adil M. Bagirov, Conny Clausen, Michael Kohler
IEEE Trans. Inf. Theory3
2008 Automatic recognition of German news focusing on future-directed beliefs and intentions
Judith Eckle-Kohler, Michael Kohler, Jens Mehnert
Comput. Speech Lang.2
2007 Nonparametric Estimation of Conditional Distributions
abstract
Estimation of conditional distributions is considered. It is assumed that the conditional distribution is either discrete or that it has a density with respect to the Lebesgue measure. Partitioning estimates of the conditional distribution are constructed and results concerning consistency and rate of convergence of the integrated total variation error of the estimates are presented.
László Györfi, Michael Kohler
IEEE Trans. Inf. Theory2
2007 On the Rate of Convergence of Local Averaging Plug-In Classification Rules Under a Margin Condition
abstract
The rates of convergence of plug-in kernel, partitioning, and nearest neighbors classification rules are analyzed. A margin condition, which measures how quickly thea posterioriprobabilities cross the decision boundary, smoothness conditions on thea posterioriprobabilities, and boundedness of the feature vector are imposed. The rates of convergence of the plug-in classifiers shown in this paper are faster than previously known.
Michael Kohler, Adam Krzyzak
IEEE Trans. Inf. Theory1
2006 Rate of convergence of local averaging plug-in classification rules under margin condition
abstract
We discuss rates of convergence of plug-in kernel, partitioning and nearest neighbors classification rules under margin condition. Margin condition characterizes the rate with which a posteriori probabilities cross the decision boundary. We show the rates of convergence of the plug-in classifiers under smoothness conditions on a posteriori probabilities and assuming that feature vectors are contained in a compact set. We obtain particularly fast rates of convergence assuming, in addition, that feature vectors distributions have densities bounded away from zero.
Michael Kohler, Adam Krzyzak
ISIT1
2005 Rates of convergence for adaptive regression estimates with multiple hidden layer feedforward neural networks
abstract
We present a general bound on the expected L2error of adaptive least squares estimates. By applying it to multiple hidden layer feedforward neural network regression function estimates we are able to obtain optimal (up to log factor) rates of convergence for Lipschitz classes and fast rates of convergence for some classes of regression functions such as additive functions
Michael Kohler, Adam Krzyzak
ISIT1
2004 Adaptive regression estimation with multilayer feedforward neural networks
abstract
We prove a general bound on the expected L/sub 2/ error of adaptive least squares estimates. By applying it to multilayer feedforward neural network regression function estimates we are able to obtain fast rates of convergence in special classes of regression functions such as additive functions.
Michael Kohler, Adam Krzyzak
ISIT1
2001 Nonparametric regression estimation using penalized least squares
abstract
We present multivariate penalized least squares regression estimates. We use Vapnik-Chervonenkis (see Statistical Learning Theory 1998) theory and bounds on the covering numbers to analyze convergence of the estimates. We show strong consistency of the truncated versions of the estimates without any conditions on the underlying distribution.
Michael Kohler, Adam Krzyzak
IEEE Trans. Inf. Theory1