VLDB 2026 Research / reviewers in the wild / expert
Clément Hongler
dblp:222/3086
· DBLP profile ↗
9ranked-venue papers
0as first author
5since 2021 · last 2024
0000-0002-6224-2035ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Deep learning architectures and training · 32% Learning theory · 28% Kernel, tree and ensemble methods · 28% | |
| Theoretical computer science
1 paper |
Information theory · 100% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Kernel, tree and ensemble methods
kernel methods |
1.7 | 4 | 2021 | Neural tangent kernel: convergence and generalization in neural networks (invited paper) · STOC 2021 Kernel Alignment Risk Estimator: Risk Prediction from Training Data · NeurIPS 2020 Implicit Regularization of Random Feature Models · ICML 2020 |
Machine learning › Deep learning architectures and training
loss landscape |
1.5 | 3 | 2022 | Feature Learning in $L_2$-regularized DNNs: Attraction/Repulsion and Sparsity · NeurIPS 2022 Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances · ICML 2021 The asymptotic spectrum of the Hessian of DNN throughout training · ICLR 2020 |
Machine learning › Learning theory
neural network theory |
0.9 | 2 | 2022 | Feature Learning in $L_2$-regularized DNNs: Attraction/Repulsion and Sparsity · NeurIPS 2022 Neural Tangent Kernel: Convergence and Generalization in Neural Networks · NeurIPS 2018 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel ridge regression |
0.9 | 2 | 2020 | Kernel Alignment Risk Estimator: Risk Prediction from Training Data · NeurIPS 2020 Implicit Regularization of Random Feature Models · ICML 2020 |
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel |
0.8 | 2 | 2021 | Neural tangent kernel: convergence and generalization in neural networks (invited paper) · STOC 2021 Neural Tangent Kernel: Convergence and Generalization in Neural Networks · NeurIPS 2018 |
Natural language and speech › Language models and text generation › neural language model
autoregressive language model |
0.8 | 1 | 2024 | Arrows of Time for Large Language Models · ICML 2024 |
Machine learning › Learning theory › neural network theory › feature learning theory
feature learning dynamics |
0.6 | 1 | 2022 | Feature Learning in $L_2$-regularized DNNs: Attraction/Repulsion and Sparsity · NeurIPS 2022 |
Machine learning › Deep learning architectures and training
overparameterized neural network |
0.5 | 1 | 2021 | Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances · ICML 2021 |
Machine learning › Deep learning architectures and training
training dynamics |
0.5 | 1 | 2021 | Neural tangent kernel: convergence and generalization in neural networks (invited paper) · STOC 2021 |
Machine learning › Learning theory
generalization bounds |
0.4 | 1 | 2020 | Kernel Alignment Risk Estimator: Risk Prediction from Training Data · NeurIPS 2020 |
Machine learning › Deep learning architectures and training › loss landscape
hessian spectrum analysis |
0.4 | 1 | 2020 | The asymptotic spectrum of the Hessian of DNN throughout training · ICLR 2020 |
Machine learning › Optimization for machine learning
implicit regularization |
0.4 | 1 | 2020 | Implicit Regularization of Random Feature Models · ICML 2020 |
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel approximation
random features |
0.4 | 1 | 2020 | Implicit Regularization of Random Feature Models · ICML 2020 |
Machine learning › Deep learning architectures and training
risk prediction |
0.4 | 1 | 2020 | Kernel Alignment Risk Estimator: Risk Prediction from Training Data · NeurIPS 2020 |
Machine learning › Learning theory › nonparametric regression
kernel regression |
0.3 | 1 | 2018 | Neural Tangent Kernel: Convergence and Generalization in Neural Networks · NeurIPS 2018 |
Machine learning › Optimization for machine learning
non-convex optimization |
0.1 | 1 | 2021 | Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances · ICML 2021 |
Machine learning › Deep learning architectures and training › training dynamics
gradient descent dynamics |
0.1 | 1 | 2018 | Neural Tangent Kernel: Convergence and Generalization in Neural Networks · NeurIPS 2018 |
Methods — techniques the papers use, named apart from their topics
perplexity analysis · 1.5random matrix theory · 0.9l2 regularization · 0.6covariance analysis · 0.6convex reformulation · 0.6permutation symmetry analysis · 0.5kernel methods · 0.5gradient descent · 0.5affine subspace decomposition · 0.5asymptotic analysis · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Arrows of Time for Large Language ModelsabstractWe study the probabilistic modeling performed by Autoregressive Large Language Models (LLMs) through the angle of time directionality, addressing a question first raised in (Shannon, 1951). For large enough models, we empirically find a time asymmetry in their ability to learn natural language: a difference in the average log-perplexity when trying to predict the next token versus when trying to predict the previous one. This difference is at the same time subtle and very consistent across various modalities (language, model size, training time, ...). Theoretically, this is surprising: from an information-theoretic point of view, there should be no such difference. We provide a theoretical framework to explain how such an asymmetry can appear from sparsity and computational complexity considerations, and outline a number of perspectives opened by our results. Vassilis Papadopoulos, Jérémie Wenger, Clément Hongler |
ICML | 3 |
| 2024 | Smart Proofs via Recursive Information Gathering: Decentralized Refereeing by Smart ContractsabstractWe introduce the SPRIG (Smart Proofs via Recursive Information Gathering) protocol. SPRIG allows agents to propose, question, and defend mathematical proofs in a decentralized fashion. A structure of stakes and bounties aims at producing debates in good faith and if those persist, they must go down to machine-level details, where they can be settled automatically. This combination of economic incentives and an oracle is designed to promote succinct and informative proofs. SPRIG can run autonomously as a smart contract on a blockchain platform, and hence it does not rely on a central trusted institution. We translate SPRIG into a general game-theoretic model and prove that the protocol satisfies two desirable properties: no spamming and monotonicity. We then characterize analytically the equilibrium of a simple two-player specification of the model: this provides important insights into the impact of the protocol’s parameters on the probabilities that it induces type I/II errors. We conclude by discussing the main attacks SPRIG’s designers will need to take into account. Sylvain Carré, Franck Gabriel, Clément Hongler, Gustavo Lacerda, Gloria Capano |
Distributed Ledger Technol. Res. Pract. | 3 |
| 2022 | Feature Learning in $L_2$-regularized DNNs: Attraction/Repulsion and SparsityabstractWe study the loss surface of DNNs with $L_{2}$ regularization. Weshow that the loss in terms of the parameters can be reformulatedinto a loss in terms of the layerwise activations $Z_{\ell}$ of thetraining set. This reformulation reveals the dynamics behind featurelearning: each hidden representations $Z_{\ell}$ are optimal w.r.t.to an attraction/repulsion problem and interpolate between the inputand output representations, keeping as little information from theinput as necessary to construct the activation of the next layer.For positively homogeneous non-linearities, the loss can be furtherreformulated in terms of the covariances of the hidden representations,which takes the form of a partially convex optimization over a convexcone.This second reformulation allows us to prove a sparsity result forhomogeneous DNNs: any local minimum of the $L_{2}$-regularized losscan be achieved with at most $N(N+1)$ neurons in each hidden layer(where $N$ is the size of the training set). We show that this boundis tight by giving an example of a local minimum that requires $N^{2}/4$hidden neurons. But we also observe numerically that in more traditionalsettings much less than $N^{2}$ neurons are required to reach theminima. Arthur Jacot, Eugene A. Golikov, Clément Hongler, Franck Gabriel |
NeurIPS | 3 |
| 2021 | Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and InvariancesabstractWe study how permutation symmetries in overparameterized multi-layer neural networks generate ‘symmetry-induced’ critical points. Assuming a network with $ L $ layers of minimal widths $ r_1^*, \ldots, r_{L-1}^* $ reaches a zero-loss minimum at $ r_1^*! \cdots r_{L-1}^*! $ isolated points that are permutations of one another, we show that adding one extra neuron to each layer is sufficient to connect all these previously discrete minima into a single manifold. For a two-layer overparameterized network of width $ r^*+ h =: m $ we explicitly describe the manifold of global minima: it consists of $ T(r^*, m) $ affine subspaces of dimension at least $ h $ that are connected to one another. For a network of width $m$, we identify the number $G(r,m)$ of affine subspaces containing only symmetry-induced critical points that are related to the critical points of a smaller network of width $r Cite this Paper BibTeX @InProceedings{pmlr-v139-simsek21a, title = {Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances}, author = {Simsek, Berfin and Ged, Fran{\c{c}}ois and Jacot, Arthur and Spadaro, Francesco and Hongler, Clement and Gerstner, Wulfram and Brea, Johanni}, booktitle = {Proceedings of the 38th International Conference on Machine Learning}, pages = {9722--9732}, year = {2021}, editor = {Meila, Marina and Zhang, Tong}, volume = {139}, series = {Proceedings of Machine Learning Research}, month = {18--24 Jul}, publisher = {PMLR}, pdf = {http://proceedings.mlr.press/v139/simsek21a/simsek21a.pdf}, url = {https://proceedings.mlr.press/v139/simsek21a.html}, abstract = {We study how permutation symmetries in overparameterized multi-layer neural networks generate ‘symmetry-induced’ critical points. Assuming a network with $ L $ layers of minimal widths $ r_1^*, \ldots, r_{L-1}^* $ reaches a zero-loss minimum at $ r_1^*! \cdots r_{L-1}^*! $ isolated points that are permutations of one another, we show that adding one extra neuron to each layer is sufficient to connect all these previously discrete minima into a single manifold. For a two-layer overparameterized network of width $ r^*+ h =: m $ we explicitly describe the manifold of global minima: it consists of $ T(r^*, m) $ affine subspaces of dimension at least $ h $ that are connected to one another. For a network of width $m$, we identify the number $G(r,m)$ of affine subspaces containing only symmetry-induced critical points that are related to the critical points of a smaller network of width $r Copy to Clipboard Download Endnote %0 Conference Paper %T Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances %A Berfin Simsek %A François Ged %A Arthur Jacot %A Francesco Spadaro %A Clement Hongler %A Wulfram Gerstner %A Johanni Brea %B Proceedings of the 38th International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2021 %E Marina Meila %E Tong Zhang %F pmlr-v139-simsek21a %I PMLR %P 9722--9732 %U https://proceedings.mlr.press/v139/simsek21a.html %V 139 %X We study how permutation symmetries in overparameterized multi-layer neural networks generate ‘symmetry-induced’ critical points. Assuming a network with $ L $ layers of minimal widths $ r_1^*, \ldots, r_{L-1}^* $ reaches a zero-loss minimum at $ r_1^*! \cdots r_{L-1}^*! $ isolated points that are permutations of one another, we show that adding one extra neuron to each layer is sufficient to connect all these previously discrete minima into a single manifold. For a two-layer overparameterized network of width $ r^*+ h =: m $ we explicitly describe the manifold of global minima: it consists of $ T(r^*, m) $ affine subspaces of dimension at least $ h $ that are connected to one another. For a network of width $m$, we identify the number $G(r,m)$ of affine subspaces containing only symmetry-induced critical points that are related to the critical points of a smaller network of width $r Copy to Clipboard Download APA Simsek, B., Ged, F., Jacot, A., Spadaro, F., Hongler, C., Gerstner, W. & Brea, J.. (2021). Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances. Proceedings of the 38th International Conference on Machine Learning, in Proceedings of Machine Learning Research 139:9722-9732 Available from https://proceedings.mlr.press/v139/simsek21a.html. Copy to Clipboard Download Related Material Download PDF Supplementary PDF This site last compiled Sun, 05 Jul 2026 14:53:09 +0000 Github Account Copyright © The authors and PMLR 2026. MLResearchPress Berfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro, Clément Hongler, Wulfram Gerstner, Johanni Brea |
ICML | 5 |
| 2021 | Neural tangent kernel: convergence and generalization in neural networks (invited paper)abstractThe Neural Tangent Kernel is a new way to understand the gradient descent in deep neural networks, connecting them with kernel methods. In this talk, I'll introduce this formalism and give a number of results on the Neural Tangent Kernel and explain how they give us insight into the dynamics of neural networks during training and into their generalization features. Arthur Jacot, Franck Gabriel, Clément Hongler |
STOC | 3 |
| 2020 | The asymptotic spectrum of the Hessian of DNN throughout training
Arthur Jacot, Franck Gabriel, Clément Hongler |
ICLR | 3 |
| 2020 | Implicit Regularization of Random Feature ModelsabstractRandom Features (RF) models are used as efficient parametric approximations of kernel methods. We investigate, by means of random matrix theory, the connection between Gaussian RF models and Kernel Ridge Regression (KRR). For a Gaussian RF model with $P$ features, $N$ data points, and a ridge $\lambda$, we show that the average (i.e. expected) RF predictor is close to a KRR predictor with an \emph{effective ridge} $\tilde{\lambda}$. We show that $\tilde{\lambda} > \lambda$ and $\tilde{\lambda} \searrow \lambda$ monotonically as $P$ grows, thus revealing the \emph{implicit regularization effect} of finite RF sampling. We then compare the risk (i.e. test error) of the $\tilde{\lambda}$-KRR predictor with the average risk of the $\lambda$-RF predictor and obtain a precise and explicit bound on their difference. Finally, we empirically find an extremely good agreement between the test errors of the average $\lambda$-RF predictor and $\tilde{\lambda}$-KRR predictor. Arthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler, Franck Gabriel |
ICML | 4 |
| 2020 | Kernel Alignment Risk Estimator: Risk Prediction from Training DataabstractWe study the risk (i.e. generalization error) of Kernel Ridge Regression (KRR) for a kernel $K$ with ridge $\lambda>0$ and i.i.d. observations. For this, we introduce two objects: the Signal Capture Threshold (SCT) and the Kernel Alignment Risk Estimator (KARE). The SCT $\vartheta_{K,\lambda}$ is a function of the data distribution: it can be used to identify the components of the data that the KRR predictor captures, and to approximate the (expected) KRR risk. This then leads to a KRR risk approximation by the KARE $\rho_{K, \lambda}$, an explicit function of the training data, agnostic of the true data distribution. We phrase the regression problem in a functional setting. The key results then follow from a finite-size adaptation of the resolvent method for general Wishart random matrices. Under a natural universality assumption (that the KRR moments depend asymptotically on the first two moments of the observations) we capture the mean and variance of the KRR predictor. We numerically investigate our findings on the Higgs and MNIST datasets for various classical kernels: the KARE gives an excellent approximation of the risk. This supports our universality hypothesis. Using the KARE, one can compare choices of Kernels and hyperparameters directly from the training set. The KARE thus provides a promising data-dependent procedure to select Kernels that generalize well. Arthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler, Franck Gabriel |
NeurIPS | 4 |
| 2018 | Neural Tangent Kernel: Convergence and Generalization in Neural NetworksabstractAt initialization, artificial neural networks (ANNs) are equivalent to Gaussian processes in the infinite-width limit, thus connecting them to kernel methods. We prove that the evolution of an ANN during training can also be described by a kernel: during gradient descent on the parameters of an ANN, the network function (which maps input vectors to output vectors) follows the so-called kernel gradient associated with a new object, which we call the Neural Tangent Kernel (NTK). This kernel is central to describe the generalization features of ANNs. While the NTK is random at initialization and varies during training, in the infinite-width limit it converges to an explicit limiting kernel and stays constant during training. This makes it possible to study the training of ANNs in function space instead of parameter space. Convergence of the training can then be related to the positive-definiteness of the limiting NTK. We then focus on the setting of least-squares regression and show that in the infinite-width limit, the network function follows a linear differential equation during training. The convergence is fastest along the largest kernel principal components of the input data with respect to the NTK, hence suggesting a theoretical motivation for early stopping. Finally we study the NTK numerically, observe its behavior for wide networks, and compare it to the infinite-width limit. Arthur Jacot, Clément Hongler, Franck Gabriel |
NeurIPS | 2 |