Tycho F. A. van der Ouderaa

dblp:236/4440 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
7since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Probabilistic and Bayesian machine learning · 27% Representation and self-supervised learning · 22% Deep learning architectures and training · 16%

Topics — the 20 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian model selection
2.032024
Noether's Razor: Learning Conserved Quantities · NeurIPS 2024
Stochastic Marginal Likelihood Gradients using Neural Tangent Kernels · ICML 2023
Invariance Learning in Deep Neural Networks with Differentiable Laplace Approximations · NeurIPS 2022
Machine learning › Representation and self-supervised learning
symmetry learning
2.032024
Noether's Razor: Learning Conserved Quantities · NeurIPS 2024
Learning Layer-wise Equivariances Automatically using Gradients · NeurIPS 2023
Relaxing Equivariance Constraints with Non-stationary Continuous Filters · NeurIPS 2022
Machine learning › Representation and self-supervised learning
equivariance
1.222023
Learning Layer-wise Equivariances Automatically using Gradients · NeurIPS 2023
Relaxing Equivariance Constraints with Non-stationary Continuous Filters · NeurIPS 2022
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › variational objective
evidence lower bound
0.812024
Noether's Razor: Learning Conserved Quantities · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model
0.812024
The LLM Surgeon · ICLR 2024
Machine learning › Efficient and distributed learning
model compression
0.812024
The LLM Surgeon · ICLR 2024
Machine learning › Efficient and distributed learning › model compression
pruning
0.812024
The LLM Surgeon · ICLR 2024
Machine learning › Optimization for machine learning
hyperparameter optimization
0.712023
Stochastic Marginal Likelihood Gradients using Neural Tangent Kernels · ICML 2023
Machine learning › Deep learning architectures and training › equivariant neural network
layer-wise equivariance
0.712023
Learning Layer-wise Equivariances Automatically using Gradients · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning
marginal likelihood
0.712023
Learning Layer-wise Equivariances Automatically using Gradients · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian model selection
marginal likelihood estimation
0.712023
Stochastic Marginal Likelihood Gradients using Neural Tangent Kernels · ICML 2023
Machine learning › Optimization for machine learning
stochastic gradient methods
0.712023
Stochastic Marginal Likelihood Gradients using Neural Tangent Kernels · ICML 2023
Machine learning › Deep learning architectures and training › equivariant neural network
approximate equivariance
0.612022
Relaxing Equivariance Constraints with Non-stationary Continuous Filters · NeurIPS 2022
Machine learning › Deep learning architectures and training
data augmentation
0.612022
Invariance Learning in Deep Neural Networks with Differentiable Laplace Approximations · NeurIPS 2022
Machine learning › Trustworthy machine learning › out-of-distribution generalization
invariant learning
0.612022
Invariance Learning in Deep Neural Networks with Differentiable Laplace Approximations · NeurIPS 2022
Machine learning › Generative modeling
generative adversarial network
0.412019
Reversible GANs for Memory-Efficient Image-To-Image Translation · CVPR 2019
Machine learning › Generative modeling › generative adversarial network
image-to-image translation
0.412019
Reversible GANs for Memory-Efficient Image-To-Image Translation · CVPR 2019
Machine learning › Deep learning architectures and training › feedforward neural network
invertible neural network
0.412019
Reversible GANs for Memory-Efficient Image-To-Image Translation · CVPR 2019
Machine learning › Deep learning architectures and training
transformer
0.212024
The LLM Surgeon · ICLR 2024
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel
0.212023
Stochastic Marginal Likelihood Gradients using Neural Tangent Kernels · ICML 2023

Methods — techniques the papers use, named apart from their topics

noether's theorem · 0.8kronecker-factored curvature approximation · 0.8approximate bayesian model selection · 0.8stochastic gradient descent · 0.7neural tangent kernel · 0.7linearized laplace approximation · 0.7laplace approximation · 0.7gradient-based learning · 0.7kronecker-factored approximation · 0.6differentiable laplace approximation · 0.6
YearPublicationVenuePosition
2024 The LLM Surgeon
abstract
State-of-the-art language models are becoming increasingly large in an effort to achieve the highest performance on large corpora of available textual data. However, the sheer size of the Transformer architectures makes it difficult to deploy models within computational, environmental or device-specific constraints. We explore data-driven compression of existing pretrained models as an alternative to training smaller models from scratch. To do so, we scale Kronecker-factored curvature approximations of the target loss landscape to large language models. In doing so, we can compute both the dynamic allocation of structures that can be removed as well as updates of remaining weights that account for the removal. We provide a general framework for unstructured, semi-structured and structured pruning and improve upon weight updates to capture more correlations between weights, while remaining computationally efficient. Experimentally, our method can prune rows and columns from a range of OPT models and Llamav2-7B by 20\%-30\%, with a negligible loss in performance, and achieve state-of-the-art results in unstructured and semi-structured pruning of large language models. We will open source our code on GitHub upon acceptance.
Tycho F. A. van der Ouderaa, Markus Nagel, Mart van Baalen, Tijmen Blankevoort
ICLR1
2024 Noether's Razor: Learning Conserved Quantities
abstract
Symmetries have proven useful in machine learning models, improving generalisation and overall performance. At the same time, recent advancements in learning dynamical systems rely on modelling the underlying Hamiltonian to guarantee the conservation of energy. These approaches can be connected via a seminal result in mathematical physics: Noether's theorem, which states that symmetries in a dynamical system correspond to conserved quantities. This work uses Noether's theorem to parameterise symmetries as learnable conserved quantities. We then allow conserved quantities and associated symmetries to be learned directly from train data through approximate Bayesian model selection, jointly with the regular training procedure. As training objective, we derive a variational lower bound to the marginal likelihood. The objective automatically embodies an Occam's Razor effect that avoids collapse of conversation laws to the trivial constant, without the need to manually add and tune additional regularisers. We demonstrate a proof-of-principle on n-harmonic oscillators and n-body systems. We find that our method correctly identifies the correct conserved quantities and U(n) and SE(n) symmetry groups, improving overall performance and predictive accuracy on test data.
Tycho F. A. van der Ouderaa, Mark van der Wilk, Pim de Haan
NeurIPS1
2023 Stochastic Marginal Likelihood Gradients using Neural Tangent Kernels
abstract
Selecting hyperparameters in deep learning greatly impacts its effectiveness but requires manual effort and expertise. Recent works show that Bayesian model selection with Laplace approximations can allow to optimize such hyperparameters just like standard neural network parameters using gradients and on the training data. However, estimating a single hyperparameter gradient requires a pass through the entire dataset, limiting the scalability of such algorithms. In this work, we overcome this issue by introducing lower bounds to the linearized Laplace approximation of the marginal likelihood. In contrast to previous estimators, these bounds are amenable to stochastic-gradient-based optimization and allow to trade off estimation accuracy against computational complexity. We derive them using the function-space form of the linearized Laplace, which can be estimated using the neural tangent kernel. Experimentally, we show that the estimators can significantly accelerate gradient-based hyperparameter optimization.
Alexander Immer, Tycho F. A. van der Ouderaa, Mark van der Wilk, Gunnar Rätsch, Bernhard Schölkopf
ICML2
2023 Learning Layer-wise Equivariances Automatically using Gradients
abstract
Convolutions encode equivariance symmetries into neural networks leading to better generalisation performance. However, symmetries provide fixed hard constraints on the functions a network can represent, need to be specified in advance, and can not be adapted. Our goal is to allow flexible symmetry constraints that can automatically be learned from data using gradients. Learning symmetry and associated weight connectivity structures from scratch is difficult for two reasons. First, it requires efficient and flexible parameterisations of layer-wise equivariances. Secondly, symmetries act as constraints and are therefore not encouraged by training losses measuring data fit. To overcome these challenges, we improve parameterisations of soft equivariance and learn the amount of equivariance in layers by optimising the marginal likelihood, estimated using differentiable Laplace approximations. The objective balances data fit and model complexity enabling layer-wise symmetry discovery in deep networks. We demonstrate the ability to automatically learn layer-wise equivariances on image classification tasks, achieving equivalent or improved performance over baselines with hard-coded symmetry.
Tycho F. A. van der Ouderaa, Alexander Immer, Mark van der Wilk
NeurIPS1
2022 Invariance Learning in Deep Neural Networks with Differentiable Laplace Approximations
abstract
Data augmentation is commonly applied to improve performance of deep learning by enforcing the knowledge that certain transformations on the input preserve the output. Currently, the data augmentation parameters are chosen by human effort and costly cross-validation, which makes it cumbersome to apply to new datasets. We develop a convenient gradient-based method for selecting the data augmentation without validation data during training of a deep neural network. Our approach relies on phrasing data augmentation as an invariance in the prior distribution on the functions of a neural network, which allows us to learn it using Bayesian model selection. This has been shown to work in Gaussian processes, but not yet for deep neural networks. We propose a differentiable Kronecker-factored Laplace approximation to the marginal likelihood as our objective, which can be optimised without human supervision or validation data. We show that our method can successfully recover invariances present in the data, and that this improves generalisation and data efficiency on image datasets.
Alexander Immer, Tycho F. A. van der Ouderaa, Gunnar Rätsch, Vincent Fortuin, Mark van der Wilk
NeurIPS2
2022 Relaxing Equivariance Constraints with Non-stationary Continuous Filters
abstract
Equivariances provide useful inductive biases in neural network modeling, with the translation equivariance of convolutional neural networks being a canonical example. Equivariances can be embedded in architectures through weight-sharing and place symmetry constraints on the functions a neural network can represent. The type of symmetry is typically fixed and has to be chosen in advance. Although some tasks are inherently equivariant, many tasks do not strictly follow such symmetries. In such cases, equivariance constraints can be overly restrictive. In this work, we propose a parameter-efficient relaxation of equivariance that can effectively interpolate between a (i) non-equivariant linear product, (ii) a strict-equivariant convolution, and (iii) a strictly-invariant mapping. The proposed parameterisation can be thought of as a building block to allow adjustable symmetry structure in neural networks. In addition, we demonstrate that the amount of equivariance can be learned from the training data using backpropagation. Gradient-based learning of equivariance achieves similar or improved performance compared to the best value found by cross-validation and outperforms baselines with partial or strict equivariance on CIFAR-10 and CIFAR-100 image classification tasks.
Tycho F. A. van der Ouderaa, David W. Romero, Mark van der Wilk
NeurIPS1
2022 Learning invariant weights in neural networks
abstract
Assumptions about invariances or symmetries in data can significantly increase the predictive power of statistical models. Many commonly used machine learning models are constraint to respect certain symmetries, such as translation equivariance in convolutional neural networks, and incorporating other symmetry types is actively being studied. Yet, learning invariances from the data itself remains an open research problem. It has been shown that the marginal likelihood offers a principled way to learn invariances in Gaussian Processes. We propose a weight-space equivalent to this approach, by minimizing a lower bound on the marginal likelihood to learn invariances in neural networks, resulting in naturally higher performing models.
Tycho F. A. van der Ouderaa, Mark van der Wilk
UAI1
2019 Reversible GANs for Memory-Efficient Image-To-Image Translation
abstract
The pix2pix and CycleGAN losses have vastly improved the qualitative and quantitative visual quality of results in image-to-image translation tasks. We extend this framework by exploring approximately invertible architectures which are well suited to these losses. These architectures are approximately invertible by design and thus partially satisfy cycle-consistency before training even begins. Furthermore, since invertible architectures have constant memory complexity in depth, these models can be built arbitrarily deep. We are able to demonstrate superior quantitative output on the Cityscapes and Maps datasets at near constant memory budget.
Tycho F. A. van der Ouderaa, Daniel E. Worrall
CVPR1