EDBT 2026 Demo / reviewers in the wild / expert
Tycho F. A. van der Ouderaa
dblp:236/4440
· DBLP profile ↗
8ranked-venue papers
6as first author
7since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Probabilistic and Bayesian machine learning · 27% Representation and self-supervised learning · 22% Deep learning architectures and training · 16% |
Topics — the 20 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian model selection |
2.0 | 3 | 2024 | Noether's Razor: Learning Conserved Quantities · NeurIPS 2024 Stochastic Marginal Likelihood Gradients using Neural Tangent Kernels · ICML 2023 Invariance Learning in Deep Neural Networks with Differentiable Laplace Approximations · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning
symmetry learning |
2.0 | 3 | 2024 | Noether's Razor: Learning Conserved Quantities · NeurIPS 2024 Learning Layer-wise Equivariances Automatically using Gradients · NeurIPS 2023 Relaxing Equivariance Constraints with Non-stationary Continuous Filters · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning
equivariance |
1.2 | 2 | 2023 | Learning Layer-wise Equivariances Automatically using Gradients · NeurIPS 2023 Relaxing Equivariance Constraints with Non-stationary Continuous Filters · NeurIPS 2022 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › variational objective
evidence lower bound |
0.8 | 1 | 2024 | Noether's Razor: Learning Conserved Quantities · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model |
0.8 | 1 | 2024 | The LLM Surgeon · ICLR 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | The LLM Surgeon · ICLR 2024 |
Machine learning › Efficient and distributed learning › model compression
pruning |
0.8 | 1 | 2024 | The LLM Surgeon · ICLR 2024 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.7 | 1 | 2023 | Stochastic Marginal Likelihood Gradients using Neural Tangent Kernels · ICML 2023 |
Machine learning › Deep learning architectures and training › equivariant neural network
layer-wise equivariance |
0.7 | 1 | 2023 | Learning Layer-wise Equivariances Automatically using Gradients · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning
marginal likelihood |
0.7 | 1 | 2023 | Learning Layer-wise Equivariances Automatically using Gradients · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian model selection
marginal likelihood estimation |
0.7 | 1 | 2023 | Stochastic Marginal Likelihood Gradients using Neural Tangent Kernels · ICML 2023 |
Machine learning › Optimization for machine learning
stochastic gradient methods |
0.7 | 1 | 2023 | Stochastic Marginal Likelihood Gradients using Neural Tangent Kernels · ICML 2023 |
Machine learning › Deep learning architectures and training › equivariant neural network
approximate equivariance |
0.6 | 1 | 2022 | Relaxing Equivariance Constraints with Non-stationary Continuous Filters · NeurIPS 2022 |
Machine learning › Deep learning architectures and training
data augmentation |
0.6 | 1 | 2022 | Invariance Learning in Deep Neural Networks with Differentiable Laplace Approximations · NeurIPS 2022 |
Machine learning › Trustworthy machine learning › out-of-distribution generalization
invariant learning |
0.6 | 1 | 2022 | Invariance Learning in Deep Neural Networks with Differentiable Laplace Approximations · NeurIPS 2022 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2019 | Reversible GANs for Memory-Efficient Image-To-Image Translation · CVPR 2019 |
Machine learning › Generative modeling › generative adversarial network
image-to-image translation |
0.4 | 1 | 2019 | Reversible GANs for Memory-Efficient Image-To-Image Translation · CVPR 2019 |
Machine learning › Deep learning architectures and training › feedforward neural network
invertible neural network |
0.4 | 1 | 2019 | Reversible GANs for Memory-Efficient Image-To-Image Translation · CVPR 2019 |
Machine learning › Deep learning architectures and training
transformer |
0.2 | 1 | 2024 | The LLM Surgeon · ICLR 2024 |
Machine learning › Learning theory › neural network theory › neural network kernels
neural tangent kernel |
0.2 | 1 | 2023 | Stochastic Marginal Likelihood Gradients using Neural Tangent Kernels · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
noether's theorem · 0.8kronecker-factored curvature approximation · 0.8approximate bayesian model selection · 0.8stochastic gradient descent · 0.7neural tangent kernel · 0.7linearized laplace approximation · 0.7laplace approximation · 0.7gradient-based learning · 0.7kronecker-factored approximation · 0.6differentiable laplace approximation · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | The LLM SurgeonabstractState-of-the-art language models are becoming increasingly large in an effort to achieve the highest performance on large corpora of available textual data. However, the sheer size of the Transformer architectures makes it difficult to deploy models within computational, environmental or device-specific constraints. We explore data-driven compression of existing pretrained models as an alternative to training smaller models from scratch. To do so, we scale Kronecker-factored curvature approximations of the target loss landscape to large language models. In doing so, we can compute both the dynamic allocation of structures that can be removed as well as updates of remaining weights that account for the removal. We provide a general framework for unstructured, semi-structured and structured pruning and improve upon weight updates to capture more correlations between weights, while remaining computationally efficient. Experimentally, our method can prune rows and columns from a range of OPT models and Llamav2-7B by 20\%-30\%, with a negligible loss in performance, and achieve state-of-the-art results in unstructured and semi-structured pruning of large language models. We will open source our code on GitHub upon acceptance. Tycho F. A. van der Ouderaa, Markus Nagel, Mart van Baalen, Tijmen Blankevoort |
ICLR | 1 |
| 2024 | Noether's Razor: Learning Conserved QuantitiesabstractSymmetries have proven useful in machine learning models, improving generalisation and overall performance. At the same time, recent advancements in learning dynamical systems rely on modelling the underlying Hamiltonian to guarantee the conservation of energy.
These approaches can be connected via a seminal result in mathematical physics: Noether's theorem, which states that symmetries in a dynamical system correspond to conserved quantities.
This work uses Noether's theorem to parameterise symmetries as learnable conserved quantities. We then allow conserved quantities and associated symmetries to be learned directly from train data through approximate Bayesian model selection, jointly with the regular training procedure. As training objective, we derive a variational lower bound to the marginal likelihood. The objective automatically embodies an Occam's Razor effect that avoids collapse of conversation laws to the trivial constant, without the need to manually add and tune additional regularisers. We demonstrate a proof-of-principle on n-harmonic oscillators and n-body systems. We find that our method correctly identifies the correct conserved quantities and U(n) and SE(n) symmetry groups, improving overall performance and predictive accuracy on test data. Tycho F. A. van der Ouderaa, Mark van der Wilk, Pim de Haan |
NeurIPS | 1 |
| 2023 | Stochastic Marginal Likelihood Gradients using Neural Tangent KernelsabstractSelecting hyperparameters in deep learning greatly impacts its effectiveness but requires manual effort and expertise. Recent works show that Bayesian model selection with Laplace approximations can allow to optimize such hyperparameters just like standard neural network parameters using gradients and on the training data. However, estimating a single hyperparameter gradient requires a pass through the entire dataset, limiting the scalability of such algorithms. In this work, we overcome this issue by introducing lower bounds to the linearized Laplace approximation of the marginal likelihood. In contrast to previous estimators, these bounds are amenable to stochastic-gradient-based optimization and allow to trade off estimation accuracy against computational complexity. We derive them using the function-space form of the linearized Laplace, which can be estimated using the neural tangent kernel. Experimentally, we show that the estimators can significantly accelerate gradient-based hyperparameter optimization. Alexander Immer, Tycho F. A. van der Ouderaa, Mark van der Wilk, Gunnar Rätsch, Bernhard Schölkopf |
ICML | 2 |
| 2023 | Learning Layer-wise Equivariances Automatically using GradientsabstractConvolutions encode equivariance symmetries into neural networks leading to better generalisation performance. However, symmetries provide fixed hard constraints on the functions a network can represent, need to be specified in advance, and can not be adapted. Our goal is to allow flexible symmetry constraints that can automatically be learned from data using gradients. Learning symmetry and associated weight connectivity structures from scratch is difficult for two reasons. First, it requires efficient and flexible parameterisations of layer-wise equivariances. Secondly, symmetries act as constraints and are therefore not encouraged by training losses measuring data fit. To overcome these challenges, we improve parameterisations of soft equivariance and learn the amount of equivariance in layers by optimising the marginal likelihood, estimated using differentiable Laplace approximations. The objective balances data fit and model complexity enabling layer-wise symmetry discovery in deep networks. We demonstrate the ability to automatically learn layer-wise equivariances on image classification tasks, achieving equivalent or improved performance over baselines with hard-coded symmetry. Tycho F. A. van der Ouderaa, Alexander Immer, Mark van der Wilk |
NeurIPS | 1 |
| 2022 | Invariance Learning in Deep Neural Networks with Differentiable Laplace ApproximationsabstractData augmentation is commonly applied to improve performance of deep learning by enforcing the knowledge that certain transformations on the input preserve the output. Currently, the data augmentation parameters are chosen by human effort and costly cross-validation, which makes it cumbersome to apply to new datasets. We develop a convenient gradient-based method for selecting the data augmentation without validation data during training of a deep neural network. Our approach relies on phrasing data augmentation as an invariance in the prior distribution on the functions of a neural network, which allows us to learn it using Bayesian model selection. This has been shown to work in Gaussian processes, but not yet for deep neural networks. We propose a differentiable Kronecker-factored Laplace approximation to the marginal likelihood as our objective, which can be optimised without human supervision or validation data. We show that our method can successfully recover invariances present in the data, and that this improves generalisation and data efficiency on image datasets. Alexander Immer, Tycho F. A. van der Ouderaa, Gunnar Rätsch, Vincent Fortuin, Mark van der Wilk |
NeurIPS | 2 |
| 2022 | Relaxing Equivariance Constraints with Non-stationary Continuous FiltersabstractEquivariances provide useful inductive biases in neural network modeling, with the translation equivariance of convolutional neural networks being a canonical example. Equivariances can be embedded in architectures through weight-sharing and place symmetry constraints on the functions a neural network can represent. The type of symmetry is typically fixed and has to be chosen in advance. Although some tasks are inherently equivariant, many tasks do not strictly follow such symmetries. In such cases, equivariance constraints can be overly restrictive. In this work, we propose a parameter-efficient relaxation of equivariance that can effectively interpolate between a (i) non-equivariant linear product, (ii) a strict-equivariant convolution, and (iii) a strictly-invariant mapping. The proposed parameterisation can be thought of as a building block to allow adjustable symmetry structure in neural networks. In addition, we demonstrate that the amount of equivariance can be learned from the training data using backpropagation. Gradient-based learning of equivariance achieves similar or improved performance compared to the best value found by cross-validation and outperforms baselines with partial or strict equivariance on CIFAR-10 and CIFAR-100 image classification tasks. Tycho F. A. van der Ouderaa, David W. Romero, Mark van der Wilk |
NeurIPS | 1 |
| 2022 | Learning invariant weights in neural networksabstractAssumptions about invariances or symmetries in data can significantly increase the predictive power of statistical models. Many commonly used machine learning models are constraint to respect certain symmetries, such as translation equivariance in convolutional neural networks, and incorporating other symmetry types is actively being studied. Yet, learning invariances from the data itself remains an open research problem. It has been shown that the marginal likelihood offers a principled way to learn invariances in Gaussian Processes. We propose a weight-space equivalent to this approach, by minimizing a lower bound on the marginal likelihood to learn invariances in neural networks, resulting in naturally higher performing models. Tycho F. A. van der Ouderaa, Mark van der Wilk |
UAI | 1 |
| 2019 | Reversible GANs for Memory-Efficient Image-To-Image TranslationabstractThe pix2pix and CycleGAN losses have vastly improved the qualitative and quantitative visual quality of results in image-to-image translation tasks. We extend this framework by exploring approximately invertible architectures which are well suited to these losses. These architectures are approximately invertible by design and thus partially satisfy cycle-consistency before training even begins. Furthermore, since invertible architectures have constant memory complexity in depth, these models can be built arbitrarily deep. We are able to demonstrate superior quantitative output on the Cityscapes and Maps datasets at near constant memory budget. Tycho F. A. van der Ouderaa, Daniel E. Worrall |
CVPR | 1 |