Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Lachlan E. MacDonald

dblp:306/7691 · also Lachlan Ewen MacDonald · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0003-3601-9777ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Deep learning architectures and training · 44% Optimization for machine learning · 20% 3D vision · 19%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
neural radiance field
1.322023
Curvature-Aware Training for Coordinate Networks · ICCV 2023
Flow Supervision for Deformable NeRF · CVPR 2023
Machine learning › Learning theory › inductive bias
spectral bias
1.222023
How much does Initialization Affect Generalization? · ICML 2023
On the Frequency-bias of Coordinate-MLPs · NeurIPS 2022
Machine learning › Deep learning architectures and training › training dynamics
edge of stability
0.912025
Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares · NeurIPS 2025
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent
0.912025
Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares · NeurIPS 2025
Machine learning › Learning theory
generalization
0.712023
How much does Initialization Affect Generalization? · ICML 2023
Computer vision › 3D vision
scene flow estimation
0.712023
Flow Supervision for Deformable NeRF · CVPR 2023
Machine learning › Optimization for machine learning
second-order optimization
0.712023
Curvature-Aware Training for Coordinate Networks · ICCV 2023
Machine learning › Deep learning architectures and training
training optimization
0.712023
On skip connections and normalisation layers in deep optimisation · NeurIPS 2023
Machine learning › Deep learning architectures and training
weight initialization
0.712023
How much does Initialization Affect Generalization? · ICML 2023
Machine learning › Deep learning architectures and training › feedforward neural network › MLP-based architecture
coordinate-MLP
0.612022
On the Frequency-bias of Coordinate-MLPs · NeurIPS 2022
Machine learning › Deep learning architectures and training
equivariant neural network
0.612022
Enabling Equivariance for Arbitrary Lie Groups · CVPR 2022
Machine learning › Deep learning architectures and training › training dynamics
frequency bias
0.612022
On the Frequency-bias of Coordinate-MLPs · NeurIPS 2022
Machine learning › Deep learning architectures and training › convolutional neural network › convolution design
group convolution
0.612022
Enabling Equivariance for Arbitrary Lie Groups · CVPR 2022
Machine learning › Optimization for machine learning
implicit regularization
0.612022
On the Frequency-bias of Coordinate-MLPs · NeurIPS 2022
Mathematical optimization
nonconvex optimization
0.312025
Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares · NeurIPS 2025
Machine learning › Deep learning architectures and training › convolutional neural network
residual network
0.212023
On skip connections and normalisation layers in deep optimisation · NeurIPS 2023
Rendering
novel view synthesis
0.212023
Flow Supervision for Deformable NeRF · CVPR 2023

Methods — techniques the papers use, named apart from their topics

riemannian manifold analysis · 1.7overparametrised least squares · 1.7second-order optimization · 1.3inverse function theorem · 1.3curvature-aware training · 1.3backward deformation field · 1.3fourier analysis · 1.2gradient flow analysis · 0.7gradient descent · 0.7curvature analysis · 0.7
YearPublicationVenuePosition
2025 Understanding the Learning Dynamics of LoRA: A Gradient Flow Perspective on Low-Rank Adaptation in Matrix Factorization
abstract
Despite the empirical success of Low-Rank Adaptation (LoRA) in fine-tuning pre-trained models, there is little theoretical understanding of how first-order methods with carefully crafted initialization adapt models to new tasks. In this work, we take the first step towards bridging this gap by theoretically analyzing the learning dynamics of LoRA for matrix factorization (MF) under gradient flow (GF), emphasizing the crucial role of initialization. For small initialization, we theoretically show that GF converges to a neighborhood of the optimal solution, with smaller initialization leading to lower final error. Our analysis shows that the final error is affected by the misalignment between the singular spaces of the pre-trained model and the target matrix, and reducing the initialization scale improves alignment. To address this misalignment, we propose a spectral initialization for LoRA in MF and theoretically prove that GF with small spectral initialization converges to the fine-tuning task with arbitrary precision. Numerical experiments from MF and image classification validate our findings.
Ziqing Xu, Hancheng Min, Lachlan E. MacDonald, Jinqi Luo, Salma Tarmoun, Enrique Mallada, René Vidal
AISTATS3
2025 Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares
abstract
Classical optimisation theory guarantees monotonic objective decrease for gradient descent (GD) when employed in a small step size, or "stable", regime. In contrast, gradient descent on neural networks is frequently performed in a large step size regime called the "edge of stability", in which the objective decreases non-monotonically with an observed implicit bias towards flat minima. In this paper, we take a step toward quantifying this phenomenon by providing convergence rates for gradient descent with large learning rates in an overparametrised least squares setting. The key insight behind our analysis is that, as a consequence of overparametrisation, the set of global minimisers forms a Riemannian manifold $M$, which enables the decomposition of the GD dynamics into components parallel and orthogonal to $M$. The parallel component corresponds to Riemannian gradient descent on the objective sharpness, while the orthogonal component corresponds to a quadratic dynamical system. This insight allows us to derive convergence rates in three regimes characterised by the learning rate size: the subcritical regime, in which transient instability is overcome in finite time before linear convergence to a suboptimally flat global minimum; the critical regime, in which instability persists for all time with a power-law convergence toward the optimally flat global minimum; the supercritical regime, in which instability persists for all time with linear convergence to an oscillation of period two centred on the optimally flat global minimum.
Lachlan E. MacDonald, Hancheng Min, Leandro Palma, Salma Tarmoun, Ziqing Xu, René Vidal
NeurIPS1
2024 D'OH: Decoder-Only Random Hypernetworks for Implicit Neural Representations
Cameron Gordon, Lachlan E. MacDonald, Hemanth Saratchandran, Simon Lucey
ACCV (7)2
2023 Flow Supervision for Deformable NeRF
abstract
In this paper we present a new method for deformable NeRF that can directly use optical flow as supervision. We overcome the major challenge with respect to the computationally inefficiency of enforcing the flow constraints to the backward deformation field, used by deformable NeRFs. Specifically, we show that inverting the backward deformation function is actually not needed for computing scene flows between frames. This insight dramatically simplifies the problem, as one is no longer constrained to deformation functions that can be analytically inverted. Instead, thanks to the weak assumptions required by our derivation based on the inverse function theorem, our approach can be extended to a broad class of commonly used backward deformation field. We present results on monocular novel view synthesis with rapid object motion, and demonstrate significant improvements over baselines without flow supervision.
Chaoyang Wang 0001, Lachlan E. MacDonald, László A. Jeni, Simon Lucey
CVPR2
2023 Curvature-Aware Training for Coordinate Networks
abstract
Coordinate networks are widely used in computer vision due to their ability to represent signals as compressed, continuous entities. However, training these networks with first-order optimizers can be slow, hindering their use in real-time applications. Recent works have opted for shallow voxel-based representations to achieve faster training, but this sacrifices memory efficiency. This work proposes a solution that leverages second-order optimization methods to significantly reduce training times for coordinate networks while maintaining their compressibility. Experiments demonstrate the effectiveness of this approach on various signal modalities, such as audio, images, videos, shape and neural radiance fields (NeRF).
Hemanth Saratchandran, Shin-Fang Ch'ng, Sameera Ramasinghe, Lachlan E. MacDonald, Simon Lucey
ICCV4
2023 How much does Initialization Affect Generalization?
abstract
Characterizing the remarkable generalization properties of over-parameterized neural networks remains an open problem. A growing body of recent literature shows that the bias of stochastic gradient descent (SGD) and architecture choice implicitly leads to better generalization. In this paper, we show on the contrary that, independently of architecture, SGD can itself be the cause of poor generalization if one does not ensure good initialization. Specifically, we prove that *any* differentiably parameterized model, trained under gradient flow, obeys a weak spectral bias law which states that sufficiently high frequencies train arbitrarily slowly. This implies that very high frequencies present at initialization will remain after training, and hamper generalization. Further, we empirically test the developed theoretical insights using practical, deep networks. Finally, we contrast our framework with that supplied by the *flat-minima* conjecture and show that Fourier analysis grants a more reliable framework for understanding the generalization of neural networks.
Sameera Ramasinghe, Lachlan E. MacDonald, Moshiur R. Farazi, Hemanth Saratchandran, Simon Lucey
ICML2
2023 On skip connections and normalisation layers in deep optimisation
abstract
We introduce a general theoretical framework, designed for the study of gradient optimisation of deep neural networks, that encompasses ubiquitous architecture choices including batch normalisation, weight normalisation and skip connections. Our framework determines the curvature and regularity properties of multilayer loss landscapes in terms of their constituent layers, thereby elucidating the roles played by normalisation layers and skip connections in globalising these properties. We then demonstrate the utility of this framework in two respects. First, we give the only proof of which we are aware that a class of deep neural networks can be trained using gradient descent to global optima even when such optima only exist at infinity, as is the case for the cross-entropy cost. Second, we identify a novel causal mechanism by which skip connections accelerate training, which we verify predictively with ResNets on MNIST, CIFAR10, CIFAR100 and ImageNet.
Lachlan E. MacDonald, Jack Valmadre, Hemanth Saratchandran, Simon Lucey
NeurIPS1
2023 On Quantizing Implicit Neural Representations
abstract
The role of quantization within implicit/coordinate neural networks is still not fully understood. We note that using a canonical fixed quantization scheme during training produces poor performance at low bit-rates due to the network weight distributions changing over the course of training. In this work, we show that a non-uniform quantization of neural weights can lead to significant improvements. Specifically, we demonstrate that a clustered quantization enables improved reconstruction. Finally, by characterising a trade-off between quantization and network capacity, we demonstrate that it is possible (while memory inefficient) to reconstruct signals using binary neural networks. We demonstrate our findings experimentally on 2D image reconstruction and 3D radiance fields; and show that simple quantization methods and architecture search can achieve compression of NeRF to less than 16kb with minimal loss in performance (323x smaller than the original NeRF).
Cameron Gordon, Shin-Fang Ch'ng, Lachlan E. MacDonald, Simon Lucey
WACV3
2022 Enabling Equivariance for Arbitrary Lie Groups
abstract
Although provably robust to translational perturbations, convolutional neural networks (CNNs) are known to suffer from extreme performance degradation when presented at test time with more general geometric transformations of inputs. Recently, this limitation has motivated a shift infocus from CNNs to Capsule Networks (CapsNets). However, CapsNets suffer from admitting relatively few theoretical guarantees of invariance. We introduce a rigourous mathematical framework to permit invariance to any Lie group of warps, exclusively using convolutions (over Lie groups), without the need for capsules. Previous work on group convolutions has been hampered by strong assumptions about the group, which precludes the application of such techniques to common warps in computer vision such as affine and homographic. Our framework enables the implementation of group convolutions over any finite-dimensional Lie group. We empirically validate our approach on the benchmark affine-invariant classification task, where we achieve ~30% improvement in accuracy against conventional CNNs while outperforming most CapsNets. As further illustration of the generality of our framework, we train a homography-convolutional model which achieves superior robustness on a homography-perturbed dataset, where CapsNet results degrade.
Lachlan E. MacDonald, Sameera Ramasinghe, Simon Lucey
CVPR1
2022 On the Frequency-bias of Coordinate-MLPs
abstract
We show that typical implicit regularization assumptions for deep neural networks (for regression) do not hold for coordinate-MLPs, a family of MLPs that are now ubiquitous in computer vision for representing high-frequency signals. Lack of such implicit bias disrupts smooth interpolations between training samples, and hampers generalizing across signal regions with different spectra. We investigate this behavior through a Fourier lens and uncover that as the bandwidth of a coordinate-MLP is enhanced, lower frequencies tend to get suppressed unless a suitable prior is provided explicitly. Based on these insights, we propose a simple regularization technique that can mitigate the above problem, which can be incorporated into existing networks without any architectural modifications.
Sameera Ramasinghe, Lachlan E. MacDonald, Simon Lucey
NeurIPS2