VLDB 2026 Research / reviewers in the wild / expert
Patrick Kidger
dblp:241/7262
· DBLP profile ↗
9ranked-venue papers
7as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 7 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Deep learning architectures and training · 49% Generative modeling · 20% Efficient and distributed learning · 10% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
High-performance computing · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 20 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
neural differential equations |
1.4 | 3 | 2021 | Efficient and Accurate Gradients for Neural SDEs · NeurIPS 2021 Neural Rough Differential Equations for Long Time Series · ICML 2021 Neural Controlled Differential Equations for Irregular Time Series · NeurIPS 2020 |
Machine learning › Generative modeling
generative adversarial network |
1.0 | 2 | 2021 | Efficient and Accurate Gradients for Neural SDEs · NeurIPS 2021 Neural SDEs as Infinite-Dimensional GANs · ICML 2021 |
Program synthesis and code generation
code generation with language models |
0.9 | 1 | 2025 | AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs? · NeurIPS 2025 |
High-performance computing › performance optimization
numerical program optimization |
0.9 | 1 | 2025 | AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs? · NeurIPS 2025 |
High-performance computing
performance optimization |
0.9 | 1 | 2025 | AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs? · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
gradient computation |
0.7 | 2 | 2021 | Efficient and Accurate Gradients for Neural SDEs · NeurIPS 2021 "Hey, that's not an ODE": Faster ODE Adjoints via Seminorms · ICML 2021 |
Machine learning › Deep learning architectures and training › neural differential equations
neural controlled differential equations |
0.6 | 2 | 2021 | Neural Controlled Differential Equations for Irregular Time Series · NeurIPS 2020 Neural SDEs as Infinite-Dimensional GANs · ICML 2021 |
Machine learning › Deep learning architectures and training › differentiable programming
differentiable algorithms |
0.5 | 1 | 2021 | Signatory: differentiable computations of the signature and logsignature transforms, on both CPU and GPU · ICLR 2021 |
Machine learning › Efficient and distributed learning › hardware acceleration
GPU acceleration |
0.5 | 1 | 2021 | Signatory: differentiable computations of the signature and logsignature transforms, on both CPU and GPU · ICLR 2021 |
Machine learning › Efficient and distributed learning
hardware acceleration |
0.5 | 1 | 2021 | Signatory: differentiable computations of the signature and logsignature transforms, on both CPU and GPU · ICLR 2021 |
Machine learning › Deep learning architectures and training › neural differential equations
neural ordinary differential equations |
0.5 | 1 | 2021 | "Hey, that's not an ODE": Faster ODE Adjoints via Seminorms · ICML 2021 |
Machine learning › Probabilistic and Bayesian machine learning › continuous-time model › stochastic differential equations
neural SDE |
0.5 | 1 | 2021 | Efficient and Accurate Gradients for Neural SDEs · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › neural differential equations
neural stochastic differential equations |
0.5 | 1 | 2021 | Neural SDEs as Infinite-Dimensional GANs · ICML 2021 |
Machine learning › Time series and sequential data
time series modeling |
0.5 | 1 | 2021 | Neural Rough Differential Equations for Long Time Series · ICML 2021 |
Machine learning › Generative modeling › generative adversarial network
Wasserstein GAN |
0.5 | 1 | 2021 | Neural SDEs as Infinite-Dimensional GANs · ICML 2021 |
Machine learning › Deep learning architectures and training › neural differential equations
controlled differential equation |
0.4 | 1 | 2020 | Neural Controlled Differential Equations for Irregular Time Series · NeurIPS 2020 |
Computer vision › Video understanding and tracking › temporal modeling
temporal dynamics modeling |
0.4 | 1 | 2020 | Neural Controlled Differential Equations for Irregular Time Series · NeurIPS 2020 |
Machine learning › Learning theory › approximation theory › neural network approximation
universal approximation |
0.4 | 1 | 2020 | Universal Approximation with Deep Narrow Networks · COLT 2020 |
Machine learning › Representation and self-supervised learning
feature transformation |
0.4 | 1 | 2019 | Deep Signature Transforms · NeurIPS 2019 |
Machine learning › Deep learning architectures and training
neural network layers |
0.4 | 1 | 2019 | Deep Signature Transforms · NeurIPS 2019 |
Methods — techniques the papers use, named apart from their topics
language model agent · 1.7benchmarking · 1.7adjoint backpropagation · 0.9stochastic differential equation · 0.5seminorms · 0.5rough path theory · 0.5reversible heun method · 0.5numerical solvers · 0.5neural CDE · 0.5log-signature · 0.5brownian interval · 0.5adjoint sensitivity method · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs?abstractDespite progress in language model (LM) capabilities, evaluations have thus far focused on models' performance on tasks that humans have previously solved, including in programming (SWE-Bench) and mathematics (FrontierMath). We therefore propose testing models' ability to design and implement algorithms in an open-ended benchmark: We task LMs with writing code that efficiently solves computationally challenging problems in computer science, physics, and mathematics. Our AlgoTune benchmark consists of 120 tasks collected from domain experts and a framework for validating and timing LM-synthesized solution code, which is compared to reference implementations from popular open-source packages.In addition, we develop a baseline LM agent, AlgoTuner, and evaluate its performance across a suite of frontier models.AlgoTuner achieves an average 1.58x speedup against reference solvers, including methods from packages such as SciPy, scikit-learn and CVXPY.However, we find that current models fail to discover algorithmic innovations, instead preferring surface-level optimizations. We hope that AlgoTune catalyzes the development of LM agents exhibiting creative problem solving beyond state-of-the-art human performance. Ori Press, Brandon Amos, Yikai Wu 0001, Samuel K. Ainsworth, Dominik Krupke, Patrick Kidger, Touqir Sajed, Bartolomeo Stellato, Jisun Park 0003, Nathanael Bosch, Eli Meril, Albert Steppi, Arman Zharmagambetov, Fangzhao Zhang, David Pérez-Piñeiro, Alberto Mercurio, Ni Zhan 0002, Talor Abramovich, Kilian Lieret, Shirley Huang, Matthias Bethge, Ofir Press |
NeurIPS | 7 |
| 2021 | Signatory: differentiable computations of the signature and logsignature transforms, on both CPU and GPU
Patrick Kidger, Terry J. Lyons |
ICLR | 1 |
| 2021 | "Hey, that's not an ODE": Faster ODE Adjoints via Seminorms
Patrick Kidger, Ricky T. Q. Chen, Terry J. Lyons |
ICML | 1 |
| 2021 | Neural SDEs as Infinite-Dimensional GANsabstractStochastic differential equations (SDEs) are a staple of mathematical modelling of temporal dynamics. However, a fundamental limitation has been that such models have typically been relatively inflexible, which recent work introducing Neural SDEs has sought to solve. Here, we show that the current classical approach to fitting SDEs may be approached as a special case of (Wasserstein) GANs, and in doing so the neural and classical regimes may be brought together. The input noise is Brownian motion, the output samples are time-evolving paths produced by a numerical solver, and by parameterising a discriminator as a Neural Controlled Differential Equation (CDE), we obtain Neural SDEs as (in modern machine learning parlance) continuous-time generative time series models. Unlike previous work on this problem, this is a direct extension of the classical approach without reference to either prespecified statistics or density functions. Arbitrary drift and diffusions are admissible, so as the Wasserstein loss has a unique global minima, in the infinite data limit \textit{any} SDE may be learnt. Patrick Kidger, James Foster, Terry J. Lyons |
ICML | 1 |
| 2021 | Neural Rough Differential Equations for Long Time SeriesabstractNeural controlled differential equations (CDEs) are the continuous-time analogue of recurrent neural networks, as Neural ODEs are to residual networks, and offer a memory-efficient continuous-time way to model functions of potentially irregular time series. Existing methods for computing the forward pass of a Neural CDE involve embedding the incoming time series into path space, often via interpolation, and using evaluations of this path to drive the hidden state. Here, we use rough path theory to extend this formulation. Instead of directly embedding into path space, we instead represent the input signal over small time intervals through its \textit{log-signature}, which are statistics describing how the signal drives a CDE. This is the approach for solving \textit{rough differential equations} (RDEs), and correspondingly we describe our main contribution as the introduction of Neural RDEs. This extension has a purpose: by generalising the Neural CDE approach to a broader class of driving signals, we demonstrate particular advantages for tackling long time series. In this regime, we demonstrate efficacy on problems of length up to 17k observations and observe significant training speed-ups, improvements in model performance, and reduced memory requirements compared to existing approaches. James Morrill, Cristopher Salvi, Patrick Kidger, James Foster |
ICML | 3 |
| 2021 | Efficient and Accurate Gradients for Neural SDEsabstractNeural SDEs combine many of the best qualities of both RNNs and SDEs, and as such are a natural choice for modelling many types of temporal dynamics. They offer memory efficiency, high-capacity function approximation, and strong priors on model space. Neural SDEs may be trained as VAEs or as GANs; in either case it is necessary to backpropagate through the SDE solve. In particular this may be done by constructing a backwards-in-time SDE whose solution is the desired parameter gradients. However, this has previously suffered from severe speed and accuracy issues, due to high computational complexity, numerical errors in the SDE solve, and the cost of reconstructing Brownian motion. Here, we make several technical innovations to overcome these issues. First, we introduce the \textit{reversible Heun method}: a new SDE solver that is algebraically reversible -- which reduces numerical gradient errors to almost zero, improving several test metrics by substantial margins over state-of-the-art. Moreover it requires half as many function evaluations as comparable solvers, giving up to a $1.98\times$ speedup. Next, we introduce the \textit{Brownian interval}. This is a new and computationally efficient way of exactly sampling \textit{and reconstructing} Brownian motion; this is in contrast to previous reconstruction techniques that are both approximate and relatively slow. This gives up to a $10.6\times$ speed improvement over previous techniques. After that, when specifically training Neural SDEs as GANs (Kidger et al. 2021), we demonstrate how SDE-GANs may be trained through careful weight clipping and choice of activation function. This reduces computational cost (giving up to a $1.87\times$ speedup), and removes the truncation errors of the double adjoint required for gradient penalty, substantially improving several test metrics. Altogether these techniques offer substantial improvements over the state-of-the-art, with respect to both training speed and with respect to classification, prediction, and MMD test metrics. We have contributed implementations of all of our techniques to the \texttt{torchsde} library to help facilitate their adoption. Patrick Kidger, James Foster, Terry J. Lyons |
NeurIPS | 1 |
| 2020 | Universal Approximation with Deep Narrow NetworksabstractThe classical Universal Approximation Theorem holds for neural networks of arbitrary width and bounded depth. Here we consider the natural ‘dual’ scenario for networks of bounded width and arbitrary depth. Precisely, let $n$ be the number of inputs neurons, $m$ be the number of output neurons, and let $\rho$ be any nonaffine continuous function, with a continuous nonzero derivative at some point. Then we show that the class of neural networks of arbitrary depth, width $n + m + 2$, and activation function $\rho$, is dense in $C(K; \mathbb{R}^m)$ for $K \subseteq \mathbb{R}^n$ with $K$ compact. This covers every activation function possible to use in practice, and also includes polynomial activation functions, which is unlike the classical version of the theorem, and provides a qualitative difference between deep narrow networks and shallow wide networks. We then consider several extensions of this result. In particular we consider nowhere differentiable activation functions, density in noncompact domains with respect to the $L^p$-norm, and how the width may be reduced to just $n + m + 1$ for ‘most’ activation functions. Patrick Kidger, Terry J. Lyons |
COLT | 1 |
| 2020 | Neural Controlled Differential Equations for Irregular Time SeriesabstractNeural ordinary differential equations are an attractive option for modelling temporal dynamics. However, a fundamental issue is that the solution to an ordinary differential equation is determined by its initial condition, and there is no mechanism for adjusting the trajectory based on subsequent observations. Here, we demonstrate how this may be resolved through the well-understood mathematics of \emph{controlled differential equations}. The resulting \emph{neural controlled differential equation} model is directly applicable to the general setting of partially-observed irregularly-sampled multivariate time series, and (unlike previous work on this problem) it may utilise memory-efficient adjoint-based backpropagation even across observations. We demonstrate that our model achieves state-of-the-art performance against similar (ODE or RNN based) models in empirical studies on a range of datasets. Finally we provide theoretical results demonstrating universal approximation, and that our model subsumes alternative ODE models. Patrick Kidger, James Morrill, James Foster, Terry J. Lyons |
NeurIPS | 1 |
| 2019 | Deep Signature TransformsabstractThe signature is an infinite graded sequence of statistics known to characterise a stream of data up to a negligible equivalence class. It is a transform which has previously been treated as a fixed feature transformation, on top of which a model may be built. We propose a novel approach which combines the advantages of the signature transform with modern deep learning frameworks. By learning an augmentation of the stream prior to the signature transform, the terms of the signature may be selected in a data-dependent way. More generally, we describe how the signature transform may be used as a layer anywhere within a neural network. In this context it may be interpreted as a pooling operation. We present the results of empirical experiments to back up the theoretical justification. Code available at \texttt{github.com/patrick-kidger/Deep-Signature-Transforms}. Patrick Kidger, Patric Bonnier, Imanol Pérez Arribas, Cristopher Salvi, Terry J. Lyons |
NeurIPS | 1 |