EDBT 2026 Demo / reviewers in the wild / expert
Terry J. Lyons
dblp:165/7993 · also Terrence J. Lyons
· DBLP profile ↗
24ranked-venue papers
0as first author
16since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Dynamic Graph Embeddings With Neural Controlled Differential EquationsabstractThis paper focuses on representation learning for dynamic graphs with temporal interactions. A fundamental issue is that both the graph structure and the nodes own their own dynamics, and their blending induces intractable complexity in the temporal evolution over graphs. Drawing inspiration from the recent progress of physical dynamic models in deep neural networks, we propose Graph Neural Controlled Differential Equations (GN-CDEs), a continuous-time framework that jointly models node embeddings and structural dynamics by incorporating a graph enhanced neural network vector field with a time-varying graph path as the control signal. Our framework exhibits several desirable characteristics, including the ability to express dynamics on evolving graphs without piecewise integration, the capability to calibrate trajectories with subsequent data, and robustness to missing observations. Empirical evaluation on a range of dynamic graph representation learning tasks demonstrates the effectiveness of our proposed approach in capturing the complex dynamics of dynamic graphs. Tiexin Qin, Benjamin Walker 0001, Terry J. Lyons, Hong Yan 0001, Haoliang Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Structured Linear CDEs: Maximally Expressive and Parallel-in-Time Sequence ModelsabstractThis work introduces Structured Linear Controlled Differential Equations (SLiCEs), a unifying framework for sequence models with structured, input-dependent state-transition matrices that retain the maximal expressivity of dense matrices whilst being cheaper to compute. The framework encompasses existing architectures, such as input-dependent block-diagonal linear recurrent neural networks and DeltaNet's diagonal-plus-low-rank structure, as well as two novel variants based on sparsity and the Walsh-Hadamard transform. We prove that, unlike the diagonal state-transition matrices of S4D and Mamba, SLiCEs employing block-diagonal, sparse, or Walsh-Hadamard matrices match the maximal expressivity of dense matrices. Empirically, SLiCEs solve the $A_5$ state-tracking benchmark with a single layer, achieve best-in-class length generalisation on regular language tasks among parallel-in-time models, and match the performance of log neural controlled differential equations on six multivariate time-series classification datasets while cutting the average time per training step by a factor of twenty. Benjamin Walker 0001, Lingyi Yang, Nicola Muca Cirone, Cristopher Salvi, Terry J. Lyons |
NeurIPS | 5 |
| 2024 | Log Neural Controlled Differential Equations: The Lie Brackets Make A DifferenceabstractThe vector field of a controlled differential equation (CDE) describes the relationship between a control path and the evolution of a solution path. Neural CDEs (NCDEs) treat time series data as observations from a control path, parameterise a CDE’s vector field using a neural network, and use the solution path as a continuously evolving hidden state. As their formulation makes them robust to irregular sampling rates, NCDEs are a powerful approach for modelling real-world data. Building on neural rough differential equations (NRDEs), we introduce Log-NCDEs, a novel, effective, and efficient method for training NCDEs. The core component of Log-NCDEs is the Log-ODE method, a tool from the study of rough paths for approximating a CDE’s solution. Log-NCDEs are shown to outperform NCDEs, NRDEs, the linear recurrent unit, S5, and MAMBA on a range of multivariate time series datasets with up to $50{,}000$ observations. Benjamin Walker 0001, Andrew D. McLeod, Tiexin Qin, Yichuan Cheng, Haoliang Li, Terry J. Lyons |
ICML | 6 |
| 2024 | Theoretical Foundations of Deep Selective State-Space ModelsabstractStructured state-space models (SSMs) are gaining popularity as effective foundational architectures for sequential data, demonstrating outstanding performance across a diverse set of domains alongside desirable scalability properties. Recent developments show that if the linear recurrence powering SSMs allows for a selectivity mechanism leveraging multiplicative interactions between inputs and hidden states (e.g. Mamba, GLA, Hawk/Griffin, HGRN2), then the resulting architecture can surpass attention-powered foundation models trained on text in both accuracy and efficiency, at scales of billion parameters. In this paper, we give theoretical grounding to the selectivity mechanism, often linked to in-context learning, using tools from Rough Path Theory. We provide a framework for the theoretical analysis of generalized selective SSMs, fully characterizing their expressive power and identifying the gating mechanism as the crucial architectural choice. Our analysis provides a closed-form description of the expressive powers of modern SSMs, such as Mamba, quantifying theoretically the drastic improvement in performance from the previous generation of models, such as S4. Our theory not only motivates the success of modern selective state-space models, but also provides a solid framework to understand the expressive power of future SSM variants. In particular, it suggests cross-channel interactions could play a vital role in future improvements. Nicola Muca Cirone, Antonio Orvieto, Benjamin Walker 0001, Cristopher Salvi, Terry J. Lyons |
NeurIPS | 5 |
| 2023 | Sampling-based Nyström Approximation and Kernel QuadratureabstractWe analyze the Nyström approximation of a positive definite kernel associated with a probability measure. We first prove an improved error bound for the conventional Nyström approximation with i.i.d. sampling and singular-value decomposition in the continuous regime; the proof techniques are borrowed from statistical learning theory. We further introduce a refined selection of subspaces in Nyström approximation with theoretical guarantees that is applicable to non-i.i.d. landmark points. Finally, we discuss their application to convex kernel quadrature and give novel theoretical guarantees as well as numerical observations. Satoshi Hayakawa, Harald Oberhauser, Terry J. Lyons |
ICML | 3 |
| 2022 | Path Signatures for Non-Intrusive Load MonitoringabstractNon-intrusive load monitoring (NILM) is the analysis of electricity loads by means of a single supply wire, so avoiding separate monitors on individual appliances. Some approaches to NILM use the V-I trajectory for feature generation but they apply ad-hoc rules to generate the feature vector. This paper demonstrates a systematic method of feature generation called the path signature which has recently been applied in machine learning, often with notable success. We show how the path signature generates features from the V-I trajectory to give a test set accuracy of 98.81% on the COOLL dataset. We conclude that the path signature is easier to use and generalize than ad-hoc features, and it can be applied to many other applications which use multivariate sequential data. Paul Moore, Theodor-Mihai Iliant, Filip-Alexandru Ion, Yue Wu 0016, Terry J. Lyons |
ICASSP | 5 |
| 2022 | Positively Weighted Kernel Quadrature via SubsamplingabstractWe study kernel quadrature rules with convex weights. Our approach combines the spectral properties of the kernel with recombination results about point measures. This results in effective algorithms that construct convex quadrature rules using only access to i.i.d. samples from the underlying measure and evaluation of the kernel and that result in a small worst-case error. In addition to our theoretical results and the benefits resulting from convex weights, our experiments indicate that this construction can compete with the optimal bounds in well-known examples. Satoshi Hayakawa, Harald Oberhauser, Terry J. Lyons |
NeurIPS | 3 |
| 2021 | Distribution Regression for Sequential DataabstractDistribution regression refers to the supervised learning problem where labels are only available for groups of inputs instead of individual inputs. In this paper, we develop a rigorous mathematical framework for distribution regression where inputs are complex data streams. Leveraging properties of the expected signature and a recent signature kernel trick for sequential data from stochastic analysis, we introduce two new learning techniques, one feature-based and the other kernel-based. Each is suited to a different data regime in terms of the number of data streams and the dimensionality of the individual streams. We provide theoretical results on the universality of both approaches and demonstrate empirically their robustness to irregularly sampled multivariate time-series, achieving state-of-the-art performance on both synthetic and real-world examples from thermodynamics, mathematical finance and agricultural science. Maud Lemercier, Cristopher Salvi, Theodoros Damoulas, Edwin V. Bonilla, Terry J. Lyons |
AISTATS | 5 |
| 2021 | Logsig-RNN: a novel network for robust and efficient skeleton-based action recognition
Hao Ni 0001, Shujian Liao, Kevin Schlegel, Terry J. Lyons |
BMVC | 5 |
| 2021 | Modelling Paralinguistic Properties in Conversational Speech to Detect Bipolar Disorder and Borderline Personality DisorderabstractBipolar disorder (BD) and borderline personality disorder (BPD) are two chronic mental health conditions that clinicians find challenging to distinguish based on clinical interviews, due to their overlapping symptoms. In this work, we investigate the automatic detection of these two conditions by modelling both verbal and non-verbal cues in a set of interviews. We propose a new approach of modelling short-term features with visibility-signature transform, and compare it with widely used high-level statistical functions. We demonstrate the superior performance of our proposed signature-based model. Furthermore, we show the role of different sets of features in characterising BD and BPD. Bo Wang 0034, Yue Wu 0016, Nemanja Vaci, Maria Liakata, Terry J. Lyons, Kate Saunders |
ICASSP | 5 |
| 2021 | Signatory: differentiable computations of the signature and logsignature transforms, on both CPU and GPU
Patrick Kidger, Terry J. Lyons |
ICLR | 2 |
| 2021 | "Hey, that's not an ODE": Faster ODE Adjoints via Seminorms
Patrick Kidger, Ricky T. Q. Chen, Terry J. Lyons |
ICML | 3 |
| 2021 | Neural SDEs as Infinite-Dimensional GANsabstractStochastic differential equations (SDEs) are a staple of mathematical modelling of temporal dynamics. However, a fundamental limitation has been that such models have typically been relatively inflexible, which recent work introducing Neural SDEs has sought to solve. Here, we show that the current classical approach to fitting SDEs may be approached as a special case of (Wasserstein) GANs, and in doing so the neural and classical regimes may be brought together. The input noise is Brownian motion, the output samples are time-evolving paths produced by a numerical solver, and by parameterising a discriminator as a Neural Controlled Differential Equation (CDE), we obtain Neural SDEs as (in modern machine learning parlance) continuous-time generative time series models. Unlike previous work on this problem, this is a direct extension of the classical approach without reference to either prespecified statistics or density functions. Arbitrary drift and diffusions are admissible, so as the Wasserstein loss has a unique global minima, in the infinite data limit \textit{any} SDE may be learnt. Patrick Kidger, James Foster, Terry J. Lyons |
ICML | 4 |
| 2021 | SigGPDE: Scaling Sparse Gaussian Processes on Sequential DataabstractMaking predictions and quantifying their uncertainty when the input data is sequential is a fundamental learning challenge, recently attracting increasing attention. We develop SigGPDE, a new scalable sparse variational inference framework for Gaussian Processes (GPs) on sequential data. Our contribution is twofold. First, we construct inducing variables underpinning the sparse approximation so that the resulting evidence lower bound (ELBO) does not require any matrix inversion. Second, we show that the gradients of the GP signature kernel are solutions of a hyperbolic partial differential equation (PDE). This theoretical insight allows us to build an efficient back-propagation algorithm to optimize the ELBO. We showcase the significant computational gains of SigGPDE compared to existing methods, while achieving state-of-the-art performance for classification tasks on large datasets of up to 1 million multivariate time series. Maud Lemercier, Cristopher Salvi, Thomas Cass, Edwin V. Bonilla, Theodoros Damoulas, Terry J. Lyons |
ICML | 6 |
| 2021 | Efficient and Accurate Gradients for Neural SDEsabstractNeural SDEs combine many of the best qualities of both RNNs and SDEs, and as such are a natural choice for modelling many types of temporal dynamics. They offer memory efficiency, high-capacity function approximation, and strong priors on model space. Neural SDEs may be trained as VAEs or as GANs; in either case it is necessary to backpropagate through the SDE solve. In particular this may be done by constructing a backwards-in-time SDE whose solution is the desired parameter gradients. However, this has previously suffered from severe speed and accuracy issues, due to high computational complexity, numerical errors in the SDE solve, and the cost of reconstructing Brownian motion. Here, we make several technical innovations to overcome these issues. First, we introduce the \textit{reversible Heun method}: a new SDE solver that is algebraically reversible -- which reduces numerical gradient errors to almost zero, improving several test metrics by substantial margins over state-of-the-art. Moreover it requires half as many function evaluations as comparable solvers, giving up to a $1.98\times$ speedup. Next, we introduce the \textit{Brownian interval}. This is a new and computationally efficient way of exactly sampling \textit{and reconstructing} Brownian motion; this is in contrast to previous reconstruction techniques that are both approximate and relatively slow. This gives up to a $10.6\times$ speed improvement over previous techniques. After that, when specifically training Neural SDEs as GANs (Kidger et al. 2021), we demonstrate how SDE-GANs may be trained through careful weight clipping and choice of activation function. This reduces computational cost (giving up to a $1.87\times$ speedup), and removes the truncation errors of the double adjoint required for gradient penalty, substantially improving several test metrics. Altogether these techniques offer substantial improvements over the state-of-the-art, with respect to both training speed and with respect to classification, prediction, and MMD test metrics. We have contributed implementations of all of our techniques to the \texttt{torchsde} library to help facilitate their adoption. Patrick Kidger, James Foster, Terry J. Lyons |
NeurIPS | 4 |
| 2021 | Higher Order Kernel Mean Embeddings to Capture Filtrations of Stochastic ProcessesabstractStochastic processes are random variables with values in some space of paths. However, reducing a stochastic process to a path-valued random variable ignores its filtration, i.e. the flow of information carried by the process through time. By conditioning the process on its filtration, we introduce a family of higher order kernel mean embeddings (KMEs) that generalizes the notion of KME to capture additional information related to the filtration. We derive empirical estimators for the associated higher order maximum mean discrepancies (MMDs) and prove consistency. We then construct a filtration-sensitive kernel two-sample test able to capture information that gets missed by the standard MMD test. In addition, leveraging our higher order MMDs we construct a family of universal kernels on stochastic processes that allows to solve real-world calibration and optimal stopping problems in quantitative finance (such as the pricing of American options) via classical kernel-based regression methods. Finally, adapting existing tests for conditional independence to the case of stochastic processes, we design a causal-discovery algorithm to recover the causal graph of structural dependencies among interacting bodies solely from observations of their multidimensional trajectories. Cristopher Salvi, Maud Lemercier, Blanka Horvath, Theodoros Damoulas, Terry J. Lyons |
NeurIPS | 6 |
| 2020 | Universal Approximation with Deep Narrow NetworksabstractThe classical Universal Approximation Theorem holds for neural networks of arbitrary width and bounded depth. Here we consider the natural ‘dual’ scenario for networks of bounded width and arbitrary depth. Precisely, let $n$ be the number of inputs neurons, $m$ be the number of output neurons, and let $\rho$ be any nonaffine continuous function, with a continuous nonzero derivative at some point. Then we show that the class of neural networks of arbitrary depth, width $n + m + 2$, and activation function $\rho$, is dense in $C(K; \mathbb{R}^m)$ for $K \subseteq \mathbb{R}^n$ with $K$ compact. This covers every activation function possible to use in practice, and also includes polynomial activation functions, which is unlike the classical version of the theorem, and provides a qualitative difference between deep narrow networks and shallow wide networks. We then consider several extensions of this result. In particular we consider nowhere differentiable activation functions, density in noncompact domains with respect to the $L^p$-norm, and how the width may be reduced to just $n + m + 1$ for ‘most’ activation functions. Patrick Kidger, Terry J. Lyons |
COLT | 2 |
| 2020 | Signature features with the visibility transformationabstractIn this paper we put the visibility transformation on a clear theoretical footing and show that this transform is able to embed the effect of the absolute position of the data stream into signature features in a unified and efficient way. The generated feature set is particularly useful in pattern recognition tasks, for its simplifying role in allowing the signature feature set to accommodate nonlinear functions of absolute and relative values. Yue Wu 0016, Hao Ni 0001, Terry J. Lyons, Robin L. Hudson |
ICPR | 3 |
| 2020 | Learning to Detect Bipolar Disorder and Borderline Personality Disorder with Language and Speech in Non-Clinical InterviewsabstractBipolar disorder (BD) and borderline personality disorder (BPD) are both chronic psychiatric disorders.However, their overlapping symptoms and common comorbidity make it challenging for the clinicians to distinguish the two conditions on the basis of a clinical interview.In this work, we first present a new multi-modal dataset containing interviews involving individuals with BD or BPD being interviewed about a non-clinical topic .We investigate the automatic detection of the two conditions, and demonstrate a good linear classifier that can be learnt using a down-selected set of features from the different aspects of the interviews and a novel approach of summarising these features.Finally, we find that different sets of features characterise BD and BPD, thus providing insights into the difference between the automatic screening of the two conditions. Bo Wang 0034, Yue Wu 0016, Niall Taylor, Terry J. Lyons, Maria Liakata, Alejo J. Nevado-Holgado, Kate Saunders |
INTERSPEECH | 4 |
| 2020 | Neural Controlled Differential Equations for Irregular Time SeriesabstractNeural ordinary differential equations are an attractive option for modelling temporal dynamics. However, a fundamental issue is that the solution to an ordinary differential equation is determined by its initial condition, and there is no mechanism for adjusting the trajectory based on subsequent observations. Here, we demonstrate how this may be resolved through the well-understood mathematics of \emph{controlled differential equations}. The resulting \emph{neural controlled differential equation} model is directly applicable to the general setting of partially-observed irregularly-sampled multivariate time series, and (unlike previous work on this problem) it may utilise memory-efficient adjoint-based backpropagation even across observations. We demonstrate that our model achieves state-of-the-art performance against similar (ODE or RNN based) models in empirical studies on a range of datasets. Finally we provide theoretical results demonstrating universal approximation, and that our model subsumes alternative ODE models. Patrick Kidger, James Morrill, James Foster, Terry J. Lyons |
NeurIPS | 4 |
| 2019 | A Path Signature Approach for Speech Emotion RecognitionabstractAutomatic speech emotion recognition (SER) remains a \ndifficult task within human-computer interaction, despite increasing interest in the research community. One key challenge is how to effectively integrate short-term characterisation \nof speech segments with long-term information such as temporal variations. Motivated by the numerical approximation theory of stochastic differential equations (SDEs), we propose the \nnovel use of path signatures. The latter provide a pathwise definition to solve SDEs, for the integration of short speech frames. \nFurthermore we propose a hierarchical tree structure of path signatures, to capture both global and local information. A simple tree-based convolutional neural network (TBCNN) is used \nfor learning the structural information stemming from dyadic \npath-tree signatures. Our experimental results on a widely \nused benchmark dataset demonstrate comparable performance \nto complex neural network based systems. Bo Wang 0034, Maria Liakata, Hao Ni 0001, Terry J. Lyons, Alejo J. Nevado-Holgado, Kate Saunders |
INTERSPEECH | 4 |
| 2019 | Deep Signature TransformsabstractThe signature is an infinite graded sequence of statistics known to characterise a stream of data up to a negligible equivalence class. It is a transform which has previously been treated as a fixed feature transformation, on top of which a model may be built. We propose a novel approach which combines the advantages of the signature transform with modern deep learning frameworks. By learning an augmentation of the stream prior to the signature transform, the terms of the signature may be selected in a data-dependent way. More generally, we describe how the signature transform may be used as a layer anywhere within a neural network. In this context it may be interpreted as a pooling operation. We present the results of empirical experiments to back up the theoretical justification. Code available at \texttt{github.com/patrick-kidger/Deep-Signature-Transforms}. Patrick Kidger, Patric Bonnier, Imanol Pérez Arribas, Cristopher Salvi, Terry J. Lyons |
NeurIPS | 5 |
| 2018 | Learning Spatial-Semantic Context with Fully Convolutional Recurrent Network for Online Handwritten Chinese Text RecognitionabstractOnline handwritten Chinese text recognition (OHCTR) is a challenging problem as it involves a large-scale character set, ambiguous segmentation, and variable-length input sequences. In this paper, we exploit the outstanding capability of path signature to translate online pen-tip trajectories into informative signature feature maps, successfully capturing the analytic and geometric properties of pen strokes with strong local invariance and robustness. A multi-spatial-context fully convolutional recurrent network (MC-FCRN) is proposed to exploit the multiple spatial contexts from the signature feature maps and generate a prediction sequence while completely avoiding the difficult segmentation problem. Furthermore, an implicit language model is developed to make predictions based on semantic context within a predicting feature sequence, providing a new perspective for incorporating lexicon constraints and prior knowledge about a certain language in the recognition procedure. Experiments on two standard benchmarks, Dataset-CASIA and Dataset-ICDAR, yielded outstanding results, with correct rates of 97.50 and 96.58 percent, respectively, which are significantly better than the best result reported thus far in the literature. Zecheng Xie, Zenghui Sun, Hao Ni 0001, Terry J. Lyons |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2016 | Rotation-free online handwritten character recognition using dyadic path signature features, hanging normalization, and deep neural networkabstractThe path signature feature (PSF) which was initially introduced in rough paths theory as a branch of stochastic analysis, has recently been successfully applied to the field of pattern recognition for extracting sufficient quantity of information contained in a finite trajectory, but with potentially high dimension. In this paper, we propose a variation of path signature representation, namely the dyadic path signature feature (D-PSF), to fully characterize the trajectory using a hierarchical structure to solve the rotation-free online handwritten character recognition (OLHCR) problem. We adopt the deep neural network (DNN) as classifier, and investigate three hanging normalization methods to improve the robustness of the DNN to rotational distortions. Extensive experiments on digits, English letters, and Chinese radicals demonstrated that the proposed D-PSF, jointly with hanging normalization and DNN, achieved very promising results for rotated OLHCR, significantly outperforming previous methods. Hao Ni 0001, Terry J. Lyons |
ICPR | 4 |