Alistair Shilton

dblp:01/5564 · DBLP profile ↗
← Back
27ranked-venue papers
13as first author
9since 2021 · last 2025
0000-0002-0849-3271ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 11 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Reproducing Kernel Banach Space Models for Neural Networks with Application to Rademacher Complexity Analysis
abstract
This paper explores the use of Hermite transform based reproducing kernel Banach space methods to construct exact or un-approximated models of feedforward neural networks of arbitrary width, depth and topology, including ResNet and Transformers networks, assuming only a feedforward topology, finite energy activations and finite (spectral-) norm weights and biases. Using this model, two straightforward but surprisingly tight bounds on Rademacher complexity are derived, precisely (1) a general bound that is width-independent and scales exponentially with depth; and (2) a width- and depth-independent bound for networks with appropriately constrained (below threshold) weights and biases.
Alistair Shilton, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh
NeurIPS1
2025 Accelerated experimental design using a human-AI teaming framework
Arun Kumar Anjanapura Venkatesh, Alistair Shilton, Sunil Gupta 0001, Shannon Ryan, Majid Abdolshah, Hung Le 0002, Santu Rana, Julian Berk, Mahad Rashid, Svetha Venkatesh
Knowl. Based Syst.2
2024 Enhanced Bayesian Optimization via Preferential Modeling of Abstract Properties
A. V. Arun Kumar, Alistair Shilton, Sunil Gupta 0001, Santu Rana, Stewart Greenhill, Svetha Venkatesh
ECML/PKDD (6)2
2024 PINN-BO: A Black-Box Optimization Algorithm Using Physics-Informed Neural Networks
Dat Phan-Trong, Hung The Tran, Alistair Shilton, Sunil Gupta 0001
ECML/PKDD (2)3
2023 Gradient Descent in Neural Networks as Sequential Learning in Reproducing Kernel Banach Space
abstract
The study of Neural Tangent Kernels (NTKs) has provided much needed insight into convergence and generalization properties of neural networks in the over-parametrized (wide) limit by approximating the network using a first-order Taylor expansion with respect to its weights in the neighborhood of their initialization values. This allows neural network training to be analyzed from the perspective of reproducing kernel Hilbert spaces (RKHS), which is informative in the over-parametrized regime, but a poor approximation for narrower networks as the weights change more during training. Our goal is to extend beyond the limits of NTK toward a more general theory. We construct an exact power-series representation of the neural network in a finite neighborhood of the initial weights as an inner product of two feature maps, respectively from data and weight-step space, to feature space, allowing neural network training to be analyzed from the perspective of reproducing kernel Banach space (RKBS). We prove that, regardless of width, the training sequence produced by gradient descent can be exactly replicated by regularized sequential learning in RKBS. Using this, we present novel bound on uniform convergence where the iterations count and learning rate play a central role, giving new theoretical insight into neural network training.
Alistair Shilton, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh
ICML1
2022 TRF: Learning Kernels with Tuned Random Features
Alistair Shilton, Sunil Gupta 0001, Santu Rana, Arun Kumar Anjanapura Venkatesh, Svetha Venkatesh
AAAI1
2022 Human-AI Collaborative Bayesian Optimisation
abstract
Abstract Human-AI collaboration looks at harnessing the complementary strengths of both humans and AI. We propose a new method for human-AI collaboration in Bayesian optimisation where the optimum is mainly pursued by the Bayesian optimisation algorithm following complex computation, whilst getting occasional help from the accompanying expert having a deeper knowledge of the underlying physical phenomenon. We expect experts to have some understanding of the correlation structures of the experimental system, but not the location of the optimum. The expert provides feedback by either changing the current recommendation or providing her belief on the good and bad regions of the search space based on the current observations. Our proposed method takes such feedback to build a model that aligns with the expert’s model and then uses it for optimisation. We provide theoretical underpinning on why such an approach may be more efficient than the one without expert’s feedback. The empirical results show the robustness and superiority of our method with promising efficiency gains.
Arun Kumar A. V., Santu Rana, Alistair Shilton, Svetha Venkatesh
NeurIPS3
2021 Kernel Functional Optimisation
abstract
Traditional methods for kernel selection rely on parametric kernel functions or a combination thereof and although the kernel hyperparameters are tuned, these methods often provide sub-optimal results due to the limitations induced by the parametric forms. In this paper, we propose a novel formulation for kernel selection using efficient Bayesian optimisation to find the best fitting non-parametric kernel. The kernel is expressed using a linear combination of functions sampled from a prior Gaussian Process (GP) defined by a hyperkernel. We also provide a mechanism to ensure the positive definiteness of the Gram matrix constructed using the resultant kernels. Our experimental results on GP regression and Support Vector Machine (SVM) classification tasks involving both synthetic functions and several real-world datasets show the superiority of our approach over the state-of-the-art.
Arun Kumar Anjanapura Venkatesh, Alistair Shilton, Santu Rana, Sunil Gupta 0001, Svetha Venkatesh
NeurIPS2
2021 Fairness improvement for black-box classifiers with Gaussian process
Dang Nguyen 0002, Sunil Gupta 0001, Santu Rana, Alistair Shilton, Svetha Venkatesh
Inf. Sci.4
2020 Bayesian Optimization for Categorical and Category-Specific Continuous Inputs
abstract
Many real-world functions are defined over both categorical and category-specific continuous variables and thus cannot be optimized by traditional Bayesian optimization (BO) methods. To optimize such functions, we propose a new method that formulates the problem as a multi-armed bandit problem, wherein each category corresponds to an arm with its reward distribution centered around the optimum of the objective function in continuous variables. Our goal is to identify the best arm and the maximizer of the corresponding continuous function simultaneously. Our algorithm uses a Thompson sampling scheme that helps connecting both multi-arm bandit and BO in a unified framework. We extend our method to batch BO to allow parallel optimization when multiple resources are available. We theoretically analyze our method for convergence and prove sub-linear regret bounds. We perform a variety of experiments: optimization of several benchmark functions, hyper-parameter tuning of a neural network, and automatic selection of the best machine learning model along with its optimal hyper-parameters (a.k.a automated machine learning). Comparisons with other methods demonstrate the effectiveness of our proposed method.
Dang Nguyen 0002, Sunil Gupta 0001, Santu Rana, Alistair Shilton, Svetha Venkatesh
AAAI4
2020 Accelerated Bayesian Optimisation through Weight-Prior Tuning
abstract
Bayesian optimization (BO) is a widely-used method for optimizing expensive (to evaluate) problems. At the core of most BO methods is the modeling of the objective function using a Gaussian Process (GP) whose covariance is selected from a set of standard covariance functions. From a weight-space view, this models the objective as a linear function in a feature space implied by the given covariance $K$, with an arbitrary Gaussian weight prior ${\bf w} \sim ormdist ({\bf 0},{\bf I})$. In many practical applications there is data available that has a similar (covariance) structure to the objective, but which, having different form, cannot be used directly in standard transfer learning. In this paper we show how such auxiliary data may be used to construct a GP covariance corresponding to a more appropriate weight prior for the objective function. Building on this, we show that we may accelerate BO by modeling the objective function using this (learned) weight prior, which we demonstrate on both test functions and a practical application to short-polymer fibre manufacture.
Alistair Shilton, Sunil Gupta 0001, Santu Rana, Pratibha Vellanki, Cheng Li 0003, Svetha Venkatesh, Laurence Anthony F. Park, Alessandra Sutti, David Rubin, Thomas Dorin, Alireza Vahid, Murray Height, Teo Slezak
AISTATS1
2020 Multiclass Anomaly Detector: the CS++ Support Vector Machine
abstract
A new support vector machine (SVM) variant, called CS++-SVM, is presented combining multiclass classification and anomaly detection in a single-step process to create a trained machine that can simultaneously classify test data belonging to classes represented in the training set and label as anomalous test data belonging to classes not represented in the training set. A theoretical analysis of the properties of the new method, showing how it combines properties inherited both from the conic-segmentation SVM (CS-SVM) and the $1$-class SVM (to which the method described reduces to in the case of unlabelled training data), is given. Finally, experimental results are presented to demonstrate the effectiveness of the algorithm for both simulated and real-world data.
Alistair Shilton, Sutharshan Rajasegarar, Marimuthu Palaniswami
J. Mach. Learn. Res.1
2019 Multi-objective Bayesian optimisation with preferences over objectives
abstract
We present a multi-objective Bayesian optimisation algorithm that allows the user to express preference-order constraints on the objectives of the type objective A is more important than objective B. These preferences are defined based on the stability of the obtained solutions with respect to preferred objective functions. Rather than attempting to find a representative subset of the complete Pareto front, our algorithm selects those Pareto-optimal points that satisfy these constraints. We formulate a new acquisition function based on expected improvement in dominated hypervolume (EHI) to ensure that the subset of Pareto front satisfying the constraints is thoroughly explored. The hypervolume calculation is weighted by the probability of a point satisfying the constraints from a gradient Gaussian Process model. We demonstrate our algorithm on both synthetic and real-world problems.
Majid Abdolshah, Alistair Shilton, Santu Rana, Sunil Gupta 0001, Svetha Venkatesh
NeurIPS2
2018 Exploiting Strategy-Space Diversity for Batch Bayesian Optimization
abstract
This paper proposes a novel approach to batch Bayesian optimisation using a multi-objective optimisation framework with exploitation and exploration forming two objectives. The key advantage of this approach is that it uses a suite of strategies to balance exploration and exploitation and thus can efficiently handle the optimisation of a variety of functions with small to large number of local extrema. Another advantage is that it automatically determines the batch size within a specified budget avoiding unnecessary function evaluations. Theoretical analysis shows that the regret not only reduces sub-linearly but also by an additional reduction factor determined by the batch size. We demonstrate the efficiency of our algorithm by optimising a variety of benchmark functions, performing hyperparameter tuning of support vector regression and classification, and finally heat treatment process of an Al-Sc alloy. Comparisons with recent baseline algorithms confirm the usefulness of our algorithm.
Sunil Gupta 0001, Alistair Shilton, Santu Rana, Svetha Venkatesh
AISTATS2
2018 Expected Hypervolume Improvement with Constraints
abstract
Bayesian optimisation has become a powerful framework for global optimisation of black-box functions that are expensive to evaluate and possibly noisy. In addition to expensive evaluation of objective functions, many real-world optimisation problems deal with similarly expensive black-box constraints. However, there are few studies regarding the role of constraints in multi-objective Bayesian optimisation. In this paper, we extend the Expected Hypervolume Improvement by introducing expectation of constraints satisfaction and merging them into a new acquisition function called Expected Hypervolume Improvement with Constraints (EHVIC). We analyse the performance of our algorithm by estimating the feasible region dominated by Pareto front using 4 benchmark functions. The proposed method is also evaluated on a realworld problem of Alloy Design. We demonstrate that EHVIC is an effective algorithm that provides a promising performance by comparing to a well-known related method.
Majid Abdolshah, Alistair Shilton, Santu Rana, Sunil Gupta 0001, Svetha Venkatesh
ICPR2
2018 Multi-Target Optimisation via Bayesian Optimisation and Linear Programming
Alistair Shilton, Santu Rana, Sunil Gupta 0001, Svetha Venkatesh
UAI1
2017 Regret Bounds for Transfer Learning in Bayesian Optimisation
abstract
This paper studies the regret bound of two transfer learning algorithms in Bayesian optimisation. The first algorithm models any difference between the source and target functions as a noise process. The second algorithm proposes a new way to model the difference between the source and target as a Gaussian process which is then used to adapt the source data. We show that in both cases the regret bounds are tighter than in the no transfer case. We also experimentally compare the performance of these algorithms relative to no transfer learning and demonstrate benefits of transfer learning.
Alistair Shilton, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh
AISTATS1
2017 High Dimensional Bayesian Optimization using Dropout
abstract
Scaling Bayesian optimization to high dimensions is challenging task as the global optimization of high-dimensional acquisition function can be expensive and often infeasible. Existing methods depend either on limited “active” variables or the additive form of the objective function. We propose a new method for high-dimensional Bayesian optimization, that uses a drop-out strategy to optimize only a subset of variables at each iteration. We derive theoretical bounds for the regret and show how it can inform the derivation of our algorithm. We demonstrate the efficacy of our algorithms for optimization on two benchmark functions and two real-world applications - training cascade classifiers and optimizing alloy composition.
Cheng Li 0003, Sunil Gupta 0001, Santu Rana, Vu Nguyen 0001, Svetha Venkatesh, Alistair Shilton
IJCAI6
2012 Automatic detection of different walking conditions using inertial sensor data
abstract
Identifying different walking conditions is essential in order to monitor the activities of elderly population for active living or fast recovery of a patient following a surgery or even for prognosis and diagnosis of several conditions like Parkinson's disease. This paper looks at automatically detecting three different walking conditions (walking normally with preferred walking speed (PWS), walking while carrying a glass of water, and walking blind folded) using inertial sensor data. Tri-axial accelerometers and gyroscopes were used to acquire movement data from both feet during the three gait tasks. Five healthy young subjects undertook 10 trials per condition on a GAITRite mat. Statistical properties such as the mean, standard deviation (std), skewness (skew) and kurtosis were calculated for each trial that included several gait cycles' data. Altogether 48 features were analyzed using Fuzzy Clustering Mean (FCM) algorithm to verify the separable nature of sensor data. The results show that three clusters could be found with an almost equal number of points; however the membership was not high enough to result in complete discrete clusters. Then three different Support Vector Machine (SVM) classifiers were used to examine whether the conditions could be automatically classified based on the features that were extracted from inertial sensor data. The results indicate 83-84% of accurate classification of the three gait conditions with three SVM algorithms. The study demonstrates that the inertial sensor data could be used to classify differences in walking conditions using powerful computational intelligence techniques.
Braveena K. Santhiranayagam, Daniel T. H. Lai, Cancan Jiang, Alistair Shilton, Rezaul K. Begg
IJCNN4
2012 The conic-segmentation support vector machine - a target space method for multiclass classification
abstract
In this paper we propose a new multiclass SVM, the conic-segmentation SVM (CS-SVM), based on the direct mapping of points into a multidimensional target space segmented a-priori into conic class regions defined by generalized inequalities. We show that the CS-SVM is a natural multiclass analogue of the standard binary SVM in-so-far as it shares its motivation, simplicity of form, and many of its properties such as convexity, sparsity and kernelisation. We demonstrate that prior selection of the conic region structure can give both new and interesting multiclass formulations and also well-known multiclass formulations. Finally we present experimental results on artificial and real multiclass datasets to investigate the CS-SVM's performance.
Alistair Shilton, Daniel T. H. Lai, Marimuthu Palaniswami
IJCNN1
2012 A Note on Octonionic Support Vector Regression
abstract
This note presents an analysis of the octonionic form of the division algebraic support vector regressor (SVR) first introduced by Shilton A detailed derivation of the dual form is given, and three conditions under which it is analogous to the quaternionic case are exhibited. It is shown that, in the general case of an octonionic-valued feature map, the usual "kernel trick" breaks down. The cause of this (and its interpretation) is discussed in some detail, along with potential ways of extending kernel methods to take advantage of the distinct features present in the general case. Finally, the octonionic SVR is applied to an example gait analysis problem, and its performance is compared to that of the least squares SVR, the Clifford SVR, and the multidimensional SVR.
Alistair Shilton, Daniel T. H. Lai, Braveena K. Santhiranayagam, Marimuthu Palaniswami
IEEE Trans. Syst. Man Cybern. Part B1
2010 A Division Algebraic Framework for Multidimensional Support Vector Regression
abstract
In this paper, division algebras are proposed as an elegant basis upon which to extend support vector regression (SVR) to multidimensional targets. Using this framework, a multitarget SVR called epsilon(Z)-SVR is proposed based on an epsilon-insensitive loss function that is independent of the coordinate system or basis used. This is developed to dual form in a manner that is analogous to the standard epsilon-SVR. The epsilon(H)-SVR is compared and contrasted with the least-square SVR (LS-SVR), the Clifford SVR (C-SVR), and the multidimensional SVR (M-SVR). Three practical applications are considered: namely, 1) approximation of a complex-valued function; 2) chaotic time-series prediction in 3-D; and 3) communication channel equalization. Results show that the epsilon(H)-SVR performs significantly better than the C-SVR, the LS-SVR, and the M-SVR in terms of mean-squared error, outlier sensitivity, and support vector sparsity.
Alistair Shilton, Daniel T. H. Lai, Marimuthu Palaniswami
IEEE Trans. Syst. Man Cybern. Part B1
2007 Real Value Solvent Accessibility Prediction using Adaptive Support Vector Regression
abstract
Knowledge of the secondary structure and solvent accessibility of a protein plays a vital role in prediction of fold, and eventually the tertiary structure of the protein. This paper deals with prediction of relative solvent accessibility, given only the amino-acid sequence. In this paper, we use an improved support vector regression (SVR) and new kernels for real valued prediction of solvent accessibility. In this regard, two main issues are addressed. First we address the problem of e selection, which we found to be somewhat problematic in our earlier work (e is a parameter with significant influence on noise insensitivity and generalization of SVRs). In particular, rather than employ the standard trial and error based approach, we used an improved tube shrinking method to find e. Secondly, a novel kernel combining solvation model, electrostatic charge model and evolutionary information in the form of position specific scoring matrix (PSSM) is given. A new dataset of 472 proteins with less than 20% sequence identity is curated and used to evaluate the result. To make a more objective comparison with earlier methods, we use a standard dataset and show that the proposed scheme is better than the ones normally used in literature. We also report a lowest mean absolute error (MAE) so far of 0.12 on the standard dataset.
Jayavardhana Gubbi, Alistair Shilton, Marimuthu Palaniswami, Michael Parker
CIBCB2
2007 Iterative Fuzzy Support Vector Machine Classification
abstract
Fuzzy support vector machine (FSVM) classifiers are a class of nonlinear binary classifiers which extend Vapnik's support vector machine (SVM) formulation. In the absence of additional information, fuzzy membership values are usually selected based on the distribution of training vectors, where a number of assumptions are made about the underlying shape of this distribution. In this paper we present an alternative method of generating membership values which we call iterative FSVM (I-FSVM). Our method generates membership values iteratively based on the positions of training vectors relative to the SVM decision surface itself. We show that our algorithm is capable of generating results equivalent to an SVM with a modified (non distance based) penalty (risk) function. Experiments have been carried out on three real world binary classification problems taken from the UCI repository, namely the spambase dataset and the adult (census) dataset.
Alistair Shilton, Daniel T. H. Lai
FUZZ-IEEE1
2007 Quaternionic and complex-valued Support Vector Regression for Equalization and Function Approximation
abstract
Support vector regressors (SVRs) are a class of nonlinear regressor inspired by Vapnik's support vector (SV) method for pattern classification. The standard SVR has been successfully applied to real number regression problems such as financial prediction and weather forecasting. However in some applications the domain of the function to be estimated may be more naturally and efficiently expressed using complex numbers (eg. communications channels) or quaternions (eg. 3-dimensional geometrical problems). Since SVRs have previously been proven to be efficient and accurate regressors, the extension of this method to complex numbers and quaternions is of great interest. In the present paper the standard SVR method is extended to cover regression in complex numbers and quaternions. Our method differs from existing approaches in-so-far as the cost function applied in the output space is rotationally invariant, which is important as in most cases it is the magnitude of the error in the output which is important, not the angle. We demonstrate the practical usefulness of this new formulation by considering the problem of communications channel equalization.
Alistair Shilton, Daniel T. H. Lai
IJCNN1
2005 A convergence rate estimate for the SVM decomposition method
abstract
The training of support vector machines using the decomposition method has one drawback; namely the selection of working sets such that convergence is as fast as possible. It has been shown by Lin that the rate is linear in the worse case under the assumption that all bounded support vectors have been determined. The analysis was done based on the change in the objective function and under a SVMlight selection rule. However, the rate estimate given is independent of time and hence gives little indication as to how the linear convergence speed varies during the iteration. In this initial analysis, we provide a treatment of the convergence from a gradient contraction perspective. We propose a necessary and sufficient condition which when satisfied provides strict linear convergence of the algorithm. The condition can also be interpreted as a basic requirement for a sequence of working sets in order to achieve such a convergence rate. Based on this condition, a time dependent rate estimate is then further derived. This estimate is shown to monotonically approach unity from below.
Daniel T. H. Lai, Alistair Shilton, Nallasamy Mani, Marimuthu Palaniswami
IJCNN2
2005 Incremental training of support vector machines
abstract
We propose a new algorithm for the incremental training of support vector machines (SVMs) that is suitable for problems of sequentially arriving data and fast constraint parameter variation. Our method involves using a "warm-start" algorithm for the training of SVMs, which allows us to take advantage of the natural incremental properties of the standard active set approach to linearly constrained optimization problems. Incremental training involves quickly retraining a support vector machine after adding a small number of additional training vectors to the training set of an existing (trained) support vector machine. Similarly, the problem of fast constraint parameter variation involves quickly retraining an existing support vector machine using the same training set but different constraint parameters. In both cases, we demonstrate the computational superiority of incremental training over the usual batch retraining method.
Alistair Shilton, Marimuthu Palaniswami, Daniel Ralph, Ah Chung Tsoi
IEEE Trans. Neural Networks1