Ivan Tyukin

dblp:87/1262 · also Ivan Yu. Tyukin · DBLP profile ↗
← Back
35ranked-venue papers
12as first author
20since 2021 · last 2026
0000-0002-7359-7966ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 9 first-author · 15 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Deeper insights into the learning performance of Stochastic Configuration Networks
Xiufeng Yan, Dianhui Wang 0001, Ivan Tyukin
Knowl. Based Syst.3
2026 Theoretical Advances on Stochastic Configuration Networks
abstract
This article advances the theoretical foundations of stochastic configuration networks (SCNs) by rigorously analyzing their convergence properties, approximation guarantees, and the limitations of nonadaptive randomized methods. We introduce a principled objective function that aligns incremental training with orthogonal projection, ensuring maximal residual reduction at each iteration without recomputing output weights. Under this formulation, we derive a novel necessary and sufficient condition for strong convergence in Hilbert spaces and establish sufficient conditions for uniform geometric convergence, offering the first theoretical justification of the SCN residual constraint. To assess the feasibility of unguided random initialization, we present a probabilistic analysis showing that even small support shifts markedly reduce the likelihood of sampling effective nodes in high-dimensional settings, thereby highlighting the necessity of adaptive refinement in the sampling distribution. Motivated by these insights, we propose greedy SCNs (GSCNs) and two optimized variants-Newton-Raphson GSCN (NR-GSCN) and particle swarm optimization GSCN (PSO-GSCN)-that incorporate Newton-Raphson refinement and particle swarm-based exploration to improve node selection. Empirical results on synthetic and real-world datasets demonstrate that the proposed methods achieve faster convergence, better approximation accuracy, and more compact architectures compared to existing SCN training schemes. Collectively, this work establishes a rigorous theoretical and algorithmic framework for SCNs, laying out a principled foundation for subsequent developments in the field of randomized neural network (NN) training.
Xiufeng Yan, Dianhui Wang 0001, Ivan Tyukin
IEEE Trans. Neural Networks Learn. Syst.3
2025 Staining and Locking Computer Vision Models Without Retraining
Oliver J. Sutton, George Leete, Alexander N. Gorban, Ivan Tyukin
ICCV5
2025 Stability-aware Neuromorphic Computation
abstract
In this work, we consider the problem of stability in large and potentially heterogeneous neuromorphic circuits with both symbolic and non-symbolic computations. The type of stability we are considering here captures the system’s resilience to changes in the model’s weights. This differs from a somewhat more typical setting whereby the perturbations are limited to changes in the data. Nevertheless, the scenario considered in our work allows accounting for various relevant and important effects such as weights’ disturbances arising due to imperfect implementations of neuromorphic circuits (e.g. in memristors) or due to weights’ quantization emerging as a result of post-training memory optimization. We show, through both theoretical analysis and numerical exploration, that negative consequences of weights’ imprecision and/or noise can be alleviated by including auxiliary low-cost clipping and scaling layers into the computational graph in the model. The function of these layers is to reduce the contrast of signals flowing through the neuromorphic circuit. Remarkably, we show that introducing these simple architectural modifications enhances robustness to weights’ perturbations and enables the creation of stability certificates whose computational complexity does not depend on the dimension of the domain of the inputs to the model. Theoretical results are illustrated with numerical experiments.
Egor E. Nuzhin, Ivan Tyukin, Georgii Ovchinnikov, Nikolai V. Brilliantov
IJCNN2
2025 Situation-Based Neuromorphic Memory in Spiking Neuron-Astrocyte Network
abstract
Mammalian brains operate in very special surroundings: to survive they have to react quickly and effectively to the pool of stimuli patterns previously recognized as danger. Many learning tasks often encountered by living organisms involve a specific set-up centered around a relatively small set of patterns presented in a particular environment. For example, at a party, people recognize friends immediately, without deep analysis, just by seeing a fragment of their clothes. This set-up with reduced "ontology" is referred to as a "situation." Situations are usually local in space and time. In this work, we propose that neuron-astrocyte networks provide a network topology that is effectively adapted to accommodate situation-based memory. In order to illustrate this, we numerically simulate and analyze a well-established model of a neuron-astrocyte network, which is subjected to stimuli conforming to the situation-driven environment. Three pools of stimuli patterns are considered: external patterns, patterns from the situation associative pool regularly presented to the network and learned by the network, and patterns already learned and remembered by astrocytes. Patterns from the external world are added to and removed from the associative pool. Then, we show that astrocytes are structurally necessary for an effective function in such a learning and testing set-up. To demonstrate this we present a novel neuromorphic computational model for short-term memory implemented by a two-net spiking neural-astrocytic network. Our results show that such a system tested on synthesized data with selective astrocyte-induced modulation of neuronal activity provides an enhancement of retrieval quality in comparison to standard spiking neural networks trained via Hebbian plasticity only. We argue that the proposed set-up may offer a new way to analyze, model, and understand neuromorphic artificial intelligence systems.
Susanna Yu Gordleeva, Yuliya Tsybina, Mikhail Krivonosov, Ivan Tyukin, Victor B. Kazantsev, Alexey Zaikin, Alexander N. Gorban
IEEE Trans. Neural Networks Learn. Syst.4
2024 Weakly Supervised Learners for Correction of AI Errors with Provable Performance Guarantees
abstract
We present a new methodology for handling errors of Artificial Intelligence (AI) by introducing weakly supervised AI error correctors with a priori performance guarantees. These AI correctors are auxiliary maps whose role is to moderate the decisions of some previously constructed underlying classifier by either approving or rejecting its decisions. The rejection of a decision can be used as a signal to suggest abstaining from making a decision. A key technical focus of the work is in providing performance guarantees for these new AI correctors through bounds on the probabilities of incorrect decisions. These bounds are distribution agnostic and do not rely on assumptions on the data dimension. Our empirical example illustrates how the framework can be applied to improve the performance of an image classifier in a challenging real-world task where training data are scarce.
Ivan Tyukin, Tatiana Tyukina, Daniel van Helden, Zedong Zheng, Eugenij Moiseevich Mirkes, Oliver J. Sutton, Alexander N. Gorban, Penelope M. Allison
IJCNN1
2024 Stealth edits to large language models
abstract
We reveal the theoretical foundations of techniques for editing large language models, and present new methods which can do so without requiring retraining. Our theoretical insights show that a single metric (a measure of the intrinsic dimension of the model's features) can be used to assess a model's editability and reveals its previously unrecognised susceptibility to malicious *stealth attacks*. This metric is fundamental to predicting the success of a variety of editing approaches, and reveals new bridges between disparate families of editing methods. We collectively refer to these as *stealth editing* methods, because they directly update a model's weights to specify its response to specific known hallucinating prompts without affecting other model behaviour. By carefully applying our theoretical insights, we are able to introduce a new *jet-pack* network block which is optimised for highly selective model editing, uses only standard network operations, and can be inserted into existing networks. We also reveal the vulnerability of language models to stealth attacks: a small change to a model's weights which fixes its response to a single attacker-chosen prompt. Stealth attacks are computationally simple, do not require access to or knowledge of the model's training data, and therefore represent a potent yet previously unrecognised threat to redistributed foundation models. Extensive experimental results illustrate and support our methods and their theoretical underpinnings. Demos and source code are available at https://github.com/qinghua-zhou/stealth-edits.
Oliver J. Sutton, Wei Wang 0357, Desmond J. Higham, Alexander N. Gorban, Alexander Bastounis, Ivan Tyukin
NeurIPS7
2024 Coping with AI errors with provable guarantees
abstract
AI errors pose a significant challenge, hindering real-world applications. This work introduces a novel approach to cope with AI errors using weakly supervised error correctors that guarantee a specific level of error reduction. Our correctors have low computational cost and can be used to decide whether to abstain from making an unsafe classification. We provide new upper and lower bounds on the probability of errors in the corrected system. In contrast to existing works, these bounds are distribution agnostic, non-asymptotic, and can be efficiently computed just using the corrector training data. They also can be used in settings with concept drifts when the observed frequencies of separate classes vary. The correctors can easily be updated, removed, or replaced in response to changes in distributions within each class without retraining the underlying classifier. The application of the approach is illustrated with two relevant challenging tasks: (i) an image classification problem with scarce training data, and (ii) moderating responses of large language models without retraining or otherwise fine-tuning.
Ivan Tyukin, Tatiana Tyukina, Daniël P. van Helden, Zedong Zheng, Eugenij Moiseevich Mirkes, Oliver J. Sutton, Alexander N. Gorban, Penelope M. Allison
Inf. Sci.1
2024 How adversarial attacks can disrupt seemingly stable accurate classifiers
abstract
Adversarial attacks dramatically change the output of an otherwise accurate learning system using a seemingly inconsequential modification to a piece of input data. Paradoxically, empirical evidence indicates that even systems which are robust to large random perturbations of the input data remain susceptible to small, easily constructed, adversarial perturbations of their inputs. Here, we show that this may be seen as a fundamental feature of classifiers working with high dimensional input data. We introduce a simple generic and generalisable framework for which key behaviours observed in practical systems arise with high probability-notably the simultaneous susceptibility of the (otherwise accurate) model to easily constructed adversarial attacks, and robustness to random perturbations of the input data. We confirm that the same phenomena are directly observed in practical neural networks trained on standard image classification problems, where even large additive random noise fails to trigger the adversarial instability of the network. A surprising takeaway is that even small margins separating a classifier's decision surface from training and testing data can hide adversarial susceptibility from being detected using randomly sampled perturbations. Counter-intuitively, using additive noise during training or testing is therefore inefficient for eradicating or detecting adversarial examples, and more demanding adversarial training is required.
Oliver J. Sutton, Ivan Tyukin, Alexander N. Gorban, Alexander Bastounis, Desmond J. Higham
Neural Networks3
2024 Accelerating Finite State Machine-Based Testing Using Reinforcement Learning
abstract
Testing is a crucial phase in the development of complex systems, and this has led to interest in automated test generation techniques based on state-based models. Many approaches use models that are types of finite state machine (FSM). Corresponding test generation algorithms typically require that certain test components, such as reset sequences (RSs) and preset distinguishing sequences (PDSs), have been produced for the FSM specification. Unfortunately, the generation of RSs and PDSs is computationally expensive, and this affects the scalability of such FSM-based test generation algorithms. This paper addresses this scalability problem by introducing a reinforcement learning framework: the$\mathcal{Q}$-Graph framework for MBT. We show how this framework can be used in the generation of RSs and PDSs and consider both (potentially partial) timed and untimed models. The proposed approach was evaluated using three types of FSMs: randomly generated FSMs, FSMs from a benchmark, and an FSM of an Engine Status Manager for a printer. In experiments, the proposed approach was much faster and used much less memory than the state-of-the-art methods in computing PDSs and RSs.
Uraz Cengiz Türker, Robert M. Hierons, Khaled El-Fakih, Mohammad Reza Mousavi 0001, Ivan Tyukin
IEEE Trans. Software Eng.5
2023 The Boundaries of Verifiable Accuracy, Robustness, and Generalisation in Deep Learning
Alexander Bastounis, Alexander N. Gorban, Anders C. Hansen, Desmond J. Higham, Danil V. Prokhorov, Oliver J. Sutton, Ivan Tyukin
ICANN (1)7
2023 Relative Intrinsic Dimensionality Is Intrinsic to Learning
Oliver J. Sutton, Alexander N. Gorban, Ivan Tyukin
ICANN (1)4
2023 Agile gesture recognition for capacitive sensing devices: adapting on-the-job
abstract
Automated hand gesture recognition has been a focus of the AI community for decades. Traditionally, work in this domain revolved largely around scenarios assuming the availability of the flow of images of the operator's/user's hands. This has partly been due to the prevalence of camera-based devices and the wide availability of image data. However, there is growing demand for gesture recognition technology that can be implemented on low-power devices using limited sensor data instead of high-dimensional inputs like hand images. In this work, we demonstrate a hand gesture recognition system and method that uses signals from capacitive sensors embedded into the etee hand controller. The controller generates real-time signals from each of the wearer's five fingers. We use a machine learning technique to analyse the time-series signals and identify three features that can represent 5 fingers within 500 ms. The analysis is composed of a two-stage training strategy, including dimension reduction through principal component analysis and classification with K-nearest neighbour. Remarkably, we found that this combination showed a level of performance which was comparable to more advanced methods such as supervised variational autoencoder. The base system can also be equipped with the capability to learn from occasional errors by providing it with an additional adaptive error correction mechanism. The results showed that the error corrector improve the classification performance in the base system without compromising its performance. The system requires no more than 1 ms of computing time per input sample, and is smaller than deep neural networks, demonstrating the feasibility of agile gesture recognition systems based on this technology.
Liucheng Guo, Valeri A. Makarov, Alexander N. Gorban, Eugenij Moiseevich Mirkes, Ivan Tyukin
IJCNN7
2023 Neuromorphic tuning of feature spaces to overcome the challenge of low-sample high-dimensional data
abstract
For learning algorithms, accessing large volumes of annotated data is highly desirable but not always available, especially in real-world scenarios. Accordingly, learning in the high-dimensional and low-sample size (HDLS) domain is recognised as one of the core challenges for modern AI systems. In this work, we consider a particular but very practical scenario in the HDLS domain where the number of training samples is not limited to mere few observations but yet it is not large enough to reliably build models with high degrees of expressivity. To address the problem, we present a new neuromorphic algorithm capable of fine-tuning existing feature spaces via learning relevant associations in high dimensional data with high probability. The algorithm is based on the idea of Concept Cells [1] and mimics properties attributed to memory and learning inherent to live neural systems. We demonstrate, through numerous numerical experiments, that the algorithm can “fine-tune” and “adapt” the feature space of pre-trained neural networks for better performance on new tasks in the HDLS domain. In addition, we study the impact of this “tuning” on quasi-orthogonal measures, which correlates with classification and calibration metrics.
Oliver J. Sutton, Yudong Zhang 0001, Alexander N. Gorban, Valeri A. Makarov, Ivan Tyukin
IJCNN6
2023 GeoAI in urban analytics
abstract
We are writing this editorial piece at the peak of the current Artificial Intelligence (AI) ‘spring’ as generative models quickly cross the bridge from the confines of academic and industry labs into our everyday lives. During times like this, one might be excused from forgetting how old the application of AI approaches in geography is. Geographers have been here before. About forty years ago, Smith (1984) wrote: AI techniques, if properly applied, should also allow researchers to spend a greater proportion of their time on creative thinking and less on technical drudgery. As with any set of tools, the techniques of AI cannot replace a hard-earned understanding of some phenomenon and will almost certainly be overvalued and misused by some practitioners. [Nevertheless], if used with care, the techniques of AI will prove of great benefit to such an applied, problem solving discipline as geography. (p. 157). It is in the subsequent issue of the same journal that we find Nystuen’s (1984) comment, suggesting that ‘[b]enefit to geography from such an alliance [with AI] is questionable considering that our own directions are murky enough’ (p. 358). Smith, in Nystuen’s view, should be ‘a little more critical in his appraisal of the scope of possible applications’ (Nystuen 1984, p. 359). The debate between Smith and Nystuen unfolded during the ‘AI spring’ of the 1980s, but the same hopes and concerns around a data-driven (rather than theory-driven) geography echo through the discipline’s history. From Openshaw’s (1992, 1998) work on AI tools for spatial modelling and analysis to Miller and Goodchild (2015) discussion of data-driven geography in the wake of big data, to the emergence of GeoAI (Janowicz et al. 2022) – primarily used as a shorthand for geospatial AI, encompassing the efforts towards creating spatially-explicit models in the era of deep learning. As detailed by Miller and Goodchild (2015), these ‘waves’ are evolutionary rather than revolutionary. These approaches are founded in abductive reasoning and foster the same discussions, tensions and shifts between nomothetic (law-seeking) and idiographic (description-seeking) knowledge that can be traced back to the very origins of the discipline. Traditional AI approaches have long been part of Geographical Information Science (GIScience), including research both on unsupervised learning approaches to geographical data mining (e.g. geodemographic classification and dimensionality reduction, see e.g. Miller and Han 2009) and supervised methods of inference (e.g. spatial autocorrelation and geographically weighted regression, see e.g. O'Sullivan and Unwin 2003). At the same time, each ‘wave’ is unique, and the current AI spring has again brought new challenges and opportunities. This special issue stemmed from a session organised at the Annual International Conference of the Royal Geographical Society (with IBG) in August 2021, which aimed to explore those challenges and opportunities with a particular focus on deep learning and human geography. The previous decade had seen unprecedented advances in image processing following the seminal paper on Alexnet (Krizhevsky et al. 2012), the emergence of large language models (LLMs) based on the transformer architecture (Vaswani et al. 2017), as well as the development of graph neural networks (Bruna et al. 2013, Hamilton et al. 2017). While those approaches to deep learning have found wide use in many aspects of GIScience and remote sensing (e.g. computer vision in geospatial applications), their application to human geography has been slower (Harris et al. 2017). Complementing the special issue introduced by Janowicz et al. (2020) on ‘Artificial intelligence techniques for geographical knowledge discovery’, this special issue focuses on GeoAI as a broader geographical AI and its applications in urban analytics (Liu and Biljecki 2022). The next section introduces the articles included in this special issue, while the final section contextualises the main themes emerging from those articles in the current, fast-paced landscape shaken by the emergence of foundation models (Bommasani et al. 2021).
Stef De Sabbata, Andrea Ballatore, Harvey J. Miller, Renée Sieber, Ivan Tyukin, Godwin Yeboah
Int. J. Geogr. Inf. Sci.5
2022 Quasi-orthogonality and intrinsic dimensions as measures of learning and generalisation
abstract
Finding the best architectures for learning machines, such as deep neural networks, is a well-known technical and theoretical challenge. Recent work by Mellor et al [1] showed that there may exist correlations between the accuracies of trained networks and the values of some easily computable measures defined on randomly initialised networks which may enable the search of tens of thousands of neural architectures without training. Mellor et al [1] used the Hamming distance evaluated over all the ReLU neurons as such a measure. Motivated by these findings, in our work, we ask the question of the existence of other and perhaps more principled measures which could be used as determinants of the potential success of a given neural architecture. In particular, we examine if the dimensionality and quasi-orthogonality of neural networks' feature space could be correlated with the network's performance after training. We showed, using the setup as in Mellor et al [1], that dimensionality and quasi-orthogonality may jointly serve as networks' performance discriminants. In addition to offering new opportunities to accelerate neural architecture search, our findings suggest important relationships between the networks' final performance and properties of their randomly initialised feature spaces: data dimension and quasi-orthogonality.
Alexander N. Gorban, Eugenij Moiseevich Mirkes, Jonathan Bac, Andrei Yu. Zinovyev, Ivan Tyukin
IJCNN6
2021 Demystification of Few-shot and One-shot Learning
abstract
Few-shot and one-shot learning have been the subject of active and intensive research in recent years, with mounting evidence pointing to successful implementation and exploitation of few-shot learning algorithms in practice. Classical statistical learning theories do not fully explain why few- or one-shot learning is at all possible since traditional generalisation bounds normally require large training and testing samples to be meaningful. This sharply contrasts with numerous examples of successful one- and few-shot learning systems and applications. In this work we present mathematical foundations for a theory of one-shot and few-shot learning and reveal conditions specifying when such learning schemes are likely to succeed. Our theory is based on intrinsic properties of high-dimensional spaces. We show that if the ambient or latent decision space of a learning machine is sufficiently high-dimensional than a large class of objects in this space can indeed be easily learned from few examples provided that certain data non-concentration conditions are met. In this work we present mathematical foundations for a theory of one-shot and few-shot learning and reveal conditions specifying when such learning schemes are likely to succeed. Our theory is based on intrinsic properties of high-dimensional spaces. We show that if the ambient or latent decision space of a learning machine is sufficiently high-dimensional than a large class of objects in this space can indeed be easily learned from few examples provided that certain data non-concentration conditions are met.
Ivan Tyukin, Alexander N. Gorban, Muhammad H. Alkhudaydi
IJCNN1
2021 Efficient state synchronisation in model-based testing through reinforcement learning
abstract
Model-based testing is a structured method to test complex systems. Scaling up model-based testing to large systems requires improving the efficiency of various steps involved in testcase generation and more importantly, in test-execution. One of the most costly steps of model-based testing is to bring the system to a known state, best achieved through synchronising sequences. A synchronising sequence is an input sequence that brings a given system to a predetermined state regardless of system’s initial state. Depending on the structure, the system might be complete, i.e., all inputs are applicable at every state of the system. However, some systems are partial and in this case not all inputs are usable at every state. Derivation of synchronising sequences from complete or partial systems is a challenging task. In this paper, we introduce a novel Q-learning algorithm that can derive synchronising sequences from systems with complete or partial structures. The proposed algorithm is faster and can process larger systems than the fastest sequential algorithm that derives synchronising sequences from complete systems. Moreover, the proposed method is also faster and can process larger systems than the most recent massively parallel algorithm that derives synchronising sequences from partial systems. Furthermore, the proposed algorithm generates shorter synchronising sequences.
Uraz Cengiz Türker, Robert M. Hierons, Mohammad Reza Mousavi 0001, Ivan Tyukin
ASE4
2021 Blessing of dimensionality at the edge and geometry of few-shot learning
abstract
In this paper we present theory and algorithms enabling classes of Artificial Intelligence (AI) systems to continuously and incrementally improve with a priori quantifiable guarantees – or more specifically remove classification errors – over time. This is distinct from state-of-the-art machine learning, AI, and software approaches. The theory enables building few-shot AI correction algorithms and provides conditions justifying their successful application. Another feature of this approach is that, in the supervised setting, the computational complexity of training is linear in the number of training samples. At the time of classification, the computational complexity is bounded by few inner product calculations. Moreover, the implementation is shown to be very scalable. This makes it viable for deployment in applications where computational power and memory are limited, such as embedded environments. It enables the possibility for fast on-line optimisation using improved training samples. The approach is based on the concentration of measure effects and stochastic separation theorems and is illustrated with an example on the identification faulty processes in Computer Numerical Control (CNC) milling and with a case study on adaptive removal of false positives in an industrial video surveillance and analytics system.
Ivan Tyukin, Alexander N. Gorban, Alistair A. McEwan, Sepehr Meshkinfamfard
Inf. Sci.1
2021 General stochastic separation theorems with optimal bounds
abstract
Phenomenon of stochastic separability was revealed and used in machine learning to correct errors of Artificial Intelligence (AI) systems and analyze AI instabilities. In high-dimensional datasets under broad assumptions each point can be separated from the rest of the set by simple and robust Fisher's discriminant (is Fisher separable). Errors or clusters of errors can be separated from the rest of the data. The ability to correct an AI system also opens up the possibility of an attack on it, and the high dimensionality induces vulnerabilities caused by the same stochastic separability that holds the keys to understanding the fundamentals of robustness and adaptivity in high-dimensional data-driven AI. To manage errors and analyze vulnerabilities, the stochastic separation theorems should evaluate the probability that the dataset will be Fisher separable in given dimensionality and for a given class of distributions. Explicit and optimal estimates of these separation probabilities are required, and this problem is solved in the present work. The general stochastic separation theorems with optimal probability estimates are obtained for important classes of distributions: log-concave distribution, their convex combinations and product distributions. The standard i.i.d. assumption was significantly relaxed. These theorems and estimates can be used both for correction of high-dimensional data driven AI systems and for analysis of their vulnerabilities. The third area of application is the emergence of memories in ensembles of neurons, the phenomena of grandmother's cells and sparse coding in the brain, and explanation of unexpected effectiveness of small neural ensembles in high-dimensional brain.
Bogdan Grechuk, Alexander N. Gorban, Ivan Tyukin
Neural Networks3
2020 Neural Networks for the Retrieval of Methane from the Sentinel-5 Precursor Satellite
abstract
Methane is the second most common anthropocentric greenhouse gas, it is therefore important to accurately and quickly get concentrations globally from satellite readings. Currently, data acquired by TROPOMI on board the Sentinel5 Precursor satellite is transformed into Methane mixing ratio via a physics-based retrieval algorithm. This paper presents an alternative to the slow and complex algorithm: a neural network. A number of experiments are performed using a range of training sets and different network architectures. These experiments, not only help to chose the final network architecture but also allow discussion about the regional and seasonal correlation of methane mixing ratio in the data. These experiments conclude that there is some seasonal and regional correlation in the data which must be taken into account during training. Finally, a single hidden layer network, with 256 nodes in the hidden layer, is trained over 2500 epochs giving a respectable, but improvable, 1.36% error. This error is a slight improvement on the original experiment and the results of this network are analysed with areas of improvement suggested. Particularly, improvement of the extreme data points, which are highlighted as being the worst predicted by the network. This work presents the ground work for using neural networks to replace lengthy retrieval systems, not just for methane but other gases commonly investigated in a similar way.
Rose Fenwick, Hartmut Boesch, Ivan Tyukin
IJCNN3
2020 On Adversarial Examples and Stealth Attacks in Artificial Intelligence Systems
abstract
In this work we present a formal theoretical framework for assessing and analyzing two classes of malevolent action towards generic Artificial Intelligence (AI) systems. Our results apply to general multi-class classifiers that map from an input space into a decision space, including artificial neural networks used in deep learning applications. Two classes of attacks are considered. The first class involves adversarial examples and concerns the introduction of small perturbations of the input data that cause misclassification. The second class, introduced here for the first time and named stealth attacks, involves small perturbations to the AI system itself. Here the perturbed system produces whatever output is desired by the attacker on a specific small data set, perhaps even a single input, but performs as normal on a validation set (which is unknown to the attacker).We show that in both cases, i.e., in the case of an attack based on adversarial examples and in the case of a stealth attack, the dimensionality of the AI's decision-making space is a major contributor to the AI's susceptibility. For attacks based on adversarial examples, a second crucial parameter is the absence of local concentrations in the data probability distribution, a property known as Smeared Absolute Continuity. According to our findings, robustness to adversarial examples requires either (a) the data distributions in the AI's feature space to have concentrated probability density functions or (b) the dimensionality of the AI's decision variables to be sufficiently small. We also show how to construct stealth attacks on high-dimensional AI systems that are hard to spot unless the validation set is made exponentially large.
Ivan Tyukin, Desmond J. Higham, Alexander N. Gorban
IJCNN1
2020 Myocardial Infarction Detection and Quantification Based on a Convolution Neural Network with Online Error Correction Capabilities
abstract
Myocardial infarction (MI), more commonly known as heart attack, occurs when the blood flow to the heart decreases or stops. Over 100,000 people each year in the UK suffer from an MI according to the report by British Heart Foundation. Following an MI, there is irreversible heart muscle damage that will become scar. The amount of scar following larger heart attacks, ST segment elevation myocardial infarction, drives enlargement of the heart and is associated with worse prognosis (increased risk of death and subsequent heart failure). Cardiac Magnetic Resonance Imaging (MRI) late gadolinium enhancement (LGE) has become the "gold standard" for the visualization of MI. However, to date, no "gold standard" fully automated methods exist for the quantification of MI from MRI.In this work, we propose an approach to construct such methods using Artificial Intelligence (AI) and Machine Learning (ML) technologies, in particular, Convolutional Neural Networks (CNN). Uncertainties, variability, and a possibility of bias inherent to any data imply that data-driven systems which are intended for use in clinical research and practice must be capable of learning from mistakes on-the-job. Here we develop and test a first deep learning CNN system with error correction capabilities (CNNEC) for the detection and quantification of MI. The system could be viewed as a proof-of-principle for the technology.
Shuihua Wang, Gerry P. McCann, Ivan Tyukin
IJCNN3
2019 Kernel Stochastic Separation Theorems and Separability Characterizations of Kernel Classifiers
abstract
In this work we provide generalizations and extensions of stochastic separation theorems to kernel classifiers. A general separability result for two random sets is also established. We show that despite feature maps corresponding to a given kernel function may be infinite-dimensional, kernel separability characterizations can be expressed in terms of finite-dimensional volume integrals. These integrals allow to determine and quantify separability properties of an arbitrary kernel function. The theory is illustrated with numerical examples.
Ivan Tyukin, Alexander N. Gorban, Bogdan Grechuk, Stephen Green 0001
IJCNN1
2019 One-trial correction of legacy AI systems and stochastic separation theorems
Alexander N. Gorban, Richard Burton, Ilya V. Romanenko, Ivan Tyukin
Inf. Sci.4
2019 Fast construction of correcting ensembles for legacy Artificial Intelligence systems: Algorithms and a case study
Ivan Tyukin, Alexander N. Gorban, Stephen Green 0001, Danil V. Prokhorov
Inf. Sci.1
2018 Efficiency of Shallow Cascades for Improving Deep Learning AI Systems
abstract
This paper presents a technology for simple and non-iterative improvements of Multilayer and Deep Learning neural networks and Artificial Intelligence (AI) systems. The improvements are, in essence, shallow networks constructed on top of the existing Deep Learning architecture. Theoretical foundation of the technology is based on Stochastic Separation Theorems and the ideas of measure concentration. We show that, subject to mild technical assumptions on statistical properties of internal signals in Deep Learning AI, with probability close to one the technology enables instantaneous “learning away” of spurious and systematic errors. The method is illustrated with numerical examples.
Ivan Tyukin, Alexander N. Gorban, Danil V. Prokhorov, Stephen Green 0001
IJCNN1
2018 Correction of AI systems by linear discriminants: Probabilistic foundations
Alexander N. Gorban, A. Golubkov, Bogdan Grechuk, Eugenij Moiseevich Mirkes, Ivan Tyukin
Inf. Sci.5
2017 Stochastic separation theorems
Alexander N. Gorban, Ivan Tyukin
Neural Networks2
2016 Approximation with random bases: Pro et Contra
Alexander N. Gorban, Ivan Tyukin, Danil V. Prokhorov, Konstantin I. Sofeikov
Inf. Sci.2
2014 Learning optimization for decision tree classification of non-categorical data with information gain impurity criterion
abstract
We consider the problem of construction of decision trees in cases when data is non-categorical and is inherently high-dimensional. Using conventional tree growing algorithms that either rely on univariate splits or employ direct search methods for determining multivariate splitting conditions is computationally prohibitive. On the other hand application of standard optimization methods for finding locally optimal splitting conditions is obstructed by abundance of local minima and discontinuities of classical goodness functions such as e.g. information gain or Gini impurity. In order to avoid this limitation a method to generate smoothed replacement for measuring impurity of splits is proposed. This enables to use vast number of efficient optimization techniques for finding locally optimal splits and, at the same time, decreases the number of local minima. The approach is illustrated with examples.
Konstantin I. Sofeikov, Ivan Tyukin, Alexander N. Gorban, Eugenij Moiseevich Mirkes, Danil V. Prokhorov, Ilya V. Romanenko
IJCNN2
2010 State and Parameter Estimation for Canonic Models of Neural oscillators
abstract
We consider the problem of how to recover the state and parameter values of typical model neurons, such as Hindmarsh-Rose, FitzHugh-Nagumo, Morris-Lecar, from in-vitro measurements of membrane potentials. In control theory, in terms of observer design, model neurons qualify as locally observable. However, unlike most models traditionally addressed in control theory, no parameter-independent diffeomorphism exists, such that the original model equations can be transformed into adaptive canonic observer form. For a large class of model neurons, however, state and parameter reconstruction is possible nevertheless. We propose a method which, subject to mild conditions on the richness of the measured signal, allows model parameters and state variables to be reconstructed up to an equivalence class.
Ivan Tyukin, Erik Steur, Henk Nijmeijer, David Fairhurst, Inseon Song, Alexey V. Semyanov, Cees van Leeuwen
Int. J. Neural Syst.1
2009 Invariant template matching in systems with spatiotemporal coding: A matter of instability
Ivan Tyukin, Tatiana Tyukina, Cees van Leeuwen
Neural Networks1
2008 Adaptive Classification of Temporal Signals in Fixed-Weight Recurrent Neural Networks: An Existence Proof
abstract
Recurrent neural networks with fixed weights have been shown in practice to successfully classify adaptively signals that vary as a function of time in the presence of additive noise and parametric perturbations. We address the question: Can this ability be explained theoretically? We provide a mathematical proof that these networks have this ability even when parametric perturbations enter the signals nonlinearly. The restrictions that we impose on the signals to be classified are that they satisfy an assumption of nondegeneracy and that noise amplitude is sufficiently small. Further, we demonstrate that the recurrent neural networks may not only classify uncertain signals adaptively but also can recover the values of uncertain parameters of the signals, up to their equivalence classes.
Ivan Tyukin, Danil V. Prokhorov, Cees van Leeuwen
Neural Comput.1
2003 Parameter Estimation of Sigmoid Superpositions: Dynamical System Approach
abstract
Superposition of sigmoid function over a finite time interval is shown to be equivalent to the linear combination of the solutions of a linearly parameterized system of logistic differential equations. Due to the linearity with respect to the parameters of the system, it is possible to design an effective procedure for parameter adjustment. Stability properties of this procedure are analyzed.
Ivan Tyukin, Cees van Leeuwen, Danil V. Prokhorov
Neural Comput.1