VLDB 2026 Research / reviewers in the wild / expert
Alexander N. Gorban
dblp:45/3050
· DBLP profile ↗
43ranked-venue papers
11as first author
18since 2021 · last 2025
0000-0001-6224-1430ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 8 first-author · 15 since 2021Databases, data management, data science and information retrieval · 10 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Staining and Locking Computer Vision Models Without Retraining
Oliver J. Sutton, George Leete, Alexander N. Gorban, Ivan Tyukin |
ICCV | 4 |
| 2025 | Situation-Based Neuromorphic Memory in Spiking Neuron-Astrocyte NetworkabstractMammalian brains operate in very special surroundings: to survive they have to react quickly and effectively to the pool of stimuli patterns previously recognized as danger. Many learning tasks often encountered by living organisms involve a specific set-up centered around a relatively small set of patterns presented in a particular environment. For example, at a party, people recognize friends immediately, without deep analysis, just by seeing a fragment of their clothes. This set-up with reduced "ontology" is referred to as a "situation." Situations are usually local in space and time. In this work, we propose that neuron-astrocyte networks provide a network topology that is effectively adapted to accommodate situation-based memory. In order to illustrate this, we numerically simulate and analyze a well-established model of a neuron-astrocyte network, which is subjected to stimuli conforming to the situation-driven environment. Three pools of stimuli patterns are considered: external patterns, patterns from the situation associative pool regularly presented to the network and learned by the network, and patterns already learned and remembered by astrocytes. Patterns from the external world are added to and removed from the associative pool. Then, we show that astrocytes are structurally necessary for an effective function in such a learning and testing set-up. To demonstrate this we present a novel neuromorphic computational model for short-term memory implemented by a two-net spiking neural-astrocytic network. Our results show that such a system tested on synthesized data with selective astrocyte-induced modulation of neuronal activity provides an enhancement of retrieval quality in comparison to standard spiking neural networks trained via Hebbian plasticity only. We argue that the proposed set-up may offer a new way to analyze, model, and understand neuromorphic artificial intelligence systems. Susanna Yu Gordleeva, Yuliya Tsybina, Mikhail Krivonosov, Ivan Tyukin, Victor B. Kazantsev, Alexey Zaikin, Alexander N. Gorban |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Weakly Supervised Learners for Correction of AI Errors with Provable Performance GuaranteesabstractWe present a new methodology for handling errors of Artificial Intelligence (AI) by introducing weakly supervised AI error correctors with a priori performance guarantees. These AI correctors are auxiliary maps whose role is to moderate the decisions of some previously constructed underlying classifier by either approving or rejecting its decisions. The rejection of a decision can be used as a signal to suggest abstaining from making a decision. A key technical focus of the work is in providing performance guarantees for these new AI correctors through bounds on the probabilities of incorrect decisions. These bounds are distribution agnostic and do not rely on assumptions on the data dimension. Our empirical example illustrates how the framework can be applied to improve the performance of an image classifier in a challenging real-world task where training data are scarce. Ivan Tyukin, Tatiana Tyukina, Daniel van Helden, Zedong Zheng, Eugenij Moiseevich Mirkes, Oliver J. Sutton, Alexander N. Gorban, Penelope M. Allison |
IJCNN | 8 |
| 2024 | Stealth edits to large language modelsabstractWe reveal the theoretical foundations of techniques for editing large language models, and present new methods which can do so without requiring retraining. Our theoretical insights show that a single metric (a measure of the intrinsic dimension of the model's features) can be used to assess a model's editability and reveals its previously unrecognised susceptibility to malicious *stealth attacks*. This metric is fundamental to predicting the success of a variety of editing approaches, and reveals new bridges between disparate families of editing methods. We collectively refer to these as *stealth editing* methods, because they directly update a model's weights to specify its response to specific known hallucinating prompts without affecting other model behaviour. By carefully applying our theoretical insights, we are able to introduce a new *jet-pack* network block which is optimised for highly selective model editing, uses only standard network operations, and can be inserted into existing networks. We also reveal the vulnerability of language models to stealth attacks: a small change to a model's weights which fixes its response to a single attacker-chosen prompt. Stealth attacks are computationally simple, do not require access to or knowledge of the model's training data, and therefore represent a potent yet previously unrecognised threat to redistributed foundation models. Extensive experimental results illustrate and support our methods and their theoretical underpinnings. Demos and source code are available at https://github.com/qinghua-zhou/stealth-edits. Oliver J. Sutton, Wei Wang 0357, Desmond J. Higham, Alexander N. Gorban, Alexander Bastounis, Ivan Tyukin |
NeurIPS | 5 |
| 2024 | Coping with AI errors with provable guaranteesabstractAI errors pose a significant challenge, hindering real-world applications. This work introduces a novel approach to cope with AI errors using weakly supervised error correctors that guarantee a specific level of error reduction. Our correctors have low computational cost and can be used to decide whether to abstain from making an unsafe classification. We provide new upper and lower bounds on the probability of errors in the corrected system. In contrast to existing works, these bounds are distribution agnostic, non-asymptotic, and can be efficiently computed just using the corrector training data. They also can be used in settings with concept drifts when the observed frequencies of separate classes vary. The correctors can easily be updated, removed, or replaced in response to changes in distributions within each class without retraining the underlying classifier. The application of the approach is illustrated with two relevant challenging tasks: (i) an image classification problem with scarce training data, and (ii) moderating responses of large language models without retraining or otherwise fine-tuning. Ivan Tyukin, Tatiana Tyukina, Daniël P. van Helden, Zedong Zheng, Eugenij Moiseevich Mirkes, Oliver J. Sutton, Alexander N. Gorban, Penelope M. Allison |
Inf. Sci. | 8 |
| 2024 | How adversarial attacks can disrupt seemingly stable accurate classifiersabstractAdversarial attacks dramatically change the output of an otherwise accurate learning system using a seemingly inconsequential modification to a piece of input data. Paradoxically, empirical evidence indicates that even systems which are robust to large random perturbations of the input data remain susceptible to small, easily constructed, adversarial perturbations of their inputs. Here, we show that this may be seen as a fundamental feature of classifiers working with high dimensional input data. We introduce a simple generic and generalisable framework for which key behaviours observed in practical systems arise with high probability-notably the simultaneous susceptibility of the (otherwise accurate) model to easily constructed adversarial attacks, and robustness to random perturbations of the input data. We confirm that the same phenomena are directly observed in practical neural networks trained on standard image classification problems, where even large additive random noise fails to trigger the adversarial instability of the network. A surprising takeaway is that even small margins separating a classifier's decision surface from training and testing data can hide adversarial susceptibility from being detected using randomly sampled perturbations. Counter-intuitively, using additive noise during training or testing is therefore inefficient for eradicating or detecting adversarial examples, and more demanding adversarial training is required. Oliver J. Sutton, Ivan Tyukin, Alexander N. Gorban, Alexander Bastounis, Desmond J. Higham |
Neural Networks | 4 |
| 2023 | What is Hiding in Medicine's Dark Matter? Learning with Missing Data in Medical PracticesabstractElectronic patient records (EPRs) produce a wealth of data but contain significant missing information. Understanding and handling this missing data is an important part of clinical data analysis and if left unaddressed could result in bias in analysis and distortion in critical conclusions. Missing data may be linked to health care professional practice patterns and imputation of missing data can increase the validity of clinical decisions. This study focuses on statistical approaches for understanding and interpreting the missing data and machine learning based clinical data imputation using a single centre’s paediatric emergency data and the data from UK’s largest clinical audit for traumatic injury database (TARN). In the study of 56,961 data points related to initial vital signs and observations taken on children presenting to an Emergency Department, we have shown that missing data are likely to be non-random and how these are linked to health care professional practice patterns. We have then examined 79 TARN fields with missing values for 5,791 trauma cases. Singular Value Decomposition (SVD) and k-Nearest Neighbour (kNN) based missing data imputation methods are used and imputation results against the original dataset are compared and statistically tested. We have concluded that the INN imputer is the best imputation which indicates a usual pattern of clinical decision making: find the most similar patients and take their attributes as imputation. Neslihan Suzen, Eugenij Moiseevich Mirkes, Damian Roland, Jeremy Levesley, Alexander N. Gorban, Timothy J. Coats |
IEEE Big Data | 5 |
| 2023 | The Boundaries of Verifiable Accuracy, Robustness, and Generalisation in Deep Learning
Alexander Bastounis, Alexander N. Gorban, Anders C. Hansen, Desmond J. Higham, Danil V. Prokhorov, Oliver J. Sutton, Ivan Tyukin |
ICANN (1) | 2 |
| 2023 | Relative Intrinsic Dimensionality Is Intrinsic to Learning
Oliver J. Sutton, Alexander N. Gorban, Ivan Tyukin |
ICANN (1) | 3 |
| 2023 | Agile gesture recognition for capacitive sensing devices: adapting on-the-jobabstractAutomated hand gesture recognition has been a focus of the AI community for decades. Traditionally, work in this domain revolved largely around scenarios assuming the availability of the flow of images of the operator's/user's hands. This has partly been due to the prevalence of camera-based devices and the wide availability of image data. However, there is growing demand for gesture recognition technology that can be implemented on low-power devices using limited sensor data instead of high-dimensional inputs like hand images. In this work, we demonstrate a hand gesture recognition system and method that uses signals from capacitive sensors embedded into the etee hand controller. The controller generates real-time signals from each of the wearer's five fingers. We use a machine learning technique to analyse the time-series signals and identify three features that can represent 5 fingers within 500 ms. The analysis is composed of a two-stage training strategy, including dimension reduction through principal component analysis and classification with K-nearest neighbour. Remarkably, we found that this combination showed a level of performance which was comparable to more advanced methods such as supervised variational autoencoder. The base system can also be equipped with the capability to learn from occasional errors by providing it with an additional adaptive error correction mechanism. The results showed that the error corrector improve the classification performance in the base system without compromising its performance. The system requires no more than 1 ms of computing time per input sample, and is smaller than deep neural networks, demonstrating the feasibility of agile gesture recognition systems based on this technology. Liucheng Guo, Valeri A. Makarov, Alexander N. Gorban, Eugenij Moiseevich Mirkes, Ivan Tyukin |
IJCNN | 5 |
| 2023 | Neuromorphic tuning of feature spaces to overcome the challenge of low-sample high-dimensional dataabstractFor learning algorithms, accessing large volumes of annotated data is highly desirable but not always available, especially in real-world scenarios. Accordingly, learning in the high-dimensional and low-sample size (HDLS) domain is recognised as one of the core challenges for modern AI systems. In this work, we consider a particular but very practical scenario in the HDLS domain where the number of training samples is not limited to mere few observations but yet it is not large enough to reliably build models with high degrees of expressivity. To address the problem, we present a new neuromorphic algorithm capable of fine-tuning existing feature spaces via learning relevant associations in high dimensional data with high probability. The algorithm is based on the idea of Concept Cells [1] and mimics properties attributed to memory and learning inherent to live neural systems. We demonstrate, through numerous numerical experiments, that the algorithm can “fine-tune” and “adapt” the feature space of pre-trained neural networks for better performance on new tasks in the HDLS domain. In addition, we study the impact of this “tuning” on quasi-orthogonal measures, which correlates with classification and calibration metrics. Oliver J. Sutton, Yudong Zhang 0001, Alexander N. Gorban, Valeri A. Makarov, Ivan Tyukin |
IJCNN | 4 |
| 2022 | Quasi-orthogonality and intrinsic dimensions as measures of learning and generalisationabstractFinding the best architectures for learning machines, such as deep neural networks, is a well-known technical and theoretical challenge. Recent work by Mellor et al [1] showed that there may exist correlations between the accuracies of trained networks and the values of some easily computable measures defined on randomly initialised networks which may enable the search of tens of thousands of neural architectures without training. Mellor et al [1] used the Hamming distance evaluated over all the ReLU neurons as such a measure. Motivated by these findings, in our work, we ask the question of the existence of other and perhaps more principled measures which could be used as determinants of the potential success of a given neural architecture. In particular, we examine if the dimensionality and quasi-orthogonality of neural networks' feature space could be correlated with the network's performance after training. We showed, using the setup as in Mellor et al [1], that dimensionality and quasi-orthogonality may jointly serve as networks' performance discriminants. In addition to offering new opportunities to accelerate neural architecture search, our findings suggest important relationships between the networks' final performance and properties of their randomly initialised feature spaces: data dimension and quasi-orthogonality. Alexander N. Gorban, Eugenij Moiseevich Mirkes, Jonathan Bac, Andrei Yu. Zinovyev, Ivan Tyukin |
IJCNN | 2 |
| 2022 | Astrocytes mediate analogous memory in a multi-layer neuron-astrocyte networkabstractAbstract Modeling the neuronal processes underlying short-term working memory remains the focus of many theoretical studies in neuroscience. In this paper, we propose a mathematical model of a spiking neural network (SNN) which simulates the way a fragment of information is maintained as a robust activity pattern for several seconds and the way it completely disappears if no other stimuli are fed to the system. Such short-term memory traces are preserved due to the activation of astrocytes accompanying the SNN. The astrocytes exhibit calcium transients at a time scale of seconds. These transients further modulate the efficiency of synaptic transmission and, hence, the firing rate of neighboring neurons at diverse timescales through gliotransmitter release. We demonstrate how such transients continuously encode frequencies of neuronal discharges and provide robust short-term storage of analogous information. This kind of short-term memory can store relevant information for seconds and then completely forget it to avoid overlapping with forthcoming patterns. The SNN is inter-connected with the astrocytic layer by local inter-cellular diffusive connections. The astrocytes are activated only when the neighboring neurons fire synchronously, e.g., when an information pattern is loaded. For illustration, we took grayscale photographs of people’s faces where the shades of gray correspond to the level of applied current which stimulates the neurons. The astrocyte feedback modulates (facilitates) synaptic transmission by varying the frequency of neuronal firing. We show how arbitrary patterns can be loaded, then stored for a certain interval of time, and retrieved if the appropriate clue pattern is applied to the input. Yuliya Tsybina, Innokentiy Kastalskiy, Mikhail Krivonosov, Alexey Zaikin, Victor B. Kazantsev, Alexander N. Gorban, Susanna Yu Gordleeva |
Neural Comput. Appl. | 6 |
| 2022 | Coloring Panchromatic Nighttime Satellite Images: Comparing the Performance of Several Machine Learning MethodsabstractArtificial light-at-night (ALAN), emitted from the ground and visible from space, marks human presence on earth. Since the launch of the Suomi National Polar Partnership satellite with the Visible Infrared Imaging Radiometer Suite Day–Night Band (VIIRS/DNB) onboard, global nighttime images have significantly improved; however, they remained panchromatic. Although multispectral images are also available, they are either commercial or free of charge, but sporadic. In this article, we use several machine learning techniques, such as linear, kernel, random forest regressions, and elastic map approach, to transform panchromatic VIIRS/DBN into red, green, blue (RGB) images. To validate the proposed approach, we analyze RGB images for eight urban areas worldwide. We link RGB values, obtained from ISS photographs, to panchromatic ALAN intensities, their pixel-wise differences, and several land-use-type proxies. Each dataset is used for model training, while other datasets are used for model validation. The analysis shows that model-estimated RGB images demonstrate a high degree of correspondence with the original RGB images from the ISS database. Yet, estimates, based on linear, kernel, and random forest regressions, provide better correlations, contrast similarity, and lower WMSEs levels, while RGB images, generated using elastic map approach, provide higher consistency of predictions. Natalya A. Rybnikova, Boris A. Portnov, Eugenij Moiseevich Mirkes, Andrei Yu. Zinovyev, Anna Brook, Alexander N. Gorban |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | Modelling working memory in neuron-astrocyte networkabstractWorking memory is one of the most intriguing brain function phenomena that enables storage and recognition of several information patterns simultaneously in the form of coherent activations of specific brain circuitries. These patterns can be recalled and, if physiologically (cognitively) significant, further transferred to long term storages by cortical circuits. In the paper, we show how the working memory can be effectively organized by a multiscale network model composed of spiking neurons accompanied by an astrocytic network. The latter serves as the temporal storage of information patterns that can be manipulated (relearned, retrieved, transferred) during astrocytic calcium activation. In turn, the activation of the astrocyte network is possible when coherent firing occurs in corresponding sites of the neuronal layer. We study the role of interplay of the astrocyte-induced modulation of signal transmission in neural network and the Hebbian synaptic plasticity in the working memory organization. We show that modulation of synaptic communication caused by astrocytes does not exclude but rather complements Hebbian synaptic plasticity, and they can perfectly work in parallel. We believe this model is a significant step towards confirming the importance of non-neuron species (e.g. astrocytes) in the formation and sustainability of cognitive functions of the brain. Yuliya Tsybina, Susanna Yu Gordleeva, Mikhail Krivonosov, Innokentiy Kastalskiy, Alexey Zaikin, Alexander N. Gorban |
IJCNN | 6 |
| 2021 | Demystification of Few-shot and One-shot LearningabstractFew-shot and one-shot learning have been the subject of active and intensive research in recent years, with mounting evidence pointing to successful implementation and exploitation of few-shot learning algorithms in practice. Classical statistical learning theories do not fully explain why few- or one-shot learning is at all possible since traditional generalisation bounds normally require large training and testing samples to be meaningful. This sharply contrasts with numerous examples of successful one- and few-shot learning systems and applications. In this work we present mathematical foundations for a theory of one-shot and few-shot learning and reveal conditions specifying when such learning schemes are likely to succeed. Our theory is based on intrinsic properties of high-dimensional spaces. We show that if the ambient or latent decision space of a learning machine is sufficiently high-dimensional than a large class of objects in this space can indeed be easily learned from few examples provided that certain data non-concentration conditions are met. In this work we present mathematical foundations for a theory of one-shot and few-shot learning and reveal conditions specifying when such learning schemes are likely to succeed. Our theory is based on intrinsic properties of high-dimensional spaces. We show that if the ambient or latent decision space of a learning machine is sufficiently high-dimensional than a large class of objects in this space can indeed be easily learned from few examples provided that certain data non-concentration conditions are met. Ivan Tyukin, Alexander N. Gorban, Muhammad H. Alkhudaydi |
IJCNN | 2 |
| 2021 | Blessing of dimensionality at the edge and geometry of few-shot learningabstractIn this paper we present theory and algorithms enabling classes of Artificial Intelligence (AI) systems to continuously and incrementally improve with a priori quantifiable guarantees – or more specifically remove classification errors – over time. This is distinct from state-of-the-art machine learning, AI, and software approaches. The theory enables building few-shot AI correction algorithms and provides conditions justifying their successful application. Another feature of this approach is that, in the supervised setting, the computational complexity of training is linear in the number of training samples. At the time of classification, the computational complexity is bounded by few inner product calculations. Moreover, the implementation is shown to be very scalable. This makes it viable for deployment in applications where computational power and memory are limited, such as embedded environments. It enables the possibility for fast on-line optimisation using improved training samples. The approach is based on the concentration of measure effects and stochastic separation theorems and is illustrated with an example on the identification faulty processes in Computer Numerical Control (CNC) milling and with a case study on adaptive removal of false positives in an industrial video surveillance and analytics system. Ivan Tyukin, Alexander N. Gorban, Alistair A. McEwan, Sepehr Meshkinfamfard |
Inf. Sci. | 2 |
| 2021 | General stochastic separation theorems with optimal boundsabstractPhenomenon of stochastic separability was revealed and used in machine learning to correct errors of Artificial Intelligence (AI) systems and analyze AI instabilities. In high-dimensional datasets under broad assumptions each point can be separated from the rest of the set by simple and robust Fisher's discriminant (is Fisher separable). Errors or clusters of errors can be separated from the rest of the data. The ability to correct an AI system also opens up the possibility of an attack on it, and the high dimensionality induces vulnerabilities caused by the same stochastic separability that holds the keys to understanding the fundamentals of robustness and adaptivity in high-dimensional data-driven AI. To manage errors and analyze vulnerabilities, the stochastic separation theorems should evaluate the probability that the dataset will be Fisher separable in given dimensionality and for a given class of distributions. Explicit and optimal estimates of these separation probabilities are required, and this problem is solved in the present work. The general stochastic separation theorems with optimal probability estimates are obtained for important classes of distributions: log-concave distribution, their convex combinations and product distributions. The standard i.i.d. assumption was significantly relaxed. These theorems and estimates can be used both for correction of high-dimensional data driven AI systems and for analysis of their vulnerabilities. The third area of application is the emergence of memories in ensembles of neurons, the phenomena of grandmother's cells and sparse coding in the brain, and explanation of unexpected effectiveness of small neural ensembles in high-dimensional brain. Bogdan Grechuk, Alexander N. Gorban, Ivan Tyukin |
Neural Networks | 2 |
| 2020 | On Adversarial Examples and Stealth Attacks in Artificial Intelligence SystemsabstractIn this work we present a formal theoretical framework for assessing and analyzing two classes of malevolent action towards generic Artificial Intelligence (AI) systems. Our results apply to general multi-class classifiers that map from an input space into a decision space, including artificial neural networks used in deep learning applications. Two classes of attacks are considered. The first class involves adversarial examples and concerns the introduction of small perturbations of the input data that cause misclassification. The second class, introduced here for the first time and named stealth attacks, involves small perturbations to the AI system itself. Here the perturbed system produces whatever output is desired by the attacker on a specific small data set, perhaps even a single input, but performs as normal on a validation set (which is unknown to the attacker).We show that in both cases, i.e., in the case of an attack based on adversarial examples and in the case of a stealth attack, the dimensionality of the AI's decision-making space is a major contributor to the AI's susceptibility. For attacks based on adversarial examples, a second crucial parameter is the absence of local concentrations in the data probability distribution, a property known as Smeared Absolute Continuity. According to our findings, robustness to adversarial examples requires either (a) the data distributions in the AI's feature space to have concentrated probability density functions or (b) the dimensionality of the AI's decision variables to be sufficiently small. We also show how to construct stealth attacks on high-dimensional AI systems that are hard to spot unless the validation set is made exponentially large. Ivan Tyukin, Desmond J. Higham, Alexander N. Gorban |
IJCNN | 3 |
| 2020 | Adapting Style and Content for Attended Text Sequence RecognitionabstractIn this paper, we address the problem of learning to perform sequential OCR on photos of street name signs in a language for which no labeled data exists. Our approach leverages easily-generated synthetic data and existing labeled data in other languages to achieve reasonable performance on these unlabeled images, through a combination of a novel domain adaptation technique based on gradient reversal and a multi-task learning scheme. In order to accomplish this, we introduce and release two new datasets - Hebrew Street Name Signs (HSNS) and Synthetic Hebrew Street Name Signs (SynHSNS) - while also making use of the existing French Street Name Signs (FSNS) dataset. We demonstrate that by using a synthetic dataset of Hebrew characters and a labeled dataset of French street name signs in natural images, it is possible to achieve a significant improvement on real Hebrew street name sign transcription, where the synthetic Hebrew data and real French data each overlap with different features of the images we wish to transcribe. Steven Schwarcz, Alexander N. Gorban, Xavier Gibert Serra, Dar-Shyang Lee |
WACV | 2 |
| 2020 | Multivariate Gaussian and Student-t process regression for multi-output predictionabstractAbstract Gaussian process model for vector-valued function has been shown to be useful for multi-output prediction. The existing method for this model is to reformulate the matrix-variate Gaussian distribution as a multivariate normal distribution. Although it is effective in many cases, reformulation is not always workable and is difficult to apply to other distributions because not all matrix-variate distributions can be transformed to respective multivariate distributions, such as the case for matrix-variate Student-t distribution. In this paper, we propose a unified framework which is used not only to introduce a novel multivariate Student-t process regression model (MV-TPR) for multi-output prediction, but also to reformulate the multivariate Gaussian process regression (MV-GPR) that overcomes some limitations of the existing methods. Both MV-GPR and MV-TPR have closed-form expressions for the marginal likelihoods and predictive distributions under this unified framework and thus can adopt the same optimization approaches as used in the conventional GPR. The usefulness of the proposed methods is illustrated through several simulated and real-data examples. In particular, we verify empirically that MV-TPR has superiority for the datasets considered, including air quality prediction and bike rent prediction. At last, the proposed methods are shown to produce profitable investment strategies in the stock markets. Zexun Chen, Alexander N. Gorban |
Neural Comput. Appl. | 3 |
| 2020 | Correction to: Multivariate Gaussian and Student-t process regression for multi-output prediction
Zexun Chen, Alexander N. Gorban |
Neural Comput. Appl. | 3 |
| 2019 | Do Fractional Norms and Quasinorms Help to Overcome the Curse of Dimensionality?abstractThe curse of dimensionality causes well-known and widely discussed problems for machine learning methods. There is a hypothesis that usage of Manhattan distance and even fractional quasinorms lp (for p less than 1) can help to overcome the curse of dimensionality in classification problems. In this study, we systematically test this hypothesis for 37 binary classification problems on 25 databases. We confirm that fractional quasinorms have greater relative contrast or coefficient of variation than Euclidean norm l2, but we demonstrate also that the distance concentration shows qualitatively the same behaviour for all tested norms and quasinorms and the difference between them decays while dimension tends to infinity. Estimation of classification quality for kNN based on different norms and quasinorms shows that the greater relative contrast does not mean the better classifier performance and the worst performance for different databases was shown by the different norms (quasinorms). A systematic comparison shows that the difference in performance of kNN based on lp for p=2, 1, and 0.5 is statistically insignificant. Eugenij Moiseevich Mirkes, Jeza Allohibi, Alexander N. Gorban |
IJCNN | 3 |
| 2019 | Kernel Stochastic Separation Theorems and Separability Characterizations of Kernel ClassifiersabstractIn this work we provide generalizations and extensions of stochastic separation theorems to kernel classifiers. A general separability result for two random sets is also established. We show that despite feature maps corresponding to a given kernel function may be infinite-dimensional, kernel separability characterizations can be expressed in terms of finite-dimensional volume integrals. These integrals allow to determine and quantify separability properties of an arbitrary kernel function. The theory is illustrated with numerical examples. Ivan Tyukin, Alexander N. Gorban, Bogdan Grechuk, Stephen Green 0001 |
IJCNN | 2 |
| 2019 | One-trial correction of legacy AI systems and stochastic separation theorems
Alexander N. Gorban, Richard Burton, Ilya V. Romanenko, Ivan Tyukin |
Inf. Sci. | 1 |
| 2019 | Fast construction of correcting ensembles for legacy Artificial Intelligence systems: Algorithms and a case study
Ivan Tyukin, Alexander N. Gorban, Stephen Green 0001, Danil V. Prokhorov |
Inf. Sci. | 2 |
| 2018 | Large Scale Scene Text Verification with Guided Attention
Dafang He, Yeqing Li, Alexander N. Gorban, Derrall Heath, Julian Ibarz, Daniel Kifer, C. Lee Giles |
ACCV (5) | 3 |
| 2018 | Data analysis with arbitrary error measures approximated by piece-wise quadratic PQSQ functionsabstractDefining an error function (a measure of deviation of a model prediction from the data) is a critical step in any optimization-based data analysis method, including regression, clustering and dimension reduction. Usual quadratic error function in case of real-life high-dimensional and noisy data suffers from non-robustness to presence of outliers. Therefore, using non-quadratic error functions in data analysis and machine learning (such as L1 norm-based) is an active field of modern research but the majority of methods suggested are either slow or imprecise (use arbitrary heuristics). We suggest a flexible and highly performant approach to generalize most of existing data analysis methods to an arbitrary error function of subquadratic growth. For this purpose, we exploit PQSQ functions (piece-wise quadratic of subquadratic growth), which can be minimized by a simple and fast splitting-based iterative algorithm. The theoretical basis of the PQSQ approach is an application of min-plus (idempotent) algebra to data approximation. We introduce the general idea of the approach and illustrate it on four standard tools of machine learning: simple regression, regularized regression, k-mean clustering and principal component analysis. In all cases, PQSQ-based methods achieve better robustness with respect to the presence of strong noise in the data compared to the standard methods. Alexander N. Gorban, Eugenij Moiseevich Mirkes, Andrei Yu. Zinovyev |
IJCNN | 1 |
| 2018 | Efficiency of Shallow Cascades for Improving Deep Learning AI SystemsabstractThis paper presents a technology for simple and non-iterative improvements of Multilayer and Deep Learning neural networks and Artificial Intelligence (AI) systems. The improvements are, in essence, shallow networks constructed on top of the existing Deep Learning architecture. Theoretical foundation of the technology is based on Stochastic Separation Theorems and the ideas of measure concentration. We show that, subject to mild technical assumptions on statistical properties of internal signals in Deep Learning AI, with probability close to one the technology enables instantaneous “learning away” of spurious and systematic errors. The method is illustrated with numerical examples. Ivan Tyukin, Alexander N. Gorban, Danil V. Prokhorov, Stephen Green 0001 |
IJCNN | 2 |
| 2018 | Correction of AI systems by linear discriminants: Probabilistic foundations
Alexander N. Gorban, A. Golubkov, Bogdan Grechuk, Eugenij Moiseevich Mirkes, Ivan Tyukin |
Inf. Sci. | 1 |
| 2017 | Attention-Based Extraction of Structured Information from Street View ImageryabstractWe present a neural network model — based on Convolutional Neural Networks, Recurrent Neural Networks and a novel attention mechanism — which achieves 84.2% accuracy on the challenging French Street Name Signs (FSNS) dataset, significantly outperforming the previous state of the art (Smith'16), which achieved 72.46%. Furthermore, our new method is much simpler and more general than the previous approach. To demonstrate the generality of our model, we show that it also performs well on an even more challenging dataset derived from Google Street View, in which the goal is to extract business names from store fronts. Finally, we study the speed/accuracy tradeoff that results from using CNN feature extractors of different depths. Surprisingly, we find that deeper is not always better (in terms of accuracy, as well as speed). Our resulting model is simple, accurate and fast, allowing it to be used at scale on a variety of challenging real-world text extraction problems. Zbigniew Wojna, Alexander N. Gorban, Dar-Shyang Lee, Kevin Murphy 0002, Yeqing Li, Julian Ibarz |
ICDAR | 2 |
| 2017 | Stochastic separation theorems
Alexander N. Gorban, Ivan Tyukin |
Neural Networks | 1 |
| 2016 | Detecting Events and Key Actors in Multi-person VideosabstractMulti-person event recognition is a challenging task, often with many people active in the scene but only a small subset contributing to an actual event. In this paper, we propose a model which learns to detect events in such videos while automatically "attending" to the people responsible for the event. Our model does not use explicit annotations regarding who or where those people are during training and testing. In particular, we track people in videos and use a recurrent neural network (RNN) to represent the track features. We learn time-varying attention weights to combine these features at each time-instant. The attended features are then processed using another RNN for event detection/ classification. Since most video datasets with multiple people are restricted to a small number of videos, we also collected a new basketball dataset comprising 257 basketball games with 14K event annotations corresponding to 11 event classes. Our model outperforms state-of-the-art methods for both event classification and detection on this new dataset. Additionally, we show that the attention mechanism is able to consistently localize the relevant players. Vignesh Ramanathan, Jonathan Huang, Sami Abu-El-Haija, Alexander N. Gorban, Kevin Murphy 0002, Li Fei-Fei 0001 |
CVPR | 4 |
| 2016 | SOM: Stochastic initialization versus principal components
Ayodeji A. Akinduko, Eugenij Moiseevich Mirkes, Alexander N. Gorban |
Inf. Sci. | 3 |
| 2016 | Approximation with random bases: Pro et Contra
Alexander N. Gorban, Ivan Tyukin, Danil V. Prokhorov, Konstantin I. Sofeikov |
Inf. Sci. | 1 |
| 2016 | Piece-wise quadratic approximations of arbitrary error functions for fast and robust machine learning
Alexander N. Gorban, Eugenij Moiseevich Mirkes, Andrei Yu. Zinovyev |
Neural Networks | 1 |
| 2015 | Fast and user-friendly non-linear principal manifold learning by method of elastic mapsabstractMethod of elastic maps allows fast learning of nonlinear principal manifolds for large datasets. We present user-friendly implementation of the method in ViDaExpert software. Equipped with several dialogs for configuring data point representations (size, shape, color) and fast 3D viewer, ViDaExpert is a handy tool allowing to construct an interactive 3D-scene representing a table of data in multidimensional space and perform its quick and insightfull statistical analysis, from basic to advanced methods. We list several recent application examples of manifold learning by method of elastic maps in various fields of life sciences. Alexander N. Gorban, Andrei Yu. Zinovyev |
DSAA | 1 |
| 2015 | Im2Calories: Towards an Automated Mobile Vision Food DiaryabstractWe present a system which can recognize the contents of your meal from a single image, and then predict its nutritional contents, such as calories. The simplest version assumes that the user is eating at a restaurant for which we know the menu. In this case, we can collect images offline to train a multi-label classifier. At run time, we apply the classifier (running on your phone) to predict which foods are present in your meal, and we lookup the corresponding nutritional facts. We apply this method to a new dataset of images from 23 different restaurants, using a CNN-based classifier, significantly outperforming previous work. The more challenging setting works outside of restaurants. In this case, we need to estimate the size of the foods, as well as their labels. This requires solving segmentation and depth / volume estimation from a single image. We present CNN-based approaches to these problems, with promising preliminary results. Austin Myers, Nicholas Johnston, Vivek Rathod, Anoop Korattikara Balan, Alexander N. Gorban, Nathan Silberman, Sergio Guadarrama, George Papandreou, Jonathan Huang, Kevin Murphy 0002 |
ICCV | 5 |
| 2014 | Learning optimization for decision tree classification of non-categorical data with information gain impurity criterionabstractWe consider the problem of construction of decision trees in cases when data is non-categorical and is inherently high-dimensional. Using conventional tree growing algorithms that either rely on univariate splits or employ direct search methods for determining multivariate splitting conditions is computationally prohibitive. On the other hand application of standard optimization methods for finding locally optimal splitting conditions is obstructed by abundance of local minima and discontinuities of classical goodness functions such as e.g. information gain or Gini impurity. In order to avoid this limitation a method to generate smoothed replacement for measuring impurity of splits is proposed. This enables to use vast number of efficient optimization techniques for finding locally optimal splits and, at the same time, decreases the number of local minima. The approach is illustrated with examples. Konstantin I. Sofeikov, Ivan Tyukin, Alexander N. Gorban, Eugenij Moiseevich Mirkes, Danil V. Prokhorov, Ilya V. Romanenko |
IJCNN | 3 |
| 2010 | Principal Manifolds and Graphs in Practice: from Molecular Biology to Dynamical SystemsabstractWe present several applications of non-linear data modeling, using principal manifolds and principal graphs constructed using the metaphor of elasticity (elastic principal graph approach). These approaches are generalizations of the Kohonen's self-organizing maps, a class of artificial neural networks. On several examples we show advantages of using non-linear objects for data approximation in comparison to the linear ones. We propose four numerical criteria for comparing linear and non-linear mappings of datasets into the spaces of lower dimension. The examples are taken from comparative political science, from analysis of high-throughput data in molecular biology, from analysis of dynamical systems. Alexander N. Gorban, Andrei Yu. Zinovyev |
Int. J. Neural Syst. | 1 |
| 2007 | Branching Principal Components: Elastic Graphs, Topological Grammars and Metro MapsabstractTo approximate complex data, we propose new type of low-dimensional "principal object":principal cubic complex. This complex is a generalization of linear and nonlinear principal manifolds and includes them as a particular case. To construct such an object, we combine the method oftopological grammarswith the minimization of elastic energy defined for its embedment into multidimensional data space. The whole complex is presented as a system of nodes and springs and as a product of one-dimensional continua (represented by graphs), and the grammars describe how these continua transform during the process of optimal complex construction. The simplest case of a topological grammar ("add a node or bisect an edge") produces "principal trees" that are useful in many practical applications. We demonstrate how this can be applied to the analysis of bacterial genomes and for visualization of microarray data using "metro map" visual representation. Alexander N. Gorban, Neil R. Sumner, Andrei Yu. Zinovyev |
IJCNN | 1 |
| 2003 | Application of the method of elastic maps in analysis of genetic textsabstractMethod of elastic maps allows to construct efficiently 1D, 2D and 3D nonlinear approximations to the principal manifolds with different topology (piece of plane, sphere, torus etc.) and to project data onto it. We describe the idea of the method and demonstrate its applications in analysis of genetic sequences. Alexander N. Gorban, Andrei Yu. Zinovyev, Donald C. Wunsch II |
IJCNN | 1 |
| 1999 | Generation of explicit knowledge from empirical data through pruning of trainable neural networksabstractThis paper presents a generalized technology of extraction of explicit knowledge from data. The main ideas are: 1) maximal reduction of network complexity (not only removal of neurons or synapses, but removal all the unnecessary elements and signals and reduction of the complexity of elements); 2) using of adjustable and flexible pruning process (the user should have a possibility to prune network on his own way in order to achieve a desired network structure for the purpose of extraction of rules of desired type and form); and 3) extraction of rules not in predetermined but any desired form. Some considerations and notes about network architecture and training process and applicability of currently developed pruning techniques and rule extraction algorithms are discussed. This technology, being developed by us for more than 10 years, allowed us to create dozens of knowledge-based expert systems. Alexander N. Gorban, Eugenij Moiseevich Mirkes, Victor G. Tsaregorodtsev |
IJCNN | 1 |