EDBT 2026 Demo / reviewers in the wild / expert
Alexander N. Gorban
dblp:45/3050
· DBLP profile ↗
10ranked-venue papers in the field
4as first author
3since 2021 · last 2024
0000-0001-6224-1430ORCID · verified
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 7 (3 first)Data Mining & Knowledge Discovery · 1 (1 first)Big Data, Cloud & Distributed Data Systems · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Coping with AI errors with provable guaranteesabstractAI errors pose a significant challenge, hindering real-world applications. This work introduces a novel approach to cope with AI errors using weakly supervised error correctors that guarantee a specific level of error reduction. Our correctors have low computational cost and can be used to decide whether to abstain from making an unsafe classification. We provide new upper and lower bounds on the probability of errors in the corrected system. In contrast to existing works, these bounds are distribution agnostic, non-asymptotic, and can be efficiently computed just using the corrector training data. They also can be used in settings with concept drifts when the observed frequencies of separate classes vary. The correctors can easily be updated, removed, or replaced in response to changes in distributions within each class without retraining the underlying classifier. The application of the approach is illustrated with two relevant challenging tasks: (i) an image classification problem with scarce training data, and (ii) moderating responses of large language models without retraining or otherwise fine-tuning. Ivan Tyukin, Tatiana Tyukina, Daniël P. van Helden, Zedong Zheng, Eugenij Moiseevich Mirkes, Oliver J. Sutton, Alexander N. Gorban, Penelope M. Allison |
Inf. Sci. | 8 |
| 2023 | What is Hiding in Medicine's Dark Matter? Learning with Missing Data in Medical PracticesabstractElectronic patient records (EPRs) produce a wealth of data but contain significant missing information. Understanding and handling this missing data is an important part of clinical data analysis and if left unaddressed could result in bias in analysis and distortion in critical conclusions. Missing data may be linked to health care professional practice patterns and imputation of missing data can increase the validity of clinical decisions. This study focuses on statistical approaches for understanding and interpreting the missing data and machine learning based clinical data imputation using a single centre’s paediatric emergency data and the data from UK’s largest clinical audit for traumatic injury database (TARN). In the study of 56,961 data points related to initial vital signs and observations taken on children presenting to an Emergency Department, we have shown that missing data are likely to be non-random and how these are linked to health care professional practice patterns. We have then examined 79 TARN fields with missing values for 5,791 trauma cases. Singular Value Decomposition (SVD) and k-Nearest Neighbour (kNN) based missing data imputation methods are used and imputation results against the original dataset are compared and statistically tested. We have concluded that the INN imputer is the best imputation which indicates a usual pattern of clinical decision making: find the most similar patients and take their attributes as imputation. Neslihan Suzen, Eugenij Moiseevich Mirkes, Damian Roland, Jeremy Levesley, Alexander N. Gorban, Timothy J. Coats |
IEEE Big Data | 5 |
| 2021 | Blessing of dimensionality at the edge and geometry of few-shot learningabstractIn this paper we present theory and algorithms enabling classes of Artificial Intelligence (AI) systems to continuously and incrementally improve with a priori quantifiable guarantees – or more specifically remove classification errors – over time. This is distinct from state-of-the-art machine learning, AI, and software approaches. The theory enables building few-shot AI correction algorithms and provides conditions justifying their successful application. Another feature of this approach is that, in the supervised setting, the computational complexity of training is linear in the number of training samples. At the time of classification, the computational complexity is bounded by few inner product calculations. Moreover, the implementation is shown to be very scalable. This makes it viable for deployment in applications where computational power and memory are limited, such as embedded environments. It enables the possibility for fast on-line optimisation using improved training samples. The approach is based on the concentration of measure effects and stochastic separation theorems and is illustrated with an example on the identification faulty processes in Computer Numerical Control (CNC) milling and with a case study on adaptive removal of false positives in an industrial video surveillance and analytics system. Ivan Tyukin, Alexander N. Gorban, Alistair A. McEwan, Sepehr Meshkinfamfard |
Inf. Sci. | 2 |
| 2019 | One-trial correction of legacy AI systems and stochastic separation theorems
Alexander N. Gorban, Richard Burton, Ilya V. Romanenko, Ivan Tyukin |
Inf. Sci. | 1 |
| 2019 | Fast construction of correcting ensembles for legacy Artificial Intelligence systems: Algorithms and a case study
Ivan Tyukin, Alexander N. Gorban, Stephen Green 0001, Danil V. Prokhorov |
Inf. Sci. | 2 |
| 2018 | Correction of AI systems by linear discriminants: Probabilistic foundations
Alexander N. Gorban, A. Golubkov, Bogdan Grechuk, Eugenij Moiseevich Mirkes, Ivan Tyukin |
Inf. Sci. | 1 |
| 2017 | Attention-Based Extraction of Structured Information from Street View ImageryabstractWe present a neural network model — based on Convolutional Neural Networks, Recurrent Neural Networks and a novel attention mechanism — which achieves 84.2% accuracy on the challenging French Street Name Signs (FSNS) dataset, significantly outperforming the previous state of the art (Smith'16), which achieved 72.46%. Furthermore, our new method is much simpler and more general than the previous approach. To demonstrate the generality of our model, we show that it also performs well on an even more challenging dataset derived from Google Street View, in which the goal is to extract business names from store fronts. Finally, we study the speed/accuracy tradeoff that results from using CNN feature extractors of different depths. Surprisingly, we find that deeper is not always better (in terms of accuracy, as well as speed). Our resulting model is simple, accurate and fast, allowing it to be used at scale on a variety of challenging real-world text extraction problems. Zbigniew Wojna, Alexander N. Gorban, Dar-Shyang Lee, Kevin Murphy 0002, Yeqing Li, Julian Ibarz |
ICDAR | 2 |
| 2016 | SOM: Stochastic initialization versus principal components
Ayodeji A. Akinduko, Eugenij Moiseevich Mirkes, Alexander N. Gorban |
Inf. Sci. | 3 |
| 2016 | Approximation with random bases: Pro et Contra
Alexander N. Gorban, Ivan Tyukin, Danil V. Prokhorov, Konstantin I. Sofeikov |
Inf. Sci. | 1 |
| 2015 | Fast and user-friendly non-linear principal manifold learning by method of elastic mapsabstractMethod of elastic maps allows fast learning of nonlinear principal manifolds for large datasets. We present user-friendly implementation of the method in ViDaExpert software. Equipped with several dialogs for configuring data point representations (size, shape, color) and fast 3D viewer, ViDaExpert is a handy tool allowing to construct an interactive 3D-scene representing a table of data in multidimensional space and perform its quick and insightfull statistical analysis, from basic to advanced methods. We list several recent application examples of manifold learning by method of elastic maps in various fields of life sciences. Alexander N. Gorban, Andrei Yu. Zinovyev |
DSAA | 1 |