VLDB 2026 Research / reviewers in the wild / expert
Vladimir V. Arlazarov
dblp:213/8339
· DBLP profile ↗
58ranked-venue papers
4as first author
18since 2021 · last 2025
0000-0003-3260-9104ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 18 · 10 since 2021Databases, data management, data science and information retrieval · 10 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MIDV-UP: A Dataset of Pakistani and Iranian ID Documents
Yulia S. Chernyshova, Daniil A. Ilyukhin, Vladimir V. Arlazarov |
ICDAR (4) | 3 |
| 2025 | Template-based text field segmentation for ID documents using dynamic squeezeboxes packing
Michael Zingerenko, Elena Limonova, Vladimir V. Arlazarov |
Multim. Tools Appl. | 3 |
| 2024 | An Ultra-lightweight Approach for Machine Readable Zone Detection via Semantic Segmentation and Fast Hough Transform
Daria M. Ershova, Alexander V. Gayer, Alexander Sheshkus, Vladimir V. Arlazarov |
ICDAR (4) | 4 |
| 2024 | Fully Automatic Virtual Unwrapping Method for Documents Imaged by X-Ray Tomography
Petr Kulagin, Dmitry Polevoy, Marina V. Chukalina, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICDAR (3) | 5 |
| 2023 | MIDV-Holo: A Dataset for ID Document Hologram Detection in a Video Stream
L. I. Koliaskina, Ekaterina Emelianova, Daniil V. Tropin, V. V. Popov, Konstantin B. Bulatov, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICDAR (3) | 7 |
| 2023 | Multilanguage ID document images synthesis for testing recognition pipelinesabstractDatasets are de facto the only way to test the recognition pipelines and to compare them with each other. To avoid the manual gathering of documents and, moreover, to avoid problems with the law in the case of ID documents researchers create synthetic datasets or datasets of fake documents, but this process is also time-consuming. In this paper, we present a simple method to use when you need to test a recognition pipeline or some part of it. The method employs only the information that the developers of such pipelines use in their work and allows them to create natural-looking images. The quantitative experiments show that the recognition accuracy of the synthesized images corresponds with the recognition accuracy of the MIDV-2020 dataset. The qualitative comparison also demonstrates that such images can be helpful in recognition systems’ development. Yulia S. Chernyshova, Konstantin K. Suloev, Vladimir V. Arlazarov |
ICMV | 3 |
| 2023 | Threshold U-Net: speed up document binarization with adaptive thresholdsabstractU-Net similar architectures are widely used in the task of document image binarization. However, despite the good quality of binarization, they also have high computational complexity, which greatly limits their use on mobile and embedded devices. The performance bottleneck of U-Net architectures is the first encoder layers and the last decoder layers, which operate on high-resolution input data and contain the largest number of operations. Based on this, in this paper we propose a new Threshold U-Net model: instead of predicting the final image, Threshold U-Net predicts a low-resolution adaptive threshold map, with which the input image is binarized. The proposed architecture naturally combines the ideas of classical algorithms that calculate the binarization threshold for a specific image region with an approach based on a deep learning model with a large receptive field and context understanding. Threshold U-Net demonstrates quality of binarization of historical documents comparable to U-Net on the DIBCO-2017 dataset. At the same time, depending on the resolution of the threshold map, Threshold U-Net is up to 2 times faster, requires up to 26% less RAM and consists up to 10% fewer parameters. Konstantin E. Lihota, Alexander V. Gayer, Vladimir V. Arlazarov |
ICMV | 3 |
| 2023 | Quantization method for bipolar morphological neural networksabstractIn the paper, we present a quantization method for bipolar morphological neural networks. Bipolar morphological neural networks use only addition, subtraction, and maximum operations inside the neuron and exponent and logarithm as activation functions of the layers. These operations allow fast and compact gate implementation for FPGA and ASIC, which makes these networks a promising solution for embedded devices. Quantization allows us to reach an additional increase in computational efficiency and reduce the complexity of hardware implementation by using integer values of low bitwidth for computations. We propose an 8-bit quantization scheme based on integer maximum, addition, and lookup tables for non-linear functions and experimentally demonstrate that basic models for image classification can be quantized without noticeable accuracy loss. More advanced models still provide high recognition accuracy but would benefit from further fine-tuning. Elena Limonova, Michael Zingerenko, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICMV | 4 |
| 2023 | Enhanced multiple-instance pruning for learning soft cascade detectorsabstractObject detection is one of the most common problems solved by computer vision systems. Even though neural network methods have become a standard tool for solving the problems, these methods have many disadvantages, which include high computational power requirements both for training and inference stages and tremendous training sets. This paper considers such a classical method for object detection as the Viola and Jones method and proposes an enhanced soft cascade calibration method based on Multiple-Instance Pruning to increase detection performance. The proposed method considers a response of the classifier to an image region as a random variable and follows a statistical approach to provide robust detectors. In addition, the paper addresses the problem of non-conformity of detection parameters at training and inference stages and studies performance decline. The performance of the proposed methods is demonstrated in a variety of practical tasks, including identity document detection and document fraud detection. Daniil Matalov, Vladimir V. Arlazarov |
ICMV | 2 |
| 2023 | Fast keypoint filtering for feature-based identity documents classification on complex backgroundabstractThe initial steps of many computer vision algorithms are local feature extraction and matching. However, in the problem of recognizing objects in images with complex backgrounds, this approach has a weak point since keypoints may be found not only in the object of interest, but also in the background. This leads to redundant calculations and can cause mismatches. In this paper, we propose a keypoints filtering method applicable to the problem of classification and localization of ID documents in the wild. Using a light-weight deep learning model, keypoints are divided into ”document” and ”background” classes, after which the keypoints of the background are removed. Experimental results show that adding the proposed filtering step gives an average speedup of 3.14% on the entire MIDV-500 dataset and 14.77% on MIDV-2020. At the same time, the acceleration on target images with complex backgrounds reaches 81%. Nargiza Z. Valishina, Alexander V. Gayer, Natalya Skoryukina, Vladimir V. Arlazarov |
ICMV | 4 |
| 2023 | Reducing radiation dose for NN-based COVID-19 detection in helical chest CT using real-time monitored reconstruction
Konstantin B. Bulatov, Anastasia Ingacheva, Marat I. Gilmanov, Marina V. Chukalina, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
Expert Syst. Appl. | 6 |
| 2023 | An accurate approach to real-time machine-readable zone detection with mobile devices
Alexander V. Gayer, Daria M. Ershova, Vladimir V. Arlazarov |
Int. J. Document Anal. Recognit. | 3 |
| 2022 | From tomographic reconstruction to automatic text recognition: the next frontier task for the artificial intelligenceabstractVirtual unrolling or unfolding, digital unwrapping, flattening or unfurling - all these terms are used to describe the process of surface straightening of a tomographically reconstructed digital object. For many objects of historical heritage, tomography is the only way to obtain a hidden image of the original object without its destruction. Digital flattening is no longer considered a unique met hodology. It being applied by many research group, but AI-based methods are used insignificantly in such projects, despite the amazing success of AI in computer vision, in particular optical text recognition. It can be explained by the fact that the success of AI depends on large, broad and high quality datasets, but there are very few published CT-based datasets relevant to the task of digital flattening. Accumulation of a sufficient amount of data necessary for training models is a key point for the next technological breakthrough. In this paper, we present open and cumulative dataset CT-OCR-2022. Dataset includes 6 packages data for different model objects that help to enrich tomographic solutions and to train machine learning models. Each package contains optically scanned image of model objects, 400 measured X-ray projections, 2687 CT- reconstructed cross-sections of 3D reconstructed image, segmentation markups. We believe that CT-OCR-2022 dataset will serve as a benchmark for reconstructed object digital flattening and recognition systems, and that it will prove invaluable for advancement of the field of CT-reconstruction, symbols analysis and recognition. The data presented are openly available in Zenodo at doi:10.5281/zenodo.7123495 and linked repositories. Dmitry Polevoy, Petr Kulagin, Anastasia Ingacheva, Zh. V. Soldatova, Marina V. Chukalina, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICMV | 7 |
| 2022 | Fast matrix multiplication for binary and ternary CNNs on ARM CPUabstractLow-bit quantized neural networks (QNNs) are of great interest in practical applications because they significantly reduce the consumption of both memory and computational resources. Binary neural networks (BNNs) are memory and computationally efficient as they require only one bit per weight and activation and can be computed using Boolean logic and bit count operations. QNNs with ternary weights and activations (TNNs) and binary weights and ternary activations (TBNs) aim to improve recognition quality compared to BNNs while preserving low bit-width. However, their efficient implementation is usually considered on ASICs and FPGAs, limiting their applicability in real-life tasks. At the same time, one of the areas where efficient recognition is most in demand is recognition on mobile devices using their CPUs. However, there are no known fast implementations of TBNs and TNN, only the daBNN library for BNNs inference. In this paper, we propose novel fast algorithms of ternary, ternary-binary, and binary matrix multiplication for mobile devices with ARM architecture. In our algorithms, ternary weights are represented using 2-bit encoding and binary - using one bit. It allows us to replace matrix multiplication with Boolean logic operations that can be computed on 128-bits simultaneously, using ARM NEON SIMD extension. The matrix multiplication results are accumulated in 16-bit integer registers. We also use special reordering of values in left and right matrices. All that allows us to efficiently compute a matrix product while minimizing the number of loads and stores compared to the algorithm from daBNN. Our algorithms can be used to implement inference of convolutional and fully connected layers of TNNs, TBNs, and BNNs. We evaluate them experimentally on ARM Cortex-A73 CPU and compare their inference speed to efficient implementations of full-precision, 8-bit, and 4-bit quantized matrix multiplications. Our experiment shows our implementations of ternary and ternary-binary matrix multiplications to have almost the same inference time, and they are 3.6 times faster than full-precision, 2.5 times faster than 8-bit quantized, and 1.4 times faster than 4-bit quantized matrix multiplication but 2.9 slower than binary matrix multiplication. Anton Trusov, Elena Limonova, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICPR | 4 |
| 2021 | Determining Optimal Frame Processing Strategies for Real-Time Document Recognition Systems
Konstantin B. Bulatov, Vladimir V. Arlazarov |
ICDAR (2) | 2 |
| 2021 | MIDV-LAIT: A Challenging Dataset for Recognition of IDs with Perso-Arabic, Thai, and Indian Scripts
Yulia S. Chernyshova, Ekaterina Emelianova, Alexander Sheshkus, Vladimir V. Arlazarov |
ICDAR (2) | 4 |
| 2021 | RFDoc: Memory Efficient Local Descriptors for ID Documents Localization and Classification
Daniil Matalov, Elena Limonova, Natalya Skoryukina, Vladimir V. Arlazarov |
ICDAR (2) | 4 |
| 2021 | Character sequence prediction method for training data creation in the task of text recognitionabstractFor text line recognition, much attention is paid to augmentation of the training images. Yet the inner structure of the textual information in the images also affects the accuracy of the resulting model. In this paper, we propose an ANNbased method for textual data generation for printing in images with a background of a synthetic training sample. In our method we avoid the usage of completely random sequences as well as the dictionary-based ones. As a result, we gain the data that saves the basic properties of the target language model, such as the balance of vowels and consonants, but avoid the lexicon-based properties, like the prevalence of the specific characters. Moreover, as our method focuses only on high-levels features and does not try to generate the real words, we can use a small training sample and light-weight ANN for text generation. To check our method, we train three ANNs with same architecture, but with different training samples. We choose machine readable zones as a target field because of their structure that does not correspond with the ordinary lexicon. The results of the experiments on three public datasets of identity documents demonstrate the effectiveness of our method and allows to enhance the state-of-the art results for the target field. Pavel K. Zlobin, Alexander Sheshkus, Vladimir V. Arlazarov |
ICMV | 3 |
| 2020 | Generative approach for 1D barcode dataset population for mobile-based recognitionabstractBarcode recognition via the mobile device camera is an actual problem which occurs in various fields. To receive experimental baselines and provide methods comparison researchers require some carefully annotated data samples. For those purposes, some datasets have already been collected and published in the public domain. But they suffer from different disadvantages. In this paper, we present a novel challenging dataset designed for barcode reading quality evaluation. It is populated using the generative approach with well-established image augmentation technique application. Among them, different geometrical transformations, kind of noises, brightness and lighting variance were used. The generative approach also allowed to populate ground-truth automatically and exclude errors introduced during the manual process. The dataset consists of training (41184 images) and validation parts (10296 images). It contains seven most popular symbologies: CODABAR, CODE-39, CODE-93, CODE-128, EAN-13, UPC-A, UPC-E. As an experimental baseline, the reading quality obtained with the open-source Zxing library is provided. The dataset with all the supplementary materials is available at ftp://smartengines.com/barcode. Vladimir V. Arlazarov |
ICMV | 1 |
| 2020 | An approach to road scene text recognition with per-frame accumulation and dynamic stopping decisionabstractCamera-based road scene analysis is an important task for building driving assistance systems and autonomous vehicles. An crucial component of road scene analysis is detection, tracking, and recognition of text object. In this paper, we consider the recognition of road scene text objects in sequences of video frames, and propose an approach to per-frame recognition results accumulation with a dynamic stopping decision. Experimental evaluation on an open dataset RoadText-1K showed that the proposed approach allows to achieve mean lower recognition error for the same mean number of processed frames, and significantly reduce the number of text objects which have to be recognized in each frame, thus relieving the load on the computational unit. Vladimir V. Arlazarov |
ICMV | 1 |
| 2020 | Empirical analysis of the optimality of RSRE-based stopping rules for monitored reconstructionabstractOne of the challenges present in the field of tomographic imaging is the reduction of the radiation dose imparted to the object. Monitored reconstruction is one of the approaches to reduce the dose by means of dynamic stopping of the scanning process. In this paper, analysis was performed for the RSRE-based stopping rules for monitored reconstruction, using the previously analyzed RSRE-based reconstruction quality metrics, a widely-used PSNR metric, as well as quality metrics designed to mimic human perception, such as SSIM and ISSIM. It was shown that the unnormalized RSRE-based stopping rule performs better than the baseline for SSIM, ISSIM, and PSNR, and that for the latter the best result was achieved using the stopping rule designed for the reconstruction quality metric normalized to the Radon invariant. The stopping rules were compared with a synthetic a-posteriori “perfect” stopping rule and it was shown that the RSRE-based stopping rules closely approach the perfect stopping rule for the RSRE metric normalized on the Radon invariant, as well as for the ISSIM metric. Vladimir V. Arlazarov |
ICMV | 1 |
| 2020 | Improvement of U-Net architecture for image binarization with activation functions replacementabstractIn this work we study the effect of activation functions in a neural network. We consider how activation functions with different properties and their combination affect the final quality of the model. Due to optimization and speed performance issues with most of bounded functions that are represented by sigmoids, we propose the generalized version of SoftSign function - ratio function (rf). Its shape greatly depends on introduced degree parameter, which in theory leads to new interesting property - contraction to zero. For evaluation, we chose image binarization problem: based on UNet architecture of DIBCO-2017 winners, we conducted all experiments with replacing activation functions only. Our research has led us to the state-of-the-art results in binarization quality on DIBCO-2017 test dataset. U-Net with modified activation functions significantly outperforms all existing solutions in all metrics. Alexander V. Gayer, Alexander Sheshkus, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICMV | 4 |
| 2020 | Block convolutional layer for position dependent features calculationabstractImage recognition includes problems where special features can be found only in a specific area of an image. This fact suggests us to apply different filters to different areas of input images. Convolutional networks have only fully-connected and locally-connected layers to make it. A Fully-connected layer erases the position factor for every output and a locally connected layer storage an enormous number of parameters. We need a layer that can apply different convolution kernels for different areas of an input image and not carry so many parameters as a locally-connected layer for high scale resolution images. This is why in this paper, we introduce a new type of convolutional layer - a block layer, and a way to construct a neural network using block convolutional layers to achieve better performance in the image classification problem. The influence of block layers on the quality of the neural network classifier is shown in this paper. We also provide a comparison with neural network architecture LeNet-5 as a baseline. The research was conducted on open datasets: MNIST, CIFAR-10, Fashion MNIST. The results of our research prove that this layer can increase the accuracy of neural network classifiers without increasing the number of operations for the neural network. Sergey A. Ilyuhin, Alexander Sheshkus, Vladimir V. Arlazarov |
ICMV | 3 |
| 2020 | Distance-based online pairs generation method for metric networks trainingabstractIn this work, we consider the pairs generation algorithm based on the distances between elements in metric space. The right generation of training data is an actual issue, and its solution leads to better neural network learning. Understanding the properties of the source data, we can select pairs for training in such a way that the network will pay more attention to elements that are close in the metric space and have different classes. However, the problem arises when these properties are difficult to extract from the data and a more universal pairs generation method is needed. Our method generates pairs using the results of the network from previous iterations, in parallel with the training process itself. Thus, we do not need to evaluate the properties of elements ourselves, and we can use absolutely any data as learning objects. We demonstrate this approach using the example of Korean character recognition, and also compare it with other commonly used pair generation methods. Ivan V. Kondrashev, Alexander Sheshkus, Vladimir V. Arlazarov |
ICMV | 3 |
| 2020 | Bipolar morphological U-Net for document binarizationabstractDeep neural networks are widely used in various AI systems. Many such systems rely on the edge computing concept and try to perform computations on end devices while still being energy and memory efficient. Therefore, substantial time and memory requirements are imposed on neural networks. One way to improve neural network efficiency is to simplify computations inside a neuron. A bipolar morphological neuron uses only addition, subtraction, and maximum operations inside the neuron and exponent and logarithm as activation functions for the network layers. These operations allow fast and compact gate implementation for FPGA and ASIC. In the paper, we consider the usage of bipolar morphological (BM) networks for document binarization. We examine the DIBCO 2017 binarization challenge and train the bipolar morphological convolutional neural network of U-Net architecture. Despite some accuracy decrease for a model with all BM convolutional layers, one can flexibly control the accuracy by using the partially converted model. It should be noted that even the fully BM model is suitable for solving the problem in practice. Elena Limonova, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICMV | 3 |
| 2020 | About Viola-Jones image classifier structure in the problem of stamp detection in document imagesabstractIn this paper we explore a set of modifications of the cascade structure of the Viola-Jones detector on the example of solving stamp detection problem. The experiments on the public “SPODS” dataset for various document attributes extraction problems with extremely limited training set are presented. The positive training set is augmented by applying various image processing algorithms relevant to the stamp model to an available in a single instance image for each stamp type. We describe and analyze such structures of the Viola and Jones classifiers as the original cascade structure, tree, soft cascade, and perform the training experiments. Experimental results show that each modification of the cascade structure of the classifier has its own advantages and disadvantages, and the choice of the Viola-Jones classifier design significantly affects the quality of solving object detection problem. Daniil Matalov, Sergey A. Usilin, Vladimir V. Arlazarov |
ICMV | 3 |
| 2020 | Memory consumption reduction for identity document classification with local and global features combinationabstractIn this paper we explore possibilities of memory cost reduction without significant loss of classification accuracy in connection with the problem of the ID document type recognition on mobile devices. The studied classic approach is based on representing images using constellation of feature points and descriptors. The distortion parameters are estimated by applying RANSAC. Experimental data details the approach limitations (memory, speed and accuracy) in dependence of the descriptor type. In order to maintain accuracy when using low dimensional descriptors we suggest to modify the basic approach using additional features characteristic of the document such as straight lines and quadrangles. In addition, an early filtration of the samples and the hypotheses used in RANSAC. It was shown that the proposed modifications have a positive contribution for all types of descriptors considered. The suggested algorithm was tested using the open dataset MIDV-500. The modified approach allows to achieve an accuracy improvement and significant speed up of distortion parameters estimation in RANSAC. It was shown that using compact descriptors in conjunction with the presented method allows reduce required memory cost by more than 7 times with near-zero (0.2%) loss of accuracy, and more than 14 times with the loss of accuracy is about 18%. Natalya Skoryukina, Vladimir V. Arlazarov |
ICMV | 2 |
| 2020 | The method of search for falsifications in copies of contractual documents based on N-gramsabstractThis article is focused on methods of search for falsifications in scanned copies of business documents. This task arises from a comparison of two copies of business documents signed by two parties. The comparison should be performed to detect possible changes made by one of the parties. This problem is relevant, for instance, in the banking sector when signing agreements on paper. The method of partial search for matching flexible documents, where text attributes may be changed, and unintentional modifications of non-essential words may be made is considered. The method of comparison of two scanned images based on the recognition and analysis of N-grams word sequences is proposed. The proposed method has been tested on private dataset. The proposed method has demonstrated high quality and reliability of the search for differences in two samples of one agreement-type document. Oleg A. Slavin, Elena Andreeva 0005, Vladimir V. Arlazarov |
ICMV | 3 |
| 2020 | Improved algorithm of ID card detection by a priori knowledge of the document aspect ratioabstractIn this work, we consider a problem of quadrilateral document borders detection in images captured by a mobile device’s camera. State-of-the-art algorithms for the quadrilateral document borders detection are not designed for cases when one of the document borders is either completely out of the frame, obscured, or of low contrast. We propose the algorithm which correctly processes the image in such cases. It is built on the classical contour-based algorithm. We modify the latter using the document’s aspect ratio which is known a priori. We demonstrate that this modification reduces the number of incorrect detections by 34% on an open dataset MIDV-500. Daniil V. Tropin, Ivan A. Konovalenko, Natalya Skoryukina, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICMV | 5 |
| 2020 | Fast Approximate Modelling of the Next Combination Result for Stopping the Text Recognition in a VideoabstractIn this paper, we consider a task of stopping the video stream recognition process of a text field, in which each frame is recognized independently and the individual results are combined together. The video stream recognition stopping problem is an under-researched topic with regards to computer vision, but its relevance for building high-performance video recognition systems is clear. Firstly, we describe an existing method of optimally stopping such a process based on a modelling of the next combined result. Then, we describe approximations and assumptions which allowed us to build an optimized computation scheme and thus obtain a method with reduced computational complexity. The methods were evaluated for the tasks of document text field recognition and arbitrary text recognition in a video. The experimental comparison shows that the introduced approximations do not diminish the quality of the stopping method in terms of the achieved combined result precision, while dramatically reducing the time required to make the stopping decision. The results were consistent for both text recognition tasks. Konstantin B. Bulatov, Nadezhda Fedotova, Vladimir V. Arlazarov |
ICPR | 3 |
| 2020 | ResNet-like Architecture with Low Hardware RequirementsabstractOne of the most computationally intensive parts in modern recognition systems is an inference of deep neural networks that are used for image classification, segmentation, enhancement, and recognition. The growing popularity of edge computing makes us look for ways to reduce its time for mobile and embedded devices. One way to decrease the neural network inference time is to modify a neuron model to make it more efficient for computations on a specific device. The example of such a model is a bipolar morphological neuron model. The bipolar morphological neuron is based on the idea of replacing multiplication with addition and maximum operations. This model has been demonstrated for simple image classification with LeNet-like architectures [1]. In the paper, we introduce a bipolar morphological ResNet (BM-ResNet) model obtained from a much more complex ResNet architecture by converting its layers to bipolar morphological ones. We apply BM-ResNet to image classification on MNIST and CIFAR-10 datasets with only a moderate accuracy decrease from 99.3% to 99.1 % and from 85.3% to 85.1 %. We also estimate the computational complexity of the resulting model. We show that for the majority of ResNet layers, the considered model requires 2.1-2.9 times fewer logic gates for implementation and 15-30 % lower latency. Elena Limonova, Daniil Alfonso, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICPR | 4 |
| 2020 | Approach for Document Detection by Contours and ContrastsabstractThis paper considers arbitrary document detection performed on a mobile device. The classical contour-based approach often fails in cases featuring occlusion, complex background, or blur. The region-based approach, which relies on the contrast between object and background, does not have application limitations, however, its known implementations are highly resource-consuming. We propose a modification of the contour-based method, in which the competing contour location hypotheses are ranked according to the contrast between the areas inside and outside the border. In the experiments, such modification allows for the decrease of alternatives ordering errors by 40% and the decrease of the overall detection errors by 10%. The proposed method provides unmatched state-of-the-art performance on the open MIDV-500 dataset, and it demonstrates results comparable with state-of-the-art performance on the SmartDoc dataset. Daniil V. Tropin, Sergey A. Ilyuhin, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICPR | 4 |
| 2020 | Fast Implementation of 4-bit Convolutional Neural Networks for Mobile DevicesabstractQuantized low-precision neural networks are very popular because they require less computational resources for inference and can provide high performance, which is vital for real-time and embedded recognition systems. However, their advantages are apparent for FPGA and ASIC devices, while general-purpose processor architectures are not always able to perform low-bit integer computations efficiently. The most frequently used low-precision neural network model for mobile central processors is an 8-bit quantized network. However, in a number of cases, it is possible to use fewer bits for weights and activations, and the only problem is the difficulty of efficient implementation. We introduce an efficient implementation of 4-bit matrix multiplication for quantized neural networks and perform time measurements on a mobile ARM processor. It shows 2.9 times speedup compared to standard floating-point multiplication and is 1.5 times faster than 8-bit quantized one. We also demonstrate a 4-bit quantized neural network for OCR recognition on the MIDV-500 dataset. 4-bit quantization gives 95.0% accuracy and 48% overall inference speedup, while an 8-bit quantized network gives 95.4% accuracy and 39% speedup. The results show that 4-bit quantization perfectly suits mobile devices, yielding good enough accuracy and low inference time. Anton Trusov, Elena Limonova, Dmitry Slugin, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICPR | 5 |
| 2019 | HoughNet: Neural Network Architecture for Vanishing Points DetectionabstractIn this paper we introduce a novel neural network architecture based on Fast Hough Transform layer. The layer of this type allows our neural network to accumulate features from linear areas across the entire image instead of local areas. We demonstrate its potential by solving the problem of vanishing points detection in the images of documents. Such problem occurs when dealing with camera shots of the documents in uncontrolled conditions. In this case, the document image can suffer several specific distortions including projective transform. To train our model, we use MIDV-500 dataset and provide testing results. Strong generalization ability of the suggested method is proven with its applying to a completely different ICDAR 2011 dewarping contest. In previously published papers considering this dataset authors measured quality of vanishing point detection by counting correctly recognized words with open OCR engine Tesseract. To compare with them, we reproduce this experiment and show that our method outperforms the state-of-the-art result. Alexander Sheshkus, Anastasia Ingacheva, Vladimir V. Arlazarov, Dmitry P. Nikolaev |
ICDAR | 3 |
| 2019 | Fast Method of ID Documents Location and Type Identification for Mobile and Server ApplicationabstractIn this paper we discuss the problem of simultaneous document type recognition and projective distortion parameters estimation for the images of ID documents. There are two considered cases. In the first case a video stream captured using mobile devices is processed on the device. The second case considers photos or scanned images which are processed on a server. For each case the requirements are defined for the input data and processing speed. The universal approach is proposed, which allows solving the problem in both cases. The approach is based on representing the image as a constellation of feature points and descriptors, but in order to perform more accurate distortion parameters estimation straight lines and quadrangles are extracted from the input image and used as additional features. Techniques are described which allow to combine matched feature points, lines, and quadrangles to geometric verification using RANSAC. Best alternative selection criteria are proposed along with methods of solution accuracy estimation. The differences between methods of preliminary analysis of the input image and geometric primitives location are discussed in relation to the considered problems. For quality estimation an open dataset MIDV-500 is used, together with its extension for server-side problem version, created in scope of this work. Results show that using lines and quadrangles increase the location accuracy, and the proposed algorithm surpasses previously published works in classification precision and computational performance. Natalya Skoryukina, Vladimir V. Arlazarov, Dmitry P. Nikolaev |
ICDAR | 2 |
| 2019 | Comparison of scanned administrative document imagesabstractIn this work the methods of comparison of digitized copies of administrative documents were considered. This problem arises, for example, when comparing two copies of documents signed by two parties in order to find possible modifications made by one party, in the banking sector at the conclusion of contracts in paper form. The proposed method of document image comparison is based on a combination of several ways of image comparison of words that are descriptors of text feature points. Testing was conducted on public Payslip Dataset (French). The results showed the high quality and the reliability of finding differences in two images that are versions of the same document. Elena Andreeva 0005, Vladimir V. Arlazarov, Oleg A. Slavin, Aleksey Mishev |
ICMV | 2 |
| 2019 | MIDV-2019: challenges of the modern mobile-based document OCRabstractRecognition of identity documents using mobile devices has become a topic of a wide range of computer vision research. The portfolio of methods and algorithms for solving such tasks as face detection, document detection and rectification, text field recognition, and other, is growing, and the scarcity of datasets has become an important issue. One of the openly accessible datasets for evaluating such methods is MIDV-500, containing video clips of 50 identity document types in various conditions. However, the variability of capturing conditions in MIDV-500 did not address some of the key issues, mainly significant projective distortions and different lighting conditions. In this paper we present a MIDV-2019 dataset, containing video clips shot with modern high-resolution mobile cameras, with strong projective distortions and with low lighting conditions. The description of the added data is presented, and experimental baselines for text field recognition in different conditions. Konstantin B. Bulatov, Daniil Matalov, Vladimir V. Arlazarov |
ICMV | 3 |
| 2019 | Next integrated result modelling for stopping the text field recognition process in a video using a result model with per-character alternativesabstractIn the field of document analysis and recognition using mobile devices for capturing, and the field of object recognition in a video stream, an important problem is determining the time when the capturing process should be stopped. Efficient stopping influences not only the total time spent for performing recognition and data entry, but the expected accuracy of the result as well. This paper is directed on extending the stopping method based on next integrated recognition result modelling, in order for it to be used within a string result recognition model with per-character alternatives. The stopping method and notes on its extension are described, and experimental evaluation is performed on an open dataset MIDV-500. The method was compares with previously published methods based on input observations clustering. The obtained results indicate that the stopping method based on the next integrated result modelling allows to achieve higher accuracy, even when compared with the best achievable configuration of the competing methods. Konstantin B. Bulatov, Boris Savelyev, Vladimir V. Arlazarov |
ICMV | 3 |
| 2019 | Training the convolutional neural network with statistical dependence of the response on the input data distortionabstractThe paper proposes an approach to training a convolutional neural network using information on the level of distortion of input data. The learning process is modified with an additional layer, which is subsequently deleted, so the architecture of the original network does not change. As an example, the LeNet5 architecture network with training data based on the MNIST symbols and a distortion model as Gaussian blur with a variable level of distortion is considered. This approach does not have quality loss of the network and has a significant error-free zone in responses on the test data which is absent in the traditional approach to training. The responses are statistically dependent on the level of input image’s distortions and there is a presence of a strong relationship between them. Igor Janiszewski, Vladimir V. Arlazarov, Dmitry Slugin |
ICMV | 2 |
| 2019 | Bipolar morphological neural networks: convolution without multiplicationabstractIn the paper we introduce a novel bipolar morphological neuron and bipolar morphological layer models. The models use only such operations as addition, subtraction and maximum inside the neuron and exponent and logarithm as activation functions for the layer. The proposed models unlike previously introduced morphological neural networks approximate the classical computations and show better recognition results. We also propose layer-by-layer approach to train the bipolar morphological networks, which can be further developed to an incremental approach for separate neurons to get higher accuracy. Both these approaches do not require special training algorithms and can use a variety of gradient descent methods. To demonstrate efficiency of the proposed model we consider classical convolutional neural networks and convert the pre-trained convolutional layers to the bipolar morphological layers. Seeing that the experiments on recognition of MNIST and MRZ symbols show only moderate decrease of accuracy after conversion and training, bipolar neuron model can provide faster inference and be very useful in mobile and embedded systems. Elena Limonova, Daniil Matveev, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICMV | 4 |
| 2019 | Single-sample augmentation framework for training Viola-Jones classifiersabstractIn this paper we present a single-sample augmentation framework. The key idea of the framework consists of synthesizing a positive training set from a single natural sample using relevant geometric and pixel intensity transforms. The efficiency of the proposed framework has been demonstrated solving round seal stamp detection problem using Viola-Jones approach on the public “SPODS” dataset. The mentioned image transformations make it possible to simulate different orientation of the stamps, color differences, and distortions caused by stamping process and document aging. The proposed framework can be applied to training various machine learning algorithms for solving computer vision and computed tomography problems. Daniil Matalov, Sergey A. Usilin, Vladimir V. Arlazarov |
ICMV | 3 |
| 2019 | Impact of geometrical restrictions in RANSAC sampling on the ID document classificationabstractIn this paper we explore the impact of geometrical restrictions in RANSAC sampling on the ID document type recognition accuracy in images, as well as on the accuracy of the projective distortion parameters estimation. The studied method is based on representing images as constellations of keypoints and their descriptors. The distortion parameters are estimated by applying RANSAC on the matched keypoints. Cases are studied where the base algorithm can yield erroneous or insufficiently accurate solution. A RANSAC scheme is presented with geometrical restrictors and several restriction are proposed, limiting the samples and the computed transform parameters. An experiment was conducted on the open dataset MIDV-500 and the data is presented of the dependence of classification and localization accuracy on the considered restrictors. It was shown that the introduction of restrictors allows to achieve a accuracy improvement and significant speed up. Natalya Skoryukina, Igor Faradjev, Konstantin B. Bulatov, Vladimir V. Arlazarov |
ICMV | 4 |
| 2019 | Fast approach for QR code localization on images using Viola-Jones methodabstractIn this paper, a method for QR Code localization on images obtained under uncontrolled environment is presented. The proposed method is a modified Viola-Jones object detection method in which features are calculated over the directional edge image, and a tree classifier is used instead of cascade classifier. The experiments show that the use of the QR Code localization method described in the paper can significantly improve the quality of the existing decoding algorithms. The high performance of the developed method makes it possible to use it in various real-time recognition systems. Sergey A. Usilin, Pavel Bezmaternykh, Vladimir V. Arlazarov |
ICMV | 3 |
| 2019 | On optimal stopping strategies for text recognition in a video stream as an application of a monotone sequential decision modelabstractThe paper describes the problem of stopping the text field recognition process in a video stream, which is a novel problem, particularly relevant to real-time mobile document recognition systems. A decision-theoretic framework for this problem is provided, and similarities with existing stopping rule problems are explored. Following the theoretical works on monotone stopping rule problems, a strategy is proposed based on thresholding the estimation of the expected difference between consequent recognition results. The efficiency of this strategy is evaluated on an openly accessible dataset. The results show that this method outperforms the previously published methods based on identical results cluster size thresholding. Notes on future work include incorporation of recognition result confidence estimations in the proposed model and more precise evaluation of the observation cost. Konstantin B. Bulatov, Nikita Razumnyi, Vladimir V. Arlazarov |
Int. J. Document Anal. Recognit. | 3 |
| 2018 | Experimental modeling the flow of character recognition results in video stream for document recognitionabstractThis paper considers problems regarding the development of stochastic models consistent with the results of character image recognition in video stream. Assumptions about their structure and properties are formulated for the constructed models. The description of the model components defines the Dirichlet distribution and its generalizations. The parameters of these distributions are determined using statistical estimation methods. The Akaike information criterion is used to rank models. The verification of the agreement of the proposed theoretical distributions to the sample data is carried out. Elena Andreeva 0005, Vladimir V. Arlazarov, Oleg A. Slavin, Igor Janiszewski |
ICMV | 2 |
| 2018 | Application of dynamic saliency maps to the video stream recognition systems with image quality assessmentabstractIn this paper we propose an original approach to optical video stream recognition system design, introducing new modules for image quality assessment, image quality correction and feedback. The main novelty of the proposed approach lays in combining image quality assessment results with the global dynamic object saliency maps which indicate the importance or the informative value of the corresponding image regions. The approach is applied to the identity documents video stream recognition system, where saliency maps are initially provided by document templates and are dynamically changing over time – for example, according to per-field stopping rules. Experiments demonstrated an increase of such essential recognition systems characteristics as accuracy, reliability and performance. Timofey S. Chernov, Sergey A. Ilyuhin, Vladimir V. Arlazarov |
ICMV | 3 |
| 2018 | Modification of the Viola-Jones approach for the detection of the government seal stamp of the Russian FederationabstractIn this paper we present modification of the Viola-Jones approach for solving government seal stamp of the Russian Federation detection problem. The main contributions of the proposed modification are combining brightness and edge features as well as using L1 norm of the gradient of the image for calculating edge features. This modification allows to build classifiers which are more robust to noise, absence of a characteristic structure of contrasts and object's boundaries. The modification is experimentally compared to original Viola-Jones algorithm and showing better quality on different testing sets. Daniil Matalov, Sergey A. Usilin, Vladimir V. Arlazarov |
ICMV | 3 |
| 2018 | Viability of Viola-Jones method for the problem of image classificationabstractIn this paper we study combination of Viola-Jones classifier with deep convolutional neural network as an approach to the problem of object detection and classification. It is well known that Viola-Jones detectors are fast and accurate in detection of vast variety of different objects. On the other hand, methods based on neural network usage demonstrate high accuracy in the problems of image classification. The main goal of this paper is to study viability of Viola-Jones classifier in problem of image classification. The first part of both algorithms is the same: we will use Viola-Jones classifier to find object bounding rectangle in the image. The second part of the algorithms is different: we will compare usage of Viola-Jones classifier with convolutional neural network-based classifier. We will provide speed and accuracy comparison between these two algorithms. Alexander Sheshkus, Daniil Matalov, Vladimir V. Arlazarov, Dmitry P. Nikolaev |
ICMV | 3 |
| 2018 | 2D art recognition in uncontrolled conditions using one-shot learningabstractThe paper considers the problem of 2D art identification in photos acquired with mobile devices under the conditions of museum exhibition. The proposed approach is based on a compact description of an image with a constellation of keypoints and corresponding local descriptors. A two-step comparison scheme is described for finding the best reference image matching the query. Bag-of-features approach is used as a first step, then mutual disposition of points is analyzed. Rejection of the query is performed if no suitable matches are found. Geometrical normalization of the query image is proposed to achieve higher robustness against scale and viewpoint variations. After the normalization, mutual disposition of points is estimated using a simplified geometric model. Advantages of the described approach over state-of-the-art solutions are considered. The results of the experiments conducted on the open WikiArt dataset are presented along with processing times for different hardware platforms. Natalya Skoryukina, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICMV | 3 |
| 2017 | Segments Graph-Based Approach for Document Capture in a Smartphone Video StreamabstractThe paper is devoted to the analysis of the problem of document boundaries detection in images and in a video stream. The paper proposes an algorithm for obtaining the position of the document, consisting of very reliable segments of a document boundaries extraction and a construction of an intersection graph that satisfies the projective model of the rectangle. An online algorithm for selecting and integrating possible document positions in a video stream based on the Kalman filter is proposed. The analysis of possible modifications of the algorithm and their effect on the final result are provided. Evaluation of the quality of the document at ICDAR'15 Smartphone Document Capture competition's dataset [1] showed a mean result of 95.5% in Jaccard index of projectively corrected document quadrangles and a 3rd place in the competition. Alexander Zhukovsky, Dmitry P. Nikolaev, Vladimir V. Arlazarov, Vasiliy V. Postnikov, Dmitry Polevoy, Natalya Skoryukina, Timofey S. Chernov, Julia Shemiakina, Arseniy P. Mukovozov, Ivan A. Konovalenko, Mikhail Povolotsky |
ICDAR | 3 |
| 2017 | Comparison of the scanned pages of the contractual documentsabstractIn this paper the problem statement is given to compare the digitized pages of the official papers. Such problem appears during the comparison of two customer copies signed at different times between two parties with a view to find the possible modifications introduced on the one hand. This problem is a practically significant in the banking sector during the conclusion of contracts in a paper format. The method of comparison based on the recognition, which consists in the comparison of two bag-of-words, which are the recognition result of the master and test pages, is suggested. The described experiments were conducted using the OCR Tesseract and the siamese neural network. The advantages of the suggested method are the steady operation of the comparison algorithm and the high exacting precision, and one of the disadvantages is the dependence on the chosen OCR. Elena Andreeva 0005, Vladimir V. Arlazarov, Temudzhin Manzhikov, Oleg A. Slavin |
ICMV | 2 |
| 2017 | Method of determining the necessary number of observations for video stream documents recognitionabstractThis paper discusses a task of document recognition on a sequence of video frames. In order to optimize the processing speed an estimation is performed of stability of recognition results obtained from several video frames. Considering identity document (Russian internal passport) recognition on a mobile device it is shown that significant decrease is possible of the number of observations necessary for obtaining precise recognition result. Vladimir V. Arlazarov, Konstantin B. Bulatov, Temudzhin Manzhikov, Oleg A. Slavin, Igor Janiszewski |
ICMV | 1 |
| 2017 | Image quality assessment for video stream recognition systemsabstractRecognition and machine vision systems have long been widely used in many disciplines to automate various processes of life and industry. Input images of optical recognition systems can be subjected to a large number of different distortions, especially in uncontrolled or natural shooting conditions, which leads to unpredictable results of recognition systems, making it impossible to assess their reliability. For this reason, it is necessary to perform quality control of the input data of recognition systems, which is facilitated by modern progress in the field of image quality evaluation. In this paper, we investigate the approach to designing optical recognition systems with built-in input image quality estimation modules and feedback, for which the necessary definitions are introduced and a model for describing such systems is constructed. The efficiency of this approach is illustrated by the example of solving the problem of selecting the best frames for recognition in a video stream for a system with limited resources. Experimental results are presented for the system for identity documents recognition, showing a significant increase in the accuracy and speed of the system under simulated conditions of automatic camera focusing, leading to blurring of frames. Timofey S. Chernov, Nikita P. Razumnuy, Alexander S. Kozharinov, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICMV | 5 |
| 2017 | Performance improvement of multi-class detection using greedy algorithm for Viola-Jones cascade selectionabstractThis paper aims to study the problem of multi-class object detection in video stream with Viola-Jones cascades. An adaptive algorithm for selecting Viola-Jones cascade based on greedy choice strategy in solution of the N-armed bandit problem is proposed. The efficiency of the algorithm on the problem of detection and recognition of the bank card logos in the video stream is shown. The proposed algorithm can be effectively used in documents localization and identification, recognition of road scene elements, localization and tracking of the lengthy objects , and for solving other problems of rigid object detection in a heterogeneous data flows. The computational efficiency of the algorithm makes it possible to use it both on personal computers and on mobile devices based on processors with low power consumption. Alexander A. Tereshin, Sergey A. Usilin, Vladimir V. Arlazarov |
ICMV | 3 |
| 2016 | Fast integer approximations in convolutional neural networks using layer-by-layer trainingabstractThis paper explores method of layer-by-layer training for neural networks to train neural network, that use approximate calculations and/or low precision data types. Proposed method allows to improve recognition accuracy using standard training algorithms and tools. At the same time, it allows to speed up neural network calculations using fast-processed approximate calculations and compact data types. We consider 8-bit fixed-point arithmetic as the example of such approximation for image recognition problems. In the end, we show significant accuracy increase for considered approximation along with processing speedup. Dmitry Ilin, Elena Limonova, Vladimir V. Arlazarov, Dmitry P. Nikolaev |
ICMV | 3 |
| 2016 | Slant rectification in Russian passport OCR system using fast Hough transformabstractIn this paper, we introduce slant detection method based on Fast Hough Transform calculation and demonstrate its application in industrial system for Russian passports recognition. About 1.5% of this kind of documents appear to be slant or italic. This fact reduces recognition rate, because Optical Recognition Systems are normally designed to process normal fonts. Our method uses Fast Hough Transform to analyse vertical strokes of characters extracted with the help of x-derivative of a text line image. To improve the quality of detector we also introduce field grouping rules. The resulting algorithm allowed to reach high detection quality. Almost all errors of considered approach happen on passports of nonstandard fonts, while slant detector works in appropriate way. Elena Limonova, Pavel Bezmaternykh, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICMV | 4 |
| 2016 | Snapscreen: TV-stream frame search with projectively distorted and noisy queryabstractIn this work we describe an approach to real-time image search in large databases robust to variety of query distortions such as lighting alterations, projective distortions or digital noise. The approach is based on the extraction of keypoints and their descriptors, random hierarchical clustering trees for preliminary search and RANSAC for refining search and result scoring. The algorithm is implemented in Snapscreen system which allows determining a TV-channel and a TV-show from a picture acquired with mobile device. The implementation is enhanced using preceding localization of screen region. Results for the real-world data with different modifications of the system are presented. Natalya Skoryukina, Timofey S. Chernov, Konstantin B. Bulatov, Dmitry P. Nikolaev, Vladimir V. Arlazarov |
ICMV | 5 |
| 2015 | Segments graph-based approach for smartphone document captureabstractDocument capture with a smartphone camera is already here to stay. Interactive applications for document capture and its enhancement have filled mobile application stores. However, discounting the predictions and judging only from the experience of using such applications, they are not yet ready to compete with stationary scanners when high quality and reliability is required. This paper is devoted to analysis of the problem of document detection in the image and evaluation of the quality of existing mobile applications. Based on this analysis we present a new reliable algorithm for document capture, based on the boundary segments detection and constructing a segments graph to fit rectangular projective model. The algorithm achieves about 95% quality of document detection and outperforms all of the reviewed algorithms, implemented in mobile applications. Alexander Zhukovsky, Vladimir V. Arlazarov, Vasiliy V. Postnikov, Valeriy E. Krivtsov |
ICMV | 2 |