Dmitry P. Nikolaev

dblp:88/389 · also Dmitrii P. Nikolaev, Dmitry Nikolaev 0001 · DBLP profile ↗
← Back
94ranked-venue papers
3as first author
12since 2021 · last 2025
0000-0001-5560-7668ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 74 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 24 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021
YearPublicationVenuePosition
2025 On The Robustness Of A Rectification Algorithm For Documents With A Single Crease
abstract
Modern document digitization systems tend to widely apply recognition tasks for camera-captured images. Documents in such images may appear too distorted for an OCR system to be executed on it straightforwardly. This creates the need for an image rectification step. It is particularly essential if the document in the image is mechanically distorted. One of the most common real-life document distortions is the presence of a single arbitrary crease on the paper sheet. The accuracy of a content-independent algorithm for document rectification is heavily dependent on the document localization step. Investigating the properties for document rectification is a critical and useful task, though doing it independently from the localization mistakes can appear complicated. In this paper, we propose to simulate the imperfections of a document localization subsystem by applying Gaussian noise to the manual annotation. It is shown that an actual localization algorithm based on semantic segmentation well fits the proposed model, making it possible to utilize it for the development of a real document recognition system.
Aleksandr M. Ershov, Daniil V. Tropin, Dmitry P. Nikolaev
ECMS3
2024 Fully Automatic Virtual Unwrapping Method for Documents Imaged by X-Ray Tomography
Petr Kulagin, Dmitry Polevoy, Marina V. Chukalina, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICDAR (3)4
2023 SCA-2023: A Two-Part Dataset For Benchmarking The Methods Of Image Precompensation For Users With Refractive Errors
abstract
This paper considers the problem of precompensating images shown to users with various anomalies of refraction of the eyes (e.g. myopia or astigmatism) in situations where they are not equipped with glasses or other corrective devices. Researchers have proposed a considerable number of such precompensation methods, but to this day there has been no way to accurately compare their quality. We propose an original dataset, which we called “SCA-2023”, of images specially designed for this purpose. Its most important feature is the fact that it includes not only a set of ground-truth images to for implementing the precompensation transform, but also a separate set of images characterizing specific types and degrees of manifestation of the refractive errors. The second part of the dataset is used for computer simulation of the so-called retinal image (the distribution of light on the retina of an imaginary observer). We demonstrated the capabilities of our approach using three prior-art precompensation methods and found that not all the image comparison metrics provide adequate results when applied to precompensated retinal images.
Nafe B. Alkzir, Ilia P. Nikolaev, Dmitry P. Nikolaev
ECMS3
2023 MIDV-Holo: A Dataset for ID Document Hologram Detection in a Video Stream
L. I. Koliaskina, Ekaterina Emelianova, Daniil V. Tropin, V. V. Popov, Konstantin B. Bulatov, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICDAR (3)6
2023 Search for image quality metrics suitable for assessing images specially precompensated for users with refractive errors
abstract
Recently, we presented the SCA-2023 dataset, which had been developed specifically to evaluate the quality of various image precompensation algorithms for observers with imperfect vision. Such precompensation makes it possible to bring their image perception closer to that of an observer with the ideal vision. While experimenting with various image quality metrics, we realized that it was not so easy to evaluate the quality provided by different algorithms, since the metrics ''voted'' for different things, and their choice often seemed to contradict the human perception. This is a key motivation for our study, in which we set out to select the metric best correlated with the human perception of precompensated images. We selected a suitable subdataset from our SCA-2023 dataset and, based on it, created 90 grayscale images, which were shown to our colleagues in a pairwise comparison way. More than 2,000 pairwise comparison results were collected from 24 study participants. Further, according to our original methodology, these results were compared with the ''opinion'' of some popular quality metrics, which made it possible to rank these metrics according to their adequacy within the framework of this task. Finally, we showed how to use these results in optimization procedures aimed at improving the quality of precompensation.
Nafe B. Alkzir, Ilya P. Nikolaev, Dmitry P. Nikolaev
ICMV3
2023 Robust automatic rotation axis alignment mean projection image method in cone-beam and parallel-beam CT
abstract
The rotation axis position is an important parameter of classical reconstruction algorithms in X-ray computed tomography (CT). The use of incorrect values of the axis position parameters during the reconstruction leads to the appearance of various artifacts distorting the reconstructed image. Therefore, to obtain a reconstruction of better quality, automatic rotation axis position determination and misalignment correction methods are of use. Most of the existing high-precision automatic rotation axis position determination methods are either fast, but suitable only within a parallel-beam geometric scheme, or indifferent to the geometric scheme, but computationally laborious. In this paper, we propose a method for auto-detection of two scalar parameters of rotation axis position — axis shift and tilt in the plane parallel to the detector window plane — using a pixel-wise arithmetically averaged projection image. The described method is highly accurate within both parallel-beam and cone-beam geometric schemes whereas it is characterized by robustness to noise in projection data. The method has performed an increase in reconstruction quality when compared with some well-known and still used in practice methods both on synthetic data and on real data obtained in real laboratory conditions.
Danil D. Kazimirov, Anastasia Ingacheva, Alexey V. Buzmakov, Marina V. Chukalina, Dmitry P. Nikolaev
ICMV5
2023 Quantization method for bipolar morphological neural networks
abstract
In the paper, we present a quantization method for bipolar morphological neural networks. Bipolar morphological neural networks use only addition, subtraction, and maximum operations inside the neuron and exponent and logarithm as activation functions of the layers. These operations allow fast and compact gate implementation for FPGA and ASIC, which makes these networks a promising solution for embedded devices. Quantization allows us to reach an additional increase in computational efficiency and reduce the complexity of hardware implementation by using integer values of low bitwidth for computations. We propose an 8-bit quantization scheme based on integer maximum, addition, and lookup tables for non-linear functions and experimentally demonstrate that basic models for image classification can be quantized without noticeable accuracy loss. More advanced models still provide high recognition accuracy but would benefit from further fine-tuning.
Elena Limonova, Michael Zingerenko, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICMV3
2023 CT metal artifacts simulation under x-ray total absorption
abstract
Computed tomography (CT) is a powerful tool for reconstruction and analysis of inner structure of objects applied in various fields. Although many classes of objects of interest may have highly absorbent inclusions, leading to a certain type of distortions on reconstructed volume images (metal-like artifacts). The correction of this type of artifacts can’t be considered a solved task, despite all the efforts in this direction. The development and research of methods for suppressing CT artifacts require high-quality synthetic data which allow for numerical assessment of the accuracy of the metal-like artifacts reduction methods and training of neural networks. Although simplified methods considering only beam hardening and Poisson photon distribution are commonly used to simulate the data with type of distortions. In present work we design experiments using the tomographic scanner of the Federal Research Center “Crystallography and Photonics” of the Russian Academy of Sciences to demonstrate that in some cases beam hardening may not be the dominant reason for the arising of metal-like artifacts. These experiments are closely analyzed and modeled within different approaches. The problems in both simplified and state of the art approaches are emphasized and discussed. The provided results show the importance of paying attention to the dark current modeling for synthesized data generation under the conditions of total photon absorption.
Mikhail Shutov, Marat I. Gilmanov, Dmitry Polevoy, Alexey V. Buzmakov, Anastasia Ingacheva, Marina V. Chukalina, Dmitry P. Nikolaev
ICMV7
2023 Reducing radiation dose for NN-based COVID-19 detection in helical chest CT using real-time monitored reconstruction
Konstantin B. Bulatov, Anastasia Ingacheva, Marat I. Gilmanov, Marina V. Chukalina, Dmitry P. Nikolaev, Vladimir V. Arlazarov
Expert Syst. Appl.5
2022 Calibration Model For Perceptual Compensation Of Defective Pixels Of Self-Emitting Display
abstract
In this paper, we study compensation of defective subpixels with insufficient maximum brightness. The aim of the compensation is to minimize the perceived image non-uniformity. Compensation of the displayed image non-uniformity is based on minimizing the perceived distance between the target (ideally displayed) and the simulated image displayed by the calibrated screen. In this work, we compare the efficiency of compensation depending on color coordinates we calculate color difference in. We investigated the behavior of compensations based on two different uniform color coordinates: CIELAB and Oklab. We examine the efficiency of the compensation on natural scenery images. It was found that Oklab shows better performance than CIELAB in terms of uniformity of perceived compensated image. However, taking into account the spatial properties of the human visual system using S-CIELAB preprocessing almost eliminates the difference between the color coordinates.
Olga A. Basova, Anton Grigoryev, Dmitry P. Nikolaev
ECMS3
2022 From tomographic reconstruction to automatic text recognition: the next frontier task for the artificial intelligence
abstract
Virtual unrolling or unfolding, digital unwrapping, flattening or unfurling - all these terms are used to describe the process of surface straightening of a tomographically reconstructed digital object. For many objects of historical heritage, tomography is the only way to obtain a hidden image of the original object without its destruction. Digital flattening is no longer considered a unique met hodology. It being applied by many research group, but AI-based methods are used insignificantly in such projects, despite the amazing success of AI in computer vision, in particular optical text recognition. It can be explained by the fact that the success of AI depends on large, broad and high quality datasets, but there are very few published CT-based datasets relevant to the task of digital flattening. Accumulation of a sufficient amount of data necessary for training models is a key point for the next technological breakthrough. In this paper, we present open and cumulative dataset CT-OCR-2022. Dataset includes 6 packages data for different model objects that help to enrich tomographic solutions and to train machine learning models. Each package contains optically scanned image of model objects, 400 measured X-ray projections, 2687 CT- reconstructed cross-sections of 3D reconstructed image, segmentation markups. We believe that CT-OCR-2022 dataset will serve as a benchmark for reconstructed object digital flattening and recognition systems, and that it will prove invaluable for advancement of the field of CT-reconstruction, symbols analysis and recognition. The data presented are openly available in Zenodo at doi:10.5281/zenodo.7123495 and linked repositories.
Dmitry Polevoy, Petr Kulagin, Anastasia Ingacheva, Zh. V. Soldatova, Marina V. Chukalina, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICMV6
2022 Fast matrix multiplication for binary and ternary CNNs on ARM CPU
abstract
Low-bit quantized neural networks (QNNs) are of great interest in practical applications because they significantly reduce the consumption of both memory and computational resources. Binary neural networks (BNNs) are memory and computationally efficient as they require only one bit per weight and activation and can be computed using Boolean logic and bit count operations. QNNs with ternary weights and activations (TNNs) and binary weights and ternary activations (TBNs) aim to improve recognition quality compared to BNNs while preserving low bit-width. However, their efficient implementation is usually considered on ASICs and FPGAs, limiting their applicability in real-life tasks. At the same time, one of the areas where efficient recognition is most in demand is recognition on mobile devices using their CPUs. However, there are no known fast implementations of TBNs and TNN, only the daBNN library for BNNs inference. In this paper, we propose novel fast algorithms of ternary, ternary-binary, and binary matrix multiplication for mobile devices with ARM architecture. In our algorithms, ternary weights are represented using 2-bit encoding and binary - using one bit. It allows us to replace matrix multiplication with Boolean logic operations that can be computed on 128-bits simultaneously, using ARM NEON SIMD extension. The matrix multiplication results are accumulated in 16-bit integer registers. We also use special reordering of values in left and right matrices. All that allows us to efficiently compute a matrix product while minimizing the number of loads and stores compared to the algorithm from daBNN. Our algorithms can be used to implement inference of convolutional and fully connected layers of TNNs, TBNs, and BNNs. We evaluate them experimentally on ARM Cortex-A73 CPU and compare their inference speed to efficient implementations of full-precision, 8-bit, and 4-bit quantized matrix multiplications. Our experiment shows our implementations of ternary and ternary-binary matrix multiplications to have almost the same inference time, and they are 3.6 times faster than full-precision, 2.5 times faster than 8-bit quantized, and 1.4 times faster than 4-bit quantized matrix multiplication but 2.9 slower than binary matrix multiplication.
Anton Trusov, Elena Limonova, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICPR3
2020 Simulation Of Underwater Color Images Using Banded Spectral Model
Denis A. Shepelev, Valentina P. Bozhkova, Egor I. Ershov, Dmitry P. Nikolaev
ECMS4
2020 Accelerated FBP for Computed Tomography Image Reconstruction
abstract
Filtered back projection (FBP) is a commonly used technique in tomographic image reconstruction demonstrating acceptable quality. The classical direct implementations of this algorithm require the execution of Θ(N3) operations, where N is the linear size of the 2D slice. Recent approaches including reconstruction via the Fourier slice theorem require Θ(N2log N) multiplication operations. In this paper, we propose a novel approach that reduces the computational complexity of the algorithm to Θ(N2log N) addition operations avoiding Fourier space. For speeding up the convolution, ramp filter is approximated by a pair of causal and anticausal recursive filters, also known as Infinite Impulse Response filters. The back projection is performed with the fast discrete Hough transform. Experimental results on simulated data demonstrate the efficiency of the proposed approach.
Anastasiya V. Dolmatova, Marina V. Chukalina, Dmitry P. Nikolaev
ICIP3
2020 Houghencoder: Neural Network Architecture for Document Image Semantic Segmentation
abstract
In this paper, we propose a HoughEncoder neural network architecture for the semantic image segmentation task. The main feature of the proposed architecture is that it contains layers calculating direct and transposed integral operators, namely Fast Hough Transform. These layers split deep fully convolutional architecture into three blocks. Therefore, the neural network inherits a possibility to make a decision in every point using integral features along different lines. It is important, that by doing this we do not increase the complexity of the neural network in terms of the number of trainable parameters. Our experiments on the publicly available datasets MIDV-500 and MIDV-2019 (both train and test) show that the suggested modification greatly increases quality. HoughEncoder outperforms UNet which shows state-of-the-art results in many semantic image segmentation tasks even while it has a one hundred times fewer parameters.
Alexander Sheshku, Dmitry P. Nikolaev, Vladimir L. Arlazarov
ICIP2
2020 Choosing the best image of the document owner's photograph in the video stream on the mobile device
abstract
One of the business tasks of personal documents recognition using mobile devices is to obtain a high quality image of document owner’s photograph. Such photographs are used to verify and identify the owner of the document. For example, in remote self-service systems, the image of a photo can be compared to a selfie. When a document is captured with a mobile device camera in uncontrolled conditions, the photograph’s image quality varies greatly from frame to frame. In this paper, factors influencing the image quality of a photograph are considered: features of personal documents, capture and recognition processes. A method for choosing the best photograph image is proposed. The quality of the method is assessed on real data by the method of stochastic modeling.
Mikhail A. Aliev, Dmitry P. Nikolaev
ICMV2
2020 Slope detection criterion robust to sparse 2D data
abstract
The study is referred to a task of 2D data slope estimation. We consider the integral projections analysis technique and a common criterion of sum of squared values (SSV) for optimal angle detection. This criterion is dependent on the density of input data and for very sparse data its efficiency significantly decreases. We propose the alternative criteria – the sum of the inversed lengths (SIL) that preserves SSV characteristics for dense data but that is much more robust for sparse input. The experiments conducted on simulated and real datasets demonstrate better quality of slope detection using the proposed criterion.
Dmitry Bocharov, Dmitry P. Nikolaev
ICMV2
2020 Processing and understanding of images in spectral tomography
abstract
The algorithm for 3D vector image reconstruction from a set of spectral tomographic projections collected with CT set-up completed with an optical element or elements inside the optical path behind the sample is proposed. The purpose of their placement into the optical path is to divide the integral polychromatic projection into a series of monochromatic projections, i.e., to get a multi-channel image. Understanding of the reconstruction results in the monochromatic case is beyond question, the relationship between the reconstructed spatial distribution of the linear attenuation coefficient and the discrete description of the elemental structure of the probed object is linear. In difference with monochromatic case the result of the reconstruction from polychromatic projections is a spatial distribution of the so-called effective or average attenuation coefficient, its connection to a discrete description of the elemental structure is nontrivial. However, if the distribution of the averaged coefficient is supplemented by distributions of linear coefficients for several energies, then it is possible to estimate of the local composition of the object. We present a model for the formation of spectral multi-channel projection based on crystal analyzer usage and describe the steps needed to solve the tomography inverse problem.
Marina V. Chukalina, Anastasiya I. Fadeeva, Alexey V. Buzmakov, Dmitry P. Nikolaev
ICMV4
2020 Improvement of U-Net architecture for image binarization with activation functions replacement
abstract
In this work we study the effect of activation functions in a neural network. We consider how activation functions with different properties and their combination affect the final quality of the model. Due to optimization and speed performance issues with most of bounded functions that are represented by sigmoids, we propose the generalized version of SoftSign function - ratio function (rf). Its shape greatly depends on introduced degree parameter, which in theory leads to new interesting property - contraction to zero. For evaluation, we chose image binarization problem: based on UNet architecture of DIBCO-2017 winners, we conducted all experiments with replacing activation functions only. Our research has led us to the state-of-the-art results in binarization quality on DIBCO-2017 test dataset. U-Net with modified activation functions significantly outperforms all existing solutions in all metrics.
Alexander V. Gayer, Alexander Sheshkus, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICMV3
2020 Blind CT images quality assessment of cupping artifacts
abstract
In Computed tomography (CT) usage of common reconstruction algorithms to the projection data acquired with polychromatic probing radiation leads to the appearance of a cup-like distortion. CT image quality can be improved by adjusting the CT scanner or the reconstruction algorithm, but for this purpose assessment of cupping artifacts evaluation needs to be done. Existing assessment methods either rely on expert opinion or require an object binary mask, which can be unavailable. In this paper, we propose a method for blind assessment of cupping artifacts that do not require any prior information. The main idea of the proposed method is to evaluate the degree of change in intensity near automatically found edges of optically dense objects. We prove the applicability of the method on the collected dataset with cupping artifacts. The results show a monotonic dependency between the severity of cupping artifacts and the calculated with the proposed method value.
Anastasia Ingacheva, Daniil V. Tropin, Marina V. Chukalina, Dmitry P. Nikolaev
ICMV4
2020 Bipolar morphological U-Net for document binarization
abstract
Deep neural networks are widely used in various AI systems. Many such systems rely on the edge computing concept and try to perform computations on end devices while still being energy and memory efficient. Therefore, substantial time and memory requirements are imposed on neural networks. One way to improve neural network efficiency is to simplify computations inside a neuron. A bipolar morphological neuron uses only addition, subtraction, and maximum operations inside the neuron and exponent and logarithm as activation functions for the network layers. These operations allow fast and compact gate implementation for FPGA and ASIC. In the paper, we consider the usage of bipolar morphological (BM) networks for document binarization. We examine the DIBCO 2017 binarization challenge and train the bipolar morphological convolutional neural network of U-Net architecture. Despite some accuracy decrease for a model with all BM convolutional layers, one can flexibly control the accuracy by using the partially converted model. It should be noted that even the fully BM model is suitable for solving the problem in practice.
Elena Limonova, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICMV2
2020 Color correction of the document owner's photograph image during recognition on mobile device
abstract
The growing popularity of mobile services increases the risks of financial and other losses from fraudulent user actions. To reduce the number of illegal actions and comply with the law when using mobile services, it is often scheme that user presents it’s identity document. In case of remote access via a mobile device, this means receiving and analyzing a video of a document image. One of the criteria for the authenticity of the captured security document is the presence of security optically variable devices (kinegrams, holograms). Reliable determination of the presence or absence of such security elements from video stream frames is greatly complicated by changes in color between frames. The paper discusses the possibility of using a priori information about the monochrome photograph of the document owner to compensate changes in color between frames. Color distributions are investigated on the example of black-and-white photographs. A new method for automatic white balance correction is proposed. The results of the method are tested on real data obtained with a mobile device.
Ekaterina I. Panfilova, Dmitry P. Nikolaev
ICMV2
2020 Line detection via a lightweight CNN with a Hough layer
abstract
Line detection is an important computer vision task traditionally solved by Hough Transform. With the advance of deep learning, however, trainable approaches to line detection became popular. In this paper we propose a lightweight CNN for line detection with an embedded parameter-free Hough layer, which allows the network neurons to have global strip-like receptive fields. We argue that traditional convolutional networks have two inherent problems when applied to the task of line detection and show how insertion of a Hough layer into the network solves them. Additionally, we point out some major inconsistencies in the current datasets used for line detection.
Lev Teplyakov, Evgeny A. Shvets, Dmitry P. Nikolaev
ICMV3
2020 Improved algorithm of ID card detection by a priori knowledge of the document aspect ratio
abstract
In this work, we consider a problem of quadrilateral document borders detection in images captured by a mobile device’s camera. State-of-the-art algorithms for the quadrilateral document borders detection are not designed for cases when one of the document borders is either completely out of the frame, obscured, or of low contrast. We propose the algorithm which correctly processes the image in such cases. It is built on the classical contour-based algorithm. We modify the latter using the document’s aspect ratio which is known a priori. We demonstrate that this modification reduces the number of incorrect detections by 34% on an open dataset MIDV-500.
Daniil V. Tropin, Ivan A. Konovalenko, Natalya Skoryukina, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICMV4
2020 Lightweight denoising filtering neural network for FBP algorithm
abstract
In that paper, we a suggest lightweight filtering neural network, which implements the filtering stage in the Filtered Back-Projection algorithm (FBP), but good reconstruction results are achieved not only in ideal data but also in noisy data, which a usual FBP algorithm cannot achieve. Thus, our neural network is not an only variation of Ramp filter, which is usually used then FBP algorithm, but also a denoising filter. The neural network architecture was inspired with the idea of the possibility of the Ramp filtering operation’s approximation with sufficient accuracy. The efficiency of our network was shown on the synthetic data, which imitate tomographic projections collected with low exposition. In the generation of synthetic data, we have taken into account the quantum nature of X-ray radiation, exposition time of one frame, and non-linear detector response. The FBP reconstruction time with our neural network was 13 times faster than the time of reconstruction neural network from Learned Primal-Dual Reconstruction, and our reconstruction quality 0.906 by SSIM metric, which is enough to identify most significant objects.
Andrei V. Yamaev, Marina V. Chukalina, Dmitry P. Nikolaev, Alexander Sheshkus, Alexey I. Chulichkov
ICMV3
2020 ResNet-like Architecture with Low Hardware Requirements
abstract
One of the most computationally intensive parts in modern recognition systems is an inference of deep neural networks that are used for image classification, segmentation, enhancement, and recognition. The growing popularity of edge computing makes us look for ways to reduce its time for mobile and embedded devices. One way to decrease the neural network inference time is to modify a neuron model to make it more efficient for computations on a specific device. The example of such a model is a bipolar morphological neuron model. The bipolar morphological neuron is based on the idea of replacing multiplication with addition and maximum operations. This model has been demonstrated for simple image classification with LeNet-like architectures [1]. In the paper, we introduce a bipolar morphological ResNet (BM-ResNet) model obtained from a much more complex ResNet architecture by converting its layers to bipolar morphological ones. We apply BM-ResNet to image classification on MNIST and CIFAR-10 datasets with only a moderate accuracy decrease from 99.3% to 99.1 % and from 85.3% to 85.1 %. We also estimate the computational complexity of the resulting model. We show that for the majority of ResNet layers, the considered model requires 2.1-2.9 times fewer logic gates for implementation and 15-30 % lower latency.
Elena Limonova, Daniil Alfonso, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICPR3
2020 Approach for Document Detection by Contours and Contrasts
abstract
This paper considers arbitrary document detection performed on a mobile device. The classical contour-based approach often fails in cases featuring occlusion, complex background, or blur. The region-based approach, which relies on the contrast between object and background, does not have application limitations, however, its known implementations are highly resource-consuming. We propose a modification of the contour-based method, in which the competing contour location hypotheses are ranked according to the contrast between the areas inside and outside the border. In the experiments, such modification allows for the decrease of alternatives ordering errors by 40% and the decrease of the overall detection errors by 10%. The proposed method provides unmatched state-of-the-art performance on the open MIDV-500 dataset, and it demonstrates results comparable with state-of-the-art performance on the SmartDoc dataset.
Daniil V. Tropin, Sergey A. Ilyuhin, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICPR3
2020 Fast Implementation of 4-bit Convolutional Neural Networks for Mobile Devices
abstract
Quantized low-precision neural networks are very popular because they require less computational resources for inference and can provide high performance, which is vital for real-time and embedded recognition systems. However, their advantages are apparent for FPGA and ASIC devices, while general-purpose processor architectures are not always able to perform low-bit integer computations efficiently. The most frequently used low-precision neural network model for mobile central processors is an 8-bit quantized network. However, in a number of cases, it is possible to use fewer bits for weights and activations, and the only problem is the difficulty of efficient implementation. We introduce an efficient implementation of 4-bit matrix multiplication for quantized neural networks and perform time measurements on a mobile ARM processor. It shows 2.9 times speedup compared to standard floating-point multiplication and is 1.5 times faster than 8-bit quantized one. We also demonstrate a 4-bit quantized neural network for OCR recognition on the MIDV-500 dataset. 4-bit quantization gives 95.0% accuracy and 48% overall inference speedup, while an 8-bit quantized network gives 95.4% accuracy and 39% speedup. The results show that 4-bit quantization perfectly suits mobile devices, yielding good enough accuracy and low inference time.
Anton Trusov, Elena Limonova, Dmitry Slugin, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICPR4
2019 HoughNet: Neural Network Architecture for Vanishing Points Detection
abstract
In this paper we introduce a novel neural network architecture based on Fast Hough Transform layer. The layer of this type allows our neural network to accumulate features from linear areas across the entire image instead of local areas. We demonstrate its potential by solving the problem of vanishing points detection in the images of documents. Such problem occurs when dealing with camera shots of the documents in uncontrolled conditions. In this case, the document image can suffer several specific distortions including projective transform. To train our model, we use MIDV-500 dataset and provide testing results. Strong generalization ability of the suggested method is proven with its applying to a completely different ICDAR 2011 dewarping contest. In previously published papers considering this dataset authors measured quality of vanishing point detection by counting correctly recognized words with open OCR engine Tesseract. To compare with them, we reproduce this experiment and show that our method outperforms the state-of-the-art result.
Alexander Sheshkus, Anastasia Ingacheva, Vladimir V. Arlazarov, Dmitry P. Nikolaev
ICDAR4
2019 Fast Method of ID Documents Location and Type Identification for Mobile and Server Application
abstract
In this paper we discuss the problem of simultaneous document type recognition and projective distortion parameters estimation for the images of ID documents. There are two considered cases. In the first case a video stream captured using mobile devices is processed on the device. The second case considers photos or scanned images which are processed on a server. For each case the requirements are defined for the input data and processing speed. The universal approach is proposed, which allows solving the problem in both cases. The approach is based on representing the image as a constellation of feature points and descriptors, but in order to perform more accurate distortion parameters estimation straight lines and quadrangles are extracted from the input image and used as additional features. Techniques are described which allow to combine matched feature points, lines, and quadrangles to geometric verification using RANSAC. Best alternative selection criteria are proposed along with methods of solution accuracy estimation. The differences between methods of preliminary analysis of the input image and geometric primitives location are discussed in relation to the considered problems. For quality estimation an open dataset MIDV-500 is used, together with its extension for server-side problem version, created in scope of this work. Results show that using lines and quadrangles increase the location accuracy, and the proposed algorithm surpasses previously published works in classification precision and computational performance.
Natalya Skoryukina, Vladimir V. Arlazarov, Dmitry P. Nikolaev
ICDAR3
2019 A document skew detection method using fast Hough transform
abstract
The majority of document image analysis systems use a document skew detection algorithm to simplify all its further processing stages. A huge amount of such algorithms based on Hough transform (HT) analysis has already been proposed. Despite this, we managed to find only one work where the Fast Hough Transform (FHT) usage was suggested to solve the indicated problem. Unfortunately, no study of that method was provided. In this work, we propose and study a skew detection algorithm for the document images which relies on FHT analysis. To measure this algorithm quality we use the dataset from the problem oriented DISEC‘13 contest and its evaluation methodology. Obtained values for AED, T OP80, and CE criteria are equal to 0.086, 0.056, 68.80 respectively.
Pavel Bezmaternykh, Dmitry P. Nikolaev
ICMV2
2019 Orthotropic artifacts suppression for THz and x-ray images using guided filtering
abstract
The paper presents a novel method for suppression of the orthotropic stripe artifacts typical for sensitive optical detector arrays. The algorithm is based on the guided filtering technique where the guidance image is constructed from the input frame in a way that removes artifacts from local contrast structures while disregarding the low-frequency distortions. The artifact suppression procedure was applied to the images of human faces taken with the IR -- THz camera in the diagnosis of psycho-emotional states. In this case, the presence of orthotropic artifacts prevents digital image stabilization. We also demonstrated that adaptation of the alg
Anastasiya V. Dolmatova, E. E. Berlovskaya, Inna Bukreeva, Alessia Cedola, B. R. Islamov, Elena G. Kuznetsova, I. A. Ozheredov, Dmitry P. Nikolaev
ICMV8
2019 Method for numeric estimation of Cupping effect on CT images
abstract
Usage of common reconstruction algorithms like Filtered Back Projection and Algebraic Reconstruction Technique to the projection data acquired with poly-chromatic probing radiation leads to the appearance of a cup-like distortion of the value profile in reconstructed images. While many methods of the poly-chromatic probing artifacts suppression are suggested, the numerical estimation algorithm of the “Cupping effect” typically is not considered to be important. Described methods imply manual regions selection where the intensity will be compared, or just use experts’ opinion on the effect presence. In this paper, we suggest automatic estimation of the “Cupping effect” method based on utilizing the distance transform built using the objects mask. As a result, we obtain a numeric estimation of the intensity change from the border to the center of the object. As the final image index, a weighted sum of the ratings of all objects is used. While positive value shows the magnitude of the “Cupping effect”, a negative value, on the contrary, shows magnitude of the reverse “Cupping effect”. In the paper, we demonstrate the method used on simulated data and compare it with several different techniques for distortion evaluation due to poly-chromatic probing. Finally, we show method effectiveness on real data acquired with laboratory tomography.
Anastasia Ingacheva, Marina V. Chukalina, Alexey V. Buzmakov, Dmitry P. Nikolaev
ICMV4
2019 Bipolar morphological neural networks: convolution without multiplication
abstract
In the paper we introduce a novel bipolar morphological neuron and bipolar morphological layer models. The models use only such operations as addition, subtraction and maximum inside the neuron and exponent and logarithm as activation functions for the layer. The proposed models unlike previously introduced morphological neural networks approximate the classical computations and show better recognition results. We also propose layer-by-layer approach to train the bipolar morphological networks, which can be further developed to an incremental approach for separate neurons to get higher accuracy. Both these approaches do not require special training algorithms and can use a variety of gradient descent methods. To demonstrate efficiency of the proposed model we consider classical convolutional neural networks and convert the pre-trained convolutional layers to the bipolar morphological layers. Seeing that the experiments on recognition of MNIST and MRZ symbols show only moderate decrease of accuracy after conversion and training, bipolar neuron model can provide faster inference and be very useful in mobile and embedded systems.
Elena Limonova, Daniil Matveev, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICMV3
2019 A method of detecting end-to-end curves of limited curvature
abstract
In this paper we consider a method for detecting end-to-end curves of limited curvature like the k-link polylines with bending angle between adjacent segments in a given range. The approximation accuracy is achieved by maximization of the quality function in the image matrix. The method is based on a dynamic programming scheme constructed over Fast Hough Transform calculation results for image bands. The proposed method asymptotic complexity is O(h⋅(w+h/k)⋅log(h/k)), where h and w are the image size, and k is the approximating polyline links number, which is an analogue of the complexity of the fast Fourier transform or the fast Hough transform. We also show the results of the proposed method on synthetic and real data.
Ekaterina I. Panfilova, Mikhail A. Aliev, Irina A. Kunina, Vasiliy V. Postnikov, Dmitry P. Nikolaev
ICMV5
2019 Transfer of a high-level knowledge in HoughNet neural network
abstract
In this paper, we study the recently introduced neural network architecture HoughNet for the ability to accumulate transferable high-level features. The main idea of that neural network is to use convolutional layers separated with Fast Hough Transform layers to enable an analysis of complex non-linear statistics along different lines. We show that different convolutional blocks in this neural network have essentially different purposes. While initial features extracting is task-specific, the main part of the neural network operates with high-level features and do not require re-training in order to be applied to data from a different domain. To prove our statement, we two sets of the images with different origins and demonstrate Transfer Learning presence in the neural network except for the first layers which are highly task-specific.
Alexander Sheshkus, Dmitry P. Nikolaev
ICMV2
2019 Unsupervised domain adaptation for DNN-based automated harvesting
abstract
Computer vision systems based on convolutional neural networks are being rapidly introduced in the field of precision agriculture to solve the problem of scene recognition. Convolutional networks allow performing high-precision recognition, but a significant problem is the expensive process of adapting the network to new conditions. This article proposes a method of fast adaptation of the trained network to minor changes in the source domain without annotating new data. This method is known as Adversarial Domain Adaptation, in the current paper it is applied to the problem of agricultural scene recognition in automated harvesting. The initial training procedure is modified for parallel training of an additional subnet on unannotated data, which makes it possible to compensate the domain shift due to adversarial training. This approach allows us to monotonically increase the quality of all recognized classes of objects and to enhance the stability of CNN model.
Aleksandr Yu. Shkanaev, Dmitry L. Sholomov, Dmitry P. Nikolaev
ICMV3
2018 On the use of FHT, its modification for practical applications and the structure of Hough image
abstract
This work focuses on the Fast Hough Transform (FHT) algorithm proposed by M.L. Brady. We propose how to modify the standard FHT to calculate sums along lines within any given range of their inclination angles. We also describe a new way to visualise Hough-image based on regrouping of accumulator space around its center. Finally, we prove that using Brady parameterization transforms any line into a figure of type “angle”.
Mikhail A. Aliev, Egor I. Ershov, Dmitry P. Nikolaev
ICMV3
2018 A method of projective transformations graph adjustment for planar object panorama stitching
abstract
This paper proposes an improvement for an existing and widely spread approach of panorama stitching for images of planar objects. The proposed method is based on projective transformations graph adjustment. Evaluation is presented on a heterogeneous dataset which contains images of Earth’s and Mars’s surfaces, images taken using a microscope, as well as handwritten and printed text documents. Quality enhancement of panorama stitching method is illustrated on this dataset and shows more than twofold reduction in the accumulated computation error of projective transformations.
B. I. Savelyev, I. B. Mamay, Dmitry P. Nikolaev, Vladimir L. Arlazarov
ICMV3
2018 Automatic cropping of images under projective transformation
abstract
The paper considers the problem of images cropping obtained by projective transformation of source images. The problem is highly relevant to analysis of projective distorted images. We propose two cropping algorithms based on estimation of pixel stretching under the transformation. The algorithms use the ratio of pixel neighborhood areas and the ratio of their chord lengths. The methods comparison is conducted by estimation of cropped background relative areas. The experiment uses real dataset containing projective distorted images of the pages of Russian civil passports. The method based on chord lengths ratio shows better results on highly distorted images.
Julia Shemiakina, Alexander Zhukovsky, Ivan A. Konovalenko, Dmitry P. Nikolaev
ICMV4
2018 Viability of Viola-Jones method for the problem of image classification
abstract
In this paper we study combination of Viola-Jones classifier with deep convolutional neural network as an approach to the problem of object detection and classification. It is well known that Viola-Jones detectors are fast and accurate in detection of vast variety of different objects. On the other hand, methods based on neural network usage demonstrate high accuracy in the problems of image classification. The main goal of this paper is to study viability of Viola-Jones classifier in problem of image classification. The first part of both algorithms is the same: we will use Viola-Jones classifier to find object bounding rectangle in the image. The second part of the algorithms is different: we will compare usage of Viola-Jones classifier with convolutional neural network-based classifier. We will provide speed and accuracy comparison between these two algorithms.
Alexander Sheshkus, Daniil Matalov, Vladimir V. Arlazarov, Dmitry P. Nikolaev
ICMV4
2018 2D art recognition in uncontrolled conditions using one-shot learning
abstract
The paper considers the problem of 2D art identification in photos acquired with mobile devices under the conditions of museum exhibition. The proposed approach is based on a compact description of an image with a constellation of keypoints and corresponding local descriptors. A two-step comparison scheme is described for finding the best reference image matching the query. Bag-of-features approach is used as a first step, then mutual disposition of points is analyzed. Rejection of the query is performed if no suitable matches are found. Geometrical normalization of the query image is proposed to achieve higher robustness against scale and viewpoint variations. After the normalization, mutual disposition of points is estimated using a simplified geometric model. Advantages of the described approach over state-of-the-art solutions are considered. The results of the experiments conducted on the open WikiArt dataset are presented along with processing times for different hardware platforms.
Natalya Skoryukina, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICMV2
2018 Linear colour segmentation revisited
abstract
In this work we discuss the known algorithms for linear colour segmentation based on a physical approach and propose a new modification of segmentation algorithm. This algorithm is based on a region adjacency graph framework without a pre-segmentation stage. Proposed edge weight functions are defined from linear image model with normal noise. The colour space projective transform is introduced as a novel pre-processing technique for better handling of shadow and highlight areas. The resulting algorithm is tested on a benchmark dataset consisting of the images of 19 natural scenes selected from the Barnard’s DXC-930 SFU dataset and 12 natural scene images newly published for common use. The dataset is provided with pixel-by-pixel ground truth colour segmentation for every image. Using this dataset, we show that the proposed algorithm modifications lead to qualitative advantages over other model-based segmentation algorithms, and also show the positive effect of each proposed modification. The source code and datasets for this work are available for free access at http://github.com/visillect/segmentation.
Anna Smagina, Valentina P. Bozhkova, Sergey Gladilin, Dmitry P. Nikolaev
ICMV4
2018 The method of image alignment based on sharpness maximization
abstract
In this paper the method of image alignment based on average image sharpness maximization is proposed. The algorithm for global-shift model is investigated, its efficiency by applying FFT is shown. For projective model, an approach for image alignment using local shifts and RANSAC to obtain the final transform is considered. Experimental results for the system of document's reconstruction in a video stream increasing quality of output image are demonstrated.
Daniil V. Tropin, Dmitry P. Nikolaev, Dmitry Slugin
ICMV2
2017 Automatic Beam Hardening Correction For CT Reconstruction
Marina V. Chukalina, Anastasia Ingacheva, Alexey V. Buzmakov, Igor V. Polyakov, Andrey Gladkov, Ivan Yakimchuk, Dmitry P. Nikolaev
ECMS7
2017 Generation Algorithms Of Fast Generalized Hough Transform
Egor I. Ershov, Evgeny A. Shvets, Timur M. Khanipov, Dmitry P. Nikolaev
ECMS4
2017 Segments Graph-Based Approach for Document Capture in a Smartphone Video Stream
abstract
The paper is devoted to the analysis of the problem of document boundaries detection in images and in a video stream. The paper proposes an algorithm for obtaining the position of the document, consisting of very reliable segments of a document boundaries extraction and a construction of an intersection graph that satisfies the projective model of the rectangle. An online algorithm for selecting and integrating possible document positions in a video stream based on the Kalman filter is proposed. The analysis of possible modifications of the algorithm and their effect on the final result are provided. Evaluation of the quality of the document at ICDAR'15 Smartphone Document Capture competition's dataset [1] showed a mean result of 95.5% in Jaccard index of projectively corrected document quadrangles and a 3rd place in the competition.
Alexander Zhukovsky, Dmitry P. Nikolaev, Vladimir V. Arlazarov, Vasiliy V. Postnikov, Dmitry Polevoy, Natalya Skoryukina, Timofey S. Chernov, Julia Shemiakina, Arseniy P. Mukovozov, Ivan A. Konovalenko, Mikhail Povolotsky
ICDAR2
2017 Neural network-based feature point descriptors for registration of optical and SAR images
abstract
Registration of images of different nature is an important technique used in image fusion, change detection, efficient information representation and other problems of computer vision. Solving this task using feature-based approaches is usually more complex than registration of several optical images because traditional feature descriptors (SIFT, SURF, etc.) perform poorly when images have different nature. In this paper we consider the problem of registration of SAR and optical images. We train neural network to build feature point descriptors and use RANSAC algorithm to align found matches. Experimental results are presented that confirm the method’s effectiveness.
Dmitry Abulkhanov, Ivan A. Konovalenko, Dmitry P. Nikolaev, A. V. Savchik, Evgeny Shvets, D. Sidorchuk
ICMV3
2017 Textual blocks rectification method based on fast Hough transform analysis in identity documents recognition
abstract
Textual blocks rectification or slant correction is an important stage of document image processing in OCR systems. This paper considers existing methods and introduces an approach for the construction of such algorithms based on Fast Hough Transform analysis. A quality measurement technique is proposed and obtained results are shown for both printed and handwritten textual blocks processing as a part of an industrial system of identity documents recognition on mobile devices.
Pavel Bezmaternykh, Dmitry P. Nikolaev, Vladimir L. Arlazarov
ICMV2
2017 Analysis of computer images in the presence of metals
abstract
Artifacts caused by intensely absorbing inclusions are encountered in computed tomography via polychromatic scanning and may obscure or simulate pathologies in medical applications. Тo improve the quality of reconstruction if high-Z inclusions in presence, previously we proposed and tested with synthetic data an iterative technique with soft penalty mimicking linear inequalities on the photon-starved rays. This note reports a test at the tomographic laboratory set-up at the Institute of Crystallography FSRC “Crystallography and Photonics” RAS in which tomographic scans were successfully made of temporary tooth without inclusion and with Pb inclusion.
Alexey V. Buzmakov, Anastasia Ingacheva, Victor E. Prun, Dmitry P. Nikolaev, Marina V. Chukalina, Claudio Ferrero, Victor E. Asadchikov
ICMV4
2017 Overview of machine vision methods in x-ray imaging and microtomography
abstract
Digital X-ray imaging became widely used in science, medicine, non-destructive testing. This allows using modern digital images analysis for automatic information extraction and interpretation. We give short review of scientific applications of machine vision in scientific X-ray imaging and microtomography, including image processing, feature detection and extraction, images compression to increase camera throughput, microtomography reconstruction, visualization and setup adjustment.
Alexey V. Buzmakov, Denis Zolotov, Marina V. Chukalina, Dmitry P. Nikolaev, Andrey Gladkov, Anastasia Ingacheva, Ivan Yakimchuk, Victor E. Asadchikov
ICMV4
2017 Image quality assessment for video stream recognition systems
abstract
Recognition and machine vision systems have long been widely used in many disciplines to automate various processes of life and industry. Input images of optical recognition systems can be subjected to a large number of different distortions, especially in uncontrolled or natural shooting conditions, which leads to unpredictable results of recognition systems, making it impossible to assess their reliability. For this reason, it is necessary to perform quality control of the input data of recognition systems, which is facilitated by modern progress in the field of image quality evaluation. In this paper, we investigate the approach to designing optical recognition systems with built-in input image quality estimation modules and feedback, for which the necessary definitions are introduced and a model for describing such systems is constructed. The efficiency of this approach is illustrated by the example of solving the problem of selecting the best frames for recognition in a video stream for a system with limited resources. Experimental results are presented for the system for identity documents recognition, showing a significant increase in the accuracy and speed of the system under simulated conditions of automatic camera focusing, leading to blurring of frames.
Timofey S. Chernov, Nikita P. Razumnuy, Alexander S. Kozharinov, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICMV4
2017 Fast words boundaries localization in text fields for low quality document images
abstract
The paper examines the problem of word boundaries precise localization in document text zones. Document processing on a mobile device consists of document localization, perspective correction, localization of individual fields, finding words in separate zones, segmentation and recognition. While capturing an image with a mobile digital camera under uncontrolled capturing conditions, digital noise, perspective distortions or glares may occur. Further document processing gets complicated because of its specifics: layout elements, complex background, static text, document security elements, variety of text fonts. However, the problem of word boundaries localization has to be solved at runtime on mobile CPU with limited computing capabilities under specified restrictions. At the moment, there are several groups of methods optimized for different conditions. Methods for the scanned printed text are quick but limited only for images of high quality. Methods for text in the wild have an excessively high computational complexity, thus, are hardly suitable for running on mobile devices as part of the mobile document recognition system. The method presented in this paper solves a more specialized problem than the task of finding text on natural images. It uses local features, a sliding window and a lightweight neural network in order to achieve an optimal algorithm speed-precision ratio. The duration of the algorithm is 12 ms per field running on an ARM processor of a mobile device. The error rate for boundaries localization on a test sample of 8000 fields is 0.3
Dmitry Ilin, Dmitriy Novikov, Dmitry Polevoy, Dmitry P. Nikolaev
ICMV4
2017 Blur kernel estimation with algebraic tomography technique and intensity profiles of object boundaries
abstract
Motion blur caused by camera vibration is a common source of degradation in photographs. In this paper we study the problem of finding the point spread function (PSF) of a blurred image using the tomography technique. The PSF reconstruction result strongly depends on the particular tomography technique used. We present a tomography algorithm with regularization adapted specifically for this task. We use the algebraic reconstruction technique (ART algorithm) as the starting algorithm and introduce regularization. We use the conjugate gradient method for numerical implementation of the proposed approach. The algorithm is tested using a dataset which contains 9 kernels extracted from real photographs by the Adobe corporation where the point spread function is known. We also investigate influence of noise on the quality of image reconstruction and investigate how the number of projections influence the magnitude change of the reconstruction error.
Anastasia Ingacheva, Marina V. Chukalina, Timur M. Khanipov, Dmitry P. Nikolaev
ICMV4
2017 Aerial images visual localization on a vector map using color-texture segmentation
abstract
In this paper we study the problem of combining UAV obtained optical data and a coastal vector map in absence of satellite navigation data. The method is based on presenting the territory as a set of segments produced by color-texture image segmentation. We then find such geometric transform which gives the best match between these segments and land and water areas of the georeferenced vector map. We calculate transform consisting of an arbitrary shift relatively to the vector map and bound rotation and scaling. These parameters are estimated using the RANSAC algorithm which matches the segments contours and the contours of land and water areas of the vector map. To implement this matching we suggest computing shape descriptors robust to rotation and scaling. We performed numerical experiments demonstrating the practical applicability of the proposed method.
Irina A. Kunina, Lev Teplyakov, Andrey Gladkov, Timur M. Khanipov, Dmitry P. Nikolaev
ICMV5
2017 Mobile and embedded fast high resolution image stitching for long length rectangular monochromatic objects with periodic structure
abstract
In this paper we describe stitching protocol, which allows to obtain high resolution images of long length monochromatic objects with periodic structure. This protocol can be used for long length documents or human-induced objects in satellite images of uninhabitable regions like Arctic regions. The length of such objects can reach notable values, while modern camera sensors have limited resolution and are not able to provide good enough image of the whole object for further processing, e.g. using in OCR system. The idea of the proposed method is to acquire a video stream containing full object in high resolution and use image stitching. We expect the scanned object to have straight boundaries and periodic structure, which allow us to introduce regularization to the stitching problem and adapt algorithm for limited computational power of mobile and embedded CPUs. With the help of detected boundaries and structure we estimate homography between frames and use this information to reduce complexity of stitching. We demonstrate our algorithm on mobile device and show image processing speed of 2 fps on Samsung Exynos 5422 processor
Elena Limonova, Daniil V. Tropin, Boris Savelyev, Igor Mamay, Dmitry P. Nikolaev
ICMV5
2017 Modification of YAPE keypoint detection algorithm for wide local contrast range images
abstract
Keypoint detection is an important tool of image analysis, and among many contemporary keypoint detection algorithms YAPE is known for its computational performance, allowing its use in mobile and embedded systems. One of its shortcomings is high sensitivity to local contrast which leads to high detection density in high-contrast areas while missing detections in low-contrast ones. In this work we study the contrast sensitivity of YAPE and propose a modification which compensates for this property on images with wide local contrast range (Yet Another Contrast-Invariant Point Extractor, YACIPE). As a model example, we considered the traffic sign recognition problem, where some signs are well-lighted, whereas others are in shadows and thus have low contrast. We show that the number of traffic signs on the image of which has not been detected any keypoints is 40% less for the proposed modification compared to the original algorithm.
A. Lukoyanov, Dmitry P. Nikolaev, Ivan A. Konovalenko
ICMV2
2017 Establishing the correspondence between closed contours of objects in images with projective distortions
abstract
In this paper, we consider the task of finding the correspondence between closed contours of objects in an image pair with small projective distortions. Several methods are considered and their comparison is performed. The experiments results for two contour sets are provided. Sufficient conditions of the applicability of the method of selecting the nearest contour are represented and proven.
Alexey V. Savchik, Victoria A. Sablina, Dmitry P. Nikolaev
ICMV3
2017 The method for homography estimation between two planes based on lines and points
abstract
The paper considers the problem of estimating a transform connecting two images of one plane object. The method based on RANSAC is proposed for calculating the parameters of projective transform which uses points and lines correspondences simultaneously. A series of experiments was performed on synthesized data. Presented results show that the algorithm convergence rate is significantly higher when actual lines are used instead of points of lines intersection. When using both lines and feature points it is shown that the convergence rate does not depend on the ratio between lines and feature points in the input dataset.
Julia Shemiakina, Alexander Zhukovsky, Dmitry P. Nikolaev
ICMV3
2017 Vanishing points detection using combination of fast Hough transform and deep learning
abstract
In this paper we propose a novel method for vanishing points detection based on convolutional neural network (CNN) approach and fast Hough transform algorithm. We show how to determine fast Hough transform neural network layer and how to use it in order to increase usability of the neural network approach to the vanishing point detection task. Our algorithm includes CNN with consequence of convolutional and fast Hough transform layers. We are building estimator for distribution of possible vanishing points in the image. This distribution can be used to find candidates of vanishing point. We provide experimental results from tests of suggested method using images collected from videos of road trips. Our approach shows stable result on test images with different projective distortions and noise. Described approach can be effectively implemented for mobile GPU and CPU.
Alexander Sheshkus, Anastasia Ingacheva, Dmitry P. Nikolaev
ICMV3
2016 Fast 3D Hough Transform Computation
Egor I. Ershov, Arseniy P. Terekhin, Simon M. Karpenko, Dmitry P. Nikolaev, Vasiliy V. Postnikov
ECMS4
2016 To image analysis in computed tomography
abstract
The presence of errors in tomographic image may lead to misdiagnosis when computed tomography (CT) is used in medicine, or the wrong decision about parameters of technological processes when CT is used in the industrial applications. Two main reasons produce these errors. First, the errors occur on the step corresponding to the measurement, e.g. incorrect calibration and estimation of geometric parameters of the set-up. The second reason is the nature of the tomography reconstruction step. At the stage a mathematical model to calculate the projection data is created. Applied optimization and regularization methods along with their numerical implementations of the method chosen have their own specific errors. Nowadays, a lot of research teams try to analyze these errors and construct the relations between error sources. In this paper, we do not analyze the nature of the final error, but present a new approach for the calculation of its distribution in the reconstructed volume. We hope that the visualization of the error distribution will allow experts to clarify the medical report impression or expert summary given by them after analyzing of CT results. To illustrate the efficiency of the proposed approach we present both the simulation and real data processing results.
Marina V. Chukalina, Dmitry P. Nikolaev, Anastasia Ingacheva, Alexey V. Buzmakov, Ivan Yakimchuk, Victor E. Asadchikov
ICMV2
2016 Fast integer approximations in convolutional neural networks using layer-by-layer training
abstract
This paper explores method of layer-by-layer training for neural networks to train neural network, that use approximate calculations and/or low precision data types. Proposed method allows to improve recognition accuracy using standard training algorithms and tools. At the same time, it allows to speed up neural network calculations using fast-processed approximate calculations and compact data types. We consider 8-bit fixed-point arithmetic as the example of such approximation for image recognition problems. In the end, we show significant accuracy increase for considered approximation along with processing speedup.
Dmitry Ilin, Elena Limonova, Vladimir V. Arlazarov, Dmitry P. Nikolaev
ICMV4
2016 Aerial image geolocalization by matching its line structure with route map
abstract
The classic way of aerial photographs geolocation is to bind their local coordinates to a geographic coordinate system using GPS and IMU data. At the same time the possibility of geolocation in a jammed navigation field is also of interest for practical purposes. In this paper we consider one approach to visual localization relatively to a vector road map without GPS. We suggest a geolocalization algorithm which detects image line segments and looks for a geometrical transformation which provides the best mapping between the obtained segments set and line segments in the road map. We consider IMU and altimeter data still known which allows to work with orthorectified images. The problem is hence reduced to a search for a transformation which contains an arbitrary shift and bounded rotation and scaling relatively to the vector map. These parameters are estimated using RANSAC by matching straight line segments from the image to vector map segments. We also investigate how the proposed algorithm’s stability is influenced by segment coordinates (two spatial and one angular).
Irina A. Kunina, Arseniy P. Terekhin, Timur M. Khanipov, Elena G. Kuznetsova, Dmitry P. Nikolaev
ICMV5
2016 Slant rectification in Russian passport OCR system using fast Hough transform
abstract
In this paper, we introduce slant detection method based on Fast Hough Transform calculation and demonstrate its application in industrial system for Russian passports recognition. About 1.5% of this kind of documents appear to be slant or italic. This fact reduces recognition rate, because Optical Recognition Systems are normally designed to process normal fonts. Our method uses Fast Hough Transform to analyse vertical strokes of characters extracted with the help of x-derivative of a text line image. To improve the quality of detector we also introduce field grouping rules. The resulting algorithm allowed to reach high detection quality. Almost all errors of considered approach happen on passports of nonstandard fonts, while slant detector works in appropriate way.
Elena Limonova, Pavel Bezmaternykh, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICMV3
2016 Image deblurring in video stream based on two-level image model
abstract
An iterative algorithm is proposed for blind multi-image deblurring of binary images. The binarity is the only prior restriction imposed on the image. Image formation model assumes convolution with arbitrary kernel and addition of a constant value. Penalty functional is composed using binarity constraint for regularization. The algorithm estimates the original image and distortion parameters by alternate reduction of two parts of this functional. Experimental results for natural (non-synthetic) data are present.
Arseniy Mukovozov, Dmitry P. Nikolaev, Elena Limonova
ICMV2
2016 Combining convolutional neural networks and Hough Transform for classification of images containing lines
abstract
In this paper, we propose an expansion of convolutional neural network (CNN) input features based on Hough Transform. We perform morphological contrasting of source image followed by Hough Transform, and then use it as input for some convolutional filters. Thus, CNNs computational complexity and the number of units are not affected. Morphological contrasting and Hough Transform are the only additional computational expenses of introduced CNN input features expansion. Proposed approach was demonstrated on the example of CNN with very simple structure. We considered two image recognition problems, that were object classification on CIFAR-10 and printed character recognition on private dataset with symbols taken from Russian passports. Our approach allowed to reach noticeable accuracy improvement without taking much computational effort, which can be extremely important in industrial recognition systems or difficult problems utilising CNNs, like pressure ridge analysis and classification.
Alexander Sheshkus, Elena Limonova, Dmitry P. Nikolaev, Valeriy E. Krivtsov
ICMV3
2016 Snapscreen: TV-stream frame search with projectively distorted and noisy query
abstract
In this work we describe an approach to real-time image search in large databases robust to variety of query distortions such as lighting alterations, projective distortions or digital noise. The approach is based on the extraction of keypoints and their descriptors, random hierarchical clustering trees for preliminary search and RANSAC for refining search and result scoring. The algorithm is implemented in Snapscreen system which allows determining a TV-channel and a TV-show from a picture acquired with mobile device. The implementation is enhanced using preceding localization of screen region. Results for the real-world data with different modifications of the system are presented.
Natalya Skoryukina, Timofey S. Chernov, Konstantin B. Bulatov, Dmitry P. Nikolaev, Vladimir V. Arlazarov
ICMV4
2015 A Method Of Periodic Pattern Detection On Document Images
Timofey S. Chernov, Vitali M. Kliatskine, Dmitry P. Nikolaev
ECMS3
2015 A Way To Reduce The Artifacts Caused By Intensely Absorbing Areas In Computed Tomography
abstract
Artifacts caused by intensely absorbing areas are encountered in computed tomography and may obscure or simulate pathology in medical applications, hide or mimic the cracks and cavities in the devices at industrial applications. We simulated sinograms with different levels of absorption to demonstrate the artifacts dynamics. If the analysis of the measured data shows the presence of strongly absorbing areas in the object under study we propose to use quadratic programming technique for solving the inverse problem. Although the technique is time-consuming it allows us to avoid the typical artifacts. We compare the images reconstructed with different techniques including the proposed one.
Marina V. Chukalina, Anastasia Ingacheva, Victor E. Prun, Alexey V. Buzmakov, Dmitry P. Nikolaev
ECMS5
2015 Exact Fast Algorithm For Optimal Linear Separation Of 2D Distribution
abstract
The paper presents a new fast computation scheme for linear separation in two-dimensional feature space. This scheme is based on a combination of several image processing techniques: fast Hough transform, cumulative sum computation and expression of optimized criterion as a function of additive statistics. It is shown that complexity of the scheme is O(n log n) for chosen set of criteria. Two appropriate criteria are discussed, both being a 2D extension of well-known Otsu’s criterion: standard one considering covariance trace and one considering covariance second eigenvalues. Applicability of the latter criterion for the color segmentation problem is discussed.
Egor I. Ershov, Vasiliy V. Postnikov, Arseniy P. Terekhin, Dmitry P. Nikolaev
ECMS4
2015 Vision-Based Vehicle Wheel Detector And Axle Counter
Anton Grigoryev, Dmitry Bocharov, Arseniy P. Terekhin, Dmitry P. Nikolaev
ECMS4
2015 UAV Navigation On The Basis Of The Feature Points Detection On Underlying Surface
abstract
This work relates to the intelligent systems tracking such as UAV’s (unmanned aviation vehicle) navigation in GPS-denied environment. Generally it considers the tracking of the UAV path on the basis of bearing-only observations including azimuth and elevation angles. It is assumed that UAV’s cameras are able to capture the angular position of reference points and to measure the directional angles of the sight line. Such measurements involve the real position of UAV in implicit form, and therefore some of nonlinear filters such as Extended Kalman filter (EKF) or others must be used in order to implement these measurements for UAV control. Meanwhile, there is well-known method of pseudomeasurements which reduces the estimation problem to the linear settings, though these method has a bias. Recently it was shown that the application of the modified filter based on the pseudomeasurements approach provides the reliable UAV control on the basis of the observation of reference points nominated before the flight. This approach uses the known coordinates of reference points and then applies the optimal linear Kalman type filter. The principal difference with the usage of location of reference points nominated in advance is that here we use the observed reference points detected on-line during the flight. This approach permits to reduce the necessary on-board memory up to reasonable size. In this article the modified pseudomeasurement method without bias for estimation of the UAV position has been suggested. On the basis of this estimation the control algorithm which provides the tracking of reference path in case of external perturbation and the angles measurements errors has been developed. Another principal novelty of this work is the usage of RANSAC approach to detection of reference landmarks which used further for estimation of the UAV position.
Ivan A. Konovalenko, Alexander B. Miller, Boris M. Miller, Dmitry P. Nikolaev
ECMS4
2015 A method of periodic pattern localization on document images
abstract
Periodic patterns often present on document images as holograms, watermarks or guilloche elements which are mostly used for fraud protection. Localization of such patterns lets an embedded OCR system to vary its settings depending on pattern presence in particular image regions and improves the precision of pattern removal to preserve as much useful data as possible. Many document images’ noise detection and removal methods deal with unstructured noise or clutter on documents with simple background. In this paper we propose a method of periodic pattern localization on document images which uses discrete Fourier transform that works well on documents with complex background.
Timofey S. Chernov, Dmitry P. Nikolaev, Vitali M. Kliatskine
ICMV2
2015 CT metal artifact reduction by soft inequality constraints
abstract
The artifacts (known as metal-like artifacts) arising from incorrect reconstruction may obscure or simulate pathology in medical applications, hide or mimic cracks and cavities in the scanned objects in industrial tomographic scans. One of the main reasons caused such artifacts is photon starvation on the rays which go through highly absorbing regions. We indroduce a way to suppress such artifacts in the reconstructions using soft penalty mimicing linear inequalities on the photon starved rays. An efficient algorithm to use such information is provided and the effect of those inequalities on the reconstruction quality is studied.
Marina V. Chukalina, Dmitry P. Nikolaev, Valerii Sokolov, Anastasia Ingacheva, Alexey V. Buzmakov, Victor E. Prun
ICMV2
2015 Fast Hough transform analysis: pattern deviation from line segment
abstract
In this paper, we analyze properties of dyadic patterns. These pattern were proposed to approximate line segments in the fast Hough transform (FHT). Initially, these patterns only had recursive computational scheme. We provide simple closed form expression for calculating point coordinates and their deviation from corresponding ideal lines.
Egor I. Ershov, Arseniy P. Terekhin, Dmitry P. Nikolaev, Vasiliy V. Postnikov, Simon M. Karpenko
ICMV3
2015 Building a robust vehicle detection and classification module
abstract
The growing adoption of intelligent transportation systems (ITS) and autonomous driving requires robust real-time solutions for various event and object detection problems. Most of real-world systems still cannot rely on computer vision algorithms and employ a wide range of costly additional hardware like LIDARs. In this paper we explore engineering challenges encountered in building a highly robust visual vehicle detection and classification module that works under broad range of environmental and road conditions. The resulting technology is competitive to traditional non-visual means of traffic monitoring. The main focus of the paper is on software and hardware architecture, algorithm selection and domain-specific heuristics that help the computer vision system avoid implausible answers.
Anton Grigoryev, Timur M. Khanipov, Ivan Koptelov, Dmitry Bocharov, Vasiliy V. Postnikov, Dmitry P. Nikolaev
ICMV6
2015 Visual navigation of the UAVs on the basis of 3D natural landmarks
abstract
This work considers the tracking of the UAV (unmanned aviation vehicle) on the basis of onboard observations of natural landmarks including azimuth and elevation angles. It is assumed that UAV's cameras are able to capture the angular position of reference points and to measure the angles of the sight line. Such measurements involve the real position of UAV in implicit form, and therefore some of nonlinear filters such as Extended Kalman filter (EKF) or others must be used in order to implement these measurements for UAV control. Recently it was shown that modified pseudomeasurement method may be used to control UAV on the basis of the observation of reference points assigned along the UAV path in advance. However, the use of such set of points needs the cumbersome recognition procedure with the huge volume of on-board memory. The natural landmarks serving as such reference points which may be determined on-line can significantly reduce the on-board memory and the computational difficulties. The principal difference of this work is the usage of the 3D reference points coordinates which permits to determine the position of the UAV more precisely and thereby to guide along the path with higher accuracy which is extremely important for successful performance of the autonomous missions. The article suggests the new RANSAC for ISOMETRY algorithm and the use of recently developed estimation and control algorithms for tracking of given reference path under external perturbation and noised angular measurements.
Simon M. Karpenko, Ivan A. Konovalenko, Alexander B. Miller, Boris M. Miller, Dmitry P. Nikolaev
ICMV5
2015 Demosaicing as the problem of regularization
abstract
Demosaicing is the process of reconstruction of a full-color image from Bayer mosaic, which is used in digital cameras for image formation. This problem is usually considered as an interpolation problem. In this paper, we propose to consider the demosaicing problem as a problem of solving an underdetermined system of algebraic equations using regularization methods. We consider regularization with standard l1/2-, l1 -, l2- norms and their effect on quality image reconstruction. The experimental results showed that the proposed technique can both be used in existing methods and become the base for new ones
Irina A. Kunina, Aleksey Volkov, Sergey Gladilin, Dmitry P. Nikolaev
ICMV4
2015 Viola-Jones based hybrid framework for real-time object detection in multispectral images
abstract
This paper describes a method for real-time object detection based on a hybrid of a Viola-Jones cascade with a convolutional neural network. This scheme allows flexible trade-offs between detection quality and computational performance. We also propose a generalization of this method to multispectral images that effectively and efficiently utilizes information from each spectral channel. The new scheme is experimentally compared to traditional Viola-Jones, showing improved detection quality with adjustable performance.
Elena G. Kuznetsova, Evgeny A. Shvets, Dmitry P. Nikolaev
ICMV3
2015 Improving neural network performance on SIMD architectures
abstract
Neural network calculations for the image recognition problems can be very time consuming. In this paper we propose three methods of increasing neural network performance on SIMD architectures. The usage of SIMD extensions is a way to speed up neural network processing available for a number of modern CPUs. In our experiments, we use ARM NEON as SIMD architecture example. The first method deals with half float data type for matrix computations. The second method describes fixed-point data type for the same purpose. The third method considers vectorized activation functions implementation. For each method we set up a series of experiments for convolutional and fully connected networks designed for image recognition task.
Elena Limonova, Dmitry Ilin, Dmitry P. Nikolaev
ICMV3
2015 Approach to recognition of flexible form for credit card expiration date recognition as example
abstract
In this paper we consider a task of finding information fields within document with flexible form for credit card expiration date field as example. We discuss main difficulties and suggest possible solutions. In our case this task is to be solved on mobile devices therefore computational complexity has to be as low as possible. In this paper we provide results of the analysis of suggested algorithm. Error distribution of the recognition system shows that suggested algorithm solves the task with required accuracy.
Alexander Sheshkus, Dmitry P. Nikolaev, Anastasia Ingacheva, Natalya Skoryukina
ICMV2
2015 Complex approach to long-term multi-agent mapping in low dynamic environments
abstract
In the paper we consider the problem of multi-agent continuous mapping of a changing, low dynamic environment. The mapping problem is a well-studied one, however usage of multiple agents and operation in a non-static environment complicate it and present a handful of challenges (e.g. double-counting, robust data association, memory and bandwidth limits). All these problems are interrelated, but are very rarely considered together, despite the fact that each has drawn attention of the researches. In this paper we devise an architecture that solves the considered problems in an internally consistent manner.
Evgeny A. Shvets, Dmitry P. Nikolaev
ICMV2
2014 Vision-based industrial automatic vehicle classifier
abstract
The paper describes the automatic motor vehicle video stream based classification system. The system determines vehicle type at payment collection plazas on toll roads. Classification is performed in accordance with a preconfigured set of rules which determine type by number of wheel axles, vehicle length, height over the first axle and full height. These characteristics are calculated using various computer vision algorithms: contour detectors, correlational analysis, fast Hough transform, Viola-Jones detectors, connected components analysis, elliptic shapes detectors and others. Input data contains video streams and induction loop signals. Output signals are vehicle enter and exit events, vehicle type, motion direction, speed and the above mentioned features.
Timur M. Khanipov, Ivan Koptelov, Anton Grigoryev, Elena G. Kuznetsova, Dmitry P. Nikolaev
ICMV5
2014 Method of center localization for objects containing concentric arcs
abstract
This paper proposes a method for automatic center location of objects containing concentric arcs. The method utilizes structure tensor analysis and voting scheme optimized with Fast Hough Transform. Two applications of the proposed method are considered: (i) wheel tracking in video-based system for automatic vehicle classification and (ii) tree growth rings analysis on a tree cross cut image.
Elena G. Kuznetsova, Evgeny A. Shvets, Dmitry P. Nikolaev
ICMV3
2014 Generalization of the Viola-Jones method as a decision tree of strong classifiers for real-time object recognition in video stream
abstract
In this paper, we present a new modification of Viola-Jones complex classifiers. We describe a complex classifier in the form of a decision tree and provide a method of training for such classifiers. Performance impact of the tree structure is analyzed. Comparison is carried out of precision and performance of the presented method with that of the classical cascade. Various tree architectures are experimentally studied. The task of vehicle wheels detection on images obtained from an automatic vehicle classification system is taken as an example.
A. Minkina, Dmitry P. Nikolaev, Sergey A. Usilin, V. Kozyrev
ICMV2
2014 X-ray fluorescence tomography: Jacobin matrix and confidence of the reconstructed images
abstract
The goal of the X-ray Fluorescence Computed Tomography (XFCT) is to give the quantitative description of an object under investigation (sample) in terms of the element composition. However, light and heavy elements inside the object give different contribution to the attenuation of the X-ray probe and of the fluorescence. It leads to the elements got in the shadow area do not give any contribution to the registered spectrum. Iterative reconstruction procedures will try to set to zero the variables describing the element content in composition of corresponding unit volumes as these variables do not change system's condition number. Inversion of the XFCT Radon transform gives random values in these areas. To evaluate the confidence of the reconstructed images we first propose, in addition to the reconstructed images, to calculate a generalized image based on Jacobian matrix. This image highlights the areas of doubt in case if there are exist. In the work we have attempted to prove the advisability of such an approach. For this purpose, we analyzed in detail the process of tomographic projection formation.
Dmitry P. Nikolaev, Marina V. Chukalina
ICMV1
2014 Diamond recognition algorithm using two-channel x-ray radiographic separator
abstract
In this paper real time classification method for two-channel X-ray radiographic diamond separation is discussed. Proposed method does not require direct hardware calibration but uses sample images as a train dataset. It includes online dynamic time warping algorithm for inter-channel synchronization. Additionally, algorithms of online source signal control are discussed, including X-ray intensity control, optical noise detection and sensor occlusion detection.
Dmitry P. Nikolaev, Andrey Gladkov, Timofey S. Chernov, Konstantin B. Bulatov
ICMV1
2014 Comparison of two algorithms modifications of projective-invariant recognition of the plane boundaries with the one concavity
abstract
In this paper we present two algorithms modifications of projective-recognition of the plane boundaries with one concavity. The input images are created with orthographic pinhole camera with a fixed focal length. Thus variety of the possible projective transformations is limited. The first modification considers the task more generally, the other uses prior information about camera model. A hypothesis that the second modification has better accuracy is being checked. Results of around 20000 numeral experiments that confirm the hypothesis are included.
Natalia Pritula, Dmitry P. Nikolaev, Alexander Sheshkus, Mikhail Pritula, Petr P. Nikolayev
ICMV2
2014 Real time rectangular document detection on mobile devices
abstract
In this paper we propose an algorithm for real-time rectangular document borders detection in mobile device based applications. The proposed algorithm is based on combinatorial assembly of possible quadrangle candidates from a set of line segments and projective document reconstruction using the known focal length. Fast Hough Transform is used for line detection. 1D modification of edge detector is proposed for the algorithm.
Natalya Skoryukina, Dmitry P. Nikolaev, Alexander Sheshkus, Dmitry Polevoy
ICMV2
2014 On improvements of neural network accuracy with fixed number of active neurons
abstract
In this paper an improvement possibility of multilayer perceptron based classifiers with using composite classifier scheme with predictor function was exploited. Recognition of embossed number characters on plastic cards in the image taken by mobile camera was used as a model problem.
Natalia Sokolova, Dmitry P. Nikolaev, Dmitry Polevoy
ICMV2
2010 Structural Compression Of Document Images With PDF/A
abstract
This paper describes a new compression algorithm of document images based on separating the text layer from the graphics one on the initial image and compression of each layer by the most suitable common algorithm. Then compressed layers are placed into PDF/A, a standardizated file format for long-term archiving of electronic documents. Using the individual separation algorithm for each type of document makes it possible to save the image to the best advantage. Moreover, the text layer can be processed by an OCR system and the recognized text can also be placed into the same PDF/A file for making it easy to perform cut and paste and text search operations.
Sergey A. Usilin, Dmitry P. Nikolaev, Vasiliy V. Postnikov
ECMS2
2010 Visual appearance based document image classification
abstract
In the paper, we present a new method for classifying documents with rigid geometry. Our approach is based on the fast and robust Viola-Jones object detection algorithm. The advantages of our proposed method are high speed, the possibility of automatic model construction using a training set, and processing of raw source images without any pre-processing steps such as draft recognition, layout analysis or binarisation. Furthermore, our algorithm allows not only to classify documents, but also to detect the placement and orientation of documents within an image.
Sergey A. Usilin, Dmitry P. Nikolaev, Vasiliy V. Postnikov, Gerald Schaefer
ICIP2
2004 Linear color segmentation and its implementation
Dmitry P. Nikolaev, Petr P. Nikolayev
Comput. Vis. Image Underst.1