EDBT 2026 Demo / reviewers in the wild / expert
Palaiahnakote Shivakumara
dblp:83/1065 · also P. Shivakumara, Palaiahankote Shivakumara, Shivakumara Palaiahnakote
· DBLP profile ↗
42ranked-venue papers in the field
10as first author
7since 2021 · last 2026
0000-0001-9026-4613ORCID · verified
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 42 (10 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MVR: Diffusion-Based Multi-View Reasoning for Scene Text Detection
Debayan Das Gupta, Palaiahnakote Shivakumara, Palash Ghosal, Umapada Pal 0001, Cheng-Lin Liu 0001 |
ICDAR (2) | 2 |
| 2025 | Personality Trait Prediction from Twitter Data Using Text and Image Features
Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Daniel P. Lopresti, Tong Lu 0002 |
ICDAR (1) | 2 |
| 2025 | A Lightweight Context-Driven Training-Free Network for Scene Text Segmentation and Recognition
Ritabrata Chakraborty, Palaiahnakote Shivakumara, Umapada Pal 0001, Cheng-Lin Liu 0001 |
ICDAR (2) | 2 |
| 2025 | A New Fourier-Attention Guided Approach for Domain-Agnostic Text Localization
Arnab Halder, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Yue Lu 0001 |
ICDAR (3) | 2 |
| 2024 | A New Unsupervised Approach for Text Localization in Shaky and Non-shaky Scene Video
Arnab Halder, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Cheng-Lin Liu 0001 |
ICDAR (5) | 2 |
| 2023 | Gaussian Kernels Based Network for Multiple License Plate Number Detection in Day-Night Images
Soumi Das, Palaiahnakote Shivakumara, Umapada Pal 0001, Ramachandra Raghavendra |
ICDAR (5) | 2 |
| 2021 | DCINN: Deformable Convolution and Inception Based Neural Network for Tattoo Text Detection Through Skin Region
Tamal Chowdhury, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Ramachandra Raghavendra, Sukalpa Chanda |
ICDAR (2) | 2 |
| 2020 | A New Context-Based Method for Restoring Occluded Text in Natural Scene Images
Ayush Mittal, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein, Daniel P. Lopresti |
DAS | 2 |
| 2020 | A New Common Points Detection Method for Classification of 2D and 3D Texts in Video/Scene Images
Lokesh Nandanwar, Palaiahnakote Shivakumara, Ahlad Kumar, Tong Lu 0002, Umapada Pal 0001, Daniel P. Lopresti |
DAS | 2 |
| 2019 | Age Estimation using Disconnectedness Features in HandwritingabstractReal-time applications of handwriting analysis have increased drastically in the fields of forensic and information security because of accurate cues. One of such applications is human age estimation based on handwriting for the purpose of immigrant checking. In this paper, we have proposed a new method for age estimation using handwriting analysis using Hu invariant moments and disconnectedness features. To make the proposed method robust to both ruled and un-ruled documents, we propose to explore intersection point detection in Canny edge images of each input document, which results in text components. For each text component pair, we propose Hu invariant moments for extracting disconnectedness features, which in fact measure multi-shape components based on distance, shape and mutual position analysis of components. Furthermore, iterative k-means clustering is proposed for the classification of different age groups. Experimental results on our dataset and some standard datasets, namely, IAM and KHATT, show that the proposed method is effective and outperforms the state-of-the-art methods. V. Basavaraja, Palaiahnakote Shivakumara, D. S. Guru, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
ICDAR | 2 |
| 2019 | CRNN Based Jersey-Bib Number/Text Recognition in Sports and Marathon ImagesabstractThe primary challenge in tracing the participants in sports and marathon video or images is to detect and localize the jersey/Bib number that may present in different regions of their outfit captured in cluttered environment conditions. In this work, we proposed a new framework based on detecting the human body parts such that both Jersey Bib number and text is localized reliably. To achieve this, the proposed method first detects and localize the human in a given image using Single Shot Multibox Detector (SSD). In the next step, different human body parts namely, Torso, Left Thigh, Right Thigh, that generally contain a Bib number or text region is automatically extracted. These detected individual parts are processed individually to detect the Jersey Bib number/text using a deep CNN network based on the 2-channel architecture based on the novel adaptive weighting loss function. Finally, the detected text is cropped out and fed to a CNN-RNN based deep model abbreviated as CRNN for recognizing jersey/Bib/text. Extensive experiments are carried out on the four different datasets including both bench-marking dataset and a new dataset. The performance of the proposed method is compared with the state-of-the-art methods on all four datasets that indicates the improved performance of the proposed method on all four datasets. Sauradip Nag, Ramachandra Raghavendra, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Mohan Kankanhalli |
ICDAR | 3 |
| 2019 | A Text-Context-Aware CNN Network for Multi-oriented and Multi-language Scene Text DetectionabstractThe existing deep learning based state-of-theart scene text detection methods treat scene texts a type of general objects, or segment text regions directly. The latter category achieves remarkable detection results on arbitrary orientation and large aspect ratios of scene texts based on instance segmentation algorithms. However, due to the lack of context information with consideration of scene text unique characteristics, directly applying instance segmentation to text detection task is prone to result in low accuracy, especially producing false positive detection results. To ease this problem, we propose a novel text-context-aware scene text detection CNN structure, which appropriately encodes channel and spatial attention information to construct context-aware and discriminative feature map for multi-oriented and multi-language text detection tasks. With high representation ability of text context-aware feature map, the proposed instance segmentation based method can not only robustly detect multi-oriented and multi-language text from natural scene images, but also produce better text detection results by greatly reducing false positives. Experiments on ICDAR2015 and ICDAR2017-MLT datasets show that the proposed method has achieved superior performances in precision, recall and F-measure than most of the existing studies. Minglong Xue, Tong Lu 0002, Yirui Wu, Palaiahnakote Shivakumara |
ICDAR | 5 |
| 2017 | New Fuzzy-Mass Based Features for Video Image Type CategorizationabstractDue to the large variety of video type collections, it becomes difficult to achieve good text detection and recognition accuracy. We propose a new fuzzy-mass based method for classifying (categorizing) text frames from different types of video. For each frame of a video type, we formulate Fuzzy logic to identify straight and curved edge components from edge images. We then estimate mass locally and globally by drawing consecutive ellipses over edge images with respect to straight and curved edge components. Further, we extract features based on spatial proximity between centroid of classified straight/curved edge components and that of the whole image. This results local features. Next, the features are extracted for the whole image without ellipse drawing, which results in global features. The combination of both local and global features is then fed to an SVM classifier for video type classification. Experimental results on the proposed and existing classification methods show that the proposed classification outperforms the stat of art methods. Furthermore, experiments on before and after classification with several text detection and binarization methods show that the proposed classification is significant in improving text detection and recognition performance. Sangheeta Roy, Palaiahnakote Shivakumara, Namita Jain, Vijeta Khare, Umapada Pal 0001, Tong Lu 0002 |
ICDAR | 2 |
| 2017 | Temporal Integration for Word-Wise Caption and Scene Text IdentificationabstractGenerally video consists of edited text (i.e., caption text) and natural text (i.e., scene text), and these two texts differ from one another in nature as well as characteristics. Such different behaviors of caption and scene texts lead to poor accuracy for text recognition in video. In this paper, we explore wavelet decomposition and temporal coherency for the classification of caption and scene text. We propose wavelet of high frequency sub-bands to separate text candidates that are represented by high frequency coefficients in an input word. The proposed method studies the distribution of text candidates over word images based on the fact that the standard deviation of text candidates is high at the first zone, low at the middle zone and high at the third zone. This is extracted by mapping standard deviation values to 8 equal sized bins formed based on the range of standard deviation values. The correlation among bins at the first and second levels of wavelets is explored to differentiate caption and scene text and for determining the number of temporal frames to be analyzed. The properties of caption and scene texts are validated with the chosen temporal frames to find the stable property for classification. Experimental results on three standard datasets (ICDAR 2015, YVT and License Plate Video) show that the proposed method outperforms the existing methods in terms of classification rate and improves recognition rate significantly based on classification results. Sangheeta Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Ainuddin Wahid Abdul Wahab |
ICDAR | 2 |
| 2017 | Fourier-Residual for Printer IdentificationabstractPrinter identification is challenging due to advanced software technologies in the field of forgery detection. This paper presents a new idea of using the Fourier transform residual for the identification of documents printed by different printers. The proposed approach first convolves a Laplacian mask with a Fourier transform in the frequency domain to smoothen the edges. Next, we apply an inverse Fourier transform to reconstruct images from smoothed information (RFL). Similarly, the proposed approach reconstructs images using gray information of the input image (RFG). Then the residual is calculated by subtracting RFG from RFL. The set of statistical features, texture and spatial features are extracted from residual images for printer identification. Experimental results with the existing method on our dataset and a standard dataset show that the proposed approach outperforms the existing approach on both the datasets in terms of classification rate, recall, precision and F-measure. Palaiahnakote Shivakumara, Tong Lu 0002, M. Basavanna, Umapada Pal 0001, Michael Blumenstein |
ICDAR | 2 |
| 2017 | A Robust Symmetry-Based Method for Scene/Video Text Detection through Neural NetworkabstractText detection in video/scene images has gained a significant attention in the field of image processing and document analysis due to the inherent challenges caused by variations in contrast, orientation, background, text type, font type, non-uniform illumination and so on. In this paper, we propose a novel text detection method to explore symmetry property and appearance features of text for improved accuracy and robustness. First, the proposed method explores Extremal Regions (ER) for detecting text candidates in images. Then we propose a novel feature named as Multi-domain Strokes Symmetry Histogram (MSSH) for each text candidate, which describes the inherent symmetry property of stroke pixel pairs in gray, gradient and frequency domains. Furthermore, deep convolutional features are extracted to describe the appearance for each text candidate. We further fuse them by Auto-Encoder network to define a more discriminative text descriptor for classification. Finally, the proposed method constructs text lines based on the classification results. We demonstrate the effectiveness and robustness detection results of our proposed method by testing on four different benchmark databases. Yirui Wu, Wenhai Wang, Palaiahnakote Shivakumara, Tong Lu 0002 |
ICDAR | 3 |
| 2016 | New Sharpness Features for Image Type Classification Based on Textual InformationabstractAchieving good recognition results from a single method for text lines in video/natural scene images captured by high resolution cameras or low resolution mobile cameras, and images in web pages, is often hard. In this paper, we propose new sharpness based features of textual portion of each input text line image using HSI color space for the classification of an input image into one of the four classes (video, scene, mobile or born digital). This helps in choosing an appropriate method based on the class type of the input text for its improved recognition rate. For a given input text line image, the proposed method obtains H, S and I images. Then Canny edge images are obtained for H, S and I spaces, which results in text candidates. We perform sliding window operation over the text candidate image of each text line of each color space to estimate new sharpness by calculating stroke width and gradient information. The sharpness values of the text lines of the three color spaces are then fed to k-means clustering with maximum, minimum and average guesses, which results in three respective clusters. The mean of each cluster for respective color spaces outputs a feature vector having nine feature values for image classification with the help of an SVM classifier. Experimental results on standard datasets, namely, ICDAR 2013, ICDAR 2015 video, ICDAR 2015 natural scene data, ICDAR 2013 born digital data and the images captured by a mobile camera (our own data) show that the proposed classification method helps in improving recognition results. Raghunandan K. Srinivas, Palaiahnakote Shivakumara, G. Hemantha Kumar 0001, Umapada Pal 0001, Tong Lu 0002 |
DAS | 2 |
| 2015 | A new method based on bag of filters for character recognition in scene images by learningabstractAchieving a good recognition rate for scene characters is a big challenge due to non-uniform illumination effects, perspective distortions, multiple colors or contrasts, different fonts and their various sizes, background or orientation variations, etc. Unlike the existing recognition methods that use binary information or the features extracted from different domains, the proposed method explores gray information in the form of a filter bank to extract the discriminative power for all the 62 scene character classes. We propose a sliding window (patch) operation over a character image for learning the global features, which represent the structures of character images of all the classes by reconstructing a filter bank from the original data. We introduce shareable constrains to activate class-specific filters from the filter bank. Further, we propose constraints by studying the nearest neighbor patches and exemplar selection to maximize the gap between inter-classes and minimize the gap between intra-classes. The method is evaluated and compared with several existing recognition methods in terms of character recognition rate. Experimental results show that the proposed method outperforms the existing methods. Qisu Li, Tong Lu 0002, Palaiahnakote Shivakumara, Umapada Pal 0001, Chew Lim Tan |
ICDAR | 3 |
| 2015 | A new wavelet-Laplacian method for arbitrarily-oriented character segmentation in video text linesabstractCharacter segmentation is an important topic to improve the overall performance of text recognition methods due to low resolution, complex background and lots of visual variations in video. This paper presents a novel idea for segmenting characters from arbitrarily-oriented text lines based on wavelet and Laplacian combination. Firstly, we explore wavelet which decomposes a given input image into sub-levels like a pyramid structure for segmenting words based on the fact that as decomposition level increases, the gap between characters decreases due to the reduction in the size of the input image, which results in a single component for each word. Secondly, for each segmented word, we propose Laplacian wavelet combination in a new way to extract text candidates. Thirdly, we propose horizontal and vertical sampling for character segmentation from words. The proposed method is tested on curved, non-horizontal and horizontal text lines of video and the ICDAR 2005 natural scene dataset to evaluate its performance. A comparative study with an existing method shows that the proposed method outperforms it in terms of precision and f-measure. Guozhu Liang, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan |
ICDAR | 2 |
| 2014 | Separation of Graphics (Superimposed) and Scene Text in Video FramesabstractThe presence of both graphics and scene text in video frames makes text detection and recognition problem more challenging because the nature of the two texts differs significantly. This paper aims to propose a novel method for separation of graphics and scene text to achieve good recognition rate based on the fact that Canny and Sobel edge pattern share common property for text. We propose to use Ring Radius Transform to identify the radius that represents the medial axis in the edge image. We study the intra relationship between bins of the histograms over respective radius values, resulting in intra line graphs. In this way, the method finds intra line graphs for both Canny and Sobel edge images of the input text lines. To identify the unique distribution for separation of graphics and scene texts, we explore the inter relationship between intra line graphs of Canny and Sobel edge image with respective medial axes values. This results in Gaussian distribution for graphics and non-Gaussian for scene text. Experimental results on horizontal, non-horizontal, different scripts etc. show that the proposed method is effective for classification and the results of baseline recognition methods show that recognition rate is significantly improved after classification. Palaiahnakote Shivakumara, N. Vinay Kumar, D. S. Guru, Chew Lim Tan |
Document Analysis Systems | 1 |
| 2014 | A New Laplacian Method for Arbitrarily-Oriented Word Segmentation in VideoabstractWord segmentation from video text line is challenging because video poses several challenges, such as complex background, low resolution, arbitrary orientation, etc. Besides, word segmentation is essential for improving text recognition accuracy. Therefore, we propose a novel method for segmenting words by exploring zero crossing points for each sliding window over text line. The candidate zero crossing pointes are defined based on characteristics of positive and negative Laplacian values at text region and non-text region. The percentage of candidate zero crossing points is calculated for each sliding window and is used for identifying the seed window that represents space between words. For the seed window, we propose a novel idea of horizontal and vertical sampling based on the percentage values to estimate the width and the height of the word spacing. Then the width and the height of the word spacing are used to validate the actual word spacing. Experimental results comparing with an existing method show that the proposed method is better than the existing method in terms of recall, precision and f-measure on curved, horizontal, non-horizontal, Hua's video data, as well as ICDAR data. We also test it on our own data containing multiscript text lines to show the robustness of the proposed method. Palaiahnakote Shivakumara, Mahamad Suhil, D. S. Guru, Chew Lim Tan |
Document Analysis Systems | 1 |
| 2014 | Text Detection Using Delaunay Triangulation in Video SequenceabstractText detection and tracking in video sequence is gaining interest due to the challenges posed by low resolution and complex background. This paper proposes a new method for text detection by estimating trajectories between the corners of texts in video sequence over time. Each trajectory is considered as one node to form a graph for all trajectories and Delaunay triangulation is used to obtain edges to connect nodes of the graph. In order to identify the edges that represent text regions, we propose four pruning criteria based on spatial proximity, motion coherence, local appearance and canny rate. This results in several sub-graphs. Then we use depth first search to collect corner points, which essentially represent text candidates. False positives are eliminated using heuristics and missing trajectories will be obtained by tracking the corners in temporal frames. We test the method on different videos and evaluate the method in terms of recall, precision, f-measure with existing results. Experimental result shows that the proposed method is superior to existing method. Liang Wu 0009, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan |
Document Analysis Systems | 2 |
| 2013 | Scene Character Detection by an Edge-Ray FilterabstractEdge is a type of valuable clues for scene character detection task. Generally, the existing edge-based methods rely on the assumption of straight text line to prune away the non-character candidates. This paper proposes a new edge-based method, called edge-ray filter, to detect the scene character. The main contribution of the proposed method lies in filtering out complex backgrounds by fully utilizing the essential spatial layout of edges instead of the assumption of straight text line. Edges are extracted by a combination of Canny and Edge Preserving Smoothing Filter (EPSF). To effectively boost the filtering strength of the designed edge-ray filter, we employ a new Edge Quasi-Connectivity Analysis (EQCA) to unify complex edges as well as contour of broken character. Label Histogram Analysis (LHA) then filters out non-character edges and redundant rays through setting proper thresholds. Finally, two frequently-used heuristic rules, namely aspect ratio and occupation, are exploited to wipe off distinct false alarms. In addition to have the ability to handle special scenarios, the proposed method can accommodate dark-on-bright and bright-on-dark characters simultaneously, and provides accurate character segmentation masks. We perform experiments on the benchmark ICDAR 2011 Robust Reading Competition dataset as well as scene images with special scenarios. The experimental results demonstrate the validity of our proposal. Rong Huang 0003, Palaiahnakote Shivakumara, Seiichi Uchida |
ICDAR | 2 |
| 2013 | Recognition of Video Text through Temporal IntegrationabstractThis paper presents a method for temporal integration, which can be used to improve the recognition accuracy of video texts. Given a word detected in a video frame, we use a combination of Stroke Width Transform and SIFT (Scale Invariant Feature Transform) to track it both backward and forward in time. The text instances within the word's frame span are then extracted and aligned at pixel level. In the second step, we integrate these instances into a text probability map. By thresholding this map, we obtain an initial binarization of the word. In the final step, the shapes of the characters are refined using the intensity values. This helps to preserve the distinctive character features (e.g., sharp edges and holes), which are useful for OCR engines to distinguish between the different character classes. Experiments on English and German videos show that the proposed method outperforms existing ones in terms of recognition accuracy. Trung Quy Phan, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan |
ICDAR | 2 |
| 2013 | A New Method for Character Segmentation from Multi-oriented Video WordsabstractThis paper presents a two-stage method for multi-oriented video character segmentation. Words segmented from video text lines are considered for character segmentation in the present work. Words can contain isolated or non-touching characters, as well as touching characters. Therefore, the character segmentation problem can be viewed as a two stage problem. In the first stage, text cluster is identified and isolated (non-touching) characters are segmented. The orientation of each word is computed and the segmentation paths are found in the direction perpendicular to the orientation. Candidate segmentation points computed using the top distance profile are used to find the segmentation path between the characters considering the background cluster. In the second stage, the segmentation results are verified and a check is performed to ascertain whether the word component contains touching characters or not. The average width of the components is used to find the touching character components. For segmentation of the touching characters, segmentation points are then found using average stroke width information, along with the top and bottom distance profiles. The proposed method was tested on a large dataset and was evaluated in terms of precision, recall and f-measure. A comparative study with existing methods reveals the superiority of the proposed method. Nabin Sharma, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Chew Lim Tan |
ICDAR | 2 |
| 2013 | Detection of Curved Text in Video: Quad Tree Based MethodabstractIn this paper, we address curved text detection in video through a new enhancement criterion and the use of quad tree. The proposed method makes use of the quad tree to simplify the task of handling the entire frame at each stage. The proposed method employs a novel criterion for grouping of pixels based on their R, G and B values to enhance text information. As generally, a text detection problem is a two class problem, we used k-means with k=2 to identify potential text candidate pixels. From these potential candidates, connected components are then extracted and subjected to further analysis, where symmetry property based on stroke width is used for further authentication of the text representatives. These authenticated text representatives are then exploited as seed points to restore the text information with reference to the Sobel edge frame of the original input frame. To preserve the spatial information of text pixels the concept of quad tree is applied. From these seed blocks, text lines are extracted by the use of a region growing approach driven completely based on Sobel edge map. The proposed method is tested on curved video data and Hua's horizontal video text data in terms of recall, precision, f-measure, misdetection rate and processing time. The results are compared and analyzed to show that the proposed method outperforms several existing methods in terms of accuracy and efficiency. Palaiahnakote Shivakumara, H. T. Basavaraju, D. S. Guru, Chew Lim Tan |
ICDAR | 1 |
| 2013 | Scene Character Reconstruction through Medial AxisabstractCharacter shape reconstruction for the scene character is challenging and interesting because scene character usually suffers from uneven illumination, complex background, perspective distortion. To address such ill conditions, we propose to utilize Histogram Gradient Division (HGD) and Reverse Gradient Orientation (RGO) to select Candidate Text Pixels (CTPs) for a given input character. Ring Radius Transform is applied on each pixel in a CTP image to obtain radius map where each pixel is assigned a value which is the radius to the nearest CTP. Candidate medial axis pixels are those having maximum radius values in their neighborhoods. We find such pixels on horizontal, vertical, principal diagonal and secondary diagonal directions to determine the respective medial axis pixels. The union of all medial axis pixels at each pixel location is considered as a candidate medial axis pixel of the character. Then color difference and k-means clustering are employed to eliminate false candidate medial axis. The potential medial axis values are used to reconstruct the shape of the character. The method is tested on 1025 characters of complex foreground and background from ICDAR 2003 dataset in terms of shape reconstruction and recognition rate. Experimental results demonstrate the effectiveness of our proposed method for complex foreground and background characters in terms of character recognition rate and reconstruction error. Shangxuan Tian, Palaiahnakote Shivakumara, Trung Quy Phan, Chew Lim Tan |
ICDAR | 2 |
| 2012 | A New Method for Arbitrarily-Oriented Text Detection in VideoabstractText detection in video frames plays a vital role in enhancing the performance of information extraction systems because the text in video frames helps in indexing and retrieving video efficiently and accurately. This paper presents a new method for arbitrarily-oriented text detection in video, based on dominant text pixel selection, text representatives and region growing. The method uses gradient pixel direction and magnitude corresponding to Sobel edge pixels of the input frame to obtain dominant text pixels. Edge components in the Sobel edge map corresponding to dominant text pixels are then extracted and we call them text representatives. We eliminate broken segments of each text representatives to get candidate text representatives. Then the perimeter of candidate text representatives grows along the text direction in the Sobel edge map to group the neighboring text components which we call word patches. The word patches are used for finding the direction of text lines and then the word patches are expanded in the same direction in the Sobel edge map to group the neighboring word patches and to restore missing text information. This results in extraction of arbitrarily-oriented text from the video frame. To evaluate the method, we considered arbitrarily-oriented data, non-horizontal data, horizontal data, Hua's data and ICDAR-2003 competition data (Camera images). The experimental results show that the proposed method outperforms the existing method in terms of recall and f-measure. Nabin Sharma, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Chew Lim Tan |
Document Analysis Systems | 2 |
| 2012 | New Spatial-Gradient-Features for Video Script IdentificationabstractIn this paper, we present new features based on Spatial-Gradient-Features (SGF) at block level for identifying six video scripts namely, Arabic, Chinese, English, Japanese, Korean and Tamil. This works helps in enhancing the capability of the current OCR on video text recognition by choosing an appropriate OCR engine when video contains multi-script frames. The input for script identification is the text blocks obtained by our text frame classification method. For each text block, we obtain horizontal and vertical gradient information to enhance the contrast of the text pixels. We divide the horizontal gradient block into two equal parts as upper and lower at the centroid in the horizontal direction. Histogram on the horizontal gradient values of the upper and the lower part is performed to select dominant text pixels. In the same way, the method selects dominant pixels from the right and the left parts obtained by dividing the vertical gradient block vertically. The method combines the horizontal and the vertical dominant pixels to obtain text components. Skeleton concept is used to reduce pixel width to a single pixel to extract spatial features. We extract four features based on proximity between end points, junction points, intersection points and pixels. The method is evaluated on 770 frames of six scripts in terms of classification rate and is compared with an existing method. We have achieved 82.1% average classification rate. Danni Zhao, Palaiahnakote Shivakumara, Shijian Lu, Chew Lim Tan |
Document Analysis Systems | 2 |
| 2011 | Video Script Identification Based on Text LinesabstractIn this paper, we present a new method for video script identification which is essential before choosing an appropriate OCR engine for identifying text lines when a video frame contains more than one language. The input for script identification is the text lines obtained by our text detection method. We extract upper and lower extreme points for each connected component of Canny edges of text lines. The extracted points are connected to study the behavior of upper and lower lines. The direction of each 10-pixel segment of the lines is determined using PCA. The average angle of the segments of the upper and lower lines is computed to study the smoothness and cursiveness of the lines. In addition, to discriminate the scripts accurately, the method divides a text line into five equal zones horizontally to study the smoothness and cursiveness of the upper and lower lines of each zone. We evaluate the method by conducting experiments on different combinations of languages such as English and Chinese, English and Tamil, Chinese and Tamil, and English, Chinese and Tamil. Trung Quy Phan, Palaiahnakote Shivakumara, Zhang Ding, Shijian Lu, Chew Lim Tan |
ICDAR | 2 |
| 2011 | A Gradient Vector Flow-Based Method for Video Character SegmentationabstractIn this paper, we propose a method based on gradient vector flow for video character segmentation. By formulating character segmentation as a minimum cost path finding problem, the proposed method allows curved segmentation paths and thus it is able to segment overlapping characters and touching characters due to low contrast and complex background. Gradient vector flow is used in a new way to identify candidate cut pixels. A two-pass path finding algorithm is then applied where the forward direction helps to locate potential cuts and the backward direction serves to remove the false cuts, i.e. those that go through the characters, while retaining the true cuts. Experimental results show that the proposed method outperforms an existing method on multi-oriented English and Chinese video text lines. The proposed method also helps to improve binarization results, which lead to a better character recognition rate. Trung Quy Phan, Palaiahnakote Shivakumara, Bolan Su, Chew Lim Tan |
ICDAR | 2 |
| 2011 | A New Fourier-Moments Based Video Word and Character Extraction Method for RecognitionabstractThis paper presents a new method based on Fourier and moments features to extract words and characters from a video text line in any direction for recognition. Unlike existing methods which output the entire text line to the ensuing recognition algorithm, the proposed method obtains each extracted character from the text line as input to the recognition algorithm because the background of a single character is relatively simple compared to the text line and words. Max-Min clustering criterion is introduced to obtain text cluster from the extracted Fourier and moments feature set. Union of the text cluster with Canny operation of the input video text line is proposed to obtain missing text candidates. Then a run length criterion is used for extraction of words. From the words, we propose a new idea for extracting characters from the text candidates of each word image based on the fact that the text height difference at the character boundary column is smaller than that at other columns of the word image. We evaluate the method on a large dataset at three levels namely text line, words and characters in terms of recall, precision and f-measure. In addition to this, we show that the recognition result for the extracted character is better than words and lines. Our experimental set up involves 3527 characters including Chinese. The dataset is selected from TRECVID database of 2005 and 2006. Deepak Rajendran, Palaiahnakote Shivakumara, Bolan Su, Shijian Lu, Chew Lim Tan |
ICDAR | 2 |
| 2011 | A New Gradient Based Character Segmentation Method for Video Text RecognitionabstractThe current OCR cannot segment words and characters from video images due to complex background as well as low resolution of video images. To have better accuracy, this paper presents a new gradient based method for words and character segmentation from text line of any orientation in video frames for recognition. We propose a Max-Min clustering concept to obtain text cluster from the normalized absolute gradient feature matrix of the video text line image. Union of the text cluster with the output of Canny operation of the input video text line is proposed to restore missing text candidates. Then a run length algorithm is applied on the text candidate image for identifying word gaps. We propose a new idea for segmenting characters from the restored word image based on the fact that the text height difference at the character boundary column is smaller than that of the other columns of the word image. We have conducted experiments on a large dataset at two levels (word and character level) in terms of recall, precision and f-measure. Our experimental setup involves 3527 characters of English and Chinese, and this dataset is selected from TRECVID database of 2005 and 2006. Palaiahnakote Shivakumara, Souvik Bhowmick, Bolan Su, Chew Lim Tan, Umapada Pal 0001 |
ICDAR | 1 |
| 2011 | Video Character Recognition through Hierarchical ClassificationabstractWe present a new video character recognition method based on hierarchical classification. In the first step, we propose a method for character segmentation of the text line detected by the text detection method. The segmentation algorithm uses dynamic programming to find least-cost paths in the gray domain to identify the spaces between characters. For the segmented characters, we get a Canny edge image as input for the character recognition step. We introduce hierarchical classification based on voting criteria with structural features to classify 62 character classes into different smaller classes. We divide the perimeter of a character into 8 segments according to 8 directions at the centroid. Then the shape of each segment is studied to recognize the characters based on distances between the centroid and end points, and distances between the midpoint and end points. Our experiments on 1462 characters of upper case, lower case and numerals shows that 10% samples per class for training is enough to obtain 94.5% recognition accuracy. The dataset is chosen from TRECVID database of 2005 and 2006. Palaiahnakote Shivakumara, Trung Quy Phan, Shijian Lu, Chew Lim Tan |
ICDAR | 1 |
| 2010 | An eigen value based approach for text detection in videoabstractIn this paper, a novel approach for detection of text and non-text regions in video frames is proposed. The proposed approach performs block wise eigen analysis on the gradient image of the video frame. For each block of the gradient frame, the dominant eigen value is computed to decide if the block could be a candidate text block. The K-means clustering is then applied to further identify text blocks among the candidate blocks. From each of the identified candidate text blocks edges are extracted using the sobel operator, and then by the use of horizontal and vertical profiles a bounding rectangle is fixed up. Further, geometric properties of the identified text regions are studied to eliminate false text regions. In order to validate the efficacy of the proposed approach, experimentation on a dataset containing 800 video frames has been carried out. The obtained results ensure that the proposed approach is with increased text detection rate with very low false and misdetection rates when compared to the other existing state of the art techniques. D. S. Guru, S. Manjunath, Palaiahnakote Shivakumara, Chew Lim Tan |
Document Analysis Systems | 3 |
| 2010 | A skeleton-based method for multi-oriented video text detectionabstractIn this paper, we propose a method based on the skeletonization operation for multi-oriented video text detection. The first step uses our existing Laplacian-based method to identify candidate text regions. In the second step, each region is classified as either a simple connected component (a single text string) or a complex connected component (multiple text strings that are connected to each other) depending on the number of intersection points in its skeleton. Complex connected components are then segmented into constituent parts based on the skeleton segments in order to separate the text strings from each other. Finally, text string straightness and edge density are used for false positive elimination. Experimental results show that the proposed method is able to detect multi-oriented graphics text and scene text. Trung Quy Phan, Palaiahnakote Shivakumara, Chew Lim Tan |
Document Analysis Systems | 2 |
| 2010 | A new wavelet-median-moment based method for multi-oriented video text detectionabstractIn this paper, we present a new method based on wavelet-median-moments and a novel idea of angle projection for detecting multi-oriented text in video. The proposed method uses wavelet decomposition first to obtain three high frequency sub-bands (LH, HL and HH) and then median moments are computed on the average sub-bands of the three high frequency sub-bands to brighten the text pixels. K-means clustering (K=2) is used for obtaining text pixels from the wavelet-median-moments features (WMMF). Text candidates are obtained by mapping the output of K-means on Sobel edge map of the original input frame. To deal with multi-oriented text, we introduce a new idea of Angle Projection (AP) based on boundary growing and nearest neighbor concepts from the text candidates instead of conventional projection profiles. The proposed method is experimented on horizontal text data, non-horizontal text data, temporal data, non-text data and camera based images (scene text data of ICDAR 2003 competition) to show that the proposed method is superior to existing methods. Palaiahnakote Shivakumara, Anjan Dutta 0001, Chew Lim Tan, Umapada Pal 0001 |
Document Analysis Systems | 1 |
| 2009 | A Laplacian Method for Video Text DetectionabstractIn this paper, we propose an efficient text detection method based on the Laplacian operator. The maximum gradient difference value is computed for each pixel in the Laplacian-filtered image. K-means is then used to classify all the pixels into two clusters: text and non-text. For each candidate text region, the corresponding region in the Sobel edge map of the input image undergoes projection profile analysis to determine the boundary of the text blocks. Finally, we employ empirical rules to eliminate false positives based on geometrical properties. Experimental results show that the proposed method is able to detect text of different fonts, contrast and backgrounds. Moreover, it outperforms three existing methods in terms of detection and false positive rates. Trung Quy Phan, Palaiahnakote Shivakumara, Chew Lim Tan |
ICDAR | 2 |
| 2009 | A Gradient Difference Based Technique for Video Text DetectionabstractText detection in video images has received increasing attention, particularly in scene text detection in video images, as it plays a vital role in video indexing and information retrieval. This paper proposes a new and robust gradient difference technique for detecting both graphics and scene text in video images. The technique introduces the concept of zero crossing to determine the bounding boxes for the detected text lines in video images, rather than using the conventional projection profiles based method which fails to fix bounding boxes when there is no proper spacing between the detected text lines. We demonstrate the capability of the proposed technique by conducting experiments on video images containing both graphics text and scene text with different font shapes and sizes, languages, text directions, background and contrasts. Our experimental results show that the proposed technique outperforms existing methods in terms of detection rate for large video image database. Palaiahnakote Shivakumara, Trung Quy Phan, Chew Lim Tan |
ICDAR | 1 |
| 2009 | A Robust Wavelet Transform Based Technique for Video Text DetectionabstractIn this paper, we propose a new method based on wavelet transform, statistical features and central moments for both graphics and scene text detection in video images. The method uses wavelet single level decomposition LH, HL and HH subbands for computing features and the computed features are fed to k means clustering to classify the text pixel from the background of the image. The average of wavelet subbands and the output of k means clustering helps in classifying true text pixel in the image. The text blocks are detected based on analysis of projection profiles. Finally, we introduce a few heuristics to eliminate false positives from the image. The robustness of the proposed method is tested by conducting experiments on a variety of images of low contrast, complex background, different fonts, and size of text in the image. The experimental results show that the proposed method outperforms the existing methods in terms of detection rate, false positive rate and misdetection rate. Palaiahnakote Shivakumara, Trung Quy Phan, Chew Lim Tan |
ICDAR | 1 |
| 2008 | An Efficient Edge Based Technique for Text Detection in Video FramesabstractBoth graphic text and scene text detection in video images with complex background and low resolution is still a challenging and interesting problem for researchers in the field of image processing and computer vision. In this paper, we present a novel technique for detecting both graphic text and scene text in video images by finding segments containing text in an input image and then using statistical features such as vertical and horizontal bars for edges in the segments for detecting true text blocks efficiently. To identify a segment containing text, heuristic rules are formed based on combination of filters and edge analysis. Furthermore, the same rules are extended to grow the boundaries of a candidate segment in order to include complete text in the input image. The experimental results of the proposed method show that the technique performs better than existing methods in terms of a number of metrics. Palaiahnakote Shivakumara, Weihua Huang, Chew Lim Tan |
Document Analysis Systems | 1 |
| 2005 | A New Moments based Skew Estimation Technique using Pixels in the Word for Binary Document ImagesabstractAccurate skew angle estimation is an essential component in document analysis system to enhance the performance of the optical character recognition (OCR). In this paper, a new and efficient moments based method to estimate skew angle of a pixels in the word in the scanned document image is proposed. The proposed technique has two stages. In the first stage, using boundary-growing method, pixels in the words of skewed text are extracted. The pixels in the words extracted are given as input to moments based method. It results in a skew angle in the second stage. Extensive experiments have been conducted on various types of documents such as documents containing different languages and different fonts to reveal the robustness of the proposed method. Comparative studies with the well-known methods are presented to show that the proposed method is superior in terms of accuracy and computational efficiency. Palaiahnakote Shivakumara, G. Hemantha Kumar 0001, H. S. Varsha, S. Rekha, M. R. Rashmi Nayaka |
ICDAR | 1 |