Tong Lu 0002

dblp:33/4058-2 · DBLP profile ↗
← Back
17ranked-venue papers in the field
0as first author
3since 2021 · last 2025
0000-0002-7051-5347ORCID · conflict

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 17
YearPublicationVenuePosition
2025 Personality Trait Prediction from Twitter Data Using Text and Image Features
Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Daniel P. Lopresti, Tong Lu 0002
ICDAR (1)5
2023 ICDAR 2023 Competition on Born Digital Video Text Question Answering
Zhibo Yang 0003, Xiaoge Song, Sibo Song, Tong Lu 0002, Xiang Bai, Cheng-Lin Liu 0001, Fei Huang 0002, Cong Yao
ICDAR (2)4
2021 DCINN: Deformable Convolution and Inception Based Neural Network for Tattoo Text Detection Through Skin Region
Tamal Chowdhury, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Ramachandra Raghavendra, Sukalpa Chanda
ICDAR (2)4
2020 A New Context-Based Method for Restoring Occluded Text in Natural Scene Images
Ayush Mittal, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein, Daniel P. Lopresti
DAS4
2020 A New Common Points Detection Method for Classification of 2D and 3D Texts in Video/Scene Images
Lokesh Nandanwar, Palaiahnakote Shivakumara, Ahlad Kumar, Tong Lu 0002, Umapada Pal 0001, Daniel P. Lopresti
DAS4
2019 Age Estimation using Disconnectedness Features in Handwriting
abstract
Real-time applications of handwriting analysis have increased drastically in the fields of forensic and information security because of accurate cues. One of such applications is human age estimation based on handwriting for the purpose of immigrant checking. In this paper, we have proposed a new method for age estimation using handwriting analysis using Hu invariant moments and disconnectedness features. To make the proposed method robust to both ruled and un-ruled documents, we propose to explore intersection point detection in Canny edge images of each input document, which results in text components. For each text component pair, we propose Hu invariant moments for extracting disconnectedness features, which in fact measure multi-shape components based on distance, shape and mutual position analysis of components. Furthermore, iterative k-means clustering is proposed for the classification of different age groups. Experimental results on our dataset and some standard datasets, namely, IAM and KHATT, show that the proposed method is effective and outperforms the state-of-the-art methods.
V. Basavaraja, Palaiahnakote Shivakumara, D. S. Guru, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein
ICDAR5
2019 CRNN Based Jersey-Bib Number/Text Recognition in Sports and Marathon Images
abstract
The primary challenge in tracing the participants in sports and marathon video or images is to detect and localize the jersey/Bib number that may present in different regions of their outfit captured in cluttered environment conditions. In this work, we proposed a new framework based on detecting the human body parts such that both Jersey Bib number and text is localized reliably. To achieve this, the proposed method first detects and localize the human in a given image using Single Shot Multibox Detector (SSD). In the next step, different human body parts namely, Torso, Left Thigh, Right Thigh, that generally contain a Bib number or text region is automatically extracted. These detected individual parts are processed individually to detect the Jersey Bib number/text using a deep CNN network based on the 2-channel architecture based on the novel adaptive weighting loss function. Finally, the detected text is cropped out and fed to a CNN-RNN based deep model abbreviated as CRNN for recognizing jersey/Bib/text. Extensive experiments are carried out on the four different datasets including both bench-marking dataset and a new dataset. The performance of the proposed method is compared with the state-of-the-art methods on all four datasets that indicates the improved performance of the proposed method on all four datasets.
Sauradip Nag, Ramachandra Raghavendra, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Mohan Kankanhalli
ICDAR5
2019 A Text-Context-Aware CNN Network for Multi-oriented and Multi-language Scene Text Detection
abstract
The existing deep learning based state-of-theart scene text detection methods treat scene texts a type of general objects, or segment text regions directly. The latter category achieves remarkable detection results on arbitrary orientation and large aspect ratios of scene texts based on instance segmentation algorithms. However, due to the lack of context information with consideration of scene text unique characteristics, directly applying instance segmentation to text detection task is prone to result in low accuracy, especially producing false positive detection results. To ease this problem, we propose a novel text-context-aware scene text detection CNN structure, which appropriately encodes channel and spatial attention information to construct context-aware and discriminative feature map for multi-oriented and multi-language text detection tasks. With high representation ability of text context-aware feature map, the proposed instance segmentation based method can not only robustly detect multi-oriented and multi-language text from natural scene images, but also produce better text detection results by greatly reducing false positives. Experiments on ICDAR2015 and ICDAR2017-MLT datasets show that the proposed method has achieved superior performances in precision, recall and F-measure than most of the existing studies.
Minglong Xue, Tong Lu 0002, Yirui Wu, Palaiahnakote Shivakumara
ICDAR3
2017 New Fuzzy-Mass Based Features for Video Image Type Categorization
abstract
Due to the large variety of video type collections, it becomes difficult to achieve good text detection and recognition accuracy. We propose a new fuzzy-mass based method for classifying (categorizing) text frames from different types of video. For each frame of a video type, we formulate Fuzzy logic to identify straight and curved edge components from edge images. We then estimate mass locally and globally by drawing consecutive ellipses over edge images with respect to straight and curved edge components. Further, we extract features based on spatial proximity between centroid of classified straight/curved edge components and that of the whole image. This results local features. Next, the features are extracted for the whole image without ellipse drawing, which results in global features. The combination of both local and global features is then fed to an SVM classifier for video type classification. Experimental results on the proposed and existing classification methods show that the proposed classification outperforms the stat of art methods. Furthermore, experiments on before and after classification with several text detection and binarization methods show that the proposed classification is significant in improving text detection and recognition performance.
Sangheeta Roy, Palaiahnakote Shivakumara, Namita Jain, Vijeta Khare, Umapada Pal 0001, Tong Lu 0002
ICDAR6
2017 Temporal Integration for Word-Wise Caption and Scene Text Identification
abstract
Generally video consists of edited text (i.e., caption text) and natural text (i.e., scene text), and these two texts differ from one another in nature as well as characteristics. Such different behaviors of caption and scene texts lead to poor accuracy for text recognition in video. In this paper, we explore wavelet decomposition and temporal coherency for the classification of caption and scene text. We propose wavelet of high frequency sub-bands to separate text candidates that are represented by high frequency coefficients in an input word. The proposed method studies the distribution of text candidates over word images based on the fact that the standard deviation of text candidates is high at the first zone, low at the middle zone and high at the third zone. This is extracted by mapping standard deviation values to 8 equal sized bins formed based on the range of standard deviation values. The correlation among bins at the first and second levels of wavelets is explored to differentiate caption and scene text and for determining the number of temporal frames to be analyzed. The properties of caption and scene texts are validated with the chosen temporal frames to find the stable property for classification. Experimental results on three standard datasets (ICDAR 2015, YVT and License Plate Video) show that the proposed method outperforms the existing methods in terms of classification rate and improves recognition rate significantly based on classification results.
Sangheeta Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Ainuddin Wahid Abdul Wahab
ICDAR4
2017 Fourier-Residual for Printer Identification
abstract
Printer identification is challenging due to advanced software technologies in the field of forgery detection. This paper presents a new idea of using the Fourier transform residual for the identification of documents printed by different printers. The proposed approach first convolves a Laplacian mask with a Fourier transform in the frequency domain to smoothen the edges. Next, we apply an inverse Fourier transform to reconstruct images from smoothed information (RFL). Similarly, the proposed approach reconstructs images using gray information of the input image (RFG). Then the residual is calculated by subtracting RFG from RFL. The set of statistical features, texture and spatial features are extracted from residual images for printer identification. Experimental results with the existing method on our dataset and a standard dataset show that the proposed approach outperforms the existing approach on both the datasets in terms of classification rate, recall, precision and F-measure.
Palaiahnakote Shivakumara, Tong Lu 0002, M. Basavanna, Umapada Pal 0001, Michael Blumenstein
ICDAR3
2017 A Robust Symmetry-Based Method for Scene/Video Text Detection through Neural Network
abstract
Text detection in video/scene images has gained a significant attention in the field of image processing and document analysis due to the inherent challenges caused by variations in contrast, orientation, background, text type, font type, non-uniform illumination and so on. In this paper, we propose a novel text detection method to explore symmetry property and appearance features of text for improved accuracy and robustness. First, the proposed method explores Extremal Regions (ER) for detecting text candidates in images. Then we propose a novel feature named as Multi-domain Strokes Symmetry Histogram (MSSH) for each text candidate, which describes the inherent symmetry property of stroke pixel pairs in gray, gradient and frequency domains. Furthermore, deep convolutional features are extracted to describe the appearance for each text candidate. We further fuse them by Auto-Encoder network to define a more discriminative text descriptor for classification. Finally, the proposed method constructs text lines based on the classification results. We demonstrate the effectiveness and robustness detection results of our proposed method by testing on four different benchmark databases.
Yirui Wu, Wenhai Wang, Palaiahnakote Shivakumara, Tong Lu 0002
ICDAR4
2016 New Sharpness Features for Image Type Classification Based on Textual Information
abstract
Achieving good recognition results from a single method for text lines in video/natural scene images captured by high resolution cameras or low resolution mobile cameras, and images in web pages, is often hard. In this paper, we propose new sharpness based features of textual portion of each input text line image using HSI color space for the classification of an input image into one of the four classes (video, scene, mobile or born digital). This helps in choosing an appropriate method based on the class type of the input text for its improved recognition rate. For a given input text line image, the proposed method obtains H, S and I images. Then Canny edge images are obtained for H, S and I spaces, which results in text candidates. We perform sliding window operation over the text candidate image of each text line of each color space to estimate new sharpness by calculating stroke width and gradient information. The sharpness values of the text lines of the three color spaces are then fed to k-means clustering with maximum, minimum and average guesses, which results in three respective clusters. The mean of each cluster for respective color spaces outputs a feature vector having nine feature values for image classification with the help of an SVM classifier. Experimental results on standard datasets, namely, ICDAR 2013, ICDAR 2015 video, ICDAR 2015 natural scene data, ICDAR 2013 born digital data and the images captured by a mobile camera (our own data) show that the proposed classification method helps in improving recognition results.
Raghunandan K. Srinivas, Palaiahnakote Shivakumara, G. Hemantha Kumar 0001, Umapada Pal 0001, Tong Lu 0002
DAS5
2015 A new method based on bag of filters for character recognition in scene images by learning
abstract
Achieving a good recognition rate for scene characters is a big challenge due to non-uniform illumination effects, perspective distortions, multiple colors or contrasts, different fonts and their various sizes, background or orientation variations, etc. Unlike the existing recognition methods that use binary information or the features extracted from different domains, the proposed method explores gray information in the form of a filter bank to extract the discriminative power for all the 62 scene character classes. We propose a sliding window (patch) operation over a character image for learning the global features, which represent the structures of character images of all the classes by reconstructing a filter bank from the original data. We introduce shareable constrains to activate class-specific filters from the filter bank. Further, we propose constraints by studying the nearest neighbor patches and exemplar selection to maximize the gap between inter-classes and minimize the gap between intra-classes. The method is evaluated and compared with several existing recognition methods in terms of character recognition rate. Experimental results show that the proposed method outperforms the existing methods.
Qisu Li, Tong Lu 0002, Palaiahnakote Shivakumara, Umapada Pal 0001, Chew Lim Tan
ICDAR2
2015 A new wavelet-Laplacian method for arbitrarily-oriented character segmentation in video text lines
abstract
Character segmentation is an important topic to improve the overall performance of text recognition methods due to low resolution, complex background and lots of visual variations in video. This paper presents a novel idea for segmenting characters from arbitrarily-oriented text lines based on wavelet and Laplacian combination. Firstly, we explore wavelet which decomposes a given input image into sub-levels like a pyramid structure for segmenting words based on the fact that as decomposition level increases, the gap between characters decreases due to the reduction in the size of the input image, which results in a single component for each word. Secondly, for each segmented word, we propose Laplacian wavelet combination in a new way to extract text candidates. Thirdly, we propose horizontal and vertical sampling for character segmentation from words. The proposed method is tested on curved, non-horizontal and horizontal text lines of video and the ICDAR 2005 natural scene dataset to evaluate its performance. A comparative study with an existing method shows that the proposed method outperforms it in terms of precision and f-measure.
Guozhu Liang, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan
ICDAR3
2014 Text Detection Using Delaunay Triangulation in Video Sequence
abstract
Text detection and tracking in video sequence is gaining interest due to the challenges posed by low resolution and complex background. This paper proposes a new method for text detection by estimating trajectories between the corners of texts in video sequence over time. Each trajectory is considered as one node to form a graph for all trajectories and Delaunay triangulation is used to obtain edges to connect nodes of the graph. In order to identify the edges that represent text regions, we propose four pruning criteria based on spatial proximity, motion coherence, local appearance and canny rate. This results in several sub-graphs. Then we use depth first search to collect corner points, which essentially represent text candidates. False positives are eliminated using heuristics and missing trajectories will be obtained by tracking the corners in temporal frames. We test the method on different videos and evaluate the method in terms of recall, precision, f-measure with existing results. Experimental result shows that the proposed method is superior to existing method.
Liang Wu 0009, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan
Document Analysis Systems3
2013 Recognition of Video Text through Temporal Integration
abstract
This paper presents a method for temporal integration, which can be used to improve the recognition accuracy of video texts. Given a word detected in a video frame, we use a combination of Stroke Width Transform and SIFT (Scale Invariant Feature Transform) to track it both backward and forward in time. The text instances within the word's frame span are then extracted and aligned at pixel level. In the second step, we integrate these instances into a text probability map. By thresholding this map, we obtain an initial binarization of the word. In the final step, the shapes of the characters are refined using the intensity values. This helps to preserve the distinctive character features (e.g., sharp edges and holes), which are useful for OCR engines to distinguish between the different character classes. Experiments on English and German videos show that the proposed method outperforms existing ones in terms of recognition accuracy.
Trung Quy Phan, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan
ICDAR3