VLDB 2026 Research / reviewers in the wild / expert
Palaiahnakote Shivakumara
dblp:83/1065 · also P. Shivakumara, Palaiahankote Shivakumara, Shivakumara Palaiahnakote
· DBLP profile ↗
222ranked-venue papers
47as first author
95since 2021 · last 2026
0000-0001-9026-4613ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 156 · 33 first-author · 73 since 2021Graphics, computer vision, multimedia, augmented reality and games · 94 · 17 first-author · 36 since 2021Databases, data management, data science and information retrieval · 42 · 10 first-author · 7 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MVR: Diffusion-Based Multi-View Reasoning for Scene Text Detection
Debayan Das Gupta, Palaiahnakote Shivakumara, Palash Ghosal, Umapada Pal 0001, Cheng-Lin Liu 0001 |
ICDAR (2) | 2 |
| 2026 | MSTIS: Multi-views Scene Text Image Sequencing to Enhance Text Detection Performance
Debayan Das Gupta, Jayasmita Roy, Palaiahnakote Shivakumara, Umapada Pal 0001 |
ICPR (11) | 3 |
| 2026 | CPA-GNN: Contextual-Based Pattern-Aware Graph-Neural Network for Text Spotting
Anant Sinha, Palaiahnakote Shivakumara, Umapada Pal 0001, Mo Saraee |
ICPR (11) | 2 |
| 2026 | GResMark: A swin transformer-based watermarking framework with geometric attack resilience
Weitong Chen 0002, Jiale Zhang 0001, Chunpeng Ge 0001, Di Wu 0050, Willy Susilo, Palaiahnakote Shivakumara |
Expert Syst. Appl. | 7 |
| 2026 | Better utilization of illumination prior via KANs for nighttime flare removal
Aoxiang Ning, Minglong Xue, Senming Zhong, Palaiahnakote Shivakumara, Mingliang Zhou 0001 |
Neural Networks | 4 |
| 2026 | FedMLAC: Mutual learning driven heterogeneous federated audio classificationabstractFederated Learning (FL) offers a privacy-preserving framework for training audio classification (AC) models across decentralized clients without sharing raw data. However, Federated Audio Classification faces three major challenges: data heterogeneity , model heterogeneity , and data corruption , which degrade performance in real-world settings. While existing methods often address these issues separately, a unified solution remains underexplored. We propose FedMLAC, a mutual learning-based FL framework that tackles all three challenges simultaneously. Each client maintains a personalized local AC model and a lightweight, globally shared Plug-in model. These models interact via bidirectional knowledge distillation, enabling global knowledge sharing while adapting to local data distributions, thus addressing both data and model heterogeneity. To counter data corruption, we introduce a Layer-wise Pruning Aggregation (LPA) strategy that filters anomalous Plug-in updates based on parameter deviations during aggregation. Extensive experiments on four diverse AC benchmarks, including both speech and non-speech tasks, show that FedMLAC consistently outperforms state-of-the-art baselines in classification accuracy and robustness to noisy data. Rajib Rana, Di Wu 0050, Youyang Qu, Xiaohui Tao 0001, Ji Zhang 0001, Carlos Busso, Palaiahnakote Shivakumara |
Pattern Recognit. | 8 |
| 2026 | Continual relation extraction with wake-sleep memory consolidation
Tingting Hang, Jun Huang 0003, Yirui Wu, Umapada Pal 0001, Palaiahnakote Shivakumara |
Pattern Recognit. | 6 |
| 2026 | Diffusion models with spatial control and attention fusion for incremental few-shot semantic segmentation
Guangchen Shi, Yirui Wu, Palaiahnakote Shivakumara, Shirong Zou, Tong Lu 0002 |
Pattern Recognit. | 4 |
| 2026 | Explanation-guided backdoor defense for ID and OOD attacks in graph neural networks
Hao Sui 0003, Bing Chen 0002, Jiale Zhang 0001, Di Wu 0050, Palaiahnakote Shivakumara |
Pattern Recognit. | 6 |
| 2026 | UR2P-Dehaze: Learning a Simple Image Dehaze Enhancer via Unpaired Rich Physical Prior
Minglong Xue, Shuaibin Fan, Palaiahnakote Shivakumara, Mingliang Zhou 0001 |
Pattern Recognit. | 3 |
| 2026 | Unified image restoration and enhancement: Degradation calibrated cycle reconstruction diffusion model
Minglong Xue, Jinhong He, Palaiahnakote Shivakumara, Mingliang Zhou 0001 |
Pattern Recognit. | 3 |
| 2026 | Representative instance selection strategy for discriminative features
Lixin Yuan, Ningyu Du, Yirui Wu, Palaiahnakote Shivakumara, Umapada Pal 0001 |
Pattern Recognit. Lett. | 4 |
| 2026 | DFCCNet: Unified Dual-domain Fusion and Color-aware Residual Correction for Robust Single Image Dehazing
Wenchao Yan, Minglong Xue, Palaiahnakote Shivakumara, Mingliang Zhou 0001 |
Signal Process. | 3 |
| 2026 | DPMMN: A dual performer-multi-modal network for emotion recognitionabstractBackground and Objective Although emotion recognition systems have been widely advocated, their accuracy can be affected when a person’s normal facial features overlap with their expressions when in a particular emotional state. This study, therefore, explores how heatmaps of electroencephalography (EEG) signals can be integrated with facial information to improve the accuracy of emotion recognition systems. Method The key idea of the proposed work is to fuse EEG signal heatmaps and Facial information for recognizing eight different emotions. For implementing this new idea, we propose a Dual Performer Multi-Modal Network (DPMMN). For each modality, the proposed work integrates modified Vision Transformer (ViT) and Long Short-Term Memory (LSTM). The integration is achieved by concatenating the features extracted from each modality and using them to classify the different emotions. In contrast to a baseline ViT, which uses self-attention layers, the proposed work replaces self-attention layers with the Performer layers through a kernelized attention approach. This results in extracting distinct visual features from EEG signal heatmaps and facial images. Similarly, for capturing temporal features from EEG heatmaps and Facial videos, the proposed LSTM replaces a traditional feed-forward network with a recurrent structure. This step helps to learn sequential dependencies across the patches. Results A comprehensive evaluation of DPMMN with respect to current state-of-the-art systems shows favorable results, with DPMN achieving 97.02 % in identifying eight distinct emotions on the DEAP benchmark dataset. Conclusion The proposed work shows that the use of EEG signal heatmap with facial information is better than EEG signal and facial information alone. Similarly, integrating performer layers with ViT and LSTM is better than existing models for extracting distinct features to classify eight emotions. Shivanand S. Gornale, Palaiahnakote Shivakumara, Amruta Unki, Sunil Vadera |
Signal Process. Image Commun. | 2 |
| 2025 | Personality Trait Prediction from Twitter Data Using Text and Image Features
Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Daniel P. Lopresti, Tong Lu 0002 |
ICDAR (1) | 2 |
| 2025 | A Lightweight Context-Driven Training-Free Network for Scene Text Segmentation and Recognition
Ritabrata Chakraborty, Palaiahnakote Shivakumara, Umapada Pal 0001, Cheng-Lin Liu 0001 |
ICDAR (2) | 2 |
| 2025 | A New Fourier-Attention Guided Approach for Domain-Agnostic Text Localization
Arnab Halder, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Yue Lu 0001 |
ICDAR (3) | 2 |
| 2025 | Degradation-Consistent Learning via Bidirectional Diffusion for Low-Light Image EnhancementabstractLow-light image enhancement aims to improve the visibility of degraded images to better align with human visual perception. While diffusion-based methods have shown promising performance due to their strong generative capabilities. However, their unidirectional modelling of degradation often struggles to capture the complexity of real-world degradation patterns, leading to structural inconsistencies and pixel misalignments. To address these challenges, we propose a bidirectional diffusion optimization mechanism that jointly models the degradation processes of both low-light and normal-light images, enabling more precise degradation parameter matching and enhancing generation quality. Specifically, we perform bidirectional diffusion-from low-to-normal light and from normal-to-low light during training and introduce an adaptive feature interaction block (AFI) to refine feature representation. By leveraging the complementarity between these two paths, our approach imposes an implicit symmetry constraint on illumination attenuation and noise distribution, facilitating consistent degradation learning and improving the model's ability to perceive illumination and detail degradation. Additionally, we design a reflection-aware correction module (RACM) to guide color restoration post-denoising and suppress overexposed regions, ensuring content consistency and generating high-quality images that align with human visual perception. Extensive experiments on multiple benchmark datasets demonstrate that our method outperforms state-of-the-art methods in both quantitative and qualitative evaluations while generalizing effectively to diverse degradation scenarios.Code Jinhong He, Minglong Xue, Zhipu Liu, Mingliang Zhou 0001, Aoxiang Ning, Palaiahnakote Shivakumara |
ACM Multimedia | 6 |
| 2025 | A novel data-centric AI approach based on sensitivity and correlation analyses: A case study on multi-organ plant disease classificationabstractWith advancements in deep learning (DL), most research on classification problems has focused on developing or modifying DL models, known as model-centric artificial intelligence (AI) approaches. However, this approach is time-consuming and overlooks the exploration of the available resources and expertise required to address industrial problems. This study proposes a new data-centric AI-based approach by thoroughly investigating dataset complexities, using multi-organ plant disease classification as a case study. To the best of our knowledge, this study is the first to perform comprehensive sensitivity and correlation analyses to evaluate the relationship between dataset complexity exclusion and the accuracy of DL classifier. In contrast to conventional sensitivity analyses which only evaluate changes in model output with respect to input changes, this study introduces a novel Sensitivity Correlation Score (SC-score). The SC-score combines sensitivity and correlation analyses into a single metric formulated as the product of the Absolute Sensitivity Function and Pearson Correlation Coefficient which is normalised for interpretability. This formulation rewards positive sensitivity and strong correlation while neutralising the effects of negative correlation. The SC-score successfully evaluated both the responsiveness and consistency of the performance enhancement of the DL model owing to the elimination of dataset complexities. To demonstrate the robustness of this study, the proposed data-centric DL-based approach was validated on an external testing dataset from diverse agricultural environments and achieved an accuracy improvement of 10.94%. This study demonstrates the strength of data-centric AI in solving industry-oriented problems in real-world applications Muhammad Hammad Saleem, Fakhia Hammad, Muhammad Taha, Palaiahnakote Shivakumara, Sadaqat ur Rehman, Mohammad Saraee |
Expert Syst. Appl. | 4 |
| 2025 | EPAD: Ethereum phishing scam detection via graph contrastive learning
Hao Sui 0003, Jiale Zhang 0001, Bing Chen 0002, Di Wu 0050, Xiaobing Sun 0001, Palaiahnakote Shivakumara |
Expert Syst. Appl. | 6 |
| 2025 | SFFL: Self-aware fairness federated learning framework for heterogeneous data distributions
Jiale Zhang 0001, Ye Li 0041, Di Wu 0050, Yanchao Zhao, Palaiahnakote Shivakumara |
Expert Syst. Appl. | 5 |
| 2025 | A New Genetic Algorithm-Based Network for Text Localization in Degraded Social Media ImagesabstractABSTRACT This paper presents a novel model for understanding social image content through text localization. For text localization, we explore maximally stable extremal regions (MSER) for detecting components that work by clustering pixels with similar properties. The output of component detection includes several non‐text components due to the degradations of social media images. To select the best components among many, we explore the genetic algorithm by convolving different kernels with components, which results in a feature matrix that is further fed to EfficientNet for choosing actual text components. Therefore, the proposed model is called genetic algorithm based network for text localization in degraded social media images (TLDSMI). For evaluating text localization, we consider the images of the standard dataset of natural scenes by uploading and downloading from different social media platforms, namely, WhatsApp, Telegram, and Instagram. The effectiveness of our method is shown by testing on original and degraded standard datasets. For example, for the degraded images of different complexities including degradations caused by social media platforms, the proposed method performs well in almost all situations. In addition, the proposed model achieves the best F1‐Score, 0.76, 0.77, 0.70, and 0.78 for the degraded images of CUTE, ICDAR 2013, Total‐Text, and CTW1500, respectively, compared to the state‐of‐the‐art methods. Palaiahnakote Shivakumara, C. Pavan Kumar 0001, Pranjal Aggarwal, Pasupuleti Chandana, M. Basavanna, Umapada Pal 0001 |
IET Image Process. | 1 |
| 2025 | A Novel Infogain and Multi-Axial Wavelet-Based Transformer for Personality Trait Question AnsweringabstractVisual Question Answering (VQA) is one of the attractive topics in the field of multimedia, affective, and empathic computing to garner user interest. Unlike existing models which aim at addressing challenges of VQA for the scene images, this work aims at developing a new model for Personality Trait Question Answering (PQA). It uses Twitter account information, which includes shared images, profile pictures, banners, text in the images, and descriptions of the images. Motivated by the accomplishments of the transformer, for encoding visual features of the images, a new InfoGain Multi-Axial Wavelet Vision Transformer (IgMaWaViT) is explored here. For encoding textual features in the images and descriptions, a new Information Gain BERT (InfoBert) method is introduced, which can handle the variable length encoding of text by choosing the optimal discriminator. Furthermore, the model fuses encodings of images and text according to the questions on different personality traits for question answering. The model is called InfoGain Multi-Axial Wavelet Vision Transformer for Personality Traits Question Answering (IgMaWaViT-PQA). To validate the efficacy of the proposed model, a dataset has been constructed, and it is used along with standard datasets for experimentation. Comprehensive experiments show that the proposed model is better than the state-of-the-art models. The code is available at the link: https://github.com/biswaskunal29/InfoGain_MultiAxial_PQA . Kunal Biswas, Palaiahnakote Shivakumara, Saumik Bhattacharya, Umapada Pal 0001, Ram Sarkar |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2025 | Personality Traits Prediction Methods: A SurveyabstractPersonality traits prediction plays a significant role in several real-world applications, such as improving the education system, improving production in manufacturing, monitoring social media content, sentiment analysis of crowds for opinion mining, judging abnormalities in personal behavior, etc. The demand for personality traits prediction has increased drastically after COVID-19. Therefore, numerous methods have emerged to predict personality traits from a variety of sources, including handwriting, interviews, social media, text, images, and audio. This review focuses on methods developed between the years 2020 and 2024, categorizing them into Handwriting (Graphology), Vision (images and videos), Audio (speech and acoustic signals), Textual (status updates, descriptions), and Multimodal (combinations of the aforementioned) approaches. We critically analyze the methods proposed, the datasets, the scope of the work, the results obtained, and noteworthy remarks. Based on our critical analysis, we notice that increasingly methods tend to use deep learning over handcrafted features. Additionally, personality traits prediction methods are trending more toward multimodal methods because they consistently achieve the highest accuracy among the input modalities. Detailed discussions, tabular presentations, and figures facilitate easy comprehension and future reference. Then, we shed light on the challenges in this field. Many key applications are detailed. Additionally, we highlight significant limitations and offer insights into potential future directions. Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Ram Sarkar |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2025 | A Novel Topic Modeling Framework for Automated Surveying of Computer Vision Research in Prostate Cancer DetectionabstractRetrieving the relevant information or material makes a big difference in the medical field because it is necessary to acquire updates on recent findings and inferences so that the practitioners can understand the state-of-the-art topics. Thus, extracting relevant material automatically from the pool of large databases like Web of Science and IEE Explore on particular issues is important. This work focuses on extracting articles on different topics for prostate cancer detection. Due to a huge number of articles in large databases with many variations in the format, keywords, text in various forms, etc., extracting relevant articles is an open challenge. This observation motivated us to propose new Agnostic-topic modeling for retrieving articles across different categories, namely, Tasks (Image Classification, Image Segmentation, Gleason Grading, Object Detection, Image Enhancement, Motion Tracking, Image Registration, Image Synthesis, and Image Reconstruction), Models (CNN, UNet, GAN, Convolutional ML models, and Uncategorized) and Image Data Type (MRI, CT, Histological Images, Ultrasound Images, Nuclear Imaging, and Others). To address the above open challenge, the proposed work adapted ChatGPT and compared it with well-known models like LDA and BERT to show the robustness of ChatGPT. Furthermore, the extracted information is used to derive new inferences and findings. For example, the trend analysis according to the abovementioned tasks, models, and image types. To the best of our knowledge, this is the first work on retrieving articles automatically for prostate cancer detection. Razieh Fadaeidehcheshmeh, Kaveh Kiani, Palaiahnakote Shivakumara, Taha Mansouri |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2025 | Spatial-Frequency Based EEG Features for Classification of Human EmotionsabstractHuman emotion classification without bias and unfairness is challenging because most existing image-based methods are directly or indirectly affected by subjectivity. Therefore, we propose an EEG (Electroencephalogram) based model for an accurate emotion classification without the effect of subjectivity. The captured EEG signals are converted into Delta, Theta, Alpha, Beta, and Gama frequency bands. As emotions change, the frequency bands change and provide unique patterns for each emotion irrespective of different persons. With this observation, the statical features, namely, mean, standard deviation, variance, and kurtosis, and frequency-based features, namely, Power Spectral Density (PSD) and Petrosian Fractal Dimension (PFD) are extracted. To integrate the strength of spatial and frequency-based features, the features are supplied to quadratic discriminative analysis for the final classification. The experiments on the benchmark datasets, DEAP and SEED-IV, achieve 99.40% and 91.97% accuracy, respectively. A comparison with state-of-the-art methods shows that the method performs very well on some datasets. Shivanand S. Gornale, Palaiahnakote Shivakumara, Amruta Unki, Sunil Vadera |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2025 | A Newly Adopted YOLOv9 Model for Detecting Mould Regions Inside of BuildingsabstractMolds on wall and ceiling surfaces in damp indoor environments especially in houses with poor insulation and ventilation are common in the UK. Since it releases toxic chemicals as it grows, it is a serious health hazard for occupants who live in such houses. For example, eye irritation, sneezing, nose bleeds, respiratory infections, and skin irritations. Furthermore, there are chances of developing serious medical conditions like lung infections and respiratory diseases which may even lead to death. The main challenge here is that due to their irregular patterns, camouflaged with the background, it is not so easy to detect with our naked eyes in the early stage and often confused as stains. Therefore, inspired by the accomplishments of the Yolo architecture for object detection, the Yolov9 model is explored for mold detection by considering mold region as an object in this work. The overall result shows a promising 76% average classification rate. Since the mold does not have a shape, specific pattern, or color, adapting the Yolov9 for accurate mold detection is challenging. To the best of our knowledge, this is the first of its kind compared to existing methods. Since it is the first work, we constructed a dataset to perform experiments and evaluate the proposed method. To demonstrate the proposed method’s effectiveness, the results were also compared with the results of the Yolov8 and Yolov10 models. Taha Mansouri, Md Shadab Mashuk, Palaiahnakote Shivakumara, Aaron Chacko, Lawrence Sykes, Ali Alameer |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2025 | A novel domain independent scene text localizerabstract• The proposed domain independent model for scene text localization is new. • Exploring partial convolution with Yolov5-transformer for feature extraction. • Integrating the swin transformer with the novel channel attention modules. • The result of the proposed method is superior to the existing methods. Text localization across multiple domains is crucial for applications like autonomous driving and tracking marathon runners. This work introduces DIPCYT, a novel model that utilizes Domain Independent Partial Convolution and a Yolov5-based Transformer for text localization in scene images from various domains, including natural scenes, underwater, and drone images. Each domain presents unique challenges: underwater images suffer from poor quality and degradation, drone images suffer from tiny text and loss of shapes, and scene images suffer from arbitrarily oriented, shaped text. Additionally, license plates in drone images may not provide rich semantic information compared to other text types due to loss of contextual information between characters. To tackle these challenges, DIPCYT employs new partial convolution layers within Yolov5 and integrates Transformer detection heads with a novel Fourier Positional Convolutional Block Attention Module (FPCBAM). This approach leverages common text properties across domains, such as contextual (global) and spatial (local) relationships. Experimental results demonstrate that DIPCYT outperforms existing methods, achieving F-scores of 0.90, 0.90, 0.77, 0.85, 0.85, and 0.88 on Total-Text, ICDAR 2015, ICDAR 2019 MLT, CTW1500, Drone, and Underwater datasets, respectively. Ayush Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. | 2 |
| 2025 | Split-net: Dual transformer encoder with splitting scene text image for script identification
Ayush Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. Lett. | 2 |
| 2025 | Struck-out handwritten word detection and restoration for automatic descriptive answer evaluationabstract• Exploring the combination of ResNet50 and the diagonal lines for segmentation. • Proposing the combination of U-Net and Bi-LSTM for restoring text pixels. • Results show that the model is impressive for detection and restoration. Unlike objective type evaluation, descriptive answer evaluation is challenging due to unpredictable answers and free writing style of answers. Because of these, descriptive answer evaluation has received special attention from many researchers. Automatic answer evaluation is useful for the following situations. It can avoid human intervention for marking, eliminates bias marking and most important is that it can save huge manpower. To develop an efficient and accurate system, there are several open challenges. One such open challenge is cleaning the document, which includes struck-out words removal and restoring the struck-out words. In this paper, we have proposed a system for struck-out handwritten word detection and restoration for automatic descriptive answer evaluation. The work has two stages. In the first stage, we explore the combination of ResNet50 and the diagonal line (principal and secondary diagonal lines) segmentation module for detecting words and then classifying struck-out words using a classification network. In the second stage, we explore the combination of U-Net as a backbone and Bi-LSTM for predicting pixels that represent actual text information of the struck-out words based on the relationship between sequences of pixels for restoration. Experimental results on our dataset and standard datasets show that the proposed model is impressive for struck-out word detection and restoration. A comparative study with the state-of-the-art methods shows that the proposed approach outperforms the existing models in terms of struck-out word detection and restoration. Dajian Zhong, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
Signal Process. Image Commun. | 2 |
| 2025 | Zero-Shot Low-Light Image Enhancement via Joint Frequency Domain Priors Guided DiffusionabstractDue to the singularity of real-world paired datasets and the complexity of low-light environments, this leads to supervised methods lacking a degree of scene generalisation. Meanwhile, limited by poor lighting and content guidance, existing zero-shot methods cannot handle unknown severe degradation well. To address this problem, we will propose a new zero-shot low-light enhancement method to compensate for the lack of light and structural information in the diffusion sampling process by effectively combining the wavelet and Fourier frequency domains to construct rich a priori information. The key to the inspiration comes from the similarity between the wavelet and Fourier frequency domains: both light and structure information are closely related to specific frequency domain regions, respectively. Therefore, by transferring the diffusion process to the wavelet low-frequency domain and combining the wavelet and Fourier frequency domains by continuously decomposing them in the inverse process, the constructed rich illumination prior is utilised to guide the image generation enhancement process. Sufficient experiments show that the framework is robust and effective in various scenarios. Jinhong He, Palaiahnakote Shivakumara, Aoxiang Ning, Minglong Xue |
IEEE Signal Process. Lett. | 2 |
| 2024 | A New Unsupervised Approach for Text Localization in Shaky and Non-shaky Scene Video
Arnab Halder, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Cheng-Lin Liu 0001 |
ICDAR (5) | 2 |
| 2024 | A New Impressive and Expressive Features Based Model for Personality Traits Identification
Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Sukalpa Chanda, Xiaojun Wu 0001 |
ICPR (8) | 2 |
| 2024 | A New StyleGAN Latent Space Based Model for Image Style Transfer
Rakesh Dey, Palaiahnakote Shivakumara, Saumik Bhattacharya, Sukalpa Chanda, Umapada Pal 0001 |
ICPR (11) | 2 |
| 2024 | A New HourGlass Network for Detecting Text in Shaky and Non-shaky Video Frames
Arnab Halder, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Shivanand S. Gornale |
ICPR (20) | 2 |
| 2024 | Vehicle Detection Performance in Nordic Region
Hamam Mokayed, Rajkumar Saini, Oluwatosin Adewumi, Lama Alkhaled, Björn Backe, Palaiahnakote Shivakumara, Olle Hagner, Yan Chai Hum |
ICPR (22) | 6 |
| 2024 | DATR: Domain Agnostic Text Recognizer
Kunal Purkayastha, Shashwat Sarkar, Palaiahnakote Shivakumara, Umapada Pal 0001, Palash Ghosal |
ICPR (17) | 3 |
| 2024 | DITS: A New Domain Independent Text Spotter
Kunal Purkayastha, Shashwat Sarkar, Palaiahnakote Shivakumara, Umapada Pal 0001, Palash Ghosal, Xiaojun Wu 0001 |
ICPR (19) | 3 |
| 2024 | XLSI: A New Xception and Log Polar Transform Based Approach for Scene Text Script Identification
Ayush Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Apostolos Antonacopoulos, Michael Blumenstein |
ICPR (19) | 2 |
| 2024 | PCGAUNet: Pixel Correlation and Gaussian Attention Driven Network for Text Segmentation
Ayush Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Apostolos Antonacopoulos, Ramachandra Raghavendra |
ICPR (17) | 2 |
| 2024 | A New Attention Based UNet and Gated Edge Attention Network for Retinal Vessel Segmentation
Ayush Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Sukalpa Chanda |
ICPR (28) | 2 |
| 2024 | A new deep CNN for 3D text localization in the wild through shadow removal
Palaiahnakote Shivakumara, Ayan Banerjee 0002, Lokesh Nandanwar, Umapada Pal 0001, Apostolos Antonacopoulos, Tong Lu 0002, Michael Blumenstein |
Comput. Vis. Image Underst. | 1 |
| 2024 | TANet: Text region attention learning for vehicle re-identification
Wenbo Hu 0008, Hongjian Zhan, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | A robust script independent handwriting system for gender identificationabstractGender identification at the word level in a multi-script environment is challenging due to variations posed by free-style handwriting of individuals and geographical differences in writing styles. This paper presents a new approach, Multi-Orientation-Scale Gabor Response Fusion (MOSGF), for gender identification at the word level using handwritten text. Our method has two steps: (i) word segmentation from unconstrained lines and (ii) gender identification at the word level. In the first step, the method explores the number of zero crossing points and gradient information for word segmentation from handwritten text lines. In the second step, employs Gabor responses at different orientations and scales to detect fine details in female and male handwriting. For each Gabor response, the proposed model estimates the correlation between average templates of all Gabor responses and the individual Gabor response to extract global consistency in writing. To strengthen correlation features, the proposed method uses the Mahalanobis distance measure, which extracts local similarity. Further, the proposed approach fuses correlation coefficient and distance-based features in a novel way. The fused features are then fed to a Neural Network (NN) for gender identification. Experiments on our dataset, which comprises Roman (English), Chinese, Farsi (Persian), Arabic, and Indian scripts, and a benchmark dataset, namely, IAM which includes English text, KHATT which includes Arabic, and QUWI which includes both English and Arabic, show that the proposed system outperforms the existing methods in terms of word segmentation and gender identification. Palaiahnakote Shivakumara, Maryam Asadzadeh Kaljahi, Swati Kanchan, Umapada Pal 0001, Daniel P. Lopresti, Tong Lu 0002 |
Expert Syst. Appl. | 1 |
| 2024 | A new U-Net based system for multi-cultural wedding image classificationabstractUse of social media for communication, sharing or expressing views, broadcasting news, threatening and blackmailing has become an integral part of society. One such activity is understanding multi-cultural wedding images uploaded on social media. This paper presents a novel method based on the combination of U-Net, Convolutional Neural Network and Random Forest for classification of multicultural wedding images. In the case of wedding images, bride and bridegroom draw the attention of the viewers. This observation led to propose a U-Net for segmenting the region of bride and bridegroom in a novel way. Similarly, it is noted that the costumes of bride and bridegroom are vital information for differentiating different cultures. This cue motivated us to extract features using CNN for classification. Since the extracted features using CNN are capable of discriminating images of different classes, we propose a simple and effective Random-Forest for Multicultural Wedding Image Classification . The efficiency of the proposed model is demonstrated by testing it on our own dataset of six multi-cultural wedding classes and standard dataset of wedding and non-wedding images classes. Experimental results on both the datasets show that the proposed model outperforms the state-of-the-art models in terms of average classification rate. Palaiahnakote Shivakumara, C. Pavan Kumar 0001, Jagrut Nemade, Kshitiz Michael, Akash Kumar 0015, Basavaraj S. Anami, Umapada Pal 0001 |
Expert Syst. Appl. | 1 |
| 2024 | A novel autoencoder for structural anomalies detection in river tunnel operation
Xuyan Tan, Palaiahnakote Shivakumara, Ke Cheng 0003, Bowen Du 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Weakly supervised scene text generation for low-resource languages
Yangchen Xie, Hongjian Zhan, Palaiahnakote Shivakumara, Cong Liu 0006, Yue Lu 0001 |
Expert Syst. Appl. | 4 |
| 2024 | NDOrder: Exploring a novel decoding order for scene text recognition
Dajian Zhong, Hongjian Zhan, Shujing Lyu, Cong Liu 0006, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
Expert Syst. Appl. | 6 |
| 2024 | A New Symmetry-Based Transformer for Text Spotting in Person and Vehicle Re-Identification ImagesabstractText spotting in person and vehicle re-identification images is complex due to the presence of multiple views of the same person and vehicle. Most existing models focus on text spotting in natural scene images, our work focuses on spotting in person and vehicle re-identification images. The rationale behind this work is that the person and the vehicles share symmetry properties and the bib number in the torso and license plate number in the vehicle are text. The method divides the input image into patches, and it explores vision transformation for encoding the patches into linear patches. The linearly embedded patches are fed to the feature similarity index step, which involves phase congruency and gradient magnitude to detect symmetric patches. The transformer is proposed to encode and capture textual information from the symmetry patches for text detection and recognition. The decoder receives the attention features from the encoder and fetches a multi-task head with the information about the detected and recognized text. The experiments on person and vehicle image benchmark, viz. (Person) Re-ID, RBNR, UFPR-ALPR and RodoSol datasets show significant improvement in performance when compared to other text spotting models. The effectiveness of the proposed model is validated by testing on the benchmark datasets, namely, ICDAR 2015, Total-Text and CTW1500 of natural scene images. Furthermore, cross-data validation shows the proposed method is independent of domains. Aritro Pal Choudhury, Palaiahnakote Shivakumara, Umapada Pal 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2024 | A Locally Weighted Linear Regression-Based Approach for Arbitrary Moving Shaky and Nonshaky Video ClassificationabstractClassification and identification of objects are complex and challenging in pattern recognition and artificial intelligence if a shaky and nonshaky camera captures the videos at different distances during the day and nighttime. This work presents a model for classifying a given video as a static, uniform, or arbitrarily moving videos so that the complexity of the problem can be reduced. To avoid the threat of different distances between the objects and the camera, the proposed work introduces new steps for estimating the depth of the objects in the video frames. We explore locally weighted linear regression for feature extraction from depth information based on the notion that the regression line fits almost all the points for uniformity and does not fit for arbitrary moving. The extracted features are fed to a random forest classifier to classify static, uniform, or arbitrary moving video. The results on a large dataset, which includes videos captured day and night, show that the proposed method successfully classifies static, uniform and arbitrary videos with 0.86, 1.00 and 0.67 F-measures, respectively. Overall, our method obtains 87% accuracy for classification of static, uniform and arbitrary video, which is superior to the state-of-the-art methods. Arnab Halder, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Palash Ghosal |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2024 | A New Approach for Classification of Spices to Make Special Herbal Tea Using Caralluma FimbriataabstractClassification of multiple types of spice images is automatically challenging due to conflict between the texture patterns of spice images. This work aims to develop an automatic system for classifying different types of spice images so that the system can choose an appropriate spice to make herbal tea using Caralluma fimbriata. This work considers the following seven spices, namely, cinnamon, citrus peel, clove, ginger, jeera, kokum, mint, and Caralluma fimbriata as one more class for classification. Most of the existing systems need human intervention to choose different spices to make Caralluma fimbriata tea. It is observed that the pattern of different spice images represents different textures. This observation motivated us to extract features based on multi-Sobel kernels. To reduce the number of computations, the proposed work introduces a novel idea of corner detection based on Gaussian distribution. For each corner, the method performed is multi-Sobel kernels for extracting features. The features are fed to convolutional neural network layers for the classification of multiple spice images. The results of our dataset and comparative study with the state-of-the-art methods show that the proposed model is superior to existing methods in terms of classification rate. Prabhuswamy Prajwal Kumar, Palaiahnakote Shivakumara, M. Basavanna, H. S. Ravikumar Patil, V. M. Vyshali |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2024 | A New Contrastive Learning-Based Vision Transformer for Sentiment Analysis Using Scene Text ImagesabstractSentiment analysis using scene text images is complex and challenging because it has an arbitrary background, and the method should rely on only visual features. Unlike most existing methods that use either text or images or both, this study uses only scene text images for sentiment analysis. The intuition to use only scene text images is that sometimes users express their feelings and emotions or convey their messages by writing text in different shapes with diverse background designs. It is noted that the existing methods ignore such vital cues for sentiment analysis. This work explores a vision transformer to extract visual features that represent contextual information about the appearance of the text image. Further, to strengthen the visual features, the proposed work introduces contrastive learning which maximizes the gap between inter-classes and minimizes the gap between intra-classes of positive, negative, and neutral. To demonstrate the effectiveness of the proposed method, it is tested on our own constructed dataset and benchmark dataset. A comparative study of our method with the existing method shows the proposed method is superior in the classification of positive, negative, and neutral scene text images. Palaiahnakote Shivakumara, Dhruv Kapri, Muhammad Hammad Saleem, Umapada Pal 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2024 | Altered Handwritten Text Detection in Document Images Using Deep LearningabstractHandwritten documents possess immense significance in domains such as law, history, and administration. However, they are vulnerable to forgery, which can undermine their credibility and reliability. This paper aims to establish a dependable technique for identifying altered text in handwritten document images, even in scenarios with high levels of noise and blur. Our study investigates 10 distinct categories of handwritten text that have been altered through various forgery operations. The suggested approach employs the deep neural architectures VGG16 and Resnet50 as feature extractors. The architecture comprises three parts: Feature extraction using individual models, a feature fusion layer, and a classification layer. Initially, we optimize the training process and feature extraction using VGG16 and ResNet50. The feature vectors obtained from both models are then fused together in the feature fusion layer and input into the classification layer for the classification task. Experiments are conducted on a custom-created dataset as well as benchmark datasets including ICPR FDC, IMEI Forged Number, and Kundu to demonstrate that the proposed method is superior to existing approaches. Gayatri Patil, Palaiahnakote Shivakumara, Shivanand S. Gornale, Daniel P. Lopresti |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2024 | An Adaptive Xception Model for Classification of Brain TumorsabstractClassification of different brain tumors is challenging due to unpredictable variations in intra-inter-classes. Unlike existing methods which are not effective for images of complex backgrounds, the proposed work aims at accurate classification of diverse types of brain tumors such that an appropriate model can be used for disease identification. This study considers glioma, meningioma, no tumor, and pituitary tumors for classification. To achieve an accurate classification, we explore the Xception architecture layer, which involves flattening, dropout, and dense layer operations. The model extracts features based on shapes, spatial relationships, and structure of the image, discriminating between the different brain tumor images. The model is evaluated on a dataset of 7023 MRI images for classification. The results of a large dataset and comparative study with the existing methods show that the proposed method is better than state of the art in terms of classification rate. Specifically, our method achieves more than a 90% average classification rate, which is better than state of the art. The results on noisy and blurred datasets show that the proposed model is robust to noise and blur. Arastu Thakur, Surbhi Bhatia, Palaiahnakote Shivakumara, Vinoth Kumar Venkatesan 0001, Ahlam Almusharraf, Arwa A. Mashat |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2024 | Soft set-based MSER end-to-end system for occluded scene text detection, recognition and predictionabstractThe presence of unpredictable occlusions on natural scene text is a significant challenge, exacerbating the difficulties already posed on text detection and recognition by the variability of such images. Addressing the need for a robust, consistently performing approach that can effectively address the above challenges, this paper presents a new Soft Set-based end-to-end system for text detection, recognition and prediction in occluded natural scene images. This is the first approach to integrate text detection, recognition and prediction , unlike existing systems developed for end-to-end text spotting (text detection and recognition) only. For candidate text components detection, the proposed combination of Soft Sets with Maximally Stable Extremal Regions (SS-MSER) improves text detection and spotting in natural scene images, irrespectively of the presence of arbitrarily orientated and shaped text, complex backgrounds and occlusion. Furthermore, a Graph Recurrent Neural Network is proposed for grouping candidate text components into text lines and for fitting accurate bounding boxes to each word. Finally, a Convolutional Recurrent Neural Network (CRNN) is proposed for the recognition of text and for predicting missing characters due to occlusion. Experimental results on a new occluded scene text dataset (OSTD) and on the most relevant benchmark natural scene text datasets demonstrate that the proposed system outperforms the state-of-the-art in text detection, recognition and prediction. The code and dataset are available at https://github.com/alloydas/Softset-MSER-Based-Occluded-Scene-Text-Spotting/blob/master/Soft_set_MSER.ipynb Alloy Das, Palaiahnakote Shivakumara, Ayan Banerjee 0002, Apostolos Antonacopoulos, Umapada Pal 0001 |
Knowl. Based Syst. | 2 |
| 2024 | An end-to-end model for multi-view scene text recognition
Ayan Banerjee 0002, Palaiahnakote Shivakumara, Saumik Bhattacharya, Umapada Pal 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. | 2 |
| 2024 | TTS: Hilbert Transform-Based Generative Adversarial Network for Tattoo and Scene Text SpottingabstractText spotting in natural scenes is of increasing interest and significance due to its critical role in several applications, such as visual question answering, named entity recognition and event rumor detection on social media. One of the newly emerging challenging problems is Tattoo Text Spotting (TTS) in images for assisting forensic teams and for person identification. Unlike the generally simpler scene text addressed by current state-of-the-art methods, tattoo text is typically characterized by the presence of decorative backgrounds, calligraphic handwriting and several distortions due to the deformable nature of the skin. This paper describes the first approach to address TTS in a real-world application context by designing an end-to-end text spotting method employing a Hilbert transform-based Generative Adversarial Network (GAN). To reduce the complexity of the TTS task, the proposed approach first detects fine details in the image using the Hilbert transform and the Optimum Phase Congruency (OPC). To overcome the challenges of only having a relatively small number of training samples, a GAN is then used for generating suitable text samples and descriptors for text spotting (i.e., both detection and recognition). The superior performance of the proposed TTS approach, for both tattoo and general scene text, over the state-of-the-art methods is demonstrated on a new TTS-specific dataset (publicly available) as well as on the existing benchmark natural scene text datasets: Total-Text, CTW1500 and ICDAR 2015. Ayan Banerjee 0002, Palaiahnakote Shivakumara, Umapada Pal 0001, Apostolos Antonacopoulos, Tong Lu 0002, Josep Lladós 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | A Conformable Moments-Based Deep Learning System for Forged Handwriting DetectionabstractDetecting forged handwriting is important in a wide variety of machine learning applications, and it is challenging when the input images are degraded with noise and blur. This article presents a new model based on conformable moments (CMs) and deep ensemble neural networks (DENNs) for forged handwriting detection in noisy and blurry environments. Since CMs involve fractional calculus with the ability to model nonlinearities and geometrical moments as well as preserving spatial relationships between pixels, fine details in images are preserved. This motivates us to introduce a DENN classifier, which integrates stenographic kernels and spatial features to classify input images as normal (original, clean images), altered (handwriting changed through copy-paste and insertion operations), noisy (added noise to original image), blurred (added blur to original image), altered-noise (noise is added to the altered image), and altered-blurred (blur is added to the altered image). To evaluate our model, we use a newly introduced dataset, which comprises handwritten words altered at the character level, as well as several standard datasets, namely ACPR 2019, ICPR 2018-FDC, and the IMEI dataset. The first two of these datasets include handwriting samples that are altered at the character and word levels, and the third dataset comprises forged International Mobile Equipment Identity (IMEI) numbers. Experimental results demonstrate that the proposed method outperforms the existing methods in terms of classification rate. Lokesh Nandanwar, Palaiahnakote Shivakumara, Hamid Abdullah Jalab, Rabha W. Ibrahim, Ramachandra Raghavendra, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Gaussian Kernels Based Network for Multiple License Plate Number Detection in Day-Night Images
Soumi Das, Palaiahnakote Shivakumara, Umapada Pal 0001, Ramachandra Raghavendra |
ICDAR (5) | 2 |
| 2023 | A New Lightweight Script Independent Scene Text Style Transfer NetworkabstractScene text style transfer without a language barrier is an open challenge for the video and scene text recognition community because this plays a vital role in poster, web design, augmenting character images, and editing characters to improve scene text recognition performance and usability. This work presents a new model, called Script Independent Scene Text Style Transfer Network (SISTSTNet), for extracting scene characters and transferring text style simultaneously. The SISTSTNet performs mapping in language-independent feature space for transferring style. It is designed based on a Style Parameter Network and Target Encoder Network through lightweight MobileNetv3 convolutional and residual blocks to capture the style and shape to generate target characters. Similarly, a generative model is explored through the Visual Geometry Group (VGG) network for character replacement. The SISTSTNet is flexible and works on different languages and arbitrary examples in a neat and unified fashion. The experimental results on images in various languages, namely, English, Chinese, Hindi, Russian, Japanese, Arabic, Greek, and Bengali and cross-language validation demonstrate the effectiveness of the proposed method. The performance of the method is superior compared to the state-of-the-art methods in terms of quality measures, language independence, shape-preserving, and efficiency. The code and dataset will be released to the public to support reproducibility. Palaiahnakote Shivakumara, Ayush Roy, Lokesh Nandanwar, Umapada Pal 0001, Yue Lu 0001, Cheng-Lin Liu 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2023 | Classification of aesthetic natural scene images using statistical and semantic features
Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein, Josep Lladós 0001 |
Multim. Tools Appl. | 2 |
| 2023 | Writer age estimation through handwriting
Zhiheng Huang, Palaiahnakote Shivakumara, Maryam Asadzadeh Kaljahi, Ahlad Kumar, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
Multim. Tools Appl. | 2 |
| 2023 | A new robust approach for altered handwritten text detection
Gayatri Patil, Palaiahnakote Shivakumara, Shivanand S. Gornale, Umapada Pal 0001, Michael Blumenstein |
Multim. Tools Appl. | 2 |
| 2023 | VQAPT: A New visual question answering model for personality traits in social media images
Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Cheng-Lin Liu 0001, Yue Lu 0001 |
Pattern Recognit. Lett. | 2 |
| 2023 | SANet-SI: A new Self-Attention-Network for Script Identification in scene images
Hongjian Zhan, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
Pattern Recognit. Lett. | 3 |
| 2023 | A New Few-Shot Learning-Based Model for Prohibited Objects Detection in Cluttered Baggage X-Ray Images Through Edge Detection and Reverse ValidationabstractDetecting prohibited items via X-ray screening at airports and sensitive venues is essential for preventing smuggling and breaches of security. The difficulty in prohibited items inspection lies in accurately detecting prohibited items in complex X-ray images and limited access to X-ray images containing prohibited items. Few-shot detection aims at learning with limited examples and assigning a category label to each object. However, most few-shot learning methods do not focus on the edge information of the occluded object in X-ray images, which is crucial for the model to detect prohibited items in the X-ray images. In this paper, we presents a method (RVViT) for few-shot prohibited items detection tasks which fully acknowledges the significance of X-ray penetrability and increases the stability of few-shot learning model. Specifically, a Transformer encoder is firstly adopted for generating high-level semantic features that contain global information. At the same time, an edge detection module is devised for enhancing the edge information of prohibited items. Moreover, to further improve the stability of the few-shot learning model and ensure prototype consistency between the support and query samples, a reverse validation strategy is proposed to assist training. Extensive experiments demonstrate our method outperforms state-of-the-art approaches in terms of detection with a small number of samples. Shujing Lyu, Palaiahnakote Shivakumara, Michael Blumenstein, Yue Lu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2023 | A New Language-Independent Deep CNN for Scene Text Detection and Style Transfer in Social Media ImagesabstractDue to the adverse effect of quality caused by different social media and arbitrary languages in natural scenes, detecting text from social media images and transferring its style is challenging. This paper presents a novel end-to-end model for text detection and text style transfer in social media images. The key notion of the proposed work is to find dominant information, such as fine details in the degraded images (social media images), and then restore the structure of character information. Therefore, we first introduce a novel idea of extracting gradients from the frequency domain of the input image to reduce the adverse effect of different social media, which outputs text candidate points. The text candidates are further connected into components and used for text detection via a UNet++ like network with an EfficientNet backbone (EffiUNet++). Then, to deal with the style transfer issue, we devise a generative model, which comprises a target encoder and style parameter networks (TESP-Net) to generate the target characters by leveraging the recognition results from the first stage. Specifically, a series of residual mapping and a position attention module are devised to improve the shape and structure of generated characters. The whole model is trained end-to-end so as to optimize the performance. Experiments on our social media dataset, benchmark datasets of natural scene text detection and text style transfer show that the proposed model outperforms the existing text detection and style transfer methods in multilingual and cross-language scenario. Palaiahnakote Shivakumara, Ayan Banerjee 0002, Umapada Pal 0001, Lokesh Nandanwar, Tong Lu 0002, Cheng-Lin Liu 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | SGBANet: Semantic GAN and Balanced Attention Network for Arbitrarily Oriented Scene Text Recognition
Dajian Zhong, Shujing Lyu, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
ECCV (28) | 3 |
| 2022 | EAU-Net: A New Edge-Attention Based U-Net for Nationality Identification
Aritro Pal Choudhury, Palaiahnakote Shivakumara, Umapada Pal 0001, Cheng-Lin Liu 0001 |
ICFHR | 2 |
| 2022 | TWD: A New Deep E2E Model for Text Watermark/Caption and Scene Text Detection in VideoabstractText watermark detection in video images is challenging because text watermark characteristics are different from caption and scene texts in the video images. Developing a successful model for detecting text watermark, caption, and scene texts is an open challenge. This study aims at developing a new Deep End-to-End model for Text Watermark Detection (TWD), caption and scene text in video images. To standardize non-uniform contrast, quality, and resolution, we explore the U-Net3+ model for enhancing poor quality text without affecting high-quality text. Similarly, to address the challenges of arbitrary orientation, text shapes and complex background, we explore Stacked Hourglass Encoded Fourier Contour Embedding Network (SFCENet) by feeding the output of the U-Net3+ model as input. Furthermore, the proposed work integrates enhancement and detection models as an end-to-end model for detecting multi-type text in video images. To validate the proposed model, we create our own dataset (named TW-866), which provides video images containing text watermark, caption (subtitles), as well as scene text. The proposed model is also evaluated on standard natural scene text detection datasets, namely, ICDAR 2019 MLT, CTW1500, Total-Text, and DAST1500. The results show that the proposed method outperforms the existing methods. This is the first work on text watermark detection in video images to the best of our knowledge. Ayan Banerjee 0002, Palaiahnakote Shivakumara, Parikshit Acharya, Umapada Pal 0001, Josep Lladós 0001 |
ICPR | 2 |
| 2022 | DPAM: A New Deep Parallel Attention Model for Multiple License Plate Number RecognitionabstractLicense plate number recognition is challenging for complex scenes containing multiple vehicles of different types, shapes, distances etc. To recognize multiple license plate numbers in an image, we propose a new model, called Deep Parallel Attention Model (DPAM), which simultaneously extracts unique features at character levels. The proposed model exploits the observation that the combination of alphanumeric characters does not have correlation at semantic level for extracting the features. This led to the introduction of parallelism for feature extraction at character levels to make it efficient in terms of time to fit in a real time environment. To test the proposed model, we consider our own dataset consisting of Indian license plate numbers and other standard datasets to show the superiority of the proposed model over the existing methods in terms of recognition rate. Furthermore, the proposed method is tested on scene text dataset to show its ability to detect text in natural scene images. Amish Kumar, Palaiahnakote Shivakumara, Pinaki Nath Chowdhury, Umapada Pal 0001, Cheng-Lin Liu 0001 |
ICPR | 2 |
| 2022 | Text line segmentation from struck-out handwritten document images
Palaiahnakote Shivakumara, Tanmay Jain, Umapada Pal 0001, Nitish Surana, Apostolos Antonacopoulos, Tong Lu 0002 |
Expert Syst. Appl. | 1 |
| 2022 | Text proposals with location-awareness-attention network for arbitrarily shaped scene text detection and recognition
Dajian Zhong, Shujing Lyu, Palaiahnakote Shivakumara, Umapada Pal 0001, Yue Lu 0001 |
Expert Syst. Appl. | 3 |
| 2022 | A Knowledge Enforcement Network-Based Approach for Classifying a Photographer's ImagesabstractClassification of photos captured by different photographers is an important and challenging problem in knowledge-based and image processing. Monitoring and authenticating images uploaded on social media are essential, and verifying the source is one key piece of evidence. We present a novel framework for classifying photos of different photographers based on the combination of local features and deep learning models. The proposed work uses focused and defocused information in the input images to extract contextual information. The model estimates the weighted gradient and calculates entropy to strengthen context features. The focused and defocused information is fused to estimate cross-covariance and define a linear relationship between them. This relationship results in a feature matrix fed to Knowledge Enforcement Network (KEN) for obtaining representative features. Due to the strong discriminative ability of deep learning models, we employ the lightweight and accurate MobileNetV2. The output of KEN and MobileNetV2 is sent to a classifier for photographer classification. Experimental results of the proposed model on our dataset of 46 photographer classes (46234 images) and publicly available datasets of 41 photographer classes (218303 images) show that the method outperforms the existing techniques by 5%–10% on average. The dataset created for the experimental purpose will be made available upon publication. Palaiahnakote Shivakumara, Pinaki Nath Chowdhury, Umapada Pal 0001, David S. Doermann, Ramachandra Raghavendra, Tong Lu 0002, Michael Blumenstein |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2022 | New Deep Spatio-Structural Features of Handwritten Text Lines for Document Age ClassificationabstractDocument age estimation using handwritten text line images is useful for several pattern recognition and artificial intelligence applications such as forged signature verification, writer identification, gender identification, personality traits identification, and fraudulent document identification. This paper presents a novel method for document age classification at the text line level. For segmenting text lines from handwritten document images, the wavelet decomposition is used in a novel way. We explore multiple levels of wavelet decomposition, which introduce blur as the number of levels increases for detecting word components. The detected components are then used for a direction guided-driven growing approach with linearity, and nonlinearity criteria for segmenting text lines. For classification of text line images of different ages, inspired by the observation that, as the age of a document increases, the quality of its image degrades, the proposed method extracts the structural, contrast, and spatial features to study degradations at different wavelet decomposition levels. The specific advantages of DenseNet, namely, strong feature propagation, mitigation of the vanishing gradient problem, reuse of features, and the reduction of the number of parameters motivated us to use DenseNet121 along with a Multi-layer Perceptron (MLP) for the classification of text lines of different ages by feeding features and the original image as input. To demonstrate the efficacy of the proposed model, experiments were conducted on our own as well as standard datasets for both text line segmentation and document age classification. The results show that the proposed method outperforms the existing methods for text line segmentation in terms of precision, recall, F-measure, and document age classification in terms of average classification rate. Palaiahnakote Shivakumara, Alloy Das, Raghunandan K. Srinivas, Umapada Pal 0001, Michael Blumenstein |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2022 | Local Resultant Gradient Vector Difference and Inpainting for 3D Text Detection in the WildabstractThree-dimensional (3D) text appearing in natural scene images is common due to 3D cameras and the capture of text from different angles, which presents new problems for text detection. This is because of the presence of depth information, shadows, and decorative characters in the images. In this work, we consider those images where 3D text appears with depth, as well as shadow information for text detection. We propose a novel method based on local resultant gradient vector difference (LRGVD), inpainting and a deep learning model for detecting 3D as well as two-dimensional (2D) texts in natural scene images. The boundary of components that are invariant to the above challenges is detected by exploring LRGVD. The LRGVD uses gradient magnitude and direction in a novel way for detecting the boundary of the components. Further, we propose an inpainting method in a new way for restoring the character background information using boundaries. For a given region and the input image, the inpainting method divides the whole image into planes and then propagates the values in the planes into the missing region based on posterior probabilities and neighboring information. This results in text regions with false positives. Then, the differential binarization network (DB-Net) is proposed for detecting text irrespective of orientation, background, 3D or 2D, etc. Experiments conducted on our 3D text images and standard datasets of natural scene text images, namely ICDAR 2019 MLT, ICDAR 2019 ArT, DAST1500, Total-Text and SCUT-CTW1500, show that the proposed method is effective in detecting 3D and 2D texts in the images. Dajian Zhong, Palaiahnakote Shivakumara, Lokesh Nandanwar, Umapada Pal 0001, Michael Blumenstein, Yue Lu 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2022 | Fuzzy and genetic algorithm based approach for classification of personality traits oriented social media images
Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Tapabrata Chakraborti, Tong Lu 0002, Mohamad Nizam Ayub |
Knowl. Based Syst. | 2 |
| 2022 | A new deep model for family and non-family photo identification
Tapan Karnik, Palaiahnakote Shivakumara, Pinaki Nath Chowdhury, Umapada Pal 0001, Tong Lu 0002, Nor Badrul Anuar |
Multim. Tools Appl. | 2 |
| 2022 | A comprehensive scheme for tattoo text detectionabstractTattoo text detection provides a vital clue for person and crime identification. Due to the freestyle and unconstrained nature of handwritten tattoo text over skin regions, accurate tattoo text detection is very challenging. This paper proposes a comprehensive scheme for tattoo text detection which comprises (a) adaptive Deformable Convolutional Neural Network (DCNN) for skin region detection to reduce text detection complexity (b) a Decoupled Gradient Text Detector (DGTD) for tattoo text detection from skin region (c) a Deep Q-Network (DQN) to refine the bounding boxes detected by DGTD, and (d) a Term-Frequency-Inverse-Document-Frequency (TF-IDF) model to group the words into text lines based on semantic information to fix the bounding box for the line. To test the effectiveness, the proposed method is evaluated on different datasets, namely, (i) a newly developed tattoo text dataset, (ii) benchmark bib number dataset of the marathon, and (iii) person re-identification dataset. The proposed method achieves 91.2, 87.5, and 88.8 F-scores from these three respective datasets. To demonstrate its superior performance, the text detection module (without skin detection) is also compared with state-of-the-art scene text detection methods on benchmark datasets, namely, ICDAR 2019 ArT, Total-Text, and DAST1500 and the proposed method achieves 90.3, 88.5 and 89.8 F-score from these respective datasets. Ayan Banerjee 0002, Palaiahnakote Shivakumara, Umapada Pal 0001, Ramachandra Raghavendra, Cheng-Lin Liu 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | Oil palm tree counting in drone images
Pinaki Nath Chowdhury, Palaiahnakote Shivakumara, Lokesh Nandanwar, Faizal Samiron, Umapada Pal 0001, Tong Lu 0002 |
Pattern Recognit. Lett. | 2 |
| 2022 | A new method for detection and prediction of occluded text in natural scene images
Ayush Mittal, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
Signal Process. Image Commun. | 2 |
| 2022 | An Episodic Learning Network for Text Detection on Human Bodies in Sports ImagesabstractDue to the proliferation of sports-related multimedia content on the WWW, effective visual search and retrieval present interesting research challenges. These are caused by poor image quality, a wide range of possible camera points of view, pose variations on the part of athletes engaged in playing a sport, deformations of text appearing on sports person’s clothing and uniforms in motion, occlusions caused by other objects, etc. To address these challenges, this paper presents a new method for detecting text on human bodies in sports images. Unlike most existing methods, which attempt to exploit locations of a player’s torso, face, and skin, we propose an end-to-end episodic learning approach that employs inductive learning criteria for detecting clothing regions in an image, which are, in turn, then used for text detection. Our method integrates a Residual Network (ResNet) and Pyramidal Pooling Module (PPM) for generating a spatial attention map. The Progressive Scalable Expansion Algorithm (PSE) is adapted for text detection from these regions. Experimental results on our own dataset as well as several benchmarks (like RBNR and MMM which contain images of runners in marathons, and Re-ID which is a person re-identification dataset) demonstrate that the proposed method outperforms existing methods in terms of precision and F1-score. We also present results for sports images chosen from natural scene text detection datasets such as CTW1500 and MS-COCO to show the proposed method is effective and reliable across a range of inputs. Pinaki Nath Chowdhury, Palaiahnakote Shivakumara, Ramachandra Raghavendra, Sauradip Nag, Umapada Pal 0001, Tong Lu 0002, Daniel P. Lopresti |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | A New Deep Wavefront Based Model for Text Localization in 3D VideoabstractWith the evolution of electronic devices, such as 3D cameras, addressing the challenges of text localization in 3D video (e.g., for indexing) is increasingly drawing the attention of the multimedia and video processing community. Existing methods focus on 2D video and their performance in the presence of the challenges in 3D video, such as shadow areas associated with text and irregularly sized and shaped text, degrades. This paper proposes the first approach that successfully addresses the challenges of 3D video in addition to those of 2D. It employs a number of innovations, among which, the first is the Generalized Gradient Vector Flow (GGVF) for dominant points detection. The second is the Wavefront concept for text candidate point detection from those dominant points. In addition, an Adaptive B-Spline Polygon Curve Network (ABS-Net) is proposed for accurate text localization in 3D videos by constructing tight fitting bounding polygons using text candidate points. Extensive experiments on custom (3D video) and standard datasets (2D video and scene text) show that the proposed method is practical and useful, and overall outperforms existing state-of-the-art methods. Lokesh Nandanwar, Palaiahnakote Shivakumara, Ramachandra Raghavendra, Tong Lu 0002, Umapada Pal 0001, Apostolos Antonacopoulos, Yue Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | DCINN: Deformable Convolution and Inception Based Neural Network for Tattoo Text Detection Through Skin Region
Tamal Chowdhury, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Ramachandra Raghavendra, Sukalpa Chanda |
ICDAR (2) | 2 |
| 2021 | ARNet: Active-Reference Network for Few-Shot Image Semantic SegmentationabstractTo make predictions on unseen classes, few-shot segmentation becomes a research focus recently. However, most methods build on pixel-level annotation requiring quantity of manual work. Moreover, inherent information on same-category objects to guide segmentation could have large diversity in feature representation due to differences in size, appearance, layout, and so on. To tackle these problems, we present an active-reference network (ARNet) for few-shot segmentation. The proposed active-reference mechanism not only supports accurately cooccurrent objects in either support or query images, but also relaxes high constraint on pixel-level labeling, allowing for weakly boundary labeling. To extract more intrinsic feature representation, a category-modulation module (CMM) is further applied to fuse features extracted from multiple support images, thus forgetting useless and enhancing contributive information. Experiments on PASCAL-5idataset show the proposed method achieves a m-IOU score of 56.5% for 1-shot and 59.8% for 5-shot segmentation, being 0.5% and 1.3% higher than current state-of-the-art method. Guangchen Shi, Yirui Wu, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002 |
ICME | 3 |
| 2021 | DCT-phase statistics for forged IMEI numbers and air ticket detection
Lokesh Nandanwar, Palaiahnakote Shivakumara, Swati Kanchan, V. Basavaraja, D. S. Guru, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
Expert Syst. Appl. | 2 |
| 2021 | Improved Ring Radius Transform-Based Reconstruction for Video Character RecognitionabstractCharacter shape reconstruction in video is challenging due to low contrast, complex backgrounds and arbitrary orientation of characters. This work proposes an Improved Ring Radius Transform (IRRT) for reconstructing impaired characters through medial axis prediction. At first, the technique proposes a novel idea based on the Tangent Vector (TV) concept that identifies each actual pair of end pixels caused by gaps in impaired character components. Next, the actual direction to predict medial axis pixels using IRRT for each pair of end pixels is proposed with a new normal vector concept. The process of prediction repeats iteratively to find all the medial axis pixels for every gap in question. Further, medial axis pixels with their radii are used to reconstruct the shapes of impaired characters. The proposed technique is tested on benchmark datasets consisting of video, natural scenes, objects and multi-lingual data to demonstrate that it reconstructs shapes well, even for heterogeneous data. Comparative studies with different binarization and character recognition methods show that the proposed technique is effective, useful and outperforms existing methods. Zhiheng Huang, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001, Michael Blumenstein, Bhaarat Chetty, G. Hemantha Kumar 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2021 | A New Hybrid Method for Caption and Scene Text Classification in Action Video ImagesabstractAchieving a better recognition rate for text in action video images is challenging due to multiple types of text with unpredictable actions in the background. In this paper, we propose a new method for the classification of caption (which is edited text) and scene text (text that is a part of the video) in video images. This work considers five action classes, namely, Yoga, Concert, Teleshopping, Craft, and Recipes, where it is expected that both types of text play a vital role in understanding the video content. The proposed method introduces a new fusion criterion based on Discrete Cosine Transform (DCT) and Fourier coefficients to obtain the reconstructed images for caption and scene text. The fusion criterion involves computing the variances for coefficients of corresponding pixels of DCT and Fourier images, and the same variances are considered as the respective weights. This step results in Reconstructed image-1. Inspired by the special property of Chebyshev-Harmonic-Fourier-Moments (CHFM) that has the ability to reconstruct a redundancy-free image, we explore CHFM for obtaining the Reconstructed image-2. The reconstructed images along with the input image are passed to a Deep Convolutional Neural Network (DCNN) for classification of caption/scene text. Experimental results on five action classes and a comparative study with the existing methods demonstrate that the proposed method is effective. In addition, the recognition results of the before and after the classification obtained from different methods show that the recognition performance improves significantly after classification, compared to before classification. Lokesh Nandanwar, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2021 | A New Method for Detecting Altered Text in Document ImagesabstractAs more and more office documents are captured, stored, and shared in digital format, and as image editing software are becoming increasingly more powerful, there is a growing concern about document authenticity. To prevent illicit activities, this paper presents a new method for detecting altered text in document images. The proposed method explores the relationship between positive and negative coefficients of DCT to extract the effect of distortions caused by tampering by fusing reconstructed images of respective positive and negative coefficients, which results in Positive-Negative DCT coefficients Fusion (PNDF). To take advantage of spatial information, we propose to fuse R, G, and B color channels of input images, which results in RGBF (RGB Fusion). Next, the same fusion operation is used for fusing PNDF and RGBF, which results in a fused image for the original input one. We compute a histogram to extract features from the fused image, which results in a feature vector. The feature vector is then fed to a deep neural network for classifying altered text images. The proposed method is tested on our own dataset and the standard datasets from the ICPR 2018 Fraud Contest, Altered Handwriting (AH), and faked IMEI number images. The results show that the proposed method is effective and the proposed method outperforms the existing methods irrespective of image type. Lokesh Nandanwar, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Daniel P. Lopresti, Bhagesh Seraogi, Bidyut B. Chaudhuri |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2021 | A new context-based feature for classification of emotions in photographs
Divya Krishnani, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001, Daniel P. Lopresti, G. Hemantha Kumar 0001 |
Multim. Tools Appl. | 2 |
| 2021 | A survey on video content rating: taxonomy, challenges and open issues
Amin Khaksar Pour, Chaw-Seng Woo, Palaiahnakote Shivakumara, Hamid Tahaei, Nor Badrul Anuar |
Multim. Tools Appl. | 3 |
| 2021 | Deformable scene text detection using harmonic features and modified pixel aggregation network
Tanmay Jain, Palaiahnakote Shivakumara, Umapada Pal 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. Lett. | 2 |
| 2021 | A new DCT-PCM method for license plate number detection in drone images
Hamam Mokayed, Palaiahnakote Shivakumara, Hock Woon Hon, Mohan Kankanhalli, Tong Lu 0002, Umapada Pal 0001 |
Pattern Recognit. Lett. | 2 |
| 2021 | Arbitrarily-Oriented Text Detection in Low Light Natural Scene ImagesabstractText detection in low light natural scene images is challenging due to poor image quality and low contrast. Unlike most existing methods that focus on well-lit (normally daylight) images, the proposed method considers much darker natural scene images. For this task, our method first integrates spatial and frequency domain features through fusion to enhance fine details in the image. Next, we use Maximally Stable Extremal Regions (MSER) for detecting text candidates from the enhanced images. We then introduce Cloud of Line Distribution (COLD) features, which capture the distribution of pixels of text candidates in the polar domain. The extracted features are sent to a Convolution Neural Network (CNN) to correct the bounding boxes for arbitrarily oriented text lines by removing false positives. Experiments are conducted on a dataset of low light images to evaluate the proposed enhancement step. The results show our approach is more effective compared to existing methods in terms of standard quality measures, namely, BRISQE, NIQE and PIQE. In addition, experimental results on a variety of standard benchmark datasets, namely, ICDAR 2013, ICDAR 2015, SVT, Total-Text, ICDAR 2017-MLT and CTW1500, show that the proposed approach not only produces better results for low light images, at the same time it is also competitive for daylight images. Minglong Xue, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001, Daniel P. Lopresti, Zhibo Yang 0003 |
IEEE Trans. Multim. | 2 |
| 2021 | A New Foreground-Background based Method for Behavior-Oriented Social Media Image ClassificationabstractDue to various applications, research on personal traits using information on social media has become an important area. In this paper, a new method for the classification of behavior-oriented social images uploaded on various social media platforms is presented. The proposed method introduces a multimodality concept using skin of different parts of human body and background information, such as indoor and outdoor environments. For each image, the proposed method detects skin candidate components based on R, G, B color spaces and entropy features. The iterative mutual nearest neighbor approach is proposed to detect accurate skin candidate components, which result in foreground components. Next, the proposed method detects the remaining part (other than skin components) as background components based on structure tensor of R, G, B color spaces, and Maximally Stable Extremal Regions (MSER ) concept in the wavelet domain. We then explore Hanman Transform for extracting context features from foreground and background components through clustering and fusion operation. These features are then fed to an SVM classifier for the classification of behavior-oriented images. Comprehensive experiments on 10-class datasets of Normal Behavior-Oriented Social media Image (NBSI) and Abnormal Behavior-Oriented Social media Image (ABSI) show that the proposed method is effective and outperforms the existing methods in terms of average classification rate. Also, the results on the benchmark dataset of five classes of personality traits and two classes of emotions of different facial expressions (FERPlus dataset) demonstrated the robustness of the proposed method over the existing methods. Lokesh Nandanwar, Palaiahnakote Shivakumara, Divya Krishnani, Ramachandra Raghavendra, Tong Lu 0002, Umapada Pal 0001, Mohan Kankanhalli |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | A New Context-Based Method for Restoring Occluded Text in Natural Scene Images
Ayush Mittal, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein, Daniel P. Lopresti |
DAS | 2 |
| 2020 | A New Common Points Detection Method for Classification of 2D and 3D Texts in Video/Scene Images
Lokesh Nandanwar, Palaiahnakote Shivakumara, Ahlad Kumar, Tong Lu 0002, Umapada Pal 0001, Daniel P. Lopresti |
DAS | 2 |
| 2020 | Chebyshev-Harmonic-Fourier-Moments and Deep CNNs for Detecting Forged HandwritingabstractRecently developed sophisticated image processing techniques and tools have made easier the creation of high-quality forgeries of handwritten documents including financial and property records. To detect such forgeries of handwritten documents, this paper presents a new method by exploring the combination of Chebyshev-Harmonic-Fourier-Moments (CHFM) and deep Convolutional Neural Networks (D-CNNs). Unlike existing methods work based on abrupt changes due to distortion created by forgery operation, the proposed method works based on inconsistencies and irregular changes created by forgery operations. Inspired by the special properties of CHFM, such as its reconstruction ability by removing redundant information, the proposed method explores CHFM to obtain reconstructed images for the color components of the Original, Forged Noisy and Blurred classes. Motivated by the strong discriminative power of deep CNNs, for the reconstructed images of respective color components, the proposed method used deep CNNs for forged handwriting detection. Experimental results on our dataset and benchmark datasets (namely, ACPR 2019, ICPR 2018 FCD and IMEI datasets) show that the proposed method outperforms existing methods in terms of classification rate. Lokesh Nandanwar, Palaiahnakote Shivakumara, Sayani Kundu, Umapada Pal 0001, Tong Lu 0002, Daniel P. Lopresti |
ICPR | 2 |
| 2020 | Local Gradient Difference Features for Classification of 2D-3D Natural Scene Text ImagesabstractMethods developed for normal 2D text detection do not work well for text that is rendered using decorative, 3D effects, etc. This paper proposes a new method for classification of 2D and 3D natural scene text images so that an appropriate recognition method can be chosen accordingly based on the classification results for better performance. The proposed method explores local gradient differences for obtaining candidate pixels, which represent a stroke. To study the spatial distribution of candidate pixels, we propose a measure, called COLD, which is denser for pixels toward the center of strokes and scattered for non-stroke pixels. This observation leads us to introduce mass features for extracting the regular spatial pattern of COLD, which indicates a 2D text image. The extracted features are fed into a Neural Network (NN) for classification. The proposed method is tested on (i) a new dataset introduced in this work (ii) a second dataset assembled from standard natural scene datasets (iii) Non-Text Image datasets which does not contain text, rather it contains objects. Experimental results of the proposed method on images with text and non-text show that the proposed method is independent of text. The proposed approach improves text detection and recognition performance significantly after classification. Lokesh Nandanwar, Palaiahnakote Shivakumara, Ramachandra Raghavendra, Tong Lu 0002, Umapada Pal 0001, Daniel P. Lopresti, Nor Badrul Anuar |
ICPR | 2 |
| 2020 | Rotation invariant angle-density based features for an ice image classification system
Shengkai Yue, Minglei Yuan, Tong Lu 0002, Palaiahnakote Shivakumara, Michael Blumenstein, G. Hemantha Kumar 0001 |
Expert Syst. Appl. | 4 |
| 2020 | Forged text detection in video, scene, and document imagesabstractRapid advances in artificial intelligence have made it possible to produce forgeries good enough to fool an average user. As a result, there is growing interest in developing robust methods to counter such forgeries. This study presents a new Fourier spectrum‐based method for detecting forged text in video images. The authors' premise is that brightness distribution and the spectrum shape exhibit irregular patterns (inconsistencies) for forged text, while appearing more regular for original text. The method divides the spectrum of an input image into sectors and tracks to highlight these effects. Specifically, positive and negative coefficients for sectors and tracks are extracted to quantify the brightness distribution. Variations in the shape of the spectrum are analysed by determining the angular relationship between the principal axes and the sectors/tracks of the spectrum. Next, it combines these two features to detect forged text in the images of IMEI (International Mobile Equipment Identity) numbers and document. For evaluation, the following datasets are used: own video dataset and standard datasets, namely, IMEI number, ICPR 2018 Fraud Document Contest, and a natural scene text dataset. Experimental results show that the proposed method outperforms existing methods in terms of average classification rate and F ‐score. Lokesh Nandanwar, Palaiahnakote Shivakumara, Prabir Mondal, Raghunandan K. Srinivas, Umapada Pal 0001, Tong Lu 0002, Daniel P. Lopresti |
IET Image Process. | 2 |
| 2020 | A new augmentation-based method for text detection in night and day license plate images
Pinaki Nath Chowdhury, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
Multim. Tools Appl. | 2 |
| 2020 | Saliency-based bit plane detection for network applications
Maryam Asadzadeh Kaljahi, Palaiahnakote Shivakumara, Saqib Hakak, Mohd Yamani Idna Bin Idris, Mohammad Hossein Anisi, Deepu Rajan |
Multim. Tools Appl. | 2 |
| 2020 | A new unified method for detecting text from marathon runners and sports players in video (PR-D-19-01078R2)
Sauradip Nag, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
Pattern Recognit. | 2 |
| 2020 | Graph attention network for detecting license plates in crowded street scenes
Pinaki Nath Chowdhury, Palaiahnakote Shivakumara, Swati Kanchan, Ramachandra Raghavendra, Umapada Pal 0001, Tong Lu 0002, Daniel P. Lopresti |
Pattern Recognit. Lett. | 2 |
| 2020 | Delaunay triangulation based text detection from multi-view images of natural scene
Soumyadip Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, G. Hemantha Kumar 0001 |
Pattern Recognit. Lett. | 2 |
| 2020 | A new Fractal Series Expansion based enhancement model for license plate recognition
Pinaki Nath Chowdhury, Palaiahnakote Shivakumara, Hamid Abdullah Jalab, Rabha W. Ibrahim, Umapada Pal 0001, Tong Lu 0002 |
Signal Process. Image Commun. | 2 |
| 2019 | Age Estimation using Disconnectedness Features in HandwritingabstractReal-time applications of handwriting analysis have increased drastically in the fields of forensic and information security because of accurate cues. One of such applications is human age estimation based on handwriting for the purpose of immigrant checking. In this paper, we have proposed a new method for age estimation using handwriting analysis using Hu invariant moments and disconnectedness features. To make the proposed method robust to both ruled and un-ruled documents, we propose to explore intersection point detection in Canny edge images of each input document, which results in text components. For each text component pair, we propose Hu invariant moments for extracting disconnectedness features, which in fact measure multi-shape components based on distance, shape and mutual position analysis of components. Furthermore, iterative k-means clustering is proposed for the classification of different age groups. Experimental results on our dataset and some standard datasets, namely, IAM and KHATT, show that the proposed method is effective and outperforms the state-of-the-art methods. V. Basavaraja, Palaiahnakote Shivakumara, D. S. Guru, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
ICDAR | 2 |
| 2019 | CRNN Based Jersey-Bib Number/Text Recognition in Sports and Marathon ImagesabstractThe primary challenge in tracing the participants in sports and marathon video or images is to detect and localize the jersey/Bib number that may present in different regions of their outfit captured in cluttered environment conditions. In this work, we proposed a new framework based on detecting the human body parts such that both Jersey Bib number and text is localized reliably. To achieve this, the proposed method first detects and localize the human in a given image using Single Shot Multibox Detector (SSD). In the next step, different human body parts namely, Torso, Left Thigh, Right Thigh, that generally contain a Bib number or text region is automatically extracted. These detected individual parts are processed individually to detect the Jersey Bib number/text using a deep CNN network based on the 2-channel architecture based on the novel adaptive weighting loss function. Finally, the detected text is cropped out and fed to a CNN-RNN based deep model abbreviated as CRNN for recognizing jersey/Bib/text. Extensive experiments are carried out on the four different datasets including both bench-marking dataset and a new dataset. The performance of the proposed method is compared with the state-of-the-art methods on all four datasets that indicates the improved performance of the proposed method on all four datasets. Sauradip Nag, Ramachandra Raghavendra, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Mohan Kankanhalli |
ICDAR | 3 |
| 2019 | A Text-Context-Aware CNN Network for Multi-oriented and Multi-language Scene Text DetectionabstractThe existing deep learning based state-of-theart scene text detection methods treat scene texts a type of general objects, or segment text regions directly. The latter category achieves remarkable detection results on arbitrary orientation and large aspect ratios of scene texts based on instance segmentation algorithms. However, due to the lack of context information with consideration of scene text unique characteristics, directly applying instance segmentation to text detection task is prone to result in low accuracy, especially producing false positive detection results. To ease this problem, we propose a novel text-context-aware scene text detection CNN structure, which appropriately encodes channel and spatial attention information to construct context-aware and discriminative feature map for multi-oriented and multi-language text detection tasks. With high representation ability of text context-aware feature map, the proposed instance segmentation based method can not only robustly detect multi-oriented and multi-language text from natural scene images, but also produce better text detection results by greatly reducing false positives. Experiments on ICDAR2015 and ICDAR2017-MLT datasets show that the proposed method has achieved superior performances in precision, recall and F-measure than most of the existing studies. Minglong Xue, Tong Lu 0002, Yirui Wu, Palaiahnakote Shivakumara |
ICDAR | 5 |
| 2019 | An Automatic System for Generating Artificial Fake Character Images
Yisheng Yue, Palaiahnakote Shivakumara, Yirui Wu, Tong Lu 0002, Umapada Pal 0001 |
MMM (2) | 2 |
| 2019 | An automatic zone detection system for safe landing of UAVs
Maryam Asadzadeh Kaljahi, Palaiahnakote Shivakumara, Mohd Yamani Idna Bin Idris, Mohammad Hossein Anisi, Tong Lu 0002, Michael Blumenstein, Noorzaily Mohamed Noor |
Expert Syst. Appl. | 2 |
| 2019 | A novel character segmentation-reconstruction approach for license plate recognition
Vijeta Khare, Palaiahnakote Shivakumara, Chee Seng Chan, Tong Lu 0002, Kim Meng Liang, Hock Woon Hon, Michael Blumenstein |
Expert Syst. Appl. | 2 |
| 2019 | Fractional means based method for multi-oriented keyword spotting in video/scene/license plate images
Palaiahnakote Shivakumara, Sangheeta Roy, Hamid Abdullah Jalab, Rabha W. Ibrahim, Umapada Pal 0001, Tong Lu 0002, Vijeta Khare, Ainuddin Wahid Abdul Wahab |
Expert Syst. Appl. | 1 |
| 2019 | A new image size reduction model for an efficient visual sensor networkabstractImage size reduction for energy-efficient transmission without losing quality is critical in Visual Sensor Networks (VSNs). The proposed method finds overlapping regions using camera locations, which eliminate unfocussed regions from the input images. The sharpness for the overlapped regions is estimated to find the Dominant Overlapping Region (DOR). The proposed model partitions further the DOR into sub-DORs according to capacity of the cameras. To reduce noise effects from the sub-DOR, we propose to perform a Median operation, which results in a Compressed Significant Region (CSR). For non-DOR, we obtain Sobel edges, which reduces the size of the images down to ambinary form. The CSR and Sobel edges of the non-DORs are sent by a VSN. Experimental results and a comparative study with the state-of-the-art methods shows that the proposed model outperforms the existing methods in terms of quality, energy consumption and network lifetime. Maryam Asadzadeh Kaljahi, Palaiahnakote Shivakumara, Mohd Yamani Idna Bin Idris, Mohammad Hossein Anisi, Michael Blumenstein |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | A scene image classification technique for a ubiquitous visual surveillance system
Maryam Asadzadeh Kaljahi, Palaiahnakote Shivakumara, Mohammad Hossein Anisi, Mohd Yamani Idna Bin Idris, Michael Blumenstein, Muhammad Khurram Khan |
Multim. Tools Appl. | 2 |
| 2019 | Curved text detection in blurred/non-blurred video/scene images
Minglong Xue, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001 |
Multim. Tools Appl. | 2 |
| 2019 | Multi-Script-Oriented Text Detection and Recognition in Video/Scene/Born Digital ImagesabstractAchieving good text detection and recognition results for multi-script-oriented images is a challenging task. First, we explore bit plane slicing in order to utilize the advantage of the most significant bit information to identify text components. A new iterative nearest neighbor symmetry is then proposed based on shapes of convex and concave deficiencies of text components in bit planes to identify candidate planes. Further, we introduce a new concept called mutual nearest neighbor pair components based on gradient direction to identify representative pairs of texts in each candidate bit plane. The representative pairs are used to restore words with the help of edge image of the input one, which results in text detection results (words). Second, we propose a new idea by fixing window for character components of arbitrary oriented words based on angular relationship between sub-bands and a fused band. For each window, we extract features in contourlet wavelet domain to detect characters with the help of an SVM classifier. Further, we propose to explore HMM for recognizing characters and words of any orientation using the same feature vector. The proposed method is evaluated on standard databases such as ICDAR, YVT video, ICDAR, SVT, MSRA scene data, ICDAR born digital data, and multi-lingual data to show its superiority to the state of the art methods. Raghunandan K. Srinivas, Palaiahnakote Shivakumara, Sangheeta Roy, G. Hemantha Kumar 0001, Umapada Pal 0001, Tong Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | New COLD Feature Based Handwriting Analysis for Enthnicity/Nationality IdentificationabstractIdentifying crime for forensic investigating teams when crimes involve people of different nationals is challenging. This paper proposes a new method for ethnicity (nationality) identification based on Cloud of Line Distribution (COLD) features of handwriting components. The proposed method, at first, uses tangent angle of the contour pixels in each row and the mean of intensity values of each row for segmenting text lines. For segmented text lines, we use tangent angle and direction of base lines to remove rule lines in the image. We use polygonal approximation for finding dominant points for contours of edge components. Then the proposed method connects the nearest dominant points of every dominant point, which results in line segments of dominant point pairs. For each line segment, the proposed method estimates angle and length, which gives a point in polar domain. For all the line segments, the proposed method generates dense points in polar domain, which results in COLD distribution. As character component shapes change, according to nationals, the shape of the distribution changes. This observation is extracted based on distance from pixels of distribution to Principal Axis of the distribution. Then the features are subjected to an SVM classifier for identifying nationals. Experiments are conducted on a complex dataset, which show the proposed method is effective and outperforms the existing method. Sauradip Nag, Palaiahnakote Shivakumara, Yirui Wu, Umapada Pal 0001, Tong Lu 0002 |
ICFHR | 2 |
| 2018 | Adaptive Multi-Gradient Kernels for Handwritting Based Gender IdentificationabstractHandwriting based Gender identification is challenging due to unconstrained handwriting and individual differences in writing. To solve this problem, we propose a new adaptive multi-gradient of Sobel kernels for extracting Adaptive Multi-Gradient Features (AMGF). For extracted text lines, the proposed method finds dominant pixels based on directional symmetry of text pixels given by AMGF. We perform histogram operation for adaptive multi-gradient values extracted corresponding to dominant pixels. The gradient values that give the highest peak in respective histograms is chosen as features. This results in feature vector having four AMGF values. The same vector are generated for successive text lines in each image to study either consistency, which is expected for females or inconsistency, which is expected for males in writing styles. The correlation is estimated based on feature vectors of the first and the successive text lines until converging or diverging criteria is met. If convergence happens, the input document is considered as female else is considered as male. The method is tested on our own dataset, which includes large variations and standard datasets, namely, QUWI, IAM-1+IAM-2 and KHATT, to demonstrate the effectiveness of the proposed method. Experimental results show that the proposed method outperforms the existing methods. B. J. Navya, Palaiahnakote Shivakumara, G. C. Swetha, Sangheeta Roy, D. S. Guru, Umapada Pal 0001, Tong Lu 0002 |
ICFHR | 2 |
| 2018 | A New RGB Based Fusion for Forged IMEI Number Detection in Mobile ImagesabstractAs technology advances to make living comfortable for people, at the same time, different crimes also increase. One such sensitive crime is creating fake International Mobile Equipment Identity (IMEI) for smart mobile devices. In this paper, we present a new fusion based method using R, G and B color components for detecting forged IMEI numbers. To the best of our knowledge, this is the first work for forged IMEI number detection in mobile images. The proposed method first finds variances for R, G and B images of a forged input image to study local changes. The variances are used to derive weights for respective color components. The same weights are convolved with respective pixel values of R, G and B components, which results in the fused image. For the fused image, the proposed method extracts features based on sparsity, the number of connected components, and the average intensity values for edge components in respective R, G and B components, which gives six features. The proposed method finds absolute difference between fused and input images, which gives feature vector containing six difference values. The proposed method constructs templates based on samples chosen randomly. Feature vectors are compared with the templates for detecting forged IMEI numbers. Experiments are conducted on our own dataset and standard datasets to evaluate the proposed method. Furthermore, comparative studies with the related existing methods show that the proposed method outperforms the existing methods. Palaiahnakote Shivakumara, V. Basavaraja, Harsha S. Gowda, D. S. Guru, Umapada Pal 0001, Tong Lu 0002 |
ICFHR | 1 |
| 2018 | Weighted-Gradient Features for Handwritten Line SegmentationabstractText line segmentation from handwritten documents is challenging when a document image contains severe touching. In this paper, we propose a new idea based on Weighted-Gradient Features (WGF) for segmenting text lines. The proposed method finds the number of zero crossing points for every row of Canny edge image of the input one, which is considered as the weights of respective rows. The weights are then multiplied with gradient values of respective rows of the image to widen the gap between pixels in the middle portion of text and the other portions. Next, k-means clustering is performed on WGF to classify middle and other pixels of text. The method performs morphological operation to obtain word components as patches for the result of clustering. The patches in both the clusters are matched to find common patch areas, which helps in reducing touching effect. Then the proposed method checks linearity and non-linearity iteratively based on patch direction to segment text lines. The method is tested on our own and standard datasets, namely, Alaei, ICDAR 2013 robust competition on handwriting context and ICDAR 2015-HTR, to evaluate the performance. Further, the method is compared with the state of art methods to show its effectiveness and usefulness. Vijeta Khare, Palaiahnakote Shivakumara, B. J. Navya, G. C. Swetha, D. S. Guru, Umapada Pal 0001, Tong Lu 0002 |
ICPR | 2 |
| 2018 | Multi-Gradient Directional Features for Gender IdentificationabstractGender identification based on handwriting analysis has received a special attention to researchers in the field of document image analysis as it is useful for several real-time applications like forensic, population counting, etc. In this paper, we explore Multi-Gradient Directional (MGD) features, which provide direction of dominant pixels obtained by Canny edge image, and gradient direction symmetry. The proposed method further performs histogram operation for gradient angle information of dominant pixels of respective multi-gradient directional images to select angles, which contribute to the highest peak. This results in feature vectors. The process of feature vector formation continues for the segmented first, second, and third text lines in each image by male or female. Next, correlation is estimated for the vector of the first line with successive lines until converging or diverging criteria is met. If the convergence happens, a document is considered as by female, else is considered as by male. The method is tested on our own dataset, which includes images of different scripts, writers, papers, pens, and ages, and the standard database QUWI which includes Arabic and English texts, to demonstrate the efficiency of the proposed method. Comparative studies with the state of the art methods show that the proposed method is effective and useful. B. J. Navya, G. C. Swetha, Palaiahnakote Shivakumara, Sangheeta Roy, D. S. Guru, Umapada Pal 0001, Tong Lu 0002 |
ICPR | 3 |
| 2018 | Em-SLAM: a Fast and Robust Monocular SLAM Method for Embedded SystemsabstractSimultaneous Localization and Mapping (SLAM) is difficult to deploy in the embedded systems due to its high computation cost and stable input requirements. Building on excellent algorithms of recent years, we present Em-SLAM, a monocular SLAM method which is fast and robust in the embedded system. We present Em-SLAM in three stages comprising initial pose estimation, iterative pose optimization and correspondences, and mapping with nearest frame queue. During the first stage, we perform stable initial pose estimation based on the matched ORB features extracted around the selected key points. Regarding initial pose and corresponding key points as input, the second stage of Em-SLAM iteratively optimizes these inputs values by tracking key points in the new frames. At the last stage, we firstly determine keyframes with the help of the proposed nearest frame queue and then design a greedy search algorithm to find matched ORB features between keyframes, which are adopted for compact and robust map reconstruction. Due to the special designs for the embedded systems, Em-SLAM demonstrates a high accurate and fast performance on the embedded system for all SLAM tasks: tracking, mapping and loop closing. We evaluate Em-SLAM on he most popular datasets by comparing with one latest SLAM method. Yirui Wu, Zhikai Li, Palaiahnakote Shivakumara, Tong Lu 0002 |
ICPR | 3 |
| 2018 | Context-Aware Attention LSTM Network for Flood PredictionabstractTo minimize the negative impacts brought by floods, researchers from pattern recognition community utilize artificial intelligence based methods to solve the problem of flood prediction. Inspired by the significant power of Long Short-Term Memory (LSTM) networks in modeling the dynamics and dependencies of sequential data, we intend to utilize LSTM networks to predict sequential flow rate values based on a set of collected flood factors. Since not all factors are informative for flood prediction and the irrelevant factors often bring a lot of noise, we need to pay more attention to the informative ones. However, original LSTM doesn't have strong attention capability. Hence we propose an context-aware attention LSTM (CA-LSTM) network for flood prediction, which is capable to selectively focus on informative factors. During training, the local context-aware attention model is constructed by learning probability distributions between flow rate and hidden output of each LSTM cell. During testing, the learned local attention model assign weights to adjust relations between input factors and predictions at all steps of LSTM network. We conduct experiments on a flood dataset with several comparative methods to demonstrate high accuracy of the proposed method and the effectiveness of the proposed context-aware attention model. Yirui Wu, Zhaoyang Liu 0001, Weigang Xu, Jun Feng 0001, Palaiahnakote Shivakumara, Tong Lu 0002 |
ICPR | 5 |
| 2018 | Fourier Transform based Features for Clean and Polluted Water Image ClassificationabstractWater image classification is challenging because water images of ocean or river share the same properties with images of polluted water such as fungus, waste and rubbish. In this paper, we present a method for classifying clean and polluted water images. The proposed method explores Fourier transform based features for extracting texture properties of clean and polluted water images. Fourier spectrum of each input image is divided into several sub-regions based on angle and spatial information. For each region over the spectrum, the proposed method extracts mean and variance features using intensity values, which results in a feature matrix. The feature matrix is then passed to an SVM classifier for the classification of clean and polluted water images. Experimental results on classes of clean and polluted water images show that the proposed method is effective. Furthermore, a comparative study with the state-of-the-art method shows that the proposed method outperforms the existing method in terms of classification rate, recall, precision and F-measure. Xuerong Wu, Palaiahnakote Shivakumara, Hualu Zhang, Tong Lu 0002, Umapada Pal 0001, Michael Blumenstein |
ICPR | 2 |
| 2018 | Local and Global Bayesian Network based Model for Flood PredictionabstractTo minimize the negative impacts brought by floods, researchers from pattern recognition community pay special attention to the problem of flood prediction by involving technologies of machine learning. In this paper, we propose to construct hierarchical Bayesian network to predict floods for small rivers, which appropriately embed hydrology expert knowledge for high rationality and robustness. We present the construction of the hierarchical Bayesian network in two stages comprising local and global network construction. During the local network construction, we firstly divide the river watershed into small local regions. Following the idea of a famous hydrology model - the Xinanjiang model, we establish the entities and connections of the local Bayesian network to represent the variables and physical processes of the Xinanjiang model, respectively. During the global network construction, intermediate variables for local regions, computed by the local Bayesian network, are coupled to offer an estimation for time-varying values of flow rate by proper inferences of the global network. At last, we propose to improve the output of Bayesian network by utilizing former flow rate values. We demonstrate the accuracy and robustness of the proposed method by conducting experiments on a collected dataset with several comparative methods. Yirui Wu, Weigang Xu, Jun Feng 0001, Palaiahnakote Shivakumara, Tong Lu 0002 |
ICPR | 4 |
| 2018 | Cloud of Line Distribution and Random Forest Based Text Detection from Natural/Video Scene Images
Wenhai Wang, Yirui Wu, Palaiahnakote Shivakumara, Tong Lu 0002 |
MMM (2) | 3 |
| 2018 | Rough-fuzzy based scene categorization for text detection and recognition in video
Sangheeta Roy, Palaiahnakote Shivakumara, Namita Jain, Vijeta Khare, Anjan Dutta 0001, Umapada Pal 0001, Tong Lu 0002 |
Pattern Recognit. | 2 |
| 2018 | Riesz Fractional Based Model for Enhancing License Plate Detection and RecognitionabstractOne of the major causes of poor results in license plate recognition is low quality of images affected by multiple factors, such as severe illumination condition, complex background, different weather conditions, night light, and perspective distortions. In this paper, we propose a new mathematical model based on Riesz fractional operator for enhancing details of edge information in license plate images to improve the performances of text detection and recognition methods. The proposed model performs convolution operation of the Riesz fractional derivative over each input image by enhancing the edge strength in it. To test the performance of the proposed model, we conduct experiments on benchmark license plate image databases, namely, UCSD and ICDAR 2015-SR competition text image databases. Experimental results on enhancement show that the proposed model outperforms the existing baseline enhancement techniques in terms of quality measures. Furthermore, experimental results on text detection and recognition show that text detection and recognition rates are improved significantly after enhancement compared with before enhancement. Raghunandan K. Srinivas, Palaiahnakote Shivakumara, Hamid Abdullah Jalab, Rabha W. Ibrahim, G. Hemantha Kumar 0001, Umapada Pal 0001, Tong Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | New Fuzzy-Mass Based Features for Video Image Type CategorizationabstractDue to the large variety of video type collections, it becomes difficult to achieve good text detection and recognition accuracy. We propose a new fuzzy-mass based method for classifying (categorizing) text frames from different types of video. For each frame of a video type, we formulate Fuzzy logic to identify straight and curved edge components from edge images. We then estimate mass locally and globally by drawing consecutive ellipses over edge images with respect to straight and curved edge components. Further, we extract features based on spatial proximity between centroid of classified straight/curved edge components and that of the whole image. This results local features. Next, the features are extracted for the whole image without ellipse drawing, which results in global features. The combination of both local and global features is then fed to an SVM classifier for video type classification. Experimental results on the proposed and existing classification methods show that the proposed classification outperforms the stat of art methods. Furthermore, experiments on before and after classification with several text detection and binarization methods show that the proposed classification is significant in improving text detection and recognition performance. Sangheeta Roy, Palaiahnakote Shivakumara, Namita Jain, Vijeta Khare, Umapada Pal 0001, Tong Lu 0002 |
ICDAR | 2 |
| 2017 | Temporal Integration for Word-Wise Caption and Scene Text IdentificationabstractGenerally video consists of edited text (i.e., caption text) and natural text (i.e., scene text), and these two texts differ from one another in nature as well as characteristics. Such different behaviors of caption and scene texts lead to poor accuracy for text recognition in video. In this paper, we explore wavelet decomposition and temporal coherency for the classification of caption and scene text. We propose wavelet of high frequency sub-bands to separate text candidates that are represented by high frequency coefficients in an input word. The proposed method studies the distribution of text candidates over word images based on the fact that the standard deviation of text candidates is high at the first zone, low at the middle zone and high at the third zone. This is extracted by mapping standard deviation values to 8 equal sized bins formed based on the range of standard deviation values. The correlation among bins at the first and second levels of wavelets is explored to differentiate caption and scene text and for determining the number of temporal frames to be analyzed. The properties of caption and scene texts are validated with the chosen temporal frames to find the stable property for classification. Experimental results on three standard datasets (ICDAR 2015, YVT and License Plate Video) show that the proposed method outperforms the existing methods in terms of classification rate and improves recognition rate significantly based on classification results. Sangheeta Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Ainuddin Wahid Abdul Wahab |
ICDAR | 2 |
| 2017 | Fourier-Residual for Printer IdentificationabstractPrinter identification is challenging due to advanced software technologies in the field of forgery detection. This paper presents a new idea of using the Fourier transform residual for the identification of documents printed by different printers. The proposed approach first convolves a Laplacian mask with a Fourier transform in the frequency domain to smoothen the edges. Next, we apply an inverse Fourier transform to reconstruct images from smoothed information (RFL). Similarly, the proposed approach reconstructs images using gray information of the input image (RFG). Then the residual is calculated by subtracting RFG from RFL. The set of statistical features, texture and spatial features are extracted from residual images for printer identification. Experimental results with the existing method on our dataset and a standard dataset show that the proposed approach outperforms the existing approach on both the datasets in terms of classification rate, recall, precision and F-measure. Palaiahnakote Shivakumara, Tong Lu 0002, M. Basavanna, Umapada Pal 0001, Michael Blumenstein |
ICDAR | 2 |
| 2017 | A Robust Symmetry-Based Method for Scene/Video Text Detection through Neural NetworkabstractText detection in video/scene images has gained a significant attention in the field of image processing and document analysis due to the inherent challenges caused by variations in contrast, orientation, background, text type, font type, non-uniform illumination and so on. In this paper, we propose a novel text detection method to explore symmetry property and appearance features of text for improved accuracy and robustness. First, the proposed method explores Extremal Regions (ER) for detecting text candidates in images. Then we propose a novel feature named as Multi-domain Strokes Symmetry Histogram (MSSH) for each text candidate, which describes the inherent symmetry property of stroke pixel pairs in gray, gradient and frequency domains. Furthermore, deep convolutional features are extracted to describe the appearance for each text candidate. We further fuse them by Auto-Encoder network to define a more discriminative text descriptor for classification. Finally, the proposed method constructs text lines based on the classification results. We demonstrate the effectiveness and robustness detection results of our proposed method by testing on four different benchmark databases. Yirui Wu, Wenhai Wang, Palaiahnakote Shivakumara, Tong Lu 0002 |
ICDAR | 3 |
| 2017 | Robust Scene Text Detection for Multi-script Languages Using Deep Learning
Ruo-Ze Liu, Xin Sun 0009, Hailiang Xu, Palaiahnakote Shivakumara, Feng Su, Tong Lu 0002, Ruoyu Yang |
MMM (1) | 4 |
| 2017 | Script independent approach for multi-oriented text detection in scene image
Sounak Dey, Palaiahnakote Shivakumara, Raghunandan K. Srinivas, Umapada Pal 0001, Tong Lu 0002, G. Hemantha Kumar 0001, Chee Seng Chan |
Neurocomputing | 2 |
| 2017 | Arbitrarily-oriented multi-lingual text detection in video
Vijeta Khare, Palaiahnakote Shivakumara, Raveendran Paramesran, Michael Blumenstein |
Multim. Tools Appl. | 2 |
| 2017 | A new multi-modal approach to bib number/text detection and recognition in Marathon images
Palaiahnakote Shivakumara, Ramachandra Raghavendra, Longfei Qin, Kiran B. Raja, Tong Lu 0002, Umapada Pal 0001 |
Pattern Recognit. | 1 |
| 2017 | Fractals based multi-oriented text detection system for recognition in mobile video images
Palaiahnakote Shivakumara, Liang Wu 0009, Tong Lu 0002, Chew Lim Tan, Michael Blumenstein, Basavaraj S. Anami |
Pattern Recognit. | 1 |
| 2016 | New Sharpness Features for Image Type Classification Based on Textual InformationabstractAchieving good recognition results from a single method for text lines in video/natural scene images captured by high resolution cameras or low resolution mobile cameras, and images in web pages, is often hard. In this paper, we propose new sharpness based features of textual portion of each input text line image using HSI color space for the classification of an input image into one of the four classes (video, scene, mobile or born digital). This helps in choosing an appropriate method based on the class type of the input text for its improved recognition rate. For a given input text line image, the proposed method obtains H, S and I images. Then Canny edge images are obtained for H, S and I spaces, which results in text candidates. We perform sliding window operation over the text candidate image of each text line of each color space to estimate new sharpness by calculating stroke width and gradient information. The sharpness values of the text lines of the three color spaces are then fed to k-means clustering with maximum, minimum and average guesses, which results in three respective clusters. The mean of each cluster for respective color spaces outputs a feature vector having nine feature values for image classification with the help of an SVM classifier. Experimental results on standard datasets, namely, ICDAR 2013, ICDAR 2015 video, ICDAR 2015 natural scene data, ICDAR 2013 born digital data and the images captured by a mobile camera (our own data) show that the proposed classification method helps in improving recognition results. Raghunandan K. Srinivas, Palaiahnakote Shivakumara, G. Hemantha Kumar 0001, Umapada Pal 0001, Tong Lu 0002 |
DAS | 2 |
| 2016 | Fourier Coefficients for Fraud Handwritten Document Classification through Age AnalysisabstractAs new digital technologies emerge to improve living style, at the same time, it also lead to increase crimes. Unlike existing approaches that use content of handwriting for fraud/forged document identification, in this paper we propose a novel approach that explores the quality of handwritten documents by considering both foreground and background information to identify whether it is old or new. The proposed approach works based on the fact that if a fraud document is created with some gaps after the original one, the fraud document happened to be a new one and the original happened to be an old one in this work. To identify whether a given handwritten document is old or new with gaps, we propose to divide Fourier coefficients of the input image into positive and negative coefficient images, and then reconstruct respective images to conquer two reconstructed ones. The contrast of the reconstructed images obtained before and after divide-conquer is studied to analyze the ages of the document based on image quality. The proposed approach finds a unique relationship between reconstructed images, obtained before and after divide-conquer, to identify the input image as old or new. To evaluate the proposed approach, we conduct experiments on our own handwritten dataset and a standard database, namely, Google-LIFE magazine. Comparative studies with the existing approaches show that the proposed approach outperforms the existing approaches in terms of classification rate. Raghunandan K. Srinivas, Palaiahnakote Shivakumara, B. J. Navya, G. Pooja, Navya Prakash, G. Hemantha Kumar 0001, Umapada Pal 0001, Tong Lu 0002 |
ICFHR | 2 |
| 2016 | New Tampered Features for Scene and Caption Text Classification in Video FrameabstractThe presence of both caption/graphics/superimposed and scene texts in video frames is the major cause for the poor accuracy of text recognition methods. This paper proposes an approach for identifying tampered information by analyzing the spatial distribution of DCT coefficients in a new way for classifying caption and scene text. Since caption text is edited/superimposed, which results in artificially created texts comparing to scene texts that exist naturally in frames. We exploit this fact to identify the presence of caption and scene texts in video frames based on the advantage of DCT coefficients. The proposed method analyzes the distributions of both zero and non-zero coefficients (only positive values) locally by moving a window, and studies histogram operations over each input text line image. This generates line graphs for respective zero and non-zero coefficient coordinates. We further study the behavior of text lines, namely, linearity and smoothness based on centroid location analysis, and the principal axis direction of each text line for classification. Experimental results on standard datasets, namely, ICDAR 2013 video, 2015 video, YVT video and our own data, show that the performances of text recognition methods are improved significantly after-classification compared to before-classification. Sangheeta Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Chew Lim Tan |
ICFHR | 2 |
| 2016 | A quad tree based method for blurred and non-blurred video text frames classification through quality metricsabstractBlur is a common artifact in video, which adds more complexity to text detection and recognition. To achieve good accuracies for text detection and recognition, this paper suggests a new method for classifying blurred and non-blurred frames in video. We explore quality metrics, namely, BRISQUE, NRIQA, GPC and SI, in a new way for classification. We estimate the values of these metrics with the help of predefined samples called reference values. To widen the difference between metric values for better classification, we introduce scaling factors as a non-linear sigmoidal function, which considers the metric of each current frame and its reference and results in templates. Based on the characteristics of metrics, the proposed method finds a relationship between the metrics to derive rules for classification. To classify the frame containing local blur, we explore quad tree division with classification rules which divide non-blurred blocks to identify local blur. We use standard databases, namely, ICDAR 2013, ICDAR 2015 and YVT videos for experimentation, and evaluate the proposed method in terms of text detection and recognition rates given by text detection and binarization methods before and after classification. Vijeta Khare, Palaiahnakote Shivakumara, Ahlad Kumar, Chee Seng Chan, Tong Lu 0002, Michael Blumenstein |
ICPR | 2 |
| 2016 | Video scene text frames categorization for text detection and recognitionabstractDeveloping a unified text detection and recognition method is hard for different video types due to varying characteristics in video. This paper proposes a new method for categorizing different types of video text frames, namely, videos containing advertisement, signboard, license plate, front page of book or magazine, street view, and video of general items, for better text detection and recognition rate. We propose symmetry features using gradient vector flow for Canny and Sobel edge images of each input frame to identify candidate edge components. Then for a candidate edge component image, we extract both global and local features using colors from different channels in a new way. Besides, the proposed method extracts statistical and structural features from the spatial distribution of candidate pixels in a multi-scale environment. Lastly, the extracted features are fed to a logistic classifier for categorization. The features extracted locally and globally are tested both separately and altogether in terms of confusion matrix. The performance of the proposed categorization method is evaluated through several text detection and recognition experiments before and after categorization. We noted that the proposed categorization method is very useful in improving text detection and recognition performance. Longfei Qin, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001, Chew Lim Tan |
ICPR | 2 |
| 2016 | A novel forged blurred region detection system for image forensic applications
Diaa M. Uliyan, Hamid Abdullah Jalab, Ainuddin Wahid Abdul Wahab, Palaiahnakote Shivakumara, Somayeh Sadeghi |
Expert Syst. Appl. | 4 |
| 2016 | Modeling spatial layout for scene image understanding via a novel multiscale sum-product network
Ze-Huan Yuan, Limin Wang 0002, Tong Lu 0002, Palaiahnakote Shivakumara, Chew Lim Tan |
Expert Syst. Appl. | 5 |
| 2016 | Weakly-supervised region annotation for understanding scene images
Tong Lu 0002, Palaiahnakote Shivakumara, Chew Lim Tan |
Multim. Tools Appl. | 4 |
| 2016 | A blind deconvolution model for scene text detection and recognition in video
Vijeta Khare, Palaiahnakote Shivakumara, Raveendran Paramesran, Michael Blumenstein |
Pattern Recognit. | 2 |
| 2016 | Fractional poisson enhancement model for text detection and recognition in video frames
Sangheeta Roy, Palaiahnakote Shivakumara, Hamid Abdullah Jalab, Rabha W. Ibrahim, Umapada Pal 0001, Tong Lu 0002 |
Pattern Recognit. | 2 |
| 2016 | A new method for multi-oriented graphics-scene-3D text classification in video
Jiamin Xu, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan, Seiichi Uchida |
Pattern Recognit. | 2 |
| 2016 | Contour Restoration of Text Components for Recognition in Video/Scene ImagesabstractText recognition in video/natural scene images has gained significant attention in the field of image processing in many computer vision applications, which is much more challenging than recognition in plain background images. In this paper, we aim to restore complete character contours in video/scene images from gray values, in contrast to the conventional techniques that consider edge images/binary information as inputs for text detection and recognition. We explore and utilize the strengths of zero crossing points given by the Laplacian to identify stroke candidate pixels (SPC). For each SPC pair, we propose new symmetry features based on gradient magnitude and Fourier phase angles to identify probable stroke candidate pairs (PSCP). The same symmetry properties are proposed at the PSCP level to choose seed stroke candidate pairs (SSCP). Finally, an iterative algorithm is proposed for SSCP to restore complete character contours. Experimental results on benchmark databases, namely, the ICDAR family of video and natural scenes, Street View Data, and MSRA data sets, show that the proposed technique outperforms the existing techniques in terms of both quality measures and recognition rate. We also show that character contour restoration is effective for text detection in video and natural scene images. Yirui Wu, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan, Michael Blumenstein, G. Hemantha Kumar 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | A new method based on bag of filters for character recognition in scene images by learningabstractAchieving a good recognition rate for scene characters is a big challenge due to non-uniform illumination effects, perspective distortions, multiple colors or contrasts, different fonts and their various sizes, background or orientation variations, etc. Unlike the existing recognition methods that use binary information or the features extracted from different domains, the proposed method explores gray information in the form of a filter bank to extract the discriminative power for all the 62 scene character classes. We propose a sliding window (patch) operation over a character image for learning the global features, which represent the structures of character images of all the classes by reconstructing a filter bank from the original data. We introduce shareable constrains to activate class-specific filters from the filter bank. Further, we propose constraints by studying the nearest neighbor patches and exemplar selection to maximize the gap between inter-classes and minimize the gap between intra-classes. The method is evaluated and compared with several existing recognition methods in terms of character recognition rate. Experimental results show that the proposed method outperforms the existing methods. Qisu Li, Tong Lu 0002, Palaiahnakote Shivakumara, Umapada Pal 0001, Chew Lim Tan |
ICDAR | 3 |
| 2015 | A new wavelet-Laplacian method for arbitrarily-oriented character segmentation in video text linesabstractCharacter segmentation is an important topic to improve the overall performance of text recognition methods due to low resolution, complex background and lots of visual variations in video. This paper presents a novel idea for segmenting characters from arbitrarily-oriented text lines based on wavelet and Laplacian combination. Firstly, we explore wavelet which decomposes a given input image into sub-levels like a pyramid structure for segmenting words based on the fact that as decomposition level increases, the gap between characters decreases due to the reduction in the size of the input image, which results in a single component for each word. Secondly, for each segmented word, we propose Laplacian wavelet combination in a new way to extract text candidates. Thirdly, we propose horizontal and vertical sampling for character segmentation from words. The proposed method is tested on curved, non-horizontal and horizontal text lines of video and the ICDAR 2005 natural scene dataset to evaluate its performance. A comparative study with an existing method shows that the proposed method outperforms it in terms of precision and f-measure. Guozhu Liang, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan |
ICDAR | 2 |
| 2015 | New Gradient-Spatial-Structural Features for video script identification
Palaiahnakote Shivakumara, Ze-Huan Yuan, Danni Zhao, Tong Lu 0002, Chew Lim Tan |
Comput. Vis. Image Underst. | 1 |
| 2015 | A new Histogram Oriented Moments descriptor for multi-oriented moving text detection in video
Vijeta Khare, Palaiahnakote Shivakumara, Raveendran Paramesran |
Expert Syst. Appl. | 2 |
| 2015 | Bayesian classifier for multi-oriented video text recognition system
Sangheeta Roy, Palaiahnakote Shivakumara, Partha Pratim Roy 0001, Umapada Pal 0001, Chew Lim Tan, Tong Lu 0002 |
Expert Syst. Appl. | 2 |
| 2015 | A new ring radius transform-based thinning method for multi-oriented video characters
Yirui Wu, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001 |
Int. J. Document Anal. Recognit. | 2 |
| 2015 | Character shape restoration system through medial axis points in video
Shangxuan Tian, Palaiahnakote Shivakumara, Trung Quy Phan, Tong Lu 0002, Chew Lim Tan |
Neurocomputing | 2 |
| 2015 | Content-oriented multimedia document understanding through cross-media correlation
Tong Lu 0002, Yukang Jin, Feng Su, Palaiahnakote Shivakumara, Chew Lim Tan |
Multim. Tools Appl. | 4 |
| 2015 | Piece-wise linearity based method for text frame classification in video
Nabin Sharma, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Chew Lim Tan |
Pattern Recognit. | 2 |
| 2015 | Multi-Spectral Fusion Based Approach for Arbitrarily Oriented Scene Text Detection in Video ImagesabstractScene text detection from video as well as natural scene images is challenging due to the variations in background, contrast, text type, font type, font size, and so on. Besides, arbitrary orientations of texts with multi-scripts add more complexity to the problem. The proposed approach introduces a new idea of convolving Laplacian with wavelet sub-bands at different levels in the frequency domain for enhancing low resolution text pixels. Then, the results obtained from different sub-bands (spectral) are fused for detecting candidate text pixels. We explore maxima stable extreme regions along with stroke width transform for detecting candidate text regions. Text alignment is done based on the distance between the nearest neighbor clusters of candidate text regions. In addition, the approach presents a new symmetry driven nearest neighbor for restoring full text lines. We conduct experiments on our collected video data as well as several benchmark data sets, such as ICDAR 2011, ICDAR 2013, and MSRA-TD500 to evaluate the proposed method. The proposed approach is compared with the state-of-the-art methods to show its superiority to the existing methods. Guozhu Liang, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan |
IEEE Trans. Image Process. | 2 |
| 2015 | A New Technique for Multi-Oriented Scene Text Line Detection and Tracking in VideoabstractText detection and tracking in video is challenging due to contrast, resolution and background variations, and different orientations and text movements. In addition, the presence of both caption and scene texts in video aggravates the problem because these two text types differ in characteristics significantly . This paper proposes a new technique for detecting and tracking video texts of any orientation by using spatial and temporal information, respectively. The technique explores gradient directional symmetry at component level for smoothing edge components before text detection. Spatial information is preserved by forming Delaunay triangulation in a novel way at this level, which results in text candidates. Text characteristics are then proposed in a different way for eliminating false text candidates , which results in potential text candidates. Then grouping is proposed for combining potential text candidates regardless of orientation based on the nearest neighbor criterion. To tackle the problems of multi-font and multi-sized texts, we propose multi-scale integration by a pyramid structure, which helps in extracting full text lines. Then, the detected text lines are tracked in video by matching the subgraphs of triangulation. Experimental results for text detection and tracking on our video dataset, the benchmark video datasets, and the natural scene image benchmark datasets show that the proposed method is superior to the state-of-the-art methods in terms of recall, precision , and F-measure. Liang Wu 0009, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan |
IEEE Trans. Multim. | 2 |
| 2014 | Separation of Graphics (Superimposed) and Scene Text in Video FramesabstractThe presence of both graphics and scene text in video frames makes text detection and recognition problem more challenging because the nature of the two texts differs significantly. This paper aims to propose a novel method for separation of graphics and scene text to achieve good recognition rate based on the fact that Canny and Sobel edge pattern share common property for text. We propose to use Ring Radius Transform to identify the radius that represents the medial axis in the edge image. We study the intra relationship between bins of the histograms over respective radius values, resulting in intra line graphs. In this way, the method finds intra line graphs for both Canny and Sobel edge images of the input text lines. To identify the unique distribution for separation of graphics and scene texts, we explore the inter relationship between intra line graphs of Canny and Sobel edge image with respective medial axes values. This results in Gaussian distribution for graphics and non-Gaussian for scene text. Experimental results on horizontal, non-horizontal, different scripts etc. show that the proposed method is effective for classification and the results of baseline recognition methods show that recognition rate is significantly improved after classification. Palaiahnakote Shivakumara, N. Vinay Kumar, D. S. Guru, Chew Lim Tan |
Document Analysis Systems | 1 |
| 2014 | A New Laplacian Method for Arbitrarily-Oriented Word Segmentation in VideoabstractWord segmentation from video text line is challenging because video poses several challenges, such as complex background, low resolution, arbitrary orientation, etc. Besides, word segmentation is essential for improving text recognition accuracy. Therefore, we propose a novel method for segmenting words by exploring zero crossing points for each sliding window over text line. The candidate zero crossing pointes are defined based on characteristics of positive and negative Laplacian values at text region and non-text region. The percentage of candidate zero crossing points is calculated for each sliding window and is used for identifying the seed window that represents space between words. For the seed window, we propose a novel idea of horizontal and vertical sampling based on the percentage values to estimate the width and the height of the word spacing. Then the width and the height of the word spacing are used to validate the actual word spacing. Experimental results comparing with an existing method show that the proposed method is better than the existing method in terms of recall, precision and f-measure on curved, horizontal, non-horizontal, Hua's video data, as well as ICDAR data. We also test it on our own data containing multiscript text lines to show the robustness of the proposed method. Palaiahnakote Shivakumara, Mahamad Suhil, D. S. Guru, Chew Lim Tan |
Document Analysis Systems | 1 |
| 2014 | Text Detection Using Delaunay Triangulation in Video SequenceabstractText detection and tracking in video sequence is gaining interest due to the challenges posed by low resolution and complex background. This paper proposes a new method for text detection by estimating trajectories between the corners of texts in video sequence over time. Each trajectory is considered as one node to form a graph for all trajectories and Delaunay triangulation is used to obtain edges to connect nodes of the graph. In order to identify the edges that represent text regions, we propose four pruning criteria based on spatial proximity, motion coherence, local appearance and canny rate. This results in several sub-graphs. Then we use depth first search to collect corner points, which essentially represent text candidates. False positives are eliminated using heuristics and missing trajectories will be obtained by tracking the corners in temporal frames. We test the method on different videos and evaluate the method in terms of recall, precision, f-measure with existing results. Experimental result shows that the proposed method is superior to existing method. Liang Wu 0009, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan |
Document Analysis Systems | 2 |
| 2014 | A Novel Topic-Level Random Walk Framework for Scene Image Co-segmentation
Ze-Huan Yuan, Tong Lu 0002, Palaiahnakote Shivakumara |
ECCV (1) | 3 |
| 2014 | Optical flow based dynamic curved video text detectionabstractText detection in video is a challenging problem as it is useful in several real time applications in the field of video indexing and retrieval. Unlike existing methods that generally focus on horizontal caption or graphics text, the proposed method focuses on detecting dynamic curved text in video. The method explores the characteristics of the optical flow of text, namely, constant velocity, uniform magnitude distribution and unique angle distribution, to identify text candidates with the help of k-means clustering algorithm. We propose an iterative procedure which finds the standard deviation of text candidates between the first and its successive frames, and it terminates when there is a sudden decrease in the standard deviation values. The proposed method eliminates false text candidates based on the characteristics of optical flow at component level while retaining the potential text candidates. Then, direction guided boundary growing is proposed to traverse curved text lines in video. Furthermore, the characteristics of optical flow of text are utilized at block level to eliminate false positives. Experiments are conducted with various videos, including video with static text, static and dynamic text, and dynamic text only, to evaluate the proposed method. The results are benchmarked with the existing methods to verify the superiority of our method over the existing methods in terms of recall, precision, F-measure and average processing time. Palaiahnakote Shivakumara, Mohamed Lubani, Koksheik Wong, Tong Lu 0002 |
ICIP | 1 |
| 2014 | Anomaly Detection through Spatio-temporal Context Modeling in Crowded ScenesabstractA novel statistical framework for modeling the intrinsic structure of crowded scenes and detecting abnormal activities is presented in this paper. The proposed framework essentially turns the anomaly detection process into two parts, namely, motion pattern representation and crowded context modeling. During the first stage, we averagely divide the spatiotemporal volume into atomic blocks. Considering the fact that mutual interference of several human body parts potentially happen in the same block, we propose an atomic motion pattern representation using the Gaussian Mixture Model (GMM) to distinguish the motions inside each block in a refined way. Usual motion patterns can thus be defined as a certain type of steady motion activities appearing at specific scene positions. During the second stage, we further use the Markov Random Field (MRF) model to characterize the joint label distributions over all the adjacent local motion patterns inside the same crowded scene, aiming at modeling the severely occluded situations in a crowded scene accurately. By combining the determinations from the two stages, a weighted scheme is proposed to automatically detect anomaly events from crowded scenes. The experimental results on several different outdoor and indoor crowded scenes illustrate the effectiveness of the proposed algorithm. Tong Lu 0002, Liang Wu 0009, Xiaolin Ma, Palaiahnakote Shivakumara, Chew Lim Tan |
ICPR | 4 |
| 2014 | Gradient-Angular-Features for Word-wise Video Script IdentificationabstractScript identification at the word level is challenging because of complex backgrounds and low resolution of video. The presence of graphics and scene text in video makes the problem more challenging. In this paper, we employ gradient angle segmentation on words from video text lines. This paper presents new Gradient-Angular-Features (GAF) for video script identification, namely, Arabic, Chinese, English, Japanese, Korean and Tamil. This work enables us to select an appropriate OCR when the frame has words of multi-scripts. We employ gradient directional features for segmenting words from video text lines. For each segmented word, we study the gradient information in effective ways to identify text candidates. The skeleton of the text candidates is analyzed to identify Potential Text Candidates (PTC) by filtering out unwanted text candidates. We propose novel GAF for the PTC to study the structure of the components in the form of cursiveness and softness. The histogram operation on the GAF is performed in different ways to obtain discriminative features. The method is evaluated on 760 words of six scripts having low contrast, complex background, different font sizes, etc. in terms of the classification rate and is compared with an existing method to show the effectiveness of the method. We achieve 88.2% average classification rate. Palaiahnakote Shivakumara, Nabin Sharma, Umapada Pal 0001, Michael Blumenstein, Chew Lim Tan |
ICPR | 1 |
| 2014 | Graphics and Scene Text Classification in VideoabstractAchieving good accuracy for text detection and recognition is a challenging and interesting problem in the field of video document analysis because of the presences of both graphics text that has good clarity and scene text that is unpredictable in video frames. Therefore, in this paper, we present a novel method for classifying graphics texts and scene texts by exploiting temporal information and finding the relationship between them in video. The method proposes an iterative procedure to identify Probable Graphics Text Candidates (PGTC) and Probable Scene Text Candidates (PSTC) in video based on the fact that graphics texts in general do not have large movements especially compared to scene texts which are usually embedded on background. In addition to PGTC and PSTC, the iterative process automatically identifies the number of video frames with the help of a converging criterion. The method further explores the symmetry between intra and inter character components to identify graphics text candidates and scene text candidates. Boundary growing method is employed to restore the complete text line. For each segmented text line, we finally introduce Eigen value analysis to classify graphics and scene text lines based on the distribution of respective Eigen values. Experimental results with the existing methods show that the proposed method is effective and useful to improve the accuracy of text detection and recognition. Jiamin Xu, Palaiahnakote Shivakumara, Tong Lu 0002, Trung Quy Phan, Chew Lim Tan |
ICPR | 2 |
| 2014 | 2D and 3D Video Scene Text ClassificationabstractText detection and recognition is a challenging problem in document analysis due to the presence of the unpredictable nature of video texts, such as the variations of orientation, font and size, illumination effects, and even different 2D/3D text shadows. In this paper, we propose a novel horizontal and vertical symmetry feature by calculating the gradient direction and the gradient magnitude of each text candidate, which results in Potential Text Candidates (PTCs) after applying the k-means clustering algorithm on the gradient image of each input frame. To verify PTCs, we explore temporal information of video by proposing an iterative process that continuously verifies the PTCs of the first frame and the successive frames, until the process meets the converging criterion. This outputs Stable Potential Text Candidates (SPTCs). For each SPTC, the method obtains text representatives with the help of the edge image of the input frame. Then for each text representative, we divide it into four quadrants and check a new Mutual Nearest Neighbor Symmetry (MNNS) based on the dominant stroke width distances of the four quadrants. A voting method is finally proposed to classify each text block as either 2D or 3D by counting the text representatives that satisfy MNNS. Experimental results on classifying 2D and 3D text images are promising, and the results are further validated by text detection and recognition before classification and after classification with the exiting methods, respectively. Jiamin Xu, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan |
ICPR | 2 |
| 2014 | A robust arbitrary text detection system for natural scene images
Anhar Risnumawan, Palaiahnakote Shivakumara, Chee Seng Chan, Chew Lim Tan |
Expert Syst. Appl. | 2 |
| 2014 | Multi-oriented scene text detection in video based on wavelet and angle projection boundary growing
Palaiahnakote Shivakumara, Anjan Dutta 0001, Chew Lim Tan, Umapada Pal 0001 |
Multim. Tools Appl. | 1 |
| 2014 | Semiautomatic Ground Truth Generation for Text Detection and Recognition in Video ImagesabstractAlthough a large number of methods for video text detection and recognition have been proposed over the past years, it is hard to find the best state-of-the-art method because of nonavailability of standard datasets, ground truth, and common evaluation measures. Therefore, in this paper, we propose a semiautomatic system for ground truth generation for video text detection and recognition, which includes English and Chinese text of different orientation. The system has a facility to allow the user to manually correct the ground truth if the automatic method produces incorrect results. We propose eleven attributes at the word level, namely: line index, word index, coordinate values of bounding box, area, content, script type, orientation information, type of text (caption/scene), condition of text (distortion/distortion free), start frame, and end frame to evaluate the performance of the method. We also introduce a new dataset that consists of 466 video frames collected from TRECVID 2005 and 2006 databases. The video frames in our dataset contain both horizontal texts (278 frames: 181 with English texts and 97 with Chinese texts) and nonhorizontal texts (188 frames: 140 English and 48 Chinese). Furthermore, the performance of the proposed system is compared with existing text detection methods by calculating measures manually and automatically to show usefulness of our semiautomatic system. The ground truth and the semiautomatic system will be released to the public. Trung Quy Phan, Palaiahnakote Shivakumara, Souvik Bhowmick, Shimiao Li, Chew Lim Tan, Umapada Pal 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Recognizing Text with Perspective Distortion in Natural ScenesabstractThis paper presents an approach to text recognition in natural scene images. Unlike most existing works which assume that texts are horizontal and frontal parallel to the image plane, our method is able to recognize perspective texts of arbitrary orientations. For individual character recognition, we adopt a bag-of-key points approach, in which Scale Invariant Feature Transform (SIFT) descriptors are extracted densely and quantized using a pre-trained vocabulary. Following [1, 2], the context information is utilized through lexicons. We formulate word recognition as finding the optimal alignment between the set of characters and the list of lexicon words. Furthermore, we introduce a new dataset called StreetViewText-Perspective, which contains texts in street images with a great variety of viewpoints. Experimental results on public datasets and the proposed dataset show that our method significantly outperforms the state-of-the-art on perspective texts of arbitrary orientations. Trung Quy Phan, Palaiahnakote Shivakumara, Shangxuan Tian, Chew Lim Tan |
ICCV | 2 |
| 2013 | Scene Character Detection by an Edge-Ray FilterabstractEdge is a type of valuable clues for scene character detection task. Generally, the existing edge-based methods rely on the assumption of straight text line to prune away the non-character candidates. This paper proposes a new edge-based method, called edge-ray filter, to detect the scene character. The main contribution of the proposed method lies in filtering out complex backgrounds by fully utilizing the essential spatial layout of edges instead of the assumption of straight text line. Edges are extracted by a combination of Canny and Edge Preserving Smoothing Filter (EPSF). To effectively boost the filtering strength of the designed edge-ray filter, we employ a new Edge Quasi-Connectivity Analysis (EQCA) to unify complex edges as well as contour of broken character. Label Histogram Analysis (LHA) then filters out non-character edges and redundant rays through setting proper thresholds. Finally, two frequently-used heuristic rules, namely aspect ratio and occupation, are exploited to wipe off distinct false alarms. In addition to have the ability to handle special scenarios, the proposed method can accommodate dark-on-bright and bright-on-dark characters simultaneously, and provides accurate character segmentation masks. We perform experiments on the benchmark ICDAR 2011 Robust Reading Competition dataset as well as scene images with special scenarios. The experimental results demonstrate the validity of our proposal. Rong Huang 0003, Palaiahnakote Shivakumara, Seiichi Uchida |
ICDAR | 2 |
| 2013 | Recognition of Video Text through Temporal IntegrationabstractThis paper presents a method for temporal integration, which can be used to improve the recognition accuracy of video texts. Given a word detected in a video frame, we use a combination of Stroke Width Transform and SIFT (Scale Invariant Feature Transform) to track it both backward and forward in time. The text instances within the word's frame span are then extracted and aligned at pixel level. In the second step, we integrate these instances into a text probability map. By thresholding this map, we obtain an initial binarization of the word. In the final step, the shapes of the characters are refined using the intensity values. This helps to preserve the distinctive character features (e.g., sharp edges and holes), which are useful for OCR engines to distinguish between the different character classes. Experiments on English and German videos show that the proposed method outperforms existing ones in terms of recognition accuracy. Trung Quy Phan, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan |
ICDAR | 2 |
| 2013 | A New Method for Character Segmentation from Multi-oriented Video WordsabstractThis paper presents a two-stage method for multi-oriented video character segmentation. Words segmented from video text lines are considered for character segmentation in the present work. Words can contain isolated or non-touching characters, as well as touching characters. Therefore, the character segmentation problem can be viewed as a two stage problem. In the first stage, text cluster is identified and isolated (non-touching) characters are segmented. The orientation of each word is computed and the segmentation paths are found in the direction perpendicular to the orientation. Candidate segmentation points computed using the top distance profile are used to find the segmentation path between the characters considering the background cluster. In the second stage, the segmentation results are verified and a check is performed to ascertain whether the word component contains touching characters or not. The average width of the components is used to find the touching character components. For segmentation of the touching characters, segmentation points are then found using average stroke width information, along with the top and bottom distance profiles. The proposed method was tested on a large dataset and was evaluated in terms of precision, recall and f-measure. A comparative study with existing methods reveals the superiority of the proposed method. Nabin Sharma, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Chew Lim Tan |
ICDAR | 2 |
| 2013 | Detection of Curved Text in Video: Quad Tree Based MethodabstractIn this paper, we address curved text detection in video through a new enhancement criterion and the use of quad tree. The proposed method makes use of the quad tree to simplify the task of handling the entire frame at each stage. The proposed method employs a novel criterion for grouping of pixels based on their R, G and B values to enhance text information. As generally, a text detection problem is a two class problem, we used k-means with k=2 to identify potential text candidate pixels. From these potential candidates, connected components are then extracted and subjected to further analysis, where symmetry property based on stroke width is used for further authentication of the text representatives. These authenticated text representatives are then exploited as seed points to restore the text information with reference to the Sobel edge frame of the original input frame. To preserve the spatial information of text pixels the concept of quad tree is applied. From these seed blocks, text lines are extracted by the use of a region growing approach driven completely based on Sobel edge map. The proposed method is tested on curved video data and Hua's horizontal video text data in terms of recall, precision, f-measure, misdetection rate and processing time. The results are compared and analyzed to show that the proposed method outperforms several existing methods in terms of accuracy and efficiency. Palaiahnakote Shivakumara, H. T. Basavaraju, D. S. Guru, Chew Lim Tan |
ICDAR | 1 |
| 2013 | Scene Character Reconstruction through Medial AxisabstractCharacter shape reconstruction for the scene character is challenging and interesting because scene character usually suffers from uneven illumination, complex background, perspective distortion. To address such ill conditions, we propose to utilize Histogram Gradient Division (HGD) and Reverse Gradient Orientation (RGO) to select Candidate Text Pixels (CTPs) for a given input character. Ring Radius Transform is applied on each pixel in a CTP image to obtain radius map where each pixel is assigned a value which is the radius to the nearest CTP. Candidate medial axis pixels are those having maximum radius values in their neighborhoods. We find such pixels on horizontal, vertical, principal diagonal and secondary diagonal directions to determine the respective medial axis pixels. The union of all medial axis pixels at each pixel location is considered as a candidate medial axis pixel of the character. Then color difference and k-means clustering are employed to eliminate false candidate medial axis. The potential medial axis values are used to reconstruct the shape of the character. The method is tested on 1025 characters of complex foreground and background from ICDAR 2003 dataset in terms of shape reconstruction and recognition rate. Experimental results demonstrate the effectiveness of our proposed method for complex foreground and background characters in terms of character recognition rate and reconstruction error. Shangxuan Tian, Palaiahnakote Shivakumara, Trung Quy Phan, Chew Lim Tan |
ICDAR | 2 |
| 2013 | A novel ring radius transform for video character reconstruction
Palaiahnakote Shivakumara, Trung Quy Phan, Souvik Bhowmick, Chew Lim Tan, Umapada Pal 0001 |
Pattern Recognit. | 1 |
| 2013 | Gradient Vector Flow and Grouping-Based Method for Arbitrarily Oriented Scene Text Detection in Video ImagesabstractText detection in videos is challenging due to low resolution and complex background of videos. Besides, an arbitrary orientation of scene text lines in video makes the problem more complex and challenging. This paper presents a new method that extracts text lines of any orientations based on gradient vector flow (GVF) and neighbor component grouping. The GVF of edge pixels in the Sobel edge map of the input frame is explored to identify the dominant edge pixels which represent text components. The method extracts edge components corresponding to dominant pixels in the Sobel edge map, which we call text candidates (TC) of the text lines. We propose two grouping schemes. The first finds nearest neighbors based on geometrical properties of TC to group broken segments and neighboring characters which results in word patches. The end and junction points of skeleton of the word patches are considered to eliminate false positives, which output the candidate text components (CTC). The second is based on the direction and the size of the CTC to extract neighboring CTC and to restore missing CTC, which enables arbitrarily oriented text line detection in video frame. Experimental results on different datasets, including arbitrarily oriented text data, nonhorizontal and horizontal text data, Hua's data and ICDAR-03 data (camera images), show that the proposed method outperforms existing methods in terms of recall, precision and f-measure. Palaiahnakote Shivakumara, Trung Quy Phan, Shijian Lu, Chew Lim Tan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2012 | A New Method for Arbitrarily-Oriented Text Detection in VideoabstractText detection in video frames plays a vital role in enhancing the performance of information extraction systems because the text in video frames helps in indexing and retrieving video efficiently and accurately. This paper presents a new method for arbitrarily-oriented text detection in video, based on dominant text pixel selection, text representatives and region growing. The method uses gradient pixel direction and magnitude corresponding to Sobel edge pixels of the input frame to obtain dominant text pixels. Edge components in the Sobel edge map corresponding to dominant text pixels are then extracted and we call them text representatives. We eliminate broken segments of each text representatives to get candidate text representatives. Then the perimeter of candidate text representatives grows along the text direction in the Sobel edge map to group the neighboring text components which we call word patches. The word patches are used for finding the direction of text lines and then the word patches are expanded in the same direction in the Sobel edge map to group the neighboring word patches and to restore missing text information. This results in extraction of arbitrarily-oriented text from the video frame. To evaluate the method, we considered arbitrarily-oriented data, non-horizontal data, horizontal data, Hua's data and ICDAR-2003 competition data (Camera images). The experimental results show that the proposed method outperforms the existing method in terms of recall and f-measure. Nabin Sharma, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Chew Lim Tan |
Document Analysis Systems | 2 |
| 2012 | New Spatial-Gradient-Features for Video Script IdentificationabstractIn this paper, we present new features based on Spatial-Gradient-Features (SGF) at block level for identifying six video scripts namely, Arabic, Chinese, English, Japanese, Korean and Tamil. This works helps in enhancing the capability of the current OCR on video text recognition by choosing an appropriate OCR engine when video contains multi-script frames. The input for script identification is the text blocks obtained by our text frame classification method. For each text block, we obtain horizontal and vertical gradient information to enhance the contrast of the text pixels. We divide the horizontal gradient block into two equal parts as upper and lower at the centroid in the horizontal direction. Histogram on the horizontal gradient values of the upper and the lower part is performed to select dominant text pixels. In the same way, the method selects dominant pixels from the right and the left parts obtained by dividing the vertical gradient block vertically. The method combines the horizontal and the vertical dominant pixels to obtain text components. Skeleton concept is used to reduce pixel width to a single pixel to extract spatial features. We extract four features based on proximity between end points, junction points, intersection points and pixels. The method is evaluated on 770 frames of six scripts in terms of classification rate and is compared with an existing method. We have achieved 82.1% average classification rate. Danni Zhao, Palaiahnakote Shivakumara, Shijian Lu, Chew Lim Tan |
Document Analysis Systems | 2 |
| 2012 | Scene character detection and recognition based on multiple hypotheses framework
Rong Huang 0003, Shinpei Oba, Palaiahnakote Shivakumara, Seiichi Uchida |
ICPR | 3 |
| 2012 | Text detection in natural scenes using Gradient Vector Flow-Guided symmetry
Trung Quy Phan, Palaiahnakote Shivakumara, Chew Lim Tan |
ICPR | 2 |
| 2012 | Wavelet-gradient-fusion for video text binarization
Sangheeta Roy, Palaiahnakote Shivakumara, Partha Pratim Roy 0001, Chew Lim Tan |
ICPR | 2 |
| 2012 | A new Iterative-Midpoint-Method for video character gap filling
Palaiahnakote Shivakumara, Ding Bei Hong, Danni Zhao, Chew Lim Tan, Umapada Pal 0001 |
ICPR | 1 |
| 2012 | Detecting text in the real worldabstractThe problem of text detection in natural scene images is challenging because of the unconstrained sizes, colors, backgrounds and alignments of the characters. This paper proposes novel symmetry features for this task. Within a text line, the intra-character symmetry captures the correspondence between the inner contour and the outer contour of a character while the inter-character symmetry helps to extract information from the gap region between two consecutive characters. A formulation based on Gradient Vector Flow is used to detect both types of symmetry points. These points are then grouped into text lines using the consistency in sizes, colors, and stroke and gap thickness. Therefore, unlike most existing methods which use only character features, our method exploits both the text features and the gap features to improve the detection result. Experimentally, our method compares well to the state-of-the-art on public datasets for natural scenes and street-level images, an emerging category of image data. The proposed technique can be used in a wide range of multimedia applications such as content-based image/video retrieval, mobile visual search and sign translation. Trung Quy Phan, Palaiahnakote Shivakumara, Chew Lim Tan |
ACM Multimedia | 2 |
| 2012 | Multi-Oriented Text Detection in Scene ImagesabstractWe present a new run-length based method for multi-oriented text detection in scene images. We consider one ideal Sobel edge image of the horizontal text image to compute run-lengths for multi-oriented text images. Then the method proposes a Max–Min clustering to find ideal run-lengths that represents text pixel from an array of run-lengths of ideal image. The run-lengths computed for the input multi-oriented and horizontal text images are matched with the ideal run-lengths given by the Max–Min clustering to find potential run-lengths. The boundary growing method is introduced to traverse multi-oriented text lines given by the potential run-lengths and then the method eliminates false positives to clear the background using angle-proximity features of the text blocks. The non-horizontal text image is rotated to horizontal direction based on angle of the text lines to ease the implementation. The method explore new idea based on zero-crossing to separate text lines from the touching text lines given by the boundary growing method. The proposed method is tested on our own multi-oriented scene data captured by high resolution camera and mobile camera, and the benchmark database (ICDAR 2003 competition scene images) to evaluate the performance of the proposed method. The results are compared with the existing methods to show that the proposed method outperforms the existing methods in terms of measures. M. Basavanna, Palaiahnakote Shivakumara, G. Hemantha Kumar 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2012 | New Edge Characteristics for Scene and Object ClassificationabstractIn this paper, we show that simple edge characteristics in images, when judiciously combined, can result in improved scene and object classification. Unlike existing methods that require a large number of training samples and complex learning schemes, our method discovers simple edge properties. We introduce three sets of edge properties, namely, centroid, compactness and aspect ratio of edges in the image. The combinations of these edge properties are used to discriminate among images in each class. A class representative is calculated for each class according to the average percentage of edges that satisfy the property of a particular class. This percentage for an unknown image is compared to the class representative to assign a label to it. It is shown that this simple edge properties-based method outperforms some of the state-of-the-art results on scene and object classification on standard databases. Palaiahnakote Shivakumara, Deepu Rajan, Suresh Anand Sadananthan |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2012 | Multioriented Video Scene Text Detection Through Bayesian Classification and Boundary GrowingabstractMultioriented text detection in video frames is not as easy as detection of captions or graphics or overlaid texts, which usually appears in the horizontal direction and has high contrast compared to its background. Multioriented text generally refers to scene text that makes text detection more challenging and interesting due to unfavorable characteristics of scene text. Therefore, conventional text detection methods may not give good results for multioriented scene text detection. Hence, in this paper, we present a new enhancement method that includes the product of Laplacian and Sobel operations to enhance text pixels in videos. To classify true text pixels, we propose a Bayesian classifier without assuming a priori probability about the input frame but estimating it based on three probable matrices. Three different ways of clustering are performed on the output of the enhancement method to obtain the three probable matrices. Text candidates are obtained by intersecting the output of the Bayesian classifier with the Canny edge map of the input frame. A boundary growing method is introduced to traverse the multioriented scene text lines using text candidates. The boundary growing method works based on the concept of nearest neighbors. The robustness of the method has been tested on a variety of datasets that include our own created data (nonhorizontal and horizontal text data) and two publicly available data, namely, video frames of Hua and complex scene text data of ICDAR 2003 competition (camera images). Experimental results show that the performance of the proposed method is encouraging compared with results of existing methods in terms of recall, precision, F-measures, and computational times. Palaiahnakote Shivakumara, Rushi Padhuman Sreedhar, Trung Quy Phan, Shijian Lu, Chew Lim Tan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | Video Script Identification Based on Text LinesabstractIn this paper, we present a new method for video script identification which is essential before choosing an appropriate OCR engine for identifying text lines when a video frame contains more than one language. The input for script identification is the text lines obtained by our text detection method. We extract upper and lower extreme points for each connected component of Canny edges of text lines. The extracted points are connected to study the behavior of upper and lower lines. The direction of each 10-pixel segment of the lines is determined using PCA. The average angle of the segments of the upper and lower lines is computed to study the smoothness and cursiveness of the lines. In addition, to discriminate the scripts accurately, the method divides a text line into five equal zones horizontally to study the smoothness and cursiveness of the upper and lower lines of each zone. We evaluate the method by conducting experiments on different combinations of languages such as English and Chinese, English and Tamil, Chinese and Tamil, and English, Chinese and Tamil. Trung Quy Phan, Palaiahnakote Shivakumara, Zhang Ding, Shijian Lu, Chew Lim Tan |
ICDAR | 2 |
| 2011 | A Gradient Vector Flow-Based Method for Video Character SegmentationabstractIn this paper, we propose a method based on gradient vector flow for video character segmentation. By formulating character segmentation as a minimum cost path finding problem, the proposed method allows curved segmentation paths and thus it is able to segment overlapping characters and touching characters due to low contrast and complex background. Gradient vector flow is used in a new way to identify candidate cut pixels. A two-pass path finding algorithm is then applied where the forward direction helps to locate potential cuts and the backward direction serves to remove the false cuts, i.e. those that go through the characters, while retaining the true cuts. Experimental results show that the proposed method outperforms an existing method on multi-oriented English and Chinese video text lines. The proposed method also helps to improve binarization results, which lead to a better character recognition rate. Trung Quy Phan, Palaiahnakote Shivakumara, Bolan Su, Chew Lim Tan |
ICDAR | 2 |
| 2011 | A New Fourier-Moments Based Video Word and Character Extraction Method for RecognitionabstractThis paper presents a new method based on Fourier and moments features to extract words and characters from a video text line in any direction for recognition. Unlike existing methods which output the entire text line to the ensuing recognition algorithm, the proposed method obtains each extracted character from the text line as input to the recognition algorithm because the background of a single character is relatively simple compared to the text line and words. Max-Min clustering criterion is introduced to obtain text cluster from the extracted Fourier and moments feature set. Union of the text cluster with Canny operation of the input video text line is proposed to obtain missing text candidates. Then a run length criterion is used for extraction of words. From the words, we propose a new idea for extracting characters from the text candidates of each word image based on the fact that the text height difference at the character boundary column is smaller than that at other columns of the word image. We evaluate the method on a large dataset at three levels namely text line, words and characters in terms of recall, precision and f-measure. In addition to this, we show that the recognition result for the extracted character is better than words and lines. Our experimental set up involves 3527 characters including Chinese. The dataset is selected from TRECVID database of 2005 and 2006. Deepak Rajendran, Palaiahnakote Shivakumara, Bolan Su, Shijian Lu, Chew Lim Tan |
ICDAR | 2 |
| 2011 | A New Gradient Based Character Segmentation Method for Video Text RecognitionabstractThe current OCR cannot segment words and characters from video images due to complex background as well as low resolution of video images. To have better accuracy, this paper presents a new gradient based method for words and character segmentation from text line of any orientation in video frames for recognition. We propose a Max-Min clustering concept to obtain text cluster from the normalized absolute gradient feature matrix of the video text line image. Union of the text cluster with the output of Canny operation of the input video text line is proposed to restore missing text candidates. Then a run length algorithm is applied on the text candidate image for identifying word gaps. We propose a new idea for segmenting characters from the restored word image based on the fact that the text height difference at the character boundary column is smaller than that of the other columns of the word image. We have conducted experiments on a large dataset at two levels (word and character level) in terms of recall, precision and f-measure. Our experimental setup involves 3527 characters of English and Chinese, and this dataset is selected from TRECVID database of 2005 and 2006. Palaiahnakote Shivakumara, Souvik Bhowmick, Bolan Su, Chew Lim Tan, Umapada Pal 0001 |
ICDAR | 1 |
| 2011 | Video Character Recognition through Hierarchical ClassificationabstractWe present a new video character recognition method based on hierarchical classification. In the first step, we propose a method for character segmentation of the text line detected by the text detection method. The segmentation algorithm uses dynamic programming to find least-cost paths in the gray domain to identify the spaces between characters. For the segmented characters, we get a Canny edge image as input for the character recognition step. We introduce hierarchical classification based on voting criteria with structural features to classify 62 character classes into different smaller classes. We divide the perimeter of a character into 8 segments according to 8 directions at the centroid. Then the shape of each segment is studied to recognize the characters based on distances between the centroid and end points, and distances between the midpoint and end points. Our experiments on 1462 characters of upper case, lower case and numerals shows that 10% samples per class for training is enough to obtain 94.5% recognition accuracy. The dataset is chosen from TRECVID database of 2005 and 2006. Palaiahnakote Shivakumara, Trung Quy Phan, Shijian Lu, Chew Lim Tan |
ICDAR | 1 |
| 2011 | A Laplacian Approach to Multi-Oriented Text Detection in VideoabstractIn this paper, we propose a method based on the Laplacian in the frequency domain for video text detection. Unlike many other approaches which assume that text is horizontally-oriented, our method is able to handle text of arbitrary orientation. The input image is first filtered with Fourier-Laplacian. K-means clustering is then used to identify candidate text regions based on the maximum difference. The skeleton of each connected component helps to separate the different text strings from each other. Finally, text string straightness and edge density are used for false positive elimination. Experimental results show that the proposed method is able to handle graphics text and scene text of both horizontal and nonhorizontal orientation. Palaiahnakote Shivakumara, Trung Quy Phan, Chew Lim Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | A novel mutual nearest neighbor based symmetry for text frame classification in video
Palaiahnakote Shivakumara, Anjan Dutta 0001, Trung Quy Phan, Chew Lim Tan, Umapada Pal 0001 |
Pattern Recognit. | 1 |
| 2010 | An eigen value based approach for text detection in videoabstractIn this paper, a novel approach for detection of text and non-text regions in video frames is proposed. The proposed approach performs block wise eigen analysis on the gradient image of the video frame. For each block of the gradient frame, the dominant eigen value is computed to decide if the block could be a candidate text block. The K-means clustering is then applied to further identify text blocks among the candidate blocks. From each of the identified candidate text blocks edges are extracted using the sobel operator, and then by the use of horizontal and vertical profiles a bounding rectangle is fixed up. Further, geometric properties of the identified text regions are studied to eliminate false text regions. In order to validate the efficacy of the proposed approach, experimentation on a dataset containing 800 video frames has been carried out. The obtained results ensure that the proposed approach is with increased text detection rate with very low false and misdetection rates when compared to the other existing state of the art techniques. D. S. Guru, S. Manjunath, Palaiahnakote Shivakumara, Chew Lim Tan |
Document Analysis Systems | 3 |
| 2010 | A skeleton-based method for multi-oriented video text detectionabstractIn this paper, we propose a method based on the skeletonization operation for multi-oriented video text detection. The first step uses our existing Laplacian-based method to identify candidate text regions. In the second step, each region is classified as either a simple connected component (a single text string) or a complex connected component (multiple text strings that are connected to each other) depending on the number of intersection points in its skeleton. Complex connected components are then segmented into constituent parts based on the skeleton segments in order to separate the text strings from each other. Finally, text string straightness and edge density are used for false positive elimination. Experimental results show that the proposed method is able to detect multi-oriented graphics text and scene text. Trung Quy Phan, Palaiahnakote Shivakumara, Chew Lim Tan |
Document Analysis Systems | 2 |
| 2010 | A new wavelet-median-moment based method for multi-oriented video text detectionabstractIn this paper, we present a new method based on wavelet-median-moments and a novel idea of angle projection for detecting multi-oriented text in video. The proposed method uses wavelet decomposition first to obtain three high frequency sub-bands (LH, HL and HH) and then median moments are computed on the average sub-bands of the three high frequency sub-bands to brighten the text pixels. K-means clustering (K=2) is used for obtaining text pixels from the wavelet-median-moments features (WMMF). Text candidates are obtained by mapping the output of K-means on Sobel edge map of the original input frame. To deal with multi-oriented text, we introduce a new idea of Angle Projection (AP) based on boundary growing and nearest neighbor concepts from the text candidates instead of conventional projection profiles. The proposed method is experimented on horizontal text data, non-horizontal text data, temporal data, non-text data and camera based images (scene text data of ICDAR 2003 competition) to show that the proposed method is superior to existing methods. Palaiahnakote Shivakumara, Anjan Dutta 0001, Chew Lim Tan, Umapada Pal 0001 |
Document Analysis Systems | 1 |
| 2010 | A New Method for Handwritten Scene Text Detection in VideoabstractThere are many video images where hand written text may appear. Therefore handwritten scene text detection in video is essential and useful for many applications for efficient indexing, retrieval etc. Also there are many video frames where text line may be multi-oriented in nature. To the best of our knowledge there is no work on handwritten text detection in video, which is multi-oriented in nature. In this paper, we present a new method based on maximum color difference and boundary growing method for detection of multi-oriented handwritten scene text in video. The method computes maximum color difference for the average of R, G and B channels of the original frame to enhance the text information. The output of maximum color difference is fed to a K-means algorithm with K = 2 to separate text and non-text clusters. Text candidates are obtained by intersecting the text cluster with the Sobel output of the original frame. To tackle the fundamental problem of different orientations and skews of handwritten text, boundary growing method based on a nearest neighbor concept is employed. We evaluate the proposed method by testing on our own handwritten text database and publicly available video data (Hua's data). Experimental results obtained from the proposed method are promising. Palaiahnakote Shivakumara, Anjan Dutta 0001, Umapada Pal 0001, Chew Lim Tan |
ICFHR | 1 |
| 2010 | A New Symmetry Based on Proximity of Wavelet-Moments for Text Frame Classification in VideoabstractThis paper proposes the use of a new symmetry property based on proximity of the median moments in the wavelet domain. The method divides a given frame into 16 equally sized blocks to classify the true text frame. The average of high frequency subbands of a block is used for computing median moments to brighten the text pixel in a block of video frame. Then K-means clustering with K=2 is applied on the median moments of the block to classify it as a probable text block. For classified blocks, average wavelet median moments are computed for a sliding window. We introduce Max-Min cluster to classify the probable text pixel in each probable text block. The four quadrants are formed from the centroid of the probable text pixels. The new concept called symmetry is introduced to identify the true text block based on proximity between probable text pixels in each quadrant. If the frame produces at least one true text block, it is considered as a text frame otherwise a non-text frame. The method is tested on three datasets to evaluate the robustness of the method in classification of text frames in terms of recall and precision. Palaiahnakote Shivakumara, Anjan Dutta 0001, Chew Lim Tan, Umapada Pal 0001 |
ICPR | 1 |
| 2010 | New Wavelet and Color Features for Text Detection in VideoabstractAutomatic text detection in video is an important task for efficient and accurate indexing and retrieval of multimedia data such as events identification, events boundary identification etc. This paper presents a new method comprising of wavelet decomposition and color features namely R, G and B. The wavelet decomposition is applied on three color bands separately to obtain three high frequency sub bands (LH, HL and HH) and then the average of the three sub bands for each color band is computed further to enhance the text pixels in video frame. To take advantage of wavelet and color information, we again take the average of the three average images (AoA) obtained by the former step to increase the gap between text and non text pixels. Our previous Laplacian method is employed on AoA for text detection. The proposed method is evaluated by testing on a large dataset which includes publicly available data, non text data and ICDAR-03 data. Comparative study with existing methods shows that the results of the proposed method are encouraging and useful. Palaiahnakote Shivakumara, Trung Quy Phan, Chew Lim Tan |
ICPR | 1 |
| 2010 | Novel Edge Features for Text Frame Classification in VideoabstractText frame classification is needed in many applications such as event identification, exact event boundary identification, navigation, video surveillance in multimedia etc. To the best of our knowledge, there are no methods reported solely dedicated to text frame classifications so far. Hence this paper presents a new approach to text frame classification in video based on capturing local observable edge properties of text frames, by virtue of the strong presence of sharp edges, straight appearances of edges and consistent proximity between edges. The approach initially classifies the blocks of the frame into text blocks and non-text blocks. The true text block is then identified among classified text blocks to detect text frames by the proposed features. If the text frame produces one true text block then it is considered as a text frame otherwise a non-text frame. We evaluate the proposed approach on a large database containing both text and nontext frames and publicly available data at two levels, i.e., estimating recall and precision at the block level and the frame level. Palaiahnakote Shivakumara, Chew Lim Tan |
ICPR | 1 |
| 2010 | Accurate video text detection through classification of low and high contrast images
Palaiahnakote Shivakumara, Weihua Huang, Trung Quy Phan, Chew Lim Tan |
Pattern Recognit. | 1 |
| 2010 | New Fourier-Statistical Features in RGB Space for Video Text DetectionabstractIn this paper, we propose new Fourier-statistical features (FSF) in RGB space for detecting text in video frames of unconstrained background, different fonts, different scripts, and different font sizes. This paper consists of two parts namely automatic classification of text frames from a large database of text and non-text frames and FSF in RGB for text detection in the classified text frames. For text frame classification, we present novel features based on three visual cues, namely, sharpness in filter-edge maps, straightness of the edges, and proximity of the edges to identify a true text frame. For text detection in video frames, we present new Fourier transform based features in RGB space with statistical features and the computed FSF features from RGB bands are subject to K-means clustering to classify text pixels from the background of the frame. Text blocks of the classified text pixels are determined by analyzing the projection profiles. Finally, we introduce a few heuristics to eliminate false positives from the frame. The robustness of the proposed approach is tested by conducting experiments on a variety of frames of low contrast, complex background, different fonts, and sizes of text in the frame. Both our own test dataset and a publicly available dataset are used for the experiments. The experimental results show that the proposed approach is superior to existing approaches in terms of detection rate, false positive rate, and misdetection rate. Palaiahnakote Shivakumara, Trung Quy Phan, Chew Lim Tan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | A Laplacian Method for Video Text DetectionabstractIn this paper, we propose an efficient text detection method based on the Laplacian operator. The maximum gradient difference value is computed for each pixel in the Laplacian-filtered image. K-means is then used to classify all the pixels into two clusters: text and non-text. For each candidate text region, the corresponding region in the Sobel edge map of the input image undergoes projection profile analysis to determine the boundary of the text blocks. Finally, we employ empirical rules to eliminate false positives based on geometrical properties. Experimental results show that the proposed method is able to detect text of different fonts, contrast and backgrounds. Moreover, it outperforms three existing methods in terms of detection and false positive rates. Trung Quy Phan, Palaiahnakote Shivakumara, Chew Lim Tan |
ICDAR | 2 |
| 2009 | A Gradient Difference Based Technique for Video Text DetectionabstractText detection in video images has received increasing attention, particularly in scene text detection in video images, as it plays a vital role in video indexing and information retrieval. This paper proposes a new and robust gradient difference technique for detecting both graphics and scene text in video images. The technique introduces the concept of zero crossing to determine the bounding boxes for the detected text lines in video images, rather than using the conventional projection profiles based method which fails to fix bounding boxes when there is no proper spacing between the detected text lines. We demonstrate the capability of the proposed technique by conducting experiments on video images containing both graphics text and scene text with different font shapes and sizes, languages, text directions, background and contrasts. Our experimental results show that the proposed technique outperforms existing methods in terms of detection rate for large video image database. Palaiahnakote Shivakumara, Trung Quy Phan, Chew Lim Tan |
ICDAR | 1 |
| 2009 | A Robust Wavelet Transform Based Technique for Video Text DetectionabstractIn this paper, we propose a new method based on wavelet transform, statistical features and central moments for both graphics and scene text detection in video images. The method uses wavelet single level decomposition LH, HL and HH subbands for computing features and the computed features are fed to k means clustering to classify the text pixel from the background of the image. The average of wavelet subbands and the output of k means clustering helps in classifying true text pixel in the image. The text blocks are detected based on analysis of projection profiles. Finally, we introduce a few heuristics to eliminate false positives from the image. The robustness of the proposed method is tested by conducting experiments on a variety of images of low contrast, complex background, different fonts, and size of text in the image. The experimental results show that the proposed method outperforms the existing methods in terms of detection rate, false positive rate and misdetection rate. Palaiahnakote Shivakumara, Trung Quy Phan, Chew Lim Tan |
ICDAR | 1 |
| 2009 | Video text detection based on filters and edge featuresabstractText detection plays a vital role in retrieving and browsing video data efficiently and accurately. In this paper, we propose a method for detecting both graphics and scene text in video images by proposing initial text block identification, text portion segmentation and new edge features for false positive elimination. The heuristic rules based on filters and edge analysis are formed to identify the initial text block and to segment the complete text portion from the image. The new edge features such as straightness and cursiveness are explored to eliminate false positives. To evaluate the performance of the proposed method, we introduce misdetection rate and processing time in addition to detection rate and false positive rate. The experimental results show that the proposed method outperforms existing methods in terms of the above metrics. Palaiahnakote Shivakumara, Trung Quy Phan, Chew Lim Tan |
ICME | 1 |
| 2008 | An Efficient Edge Based Technique for Text Detection in Video FramesabstractBoth graphic text and scene text detection in video images with complex background and low resolution is still a challenging and interesting problem for researchers in the field of image processing and computer vision. In this paper, we present a novel technique for detecting both graphic text and scene text in video images by finding segments containing text in an input image and then using statistical features such as vertical and horizontal bars for edges in the segments for detecting true text blocks efficiently. To identify a segment containing text, heuristic rules are formed based on combination of filters and edge analysis. Furthermore, the same rules are extended to grow the boundaries of a candidate segment in order to include complete text in the input image. The experimental results of the proposed method show that the technique performs better than existing methods in terms of a number of metrics. Palaiahnakote Shivakumara, Weihua Huang, Chew Lim Tan |
Document Analysis Systems | 1 |
| 2008 | Detecting moving text in video using temporal informationabstractThis paper presents our work on automatically detecting moving rigid text in digital videos. The temporal information is obtained by dividing a video frame into sub-blocks and calculating inter-frame motion vector for each sub-block. Text blocks are then extracted through both intra-frame classification and inter-frame spatial relationship checking. Unlike previous works, our method achieves both detection and tracking of moving text at the same time. The method works very well detecting scrolling text in news clips and movies, and is robust towards low resolution and complex background. The computational efficiency of the method is also discussed. Weihua Huang, Palaiahnakote Shivakumara, Chew Lim Tan |
ICPR | 2 |
| 2008 | Efficient video text detection using edge featuresabstractIn this paper, we explore new edge features such as straightness for the elimination of non significant edges from the segmented text portion of a video frame to detect accurate boundary of the text lines in video images. To segment the complete text portions, the method introduces candidate text block selection from a given image. Heuristic rules are formed based on combination of filters and edge analysis for identifying a candidate text block in the image. Furthermore, the same rules are extended to grow boundary of candidate text block in order to segment complete text portions in the image. The experimental results of the proposed method show that the method outperforms an existing method in terms of a number of metrics. Palaiahnakote Shivakumara, Weihua Huang, Chew Lim Tan |
ICPR | 1 |
| 2008 | Image classification: Are rule-based systems effective when classes are fixed and known?abstractIn this paper, we investigate if rule-based systems are useful for image classification problems when the number of classes is fixed. The rules are derived from simple edge features such as width and straightness. A class representative is calculated for each class according to the average percentage of edges that satisfy the rule for a particular class. This percentage for an unknown image is compared to the class representative to assign a label to it. The proposed system does not require extensive feature extraction and classification techniques. It is shown that the rule based system outperforms some of the reported results on scene classification. Palaiahnakote Shivakumara, Deepu Rajan, Suresh Anand Sadananthan |
ICPR | 1 |
| 2006 | Diagonal Fisher linear discriminant analysis for efficient face recognition
Noushath Shaffi, G. Hemantha Kumar 0001, Palaiahnakote Shivakumara |
Neurocomputing | 3 |
| 2006 | Sliding window based approach for document image mosaicing
Palaiahnakote Shivakumara, G. Hemantha Kumar 0001, D. S. Guru, P. Nagabhushan |
Image Vis. Comput. | 1 |
| 2006 | (2D)2LDA: An efficient approach for face recognition
Noushath Shaffi, G. Hemantha Kumar 0001, Palaiahnakote Shivakumara |
Pattern Recognit. | 3 |
| 2006 | A novel boundary growing approach for accurate skew estimation of binary document images
Palaiahnakote Shivakumara, G. Hemantha Kumar 0001 |
Pattern Recognit. Lett. | 1 |
| 2005 | A New Moments based Skew Estimation Technique using Pixels in the Word for Binary Document ImagesabstractAccurate skew angle estimation is an essential component in document analysis system to enhance the performance of the optical character recognition (OCR). In this paper, a new and efficient moments based method to estimate skew angle of a pixels in the word in the scanned document image is proposed. The proposed technique has two stages. In the first stage, using boundary-growing method, pixels in the words of skewed text are extracted. The pixels in the words extracted are given as input to moments based method. It results in a skew angle in the second stage. Extensive experiments have been conducted on various types of documents such as documents containing different languages and different fonts to reveal the robustness of the proposed method. Comparative studies with the well-known methods are presented to show that the proposed method is superior in terms of accuracy and computational efficiency. Palaiahnakote Shivakumara, G. Hemantha Kumar 0001, H. S. Varsha, S. Rekha, M. R. Rashmi Nayaka |
ICDAR | 1 |
| 2004 | An Efficient Skew Estimation Technique for Binary Document Images Based on Boundary Growing and Linear Regression Analysis
Palaiahnakote Shivakumara, G. Hemantha Kumar 0001, D. S. Guru, P. Nagabhushan |
ICONIP | 1 |