VLDB 2026 Research / reviewers in the wild / expert
Michael Blumenstein
dblp:45/1824
· DBLP profile ↗
172ranked-venue papers
9as first author
37since 2021 · last 2025
0000-0002-9908-3744ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 113 · 6 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 50 · 1 first-author · 16 since 2021Databases, data management, data science and information retrieval · 40 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 1 since 2021Security and privacy · 5Human-computer interaction and ubiquitous computing · 4 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A New Fourier-Attention Guided Approach for Domain-Agnostic Text Localization
Arnab Halder, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Yue Lu 0001 |
ICDAR (3) | 4 |
| 2025 | FASTER: A Font-Agnostic Scene Text Editing and Rendering FrameworkabstractScene Text Editing (STE) is a challenging research prob-lem, that primarily aims towards modifying existing texts in an image while preserving the background and the font style of the original text. Despite its utility in numerous real-world applications, existing style-transfer-based approaches have shown sub-par editing performance due to (1) complex image backgrounds, (2) diverse font attributes, and (3) varying word lengths within the text. To address such limitations, in this paper, we propose a novel font-agnostic scene text editing and rendering framework, named FASTER, for simultaneously generating text in arbitrary styles and locations while preserving a natural and realistic appearance and structure. A combined fusion of target mask generation and style transfer units, with a cascaded self-attention mech-anism has been proposed to focus on multi-level text region edits to handle varying word lengths. Extensive evaluation on a real-world database withfurther subjective human eval-uation study indicates the superiority of FASTER in both scene text editing and rendering tasks, in terms of model per-formance and efficiency. The code and pre-trained models have been released in our Gi thub repo. Alloy Das, Sanket Biswas, Prasun Roy, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein, Josep Lladós 0001, Saumik Bhattacharya |
WACV | 6 |
| 2024 | Pre-training Cross-Modal Retrieval by Expansive Lexicon-Patch AlignmentabstractRecent large-scale vision-language pre-training depends on image-text global alignment by contrastive learning and is further boosted by fine-grained alignment in a weakly contrastive manner for cross-modal retrieval. Nonetheless, besides semantic matching learned by contrastive learning, cross-modal retrieval also largely relies on object matching between modalities. This necessitates fine-grained categorical discriminative learning, which however suffers from scarce data in full-supervised scenarios and information asymmetry in weakly-supervised scenarios when applied to cross-modal retrieval. To address these issues, we propose expansive lexicon-patch alignment (ELA) to align image patches with a vocabulary rather than only the words explicitly in the text for annotation-free alignment and information augmentation, thus enabling more effective fine-grained categorical discriminative learning for cross-modal retrieval. Experimental results show that ELA could effectively learn representative fine-grained information and outperform state-of-the-art methods on cross-modal retrieval. Yiyuan Yang, Guodong Long, Michael Blumenstein, Xiubo Geng, Chongyang Tao, Tao Shen 0001, Daxin Jiang |
LREC/COLING | 3 |
| 2024 | A New Unsupervised Approach for Text Localization in Shaky and Non-shaky Scene Video
Arnab Halder, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Cheng-Lin Liu 0001 |
ICDAR (5) | 4 |
| 2024 | λ-Color: Amplifying Long-Range Dependencies for Image Colorization
Subhankar Ghosh, Saumik Bhattacharya, Prasun Roy, Umapada Pal 0001, Michael Blumenstein |
ICPR (22) | 5 |
| 2024 | A New HourGlass Network for Detecting Text in Shaky and Non-shaky Video Frames
Arnab Halder, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Shivanand S. Gornale |
ICPR (20) | 4 |
| 2024 | Handwriting Intra-Variability Across Surface Transitions: Implications for Writer Identification
Kumari Priya, Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein |
ICPR (20) | 4 |
| 2024 | d-Sketch: Improving Visual Fidelity of Sketch-to-Image Translation with Pretrained Latent Diffusion Models without Retraining
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein |
ICPR (25) | 5 |
| 2024 | Semantically Consistent Person Image Generation
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein |
ICPR (25) | 5 |
| 2024 | XLSI: A New Xception and Log Polar Transform Based Approach for Scene Text Script Identification
Ayush Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Apostolos Antonacopoulos, Michael Blumenstein |
ICPR (19) | 5 |
| 2024 | Dual-Personalizing Adapter for Federated Foundation ModelsabstractRecently, foundation models, particularly large language models (LLMs), have demonstrated an impressive ability to adapt to various tasks by fine-tuning diverse instruction data. Notably, federated foundation models (FedFM) emerge as a privacy preservation method to fine-tune models collaboratively under federated learning (FL) settings by leveraging many distributed datasets with non-IID data. To alleviate communication and computation overhead, parameter-efficient methods are introduced for efficiency, and some research adapted personalization methods to FedFM for better user preferences alignment. However, a critical gap in existing research is the neglect of test-time distribution shifts in real-world applications, and conventional methods for test-time distribution shifts in personalized FL are less effective for FedFM due to their failure to adapt to complex distribution shift scenarios and the requirement to train all parameters. To bridge this gap, we refine the setting in FedFM, termed test-time personalization, which aims to learn personalized federated foundation models on clients while effectively handling test-time distribution shifts simultaneously. To address challenges in this setting, we explore a simple yet effective solution, a Federated Dual-Personalizing Adapter (FedDPA) architecture. By co-working with a foundation model, a global adapter and a local adapter jointly tackle the test-time distribution shifts and client-specific personalization. Additionally, we introduce an instance-wise dynamic weighting mechanism that dynamically integrates the global and local adapters for each test instance during inference, facilitating effective test-time personalization. The effectiveness of the proposed method has been evaluated on benchmark datasets across different NLP tasks. Yiyuan Yang, Guodong Long, Tao Shen 0001, Jing Jiang 0002, Michael Blumenstein |
NeurIPS | 5 |
| 2024 | A new deep CNN for 3D text localization in the wild through shadow removal
Palaiahnakote Shivakumara, Ayan Banerjee 0002, Lokesh Nandanwar, Umapada Pal 0001, Apostolos Antonacopoulos, Tong Lu 0002, Michael Blumenstein |
Comput. Vis. Image Underst. | 7 |
| 2024 | MASSNet: Multiscale Attention for Single-Stage Ship Instance SegmentationabstractMaritime surveillance is essential in understanding, predicting, and ensuring the security of events in the complex marine environment. In this context, we have used instance segmentation techniques, which provides an accurate and efficient method for segmenting objects (Ships) in maritime surveillance. However, prevalent two-stage algorithms have limitations, including complex models, extended training time, and high memory consumption, making them impractical for real-world application. To address these challenges, we present an efficient solution called Multiscale Attention for Single-Stage Ship Instance Segmentation, or MASSNet. MASSNet uses the power of attention mechanisms to enhance multiscale feature extraction across various dimensions, resulting in a more refined and contextually-aware representation. This approach significantly improves segmentation accuracy and overall performance. In our extensive experiments, we evaluate the effectiveness of MASSNet on three challenging datasets: MariboatS, ShipInsSeg, and ShipSG. Our proposed model achieves mask Average Precision (mask AP) scores of 55.4%, 55.5%, and 74.1% on MariboatS, ShipInsSeg, and ShipSG datasets and outperforms other models such as YOLACT, SOLO and SOLOv2 architectures. MASSNet offers a robust and efficient solution for Ship Instance Segmentation, making significant improvement in the capabilities of maritime surveillance. Rabi Sharma, Chin-Teng Lin, Michael Blumenstein |
Neurocomputing | 4 |
| 2024 | A Locally Weighted Linear Regression-Based Approach for Arbitrary Moving Shaky and Nonshaky Video ClassificationabstractClassification and identification of objects are complex and challenging in pattern recognition and artificial intelligence if a shaky and nonshaky camera captures the videos at different distances during the day and nighttime. This work presents a model for classifying a given video as a static, uniform, or arbitrarily moving videos so that the complexity of the problem can be reduced. To avoid the threat of different distances between the objects and the camera, the proposed work introduces new steps for estimating the depth of the objects in the video frames. We explore locally weighted linear regression for feature extraction from depth information based on the notion that the regression line fits almost all the points for uniformity and does not fit for arbitrary moving. The extracted features are fed to a random forest classifier to classify static, uniform, or arbitrary moving video. The results on a large dataset, which includes videos captured day and night, show that the proposed method successfully classifies static, uniform and arbitrary videos with 0.86, 1.00 and 0.67 F-measures, respectively. Overall, our method obtains 87% accuracy for classification of static, uniform and arbitrary video, which is superior to the state-of-the-art methods. Arnab Halder, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Palash Ghosal |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2024 | TIC: text-guided image colorization using conditional generative modelabstractAbstract Image colorization is a well-known problem in computer vision. However, due to the ill-posed nature of the task, image colorization is inherently challenging. Though several attempts have been made by researchers to make the colorization pipeline automatic, these processes often produce unrealistic results due to a lack of conditioning. In this work, we attempt to integrate textual descriptions as an auxiliary condition, along with the grayscale image that is to be colorized, to improve the fidelity of the colorization process. To the best of our knowledge, this is one of the first attempts to incorporate textual conditioning in the colorization pipeline. To do so, a novel deep network has been proposed that takes two inputs (the grayscale image and the respective encoded text description) and tries to predict the relevant color gamut. As the respective textual descriptions contain color information of the objects present in the scene, the text encoding helps to improve the overall quality of the predicted colors. The proposed model has been evaluated using different metrics like SSIM, PSNR, LPISPS and achieved scores of 0.917, 23.27,0.223, respectively. These quantitative metrics have shown that the proposed method outperforms the SOTA techniques in most of the cases. Subhankar Ghosh, Prasun Roy, Saumik Bhattacharya, Umapada Pal 0001, Michael Blumenstein |
Multim. Tools Appl. | 5 |
| 2024 | A Conformable Moments-Based Deep Learning System for Forged Handwriting DetectionabstractDetecting forged handwriting is important in a wide variety of machine learning applications, and it is challenging when the input images are degraded with noise and blur. This article presents a new model based on conformable moments (CMs) and deep ensemble neural networks (DENNs) for forged handwriting detection in noisy and blurry environments. Since CMs involve fractional calculus with the ability to model nonlinearities and geometrical moments as well as preserving spatial relationships between pixels, fine details in images are preserved. This motivates us to introduce a DENN classifier, which integrates stenographic kernels and spatial features to classify input images as normal (original, clean images), altered (handwriting changed through copy-paste and insertion operations), noisy (added noise to original image), blurred (added blur to original image), altered-noise (noise is added to the altered image), and altered-blurred (blur is added to the altered image). To evaluate our model, we use a newly introduced dataset, which comprises handwritten words altered at the character level, as well as several standard datasets, namely ACPR 2019, ICPR 2018-FDC, and the IMEI dataset. The first two of these datasets include handwriting samples that are altered at the character and word levels, and the third dataset comprises forged International Mobile Equipment Identity (IMEI) numbers. Experimental results demonstrate that the proposed method outperforms the existing methods in terms of classification rate. Lokesh Nandanwar, Palaiahnakote Shivakumara, Hamid Abdullah Jalab, Rabha W. Ibrahim, Ramachandra Raghavendra, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2023 | Improving Open-Domain Answer Sentence Selection by Distributed Clients with Privacy Preservation
Weikuan Wang, Tao Shen 0001, Michael Blumenstein, Guodong Long |
ADMA (5) | 3 |
| 2023 | Classification of aesthetic natural scene images using statistical and semantic features
Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein, Josep Lladós 0001 |
Multim. Tools Appl. | 5 |
| 2023 | Writer age estimation through handwriting
Zhiheng Huang, Palaiahnakote Shivakumara, Maryam Asadzadeh Kaljahi, Ahlad Kumar, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
Multim. Tools Appl. | 7 |
| 2023 | A new robust approach for altered handwritten text detection
Gayatri Patil, Palaiahnakote Shivakumara, Shivanand S. Gornale, Umapada Pal 0001, Michael Blumenstein |
Multim. Tools Appl. | 5 |
| 2023 | A New Few-Shot Learning-Based Model for Prohibited Objects Detection in Cluttered Baggage X-Ray Images Through Edge Detection and Reverse ValidationabstractDetecting prohibited items via X-ray screening at airports and sensitive venues is essential for preventing smuggling and breaches of security. The difficulty in prohibited items inspection lies in accurately detecting prohibited items in complex X-ray images and limited access to X-ray images containing prohibited items. Few-shot detection aims at learning with limited examples and assigning a category label to each object. However, most few-shot learning methods do not focus on the edge information of the occluded object in X-ray images, which is crucial for the model to detect prohibited items in the X-ray images. In this paper, we presents a method (RVViT) for few-shot prohibited items detection tasks which fully acknowledges the significance of X-ray penetrability and increases the stability of few-shot learning model. Specifically, a Transformer encoder is firstly adopted for generating high-level semantic features that contain global information. At the same time, an edge detection module is devised for enhancing the edge information of prohibited items. Moreover, to further improve the stability of the few-shot learning model and ensure prototype consistency between the support and query samples, a reverse validation strategy is proposed to assist training. Extensive experiments demonstrate our method outperforms state-of-the-art approaches in terms of detection with a small number of samples. Shujing Lyu, Palaiahnakote Shivakumara, Michael Blumenstein, Yue Lu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2022 | TIPS: Text-Induced Pose Synthesis
Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya, Umapada Pal 0001, Michael Blumenstein |
ECCV (38) | 5 |
| 2022 | Omni-Scale CNNs: a simple and effective kernel size configuration for time series classification
Wensi Tang, Guodong Long, Lu Liu 0019, Tianyi Zhou 0001, Michael Blumenstein, Jing Jiang 0002 |
ICLR | 5 |
| 2022 | Scene Aware Person Image Generation through Global Contextual ConditioningabstractPerson image generation is an intriguing yet challenging problem. However, this task becomes even more difficult under constrained situations. In this work, we propose a novel pipeline to generate and insert contextually relevant person images into an existing scene while preserving the global semantics. More specifically, we aim to insert a person such that the location, pose, and scale of the person being inserted blends in with the existing persons in the scene. Our method uses three individual networks in a sequential pipeline. At first, we predict the potential location and the skeletal structure of the new person by conditioning a Wasserstein Generative Adversarial Network (WGAN) on the existing human skeletons present in the scene. Next, the predicted skeleton is refined through a shallow linear network to achieve higher structural accuracy in the generated image. Finally, the target image is generated from the refined skeleton using another generative network conditioned on a given image of the target person. In our experiments, we achieve high-resolution photo-realistic generation results while preserving the general context of the scene. We conclude our paper with multiple qualitative and quantitative benchmarks on the results. Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya, Umapada Pal 0001, Michael Blumenstein |
ICPR | 5 |
| 2022 | Detecting and mitigating poisoning attacks in federated learning using generative adversarial networksabstractSummary In the age of the Internet of Things (IoT), large numbers of sensors and edge devices are deployed in various application scenarios; Therefore, collaborative learning is widely used in IoT to implement crowd intelligence by inviting multiple participants to complete a training task. As a collaborative learning framework, federated learning is designed to preserve user data privacy, where participants jointly train a global model without uploading their private training data to a third party server. Nevertheless, federated learning is under the threat of poisoning attacks, where adversaries can upload malicious model updates to contaminate the global model. To detect and mitigate poisoning attacks in federated learning, we propose a poisoning defense mechanism, which uses generative adversarial networks to generate auditing data in the training procedure and removes adversaries by auditing their model accuracy. Experiments conducted on two well‐known datasets, MNIST and Fashion‐MNIST, suggest that federated learning is vulnerable to the poisoning attack, and the proposed defense method can detect and mitigate the poisoning attack. Ying Zhao 0011, Jiale Zhang 0001, Di Wu 0050, Michael Blumenstein, Shui Yu 0001 |
Concurr. Comput. Pract. Exp. | 5 |
| 2022 | A Knowledge Enforcement Network-Based Approach for Classifying a Photographer's ImagesabstractClassification of photos captured by different photographers is an important and challenging problem in knowledge-based and image processing. Monitoring and authenticating images uploaded on social media are essential, and verifying the source is one key piece of evidence. We present a novel framework for classifying photos of different photographers based on the combination of local features and deep learning models. The proposed work uses focused and defocused information in the input images to extract contextual information. The model estimates the weighted gradient and calculates entropy to strengthen context features. The focused and defocused information is fused to estimate cross-covariance and define a linear relationship between them. This relationship results in a feature matrix fed to Knowledge Enforcement Network (KEN) for obtaining representative features. Due to the strong discriminative ability of deep learning models, we employ the lightweight and accurate MobileNetV2. The output of KEN and MobileNetV2 is sent to a classifier for photographer classification. Experimental results of the proposed model on our dataset of 46 photographer classes (46234 images) and publicly available datasets of 41 photographer classes (218303 images) show that the method outperforms the existing techniques by 5%–10% on average. The dataset created for the experimental purpose will be made available upon publication. Palaiahnakote Shivakumara, Pinaki Nath Chowdhury, Umapada Pal 0001, David S. Doermann, Ramachandra Raghavendra, Tong Lu 0002, Michael Blumenstein |
Int. J. Pattern Recognit. Artif. Intell. | 7 |
| 2022 | New Deep Spatio-Structural Features of Handwritten Text Lines for Document Age ClassificationabstractDocument age estimation using handwritten text line images is useful for several pattern recognition and artificial intelligence applications such as forged signature verification, writer identification, gender identification, personality traits identification, and fraudulent document identification. This paper presents a novel method for document age classification at the text line level. For segmenting text lines from handwritten document images, the wavelet decomposition is used in a novel way. We explore multiple levels of wavelet decomposition, which introduce blur as the number of levels increases for detecting word components. The detected components are then used for a direction guided-driven growing approach with linearity, and nonlinearity criteria for segmenting text lines. For classification of text line images of different ages, inspired by the observation that, as the age of a document increases, the quality of its image degrades, the proposed method extracts the structural, contrast, and spatial features to study degradations at different wavelet decomposition levels. The specific advantages of DenseNet, namely, strong feature propagation, mitigation of the vanishing gradient problem, reuse of features, and the reduction of the number of parameters motivated us to use DenseNet121 along with a Multi-layer Perceptron (MLP) for the classification of text lines of different ages by feeding features and the original image as input. To demonstrate the efficacy of the proposed model, experiments were conducted on our own as well as standard datasets for both text line segmentation and document age classification. The results show that the proposed method outperforms the existing methods for text line segmentation in terms of precision, recall, F-measure, and document age classification in terms of average classification rate. Palaiahnakote Shivakumara, Alloy Das, Raghunandan K. Srinivas, Umapada Pal 0001, Michael Blumenstein |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2022 | Local Resultant Gradient Vector Difference and Inpainting for 3D Text Detection in the WildabstractThree-dimensional (3D) text appearing in natural scene images is common due to 3D cameras and the capture of text from different angles, which presents new problems for text detection. This is because of the presence of depth information, shadows, and decorative characters in the images. In this work, we consider those images where 3D text appears with depth, as well as shadow information for text detection. We propose a novel method based on local resultant gradient vector difference (LRGVD), inpainting and a deep learning model for detecting 3D as well as two-dimensional (2D) texts in natural scene images. The boundary of components that are invariant to the above challenges is detected by exploring LRGVD. The LRGVD uses gradient magnitude and direction in a novel way for detecting the boundary of the components. Further, we propose an inpainting method in a new way for restoring the character background information using boundaries. For a given region and the input image, the inpainting method divides the whole image into planes and then propagates the values in the planes into the missing region based on posterior probabilities and neighboring information. This results in text regions with false positives. Then, the differential binarization network (DB-Net) is proposed for detecting text irrespective of orientation, background, 3D or 2D, etc. Experiments conducted on our 3D text images and standard datasets of natural scene text images, namely ICDAR 2019 MLT, ICDAR 2019 ArT, DAST1500, Total-Text and SCUT-CTW1500, show that the proposed method is effective in detecting 3D and 2D texts in the images. Dajian Zhong, Palaiahnakote Shivakumara, Lokesh Nandanwar, Umapada Pal 0001, Michael Blumenstein, Yue Lu 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2022 | A new method for detection and prediction of occluded text in natural scene images
Ayush Mittal, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
Signal Process. Image Commun. | 5 |
| 2021 | Text-line-up: Don't Worry About the Caret
Chandranath Adak, Bidyut B. Chaudhuri, Chin-Teng Lin, Michael Blumenstein |
ICDAR (3) | 4 |
| 2021 | Instance-based error correction for short reads of disease-associated genesabstractBACKGROUND: Genomic reads from sequencing platforms contain random errors. Global correction algorithms have been developed, aiming to rectify all possible errors in the reads using generic genome-wide patterns. However, the non-uniform sequencing depths hinder the global approach to conduct effective error removal. As some genes may get under-corrected or over-corrected by the global approach, we conduct instance-based error correction for short reads of disease-associated genes or pathways. The paramount requirement is to ensure the relevant reads, instead of the whole genome, are error-free to provide significant benefits for single-nucleotide polymorphism (SNP) or variant calling studies on the specific genes. RESULTS: To rectify possible errors in the short reads of disease-associated genes, our novel idea is to exploit local sequence features and statistics directly related to these genes. Extensive experiments are conducted in comparison with state-of-the-art methods on both simulated and real datasets of lung cancer associated genes (including single-end and paired-end reads). The results demonstrated the superiority of our method with the best performance on precision, recall and gain rate, as well as on sequence assembly results (e.g., N50, the length of contig and contig quality). CONCLUSION: Instance-based strategy makes it possible to explore fine-grained patterns focusing on specific genes, providing high precision error correction and convincing gene sequence assembly. SNP case studies show that errors occurring at some traditional SNP areas can be accurately corrected, providing high precision and sensitivity for investigations on disease-causing point mutations. Xuan Zhang 0010, Yuansheng Liu, Michael Blumenstein, Gyorgy Hutvagner, Jinyan Li 0001 |
BMC Bioinform. | 4 |
| 2021 | DCT-phase statistics for forged IMEI numbers and air ticket detection
Lokesh Nandanwar, Palaiahnakote Shivakumara, Swati Kanchan, V. Basavaraja, D. S. Guru, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
Expert Syst. Appl. | 8 |
| 2021 | Improved Ring Radius Transform-Based Reconstruction for Video Character RecognitionabstractCharacter shape reconstruction in video is challenging due to low contrast, complex backgrounds and arbitrary orientation of characters. This work proposes an Improved Ring Radius Transform (IRRT) for reconstructing impaired characters through medial axis prediction. At first, the technique proposes a novel idea based on the Tangent Vector (TV) concept that identifies each actual pair of end pixels caused by gaps in impaired character components. Next, the actual direction to predict medial axis pixels using IRRT for each pair of end pixels is proposed with a new normal vector concept. The process of prediction repeats iteratively to find all the medial axis pixels for every gap in question. Further, medial axis pixels with their radii are used to reconstruct the shapes of impaired characters. The proposed technique is tested on benchmark datasets consisting of video, natural scenes, objects and multi-lingual data to demonstrate that it reconstructs shapes well, even for heterogeneous data. Comparative studies with different binarization and character recognition methods show that the proposed technique is effective, useful and outperforms existing methods. Zhiheng Huang, Palaiahnakote Shivakumara, Tong Lu 0002, Umapada Pal 0001, Michael Blumenstein, Bhaarat Chetty, G. Hemantha Kumar 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2021 | A New Hybrid Method for Caption and Scene Text Classification in Action Video ImagesabstractAchieving a better recognition rate for text in action video images is challenging due to multiple types of text with unpredictable actions in the background. In this paper, we propose a new method for the classification of caption (which is edited text) and scene text (text that is a part of the video) in video images. This work considers five action classes, namely, Yoga, Concert, Teleshopping, Craft, and Recipes, where it is expected that both types of text play a vital role in understanding the video content. The proposed method introduces a new fusion criterion based on Discrete Cosine Transform (DCT) and Fourier coefficients to obtain the reconstructed images for caption and scene text. The fusion criterion involves computing the variances for coefficients of corresponding pixels of DCT and Fourier images, and the same variances are considered as the respective weights. This step results in Reconstructed image-1. Inspired by the special property of Chebyshev-Harmonic-Fourier-Moments (CHFM) that has the ability to reconstruct a redundancy-free image, we explore CHFM for obtaining the Reconstructed image-2. The reconstructed images along with the input image are passed to a Deep Convolutional Neural Network (DCNN) for classification of caption/scene text. Experimental results on five action classes and a comparative study with the existing methods demonstrate that the proposed method is effective. In addition, the recognition results of the before and after the classification obtained from different methods show that the recognition performance improves significantly after classification, compared to before classification. Lokesh Nandanwar, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2021 | A clustering solution for analyzing residential water consumption patterns
Md Shamsur Rahim, Khoi Anh Nguyen, Rodney A. Stewart, Tanvir Ahmed 0001, Damien Giurco, Michael Blumenstein |
Knowl. Based Syst. | 6 |
| 2021 | Deep learning detection of anomalous patterns from bus trajectories for traffic insight analysis
Xiaocai Zhang, Yi Zheng 0002, Zhixun Zhao, Yuansheng Liu, Michael Blumenstein, Jinyan Li 0001 |
Knowl. Based Syst. | 5 |
| 2021 | Robust Tensor Decomposition for Image Representation Based on Generalized CorrentropyabstractTraditional tensor decomposition methods, e.g., two dimensional principal component analysis and two dimensional singular value decomposition, that minimize mean square errors, are sensitive to outliers. To overcome this problem, in this paper we propose a new robust tensor decomposition method using generalized correntropy criterion (Corr-Tensor). A Lagrange multiplier method is used to effectively optimize the generalized correntropy objective function in an iterative manner. The Corr-Tensor can effectively improve the robustness of tensor decomposition with the existence of outliers without introducing any extra computational cost. Experimental results demonstrated that the proposed method significantly reduces the reconstruction error on face reconstruction and improves the accuracies on handwritten digit recognition and facial image clustering. Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein |
IEEE Trans. Image Process. | 4 |
| 2020 | A New Context-Based Method for Restoring Occluded Text in Natural Scene Images
Ayush Mittal, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein, Daniel P. Lopresti |
DAS | 5 |
| 2020 | ICFHR 2020 Competition on Short answer ASsessment and Thai student SIGnature and Name COMponents Recognition and Verification (SASIGCOM 2020)abstractThis paper describes the results of the competition on Short answer ASsessment and Thai student SIGnature and Name COMponents Recognition and Verification (SASIGCOM 2020) in conjunction with the 17th International Conference on Frontiers in Handwriting Recognition (ICFHR 2020). The competition was aimed to automate the evaluation process short answer-based examination and record the development and gain attention to such system. The proposed competition contains three elements which are short answer assessment (recognition and marking the answers to short-answer questions derived from examination papers), student name components (first and last names) and signature verification and recognition. Signatures and name components data were collected from 100 volunteers. For the Thai signature dataset, there are 30 genuine signatures, 12 skilled and 12 simple forgeries for each writer. With Thai name components dataset, there are 30 genuine and 12 skilfully forged name components for each writer. There are 104 exam papers in the short answer assessment dataset, 52 of which were written with cursive handwriting; the rest of 52 papers were written with printed handwriting. The exam papers contain ten questions, and the answers to the questions were designed to be a few words per question. Three teams from distinguished labs submitted their systems. For short answer assessment, word spotting task was also performed. This paper analysed the results produced by their algorithms using a performance measure and defines a way forward for this subject of research. Both the datasets, along with some of the accompanying ground truth/baseline mask will be made freely available for research purposes via the TC10/TC11. Abhijit Das 0001, Hemmaphan Suwanwiwat, Umapada Pal 0001, Michael Blumenstein |
ICFHR | 4 |
| 2020 | Why Not? Tell us the Reason for Writer DissimilarityabstractWriter verification has drawn significant attention over the past few decades due to its extensive applications in forensics and biometrics. In traditional writer verification, handwriting similarity/dissimilarity analysis is mostly performed by extracting two feature vectors from two respective handwritten samples, followed by comparing them in relation to their similarity. In the state-of-the-art writer verification approaches, a distance metric is usually employed in terms of the similarity between two handwritten samples. If the distance between two handwritten samples is greater than a given threshold, then the samples are assumed to be written by two different writers, otherwise, they are considered to be due to the same writer. In this paper, for the very first time, we propose a model that generates English sentences to explain reasons for writer dissimilarity/similarity. First, our proposed model obtains features from handwritten images by employing a convolutional neural network, verifies the writer using a Siamese architecture, and generates English words using a recurrent neural network. Finally, these two networks are merged using an affine transformation to produce an explanatory sentence in support of writer similarity/dissimilarity. We evaluated our model on a handwritten numeral database of 100 writers and obtained promising results. Chandranath Adak, Bidyut B. Chaudhuri, Chin-Teng Lin, Michael Blumenstein |
IJCNN | 4 |
| 2020 | FACLSTM: ConvLSTM with focused attention for scene text recognition
Wenjing Jia, Xiangjian He, Michael Blumenstein, Shujing Lyu, Yue Lu 0001 |
Sci. China Inf. Sci. | 5 |
| 2020 | Rotation invariant angle-density based features for an ice image classification system
Shengkai Yue, Minglei Yuan, Tong Lu 0002, Palaiahnakote Shivakumara, Michael Blumenstein, G. Hemantha Kumar 0001 |
Expert Syst. Appl. | 5 |
| 2020 | A new augmentation-based method for text detection in night and day license plate images
Pinaki Nath Chowdhury, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
Multim. Tools Appl. | 5 |
| 2020 | Integrating joint feature selection into subspace learning: A formulation of 2DPCA for outliers robust feature selection
Muhammad Imran Razzak, Raghib M. Abu-Saris, Michael Blumenstein, Guandong Xu |
Neural Networks | 3 |
| 2020 | A new unified method for detecting text from marathon runners and sports players in video (PR-D-19-01078R2)
Sauradip Nag, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
Pattern Recognit. | 5 |
| 2020 | A robust matching pursuit algorithm using information theoretic learning
Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein |
Pattern Recognit. | 4 |
| 2020 | Bi-Level Error Correction for PacBio Long ReadsabstractThe latest sequencing technologies such as the Pacific Biosciences (PacBio) and Oxford Nanopore machines can generate long reads at the length of thousands of nucleic bases which is much longer than the reads at the length of hundreds generated by Illumina machines. However, these long reads are prone to much higher error rates, for example 15%, making downstream analysis and applications very difficult. Error correction is a process to improve the quality of sequencing data. Hybrid correction strategies have been recently proposed to combine Illumina reads of low error rates to fix sequencing errors in the noisy long reads with good performance. In this paper, we propose a new method named Bicolor, a bi-level framework of hybrid error correction for further improving the quality of PacBio long reads. At the first level, our method uses a de Bruijn graph-based error correction idea to search paths in pairs of solid -mers iteratively with an increasing length of -mer. At the second level, we combine the processed results under different parameters from the first level. In particular, a multiple sequence alignment algorithm is used to align those similar long reads, followed by a voting algorithm which determines the final base at each position of the reads. We compare the superior performance of Bicolor with three state-of-the-art methods on three real data sets. Results demonstrate that Bicolor always achieves the highest identity ratio. Bicolor also achieves a higher alignment ratio () and a higher number of aligned reads than the current methods on two data sets. On the third data set, our method is closely competitive to the current methods in terms of number of aligned reads and genome coverage. The C++ source codes of our algorithm are freely available at https://github.com/yuansliu/Bicolor. Yuansheng Liu, Chaowang Lan, Michael Blumenstein, Jinyan Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2020 | Intra-Variable Handwriting Inspection Reinforced With Idiosyncrasy AnalysisabstractIn this paper, we work on intra-variable handwriting, where the writing samples of an individual can vary significantly. Such within-writer variation throws a challenge for automatic writer inspection, where the state-of-the-art methods do not perform well. To deal with intra-variability, we analyze the idiosyncrasy in individual handwriting. We identify/verify the writer from highly idiosyncratic text-patches. Such patches are detected using a deep recurrent reinforcement learning-based architecture. An idiosyncratic score is assigned to every patch, which is predicted by employing deep regression analysis. For writer identification, we propose a deep neural architecture, which makes the final decision by the idiosyncratic score-induced weighted average of patch-based decisions. For writer verification, we propose two algorithms for patch-fed deep feature aggregation, which assist in authentication using a triplet network. The experiments were performed on two databases, where we obtained encouraging results. Chandranath Adak, Bidyut B. Chaudhuri, Chin-Teng Lin, Michael Blumenstein |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2019 | Drone-vs-Bird Detection Challenge at IEEE AVSS2019abstractThis paper presents the second edition of the “drone-vs-bird” detection challenge, launched within the activities of the 16-th IEEE International Conference on Advanced Video and Signal-based Surveillance (AVSS). The challenge's goal is to detect one or more drones appearing at some point in video sequences where birds may be also present, together with motion in background or foreground. Submitted algorithms should raise an alarm and provide a position estimate only when a drone is present, while not issuing alarms on birds, nor being confused by the rest of the scene. This paper reports on the challenge results on the 2019 dataset, which extends the first edition dataset provided by the SafeShore project with additional footage under different conditions. Angelo Coluccia, Nabin Sharma, Michael Blumenstein, Vasileios Magoulianitis, Dimitrios Ataloglou, Anastasios Dimou, Dimitrios Zarpalas, Petros Daras, Céline Craye, Salem Ardjoune, Alessio Fascista, David De la Iglesia, Miguel Méndez, Raquel Dosil, Iago González, Arne Schumann, Lars Wilko Sommer, Marian Ghenescu, Tomas Piatrik, Geert De Cubber, Mrunalini Nalamati, Ankit Kapoor |
AVSS | 4 |
| 2019 | Drone Detection in Long-Range Surveillance VideosabstractThe usage of small drones/UAVs has significantly increased recently. Consequently, there is a rising potential of small drones being misused for illegal activities such as terrorism, smuggling of drugs, etc. posing high-security risks. Hence, tracking and surveillance of drones are essential to prevent security breaches. The similarity in the appearance of small drone and birds in complex background makes it challenging to detect drones in surveillance videos. This paper addresses the challenge of detecting small drones in surveillance videos using popular and advanced deep learning-based object detection methods. Different CNN-based architectures such as ResNet-101 and Inception with Faster-RCNN, as well as Single Shot Detector (SSD) model was used for experiments. Due to sparse data available for experiments, pre-trained models were used while training the CNNs using transfer learning. Best results were obtained from experiments using Faster-RCNN with the base architecture of ResNet-101. Experimental analysis on different CNN architectures is presented in the paper, along with the visual analysis of the test dataset. Mrunalini Nalamati, Ankit Kapoor, Nabin Sharma, Michael Blumenstein |
AVSS | 5 |
| 2019 | Detecting Named Entities in Unstructured Bengali Manuscript ImagesabstractIn this paper, we undertake a task to find named entities directly from unstructured handwritten document images without any intermediate text/character recognition. Here, we do not receive any assistance from natural language processing. Therefore, it becomes more challenging to detect the named entities. We work on Bengali script which brings some additional hurdles due to its own unique script characteristics. Here, we propose a new deep neural network-based architecture to extract the latent features from a text image. The embedding is then fed to a BLSTM (Bidirectional Long Short-Term Memory) layer. After that, the attention mechanism is adapted to an approach for named entity detection. We perform experimentation on two publicly-available offline handwriting repositories containing 420 Bengali handwritten pages in total. The experimental outcome of our system is quite impressive as it attains 95.43% balanced accuracy on overall named entity detection. Chandranath Adak, Bidyut B. Chaudhuri, Chin-Teng Lin, Michael Blumenstein |
ICDAR | 4 |
| 2019 | Age Estimation using Disconnectedness Features in HandwritingabstractReal-time applications of handwriting analysis have increased drastically in the fields of forensic and information security because of accurate cues. One of such applications is human age estimation based on handwriting for the purpose of immigrant checking. In this paper, we have proposed a new method for age estimation using handwriting analysis using Hu invariant moments and disconnectedness features. To make the proposed method robust to both ruled and un-ruled documents, we propose to explore intersection point detection in Canny edge images of each input document, which results in text components. For each text component pair, we propose Hu invariant moments for extracting disconnectedness features, which in fact measure multi-shape components based on distance, shape and mutual position analysis of components. Furthermore, iterative k-means clustering is proposed for the classification of different age groups. Experimental results on our dataset and some standard datasets, namely, IAM and KHATT, show that the proposed method is effective and outperforms the state-of-the-art methods. V. Basavaraja, Palaiahnakote Shivakumara, D. S. Guru, Umapada Pal 0001, Tong Lu 0002, Michael Blumenstein |
ICDAR | 6 |
| 2019 | DeepText: Detecting Text from the Wild with Multi-ASPP-Assembled DeepLababstractIn this paper, we address the issue of scene text detection in the way of direct regression and successfully adapt an effective semantic segmentation model, DeepLab v3+ [1], for this application. In order to handle texts with arbitrary orientations and sizes and improve the recall of small texts, we propose to extract features of multiple scales by inserting multiple Atrous Spatial Pyramid Pooling (ASPP) layers to the DeepLab after the feature maps with different resolutions. Then, we set multiple auxiliary IoU losses at the decoding stage and make auxiliary connections from the intermediate encoding layers to the decoder to assist network training and enhance the discrimination ability of lower encoding layers. Experiments conducted on the benchmark scene text dataset ICDAR2015 demonstrate the superior performance of our proposed network, named as DeepText, over the state-of-the-art approaches. Wenjing Jia, Xiangjian He, Yue Lu 0001, Michael Blumenstein, Shujing Lyu |
ICDAR | 5 |
| 2019 | Temporal Self-Attention Network for Medical Concept EmbeddingabstractIn longitudinal electronic health records (EHRs), the event records of a patient are distributed over a long period of time and the temporal relations between the events reflect sufficient domain knowledge to benefit prediction tasks such as the rate of inpatient mortality. Medical concept embedding as a feature extraction method that transforms a set of medical concepts with a specific time stamp into a vector, which will be fed into a supervised learning algorithm. The quality of the embedding significantly determines the learning performance over the medical data. In this paper, we propose a medical concept embedding method based on applying a self-attention mechanism to represent each medical concept. We propose a novel attention mechanism which captures the contextual information and temporal relationships between medical concepts. A light-weight neural net, "Temporal Self-Attention Network (TeSAN)", is then proposed to learn medical concept embedding based solely on the proposed attention mechanism. To test the effectiveness of our proposed methods, we have conducted clustering and prediction tasks on two public EHRs datasets comparing TeSAN against five state-of-the-art embedding methods. The experimental results demonstrate that the proposed TeSAN model is superior to all the compared methods. To the best of our knowledge, this work is the first to exploit temporal self-attentive relations between medical events. Xueping Peng, Guodong Long, Tao Shen 0001, Sen Wang 0001, Jing Jiang 0002, Michael Blumenstein |
ICDM | 6 |
| 2019 | Robust Sparse Learning Based on Kernel Non-Second Order MinimizationabstractPartial occlusions in face images pose a great problem for most face recognition algorithms due to the fact that most of these algorithms mainly focus on solving a second order loss function, e.g., mean square error (MSE), which will magnify the effect from occlusion parts. In this paper, we proposed a kernel non-second order loss function for sparse representation (KNS-SR) to recognize or restore partially occluded facial images, which both take the advantages of the correntropy and the non-second order statistics measurement. The resulted framework is more accurate than the MSE-based ones in locating and eliminating outliers information. Experimental results from image reconstruction and recognition tasks on publicly available databases show that the proposed method achieves better performances compared with existing methods. Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein |
ICIP | 4 |
| 2019 | Kernel Mean P Power Error Loss for Robust Two-Dimensional Singular Value DecompositionabstractTraditional matrix-based dimensional reduction methods, e.g., two-dimensional principal component analysis (2DPCA) and two-dimensional singular value decomposition (2DSVD), minimize mean square errors (MSE), which is sensitive to outliers. To overcome this problem, in this paper we propose a new robust 2DSVD method based on the kernel mean p power error loss (KMPE-2DSVD). Different from the MSE and the correntropy based ones which are second order statistics based measurements, the KMPE-2DSVD is based on the non-second order statistics in the kernel space, and thus is more flexible in controlling the representation error. Experimental results show that the proposed method significantly improves the accuracy of facial image clustering. Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein |
ICIP | 4 |
| 2019 | Predicting Household Water Consumption Events: Towards a Personalised Recommender System to Encourage Water-conscious BehaviourabstractRecommender systems assist customers to make decisions; however, the modest adoption of digital technology in the water industry means no such system exists for household water users. Such a system for the water industry would suggest to consumers the most effective ways to conserve water based on their historical data from smart water meters. The advantage for water utilities in metropolitan areas is in managing demand, such as low pressure during peak hours or water shortages during drought. For customers, effective recommendations could save them money. This paper presents a novel vision of a recommender system prototype and discusses the benefits both for the consumers and the water utility companies. The success of this type of system would depend on the ability to anticipate the time of the next major water use so as to make useful, timely recommendations. Hence, the prototype is based on a long short-term memory (LSTM) neural network that predicts significant water consumption events (i.e., showers, baths, irrigation, etc.) for 83 households. The preliminary results show that LSTM is a useful method of prediction with an average root mean square error (RMSE) of 0.403. The analysis also provides indications of the scope of further research required for developing a commercially successful recommender system. Md Shamsur Rahim, Khoi Anh Nguyen, Rodney A. Stewart, Damien Giurco, Michael Blumenstein |
IJCNN | 5 |
| 2019 | Adversarial Action Data Augmentation for Similar Gesture Action RecognitionabstractHuman gestures are unique for recognizing and describing human actions, and video-based human action recognition techniques are effective solutions to varies real-world applications, such as surveillance, video indexing, and human-computer interaction. Most existing video human action recognition approaches either using handcraft features from the frames or deep learning models such as convolutional neural networks (CNN) and recurrent neural networks (RNN); however, they have mostly overlooked the similar gestures between different actions when processing the frames into the models. The classifiers suffer from similar features extracted from similar gestures, which are unable to classify the actions in the video streams. In this paper, we propose a novel framework with generative adversarial networks (GAN) to generate the data augmentation for similar gesture action recognition. The contribution of our work is tri-fold: 1) we proposed a novel action data augmentation framework (ADAF) to enlarge the differences between the actions with very similar gestures; 2) the framework can boost the classification performance either on similar gesture action pairs or the whole dataset; 3) experiments conducted on both KTH and UCF101 datasets show that our data augmentation framework boost the performance on both similar gestures actions as well as the whole dataset compared with baseline methods such as 2DCNN and 3DCNN. Di Wu 0050, Nabin Sharma, Shirui Pan, Guodong Long, Michael Blumenstein |
IJCNN | 6 |
| 2019 | Feature-Dependent Graph Convolutional Autoencoders with Adversarial Training MethodsabstractGraphs are ubiquitous for describing and modeling complicated data structures, and graph embedding is an effective solution to learn a mapping from a graph to a low-dimensional vector space while preserving relevant graph characteristics. Most existing graph embedding approaches either embed the topological information and node features separately or learn one regularized embedding with both sources of information, however, they mostly overlook the interdependency between structural characteristics and node features when processing the graph data into the models. Moreover, existing methods only reconstruct the structural characteristics, which are unable to fully leverage the interaction between the topology and the features associated with its nodes during the encoding-decoding procedure. To address the problem, we propose a framework using autoencoder for graph embedding (GED) and its variational version (VEGD). The contribution of our work is two-fold: 1) the proposed frameworks exploit a feature-dependent graph matrix (FGM) to naturally merge the structural characteristics and node features according to their interdependency; and 2) the Graph Convolutional Network (GCN) decoder of the proposed framework reconstructs both structural characteristics and node features, which naturally possesses the interaction between these two sources of information while learning the embedding. We conducted the experiments on three real-world graph datasets such as Cora, Citeseer and PubMed to evaluate our framework and algorithms, and the results outperform baseline methods on both link prediction and graph clustering tasks. Di Wu 0050, Ruiqi Hu, Yu Zheng 0013, Jing Jiang 0002, Nabin Sharma, Michael Blumenstein |
IJCNN | 6 |
| 2019 | Detection of Anomalous Traffic Patterns and Insight Analysis from Bus Trajectory Data
Xiaocai Zhang, Xuan Zhang 0010, Sunny Verma, Yuansheng Liu, Michael Blumenstein, Jinyan Li 0001 |
PRICAI (3) | 5 |
| 2019 | A comparative study of different texture features for document image retrieval
Fahimeh Alaei, Alireza Alaei, Umapada Pal 0001, Michael Blumenstein |
Expert Syst. Appl. | 4 |
| 2019 | An automatic zone detection system for safe landing of UAVs
Maryam Asadzadeh Kaljahi, Palaiahnakote Shivakumara, Mohd Yamani Idna Bin Idris, Mohammad Hossein Anisi, Tong Lu 0002, Michael Blumenstein, Noorzaily Mohamed Noor |
Expert Syst. Appl. | 6 |
| 2019 | A novel character segmentation-reconstruction approach for license plate recognition
Vijeta Khare, Palaiahnakote Shivakumara, Chee Seng Chan, Tong Lu 0002, Kim Meng Liang, Hock Woon Hon, Michael Blumenstein |
Expert Syst. Appl. | 7 |
| 2019 | A new image size reduction model for an efficient visual sensor networkabstractImage size reduction for energy-efficient transmission without losing quality is critical in Visual Sensor Networks (VSNs). The proposed method finds overlapping regions using camera locations, which eliminate unfocussed regions from the input images. The sharpness for the overlapped regions is estimated to find the Dominant Overlapping Region (DOR). The proposed model partitions further the DOR into sub-DORs according to capacity of the cameras. To reduce noise effects from the sub-DOR, we propose to perform a Median operation, which results in a Compressed Significant Region (CSR). For non-DOR, we obtain Sobel edges, which reduces the size of the images down to ambinary form. The CSR and Sobel edges of the non-DORs are sent by a VSN. Experimental results and a comparative study with the state-of-the-art methods shows that the proposed model outperforms the existing methods in terms of quality, energy consumption and network lifetime. Maryam Asadzadeh Kaljahi, Palaiahnakote Shivakumara, Mohd Yamani Idna Bin Idris, Mohammad Hossein Anisi, Michael Blumenstein |
J. Vis. Commun. Image Represent. | 5 |
| 2019 | A scene image classification technique for a ubiquitous visual surveillance system
Maryam Asadzadeh Kaljahi, Palaiahnakote Shivakumara, Mohammad Hossein Anisi, Mohd Yamani Idna Bin Idris, Michael Blumenstein, Muhammad Khurram Khan |
Multim. Tools Appl. | 5 |
| 2019 | Bag-of-visual-words for signature-based multi-script document retrieval
Ranju Mandal, Partha Pratim Roy 0001, Umapada Pal 0001, Michael Blumenstein |
Neural Comput. Appl. | 4 |
| 2019 | Chord Bunch Walks for Recognizing Naturally Self-Overlapped and Compound LeavesabstractEffectively describing and recognizing leaf shapes under arbitrary variations, particularly from a large database, remains an unsolved problem. In this research, we attempted a new strategy of describing leaf shapes by walking and measuring along a bunch of chords that pass through the shape. A novel chord bunch walks (CBW) descriptor is developed through the chord walking behavior that effectively integrates the shape image function over the walked chord to reflect both the contour features and the inner properties of the shape. For each contour point, the chord bunch groups multiple pairs of chords to build a hierarchical framework for a coarse-to-fine description that can effectively characterize not only the subtle differences among leaf margin patterns but also the interior part of the shape contour formed inside a self-overlapped or compound leaf. Instead of using optimal correspondence based matching, a Log-Min distance that encourages one-to-one correspondences is proposed for efficient and effective CBW matching. The proposed CBW shape analysis method is invariant to rotation, scaling, translation, and mirror transforms. Five experiments, including image retrieval of compound leaves, image retrieval of naturally self-overlapped leaves, and retrieval of mixed leaves on three large scale datasets, are conducted. The proposed method achieved large accuracy increases with low computational costs over the state-of-the-art benchmarks, which indicates the research potential along this direction. Bin Wang 0041, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein, John La Salle |
IEEE Trans. Image Process. | 4 |
| 2018 | Offline Bengali Writer Verification by PDF-CNN and Siamese NetabstractAutomated handwriting analysis is a popular area of research owing to the variation of writing patterns. In this research area, writer verification is one of the most challenging branches, having direct impact on biometrics and forensics. In this paper, we deal with offline writer verification on complex handwriting patterns. Therefore, we choose a relatively complex script, i.e., Indic Abugida script Bengali (or, Bangla) containing more than 250 compound characters. From a handwritten sample, the probability distribution functions (PDFs) of some handcrafted features are obtained and input to a convolutional neural network (CNN). For such a CNN architecture, we coin the term "PDFCNN", where handcrafted feature PDFs are hybridized with auto-derived CNN features. Such hybrid features are then fed into a Siamese neural network for writer verification. The experiments are performed on a Bengali offline handwritten dataset of 100 writers. Our system achieves encouraging results, which sometimes exceed the results of state-of-the-art techniques on writer verification. Chandranath Adak, Simone Marinai, Bidyut B. Chaudhuri, Michael Blumenstein |
DAS | 4 |
| 2018 | Evaluation of Gist Operator for Document Image RetrievalabstractAs digitised documents normally contain a large variety of structures, a page segmentation- and layout-free method for document image retrieval is preferable. In this research work, therefore, wavelet transform as a transform-based approach is initially used to provide different under-sampled images from the original image. Then, Gist operator, as a feature extraction technique, is employed to extract a set of global features from the original image as well as the sub-images obtained from the wavelet transform. Moreover, the column-wise variances of the values in each sub-image are computed and they are then concatenated to obtain another set of features. Considering each feature set, locality-sensitive hashing is employed to compute similarity distances between a query and the document images in the database. Finally, a classifier fusion technique using the mean function is taken into account to provide a document image retrieval result. The combination of these features and a clustering score fusion strategy provides higher document image retrieval accuracy. Two different databases of the document image are considered for experimentation. The results obtained from the experimental study are detailed and the results are encouraging. Fahimeh Alaei, Alireza Alaei, Umapada Pal 0001, Michael Blumenstein |
DAS | 4 |
| 2018 | A Study on Idiosyncratic Handwriting with Impact on Writer IdentificationabstractIn this paper, we study handwriting idiosyncrasy in terms of its structural eccentricity. In this study, our approach is to find idiosyncratic handwritten text components and model the idiosyncrasy analysis task as a machine learning problem supervised by human cognition. We employ the Inception network for this purpose. The experiments are performed on two publicly available databases and an in-house database of Bengali offline handwritten samples. On these samples, subjective opinion scores of handwriting idiosyncrasy are collected from handwriting experts. We have analyzed the handwriting idiosyncrasy on this corpus which comprises the perceptive ground-truth opinion. We also investigate the effect of idiosyncratic text on writer identification by using the SqueezeNet. The performance of our system is promising. Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein |
ICFHR | 3 |
| 2018 | Signature and Logo Detection using Deep CNN for Document Image RetrievalabstractSignature and logo as a query are important for content-based document image retrieval from a scanned document repository. This paper deals with signature and logo detection from a repository of scanned documents, which can be used for document retrieval using signature or logo information. A large intra-category variance among signature and logo samples poses challenges to traditional hand-crafted feature extraction-based approaches. Hence, the potential of deep learning-based object detectors namely, Faster R-CNN and YOLOv2 were examined for automatic detection of signatures and logos from scanned administrative documents. Four different network models namely ZF, VGG16, VGGM, and YOLOv2 were considered for analysis and identifying their potential in document image retrieval. The experiments were conducted on the publicly available "Tobacco-800" dataset. The proposed approach detects Signatures and Logos simultaneously. The results obtained from the experiments are promising and at par with the existing methods. Nabin Sharma, Ranju Mandal, Rabi Sharma, Umapada Pal 0001, Michael Blumenstein |
ICFHR | 5 |
| 2018 | ICFHR 2018 Competition on Thai Student Signatures and Name Components Recognition and Verification (TSNCRV2018)abstractThis paper summarises the results of the competition on the 1st Thai Student Signature and Name Components Recognition and Verification (TSNCRV 2018). It was organised in the context of the 16th International Conference on Frontiers in Handwriting Recognition (ICFHR 2018). The aim of this competition was to record the development and gain attention on Thai student signatures and name component recognition and verification. Two different types of datasets were used for the competition: the first dataset contains Thai student signatures and the second dataset contains Thai student name components. Signatures and name components from 100 volunteers each were included in the competition datasets. For Thai signature dataset, there are 30 genuine signatures, 12 skilled and 12 simple forgeries for each writer. For Thai name components, there are 30 genuine and 12 skilfully forged name components for each writer. For both the datasets the individuals were asked to write their name/signature in the given space on a white piece of paper for number of time (with a pause between each 10 samples). The skilled forgers were asked practice emitting the original signature for certain number of times till they fill skilled to forge. Five teams from distinguish labs submitted their systems. This paper analysed the results produced by these algorithms/systems using a performance measure and defined a way forward for this subject of research. Both the datasets along with some of the accompanying ground truth/baseline mask will be made freely available for research purposes via the TC10/TC11. Hemmaphan Suwanwiwat, Abhijit Das 0001, Umapada Pal 0001, Michael Blumenstein |
ICFHR | 4 |
| 2018 | Matching Pursuit Based on Kernel Non-Second Order MinimizationabstractThe orthogonal matching pursuit (OMP) is an important sparse approximation algorithm to recover sparse signals from compressed measurements. However, most MP algorithms are based on the mean square error(MSE) to minimize the recovery error, which is suboptimal when there are outliers. In this paper, we present a new robust OMP algorithm based on kernel non-second order statistics (KNS-OMP), which not only takes advantages of the outlier resistance ability of correntropy but also further extends the second order statistics based correntropy to a non-second order similarity measurement to improve its robustness. The resulted framework is more accurate than the second order ones in reducing the effect of outliers. Experimental results on synthetic and real data show that the proposed method achieves better performances compared with existing methods. Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein |
ICIP | 4 |
| 2018 | Fourier Transform based Features for Clean and Polluted Water Image ClassificationabstractWater image classification is challenging because water images of ocean or river share the same properties with images of polluted water such as fungus, waste and rubbish. In this paper, we present a method for classifying clean and polluted water images. The proposed method explores Fourier transform based features for extracting texture properties of clean and polluted water images. Fourier spectrum of each input image is divided into several sub-regions based on angle and spatial information. For each region over the spectrum, the proposed method extracts mean and variance features using intensity values, which results in a feature matrix. The feature matrix is then passed to an SVM classifier for the classification of clean and polluted water images. Experimental results on classes of clean and polluted water images show that the proposed method is effective. Furthermore, a comparative study with the state-of-the-art method shows that the proposed method outperforms the existing method in terms of classification rate, recall, precision and F-measure. Xuerong Wu, Palaiahnakote Shivakumara, Hualu Zhang, Tong Lu 0002, Umapada Pal 0001, Michael Blumenstein |
ICPR | 8 |
| 2018 | More Realistic and Efficient Face-Based Mobile Authentication using CNNsabstractIn this work, we propose a more realistic and efficient face-based mobile authentication technique using CNNs. This paper discusses and explores an inevitable problem of using face images for mobile authentication, taken from varying distances with a front/selfie camera of the mobile phone. Incidentally, once an individual comes towards a certain distance from the camera, the face images get large and appear over-sized. Simultaneously sharp features of some portions of the face, such as forehead, cheek, and chin are changed completely. As a result, the face features change and the impact increases exponentially once the individual crosses a certain distance and gradually approaches towards the front camera. This work proposes a solution (achieving better accuracy and facial features, whereby face images were cropped and aligned around its close bounding box) to mitigate the aforementioned identified gap. The work investigated different frontier face detection and recognition techniques to justify the proposed solution. Among all the employed methods evaluated, CNNs worked best. For a quantitative comparison of the proposed method, manually cropped face images/annotations of the face images along with their close boundary were prepared. In turn, we have developed a database considering the above-mentioned scenario for 40 individuals, which will be publicly available for academic research purposes. The experimental results achieved indicate a successful implementation of the proposed method and the performance of the proposed technique is also found to be superior in comparison to the existing state-of-the-art. Abhijit Das 0001, Abira Sengupta, Umapada Pal 0001, Michael Blumenstein |
IJCNN | 5 |
| 2018 | Cognitive Analysis for Reading and Writing of Bengali ConjunctsabstractIn this paper, we study the difficulties arising in reading and writing of Bengali conjunct characters by human-beings. Such difficulties appear when the human cognitive system faces certain obstructions in effortlessly reading/writing. In our computer-based investigation, we consider the reading/writing difficulty analysis task as a machine learning problem supervised by human perception. To this end, we employ two distinct models: (a) an auto-derived feature-based Inception network and (b) a hand-crafted feature-based SVM (Support Vector Machine). Two commonly used Bengali printed fonts and three contemporary handwritten databases are used for collecting subjective opinion scores from human readers/writers. On this corpus, which contains the perceptive ground-truth opinion of reading/writing complications, we have undertaken to conduct the experiments. The experimental results obtained on various types of conjunct characters are promising. Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein |
IJCNN | 3 |
| 2018 | Connectivity Based Method for Clustering Microbial Communities from Metagenomics Data of Water and Soil SamplesabstractUnderstanding microbial community structure of metagenomics water and soil samples is a key process in discovering functions and impact of microorganisms on human and animal health. Evolution of Next Generation Sequencing (NGS) technology has encouraged researchers to sequence large quantity of microbial data from environmental sources. Clustering marker gene sequences into Operational Taxonomic Units (OTU) is the most significant task in microbial community analysis. Several methods have been developed over the years to improve OTU picking strategies. However, building strongly connected OTUs is a major issue in majority of these methods. Herein we present ConClust, a novel method for clustering OTUs that is based on quantifying connectivity among the sequences. Experimental analysis on two synthetic datasets and two real world datasets from water and soil samples demonstrate that our method can mine robust OTUs. Our method can be highly benelicial to study functions of known and unknown microbes and analyze their positive and negative effect on the environment as well as human and animal health. Jessica Sharmin Rahman, Jinyan Li 0001, Juanying Xie, Shoshana Fogelman, Michael Blumenstein |
IJCNN | 5 |
| 2018 | Robust 2D Joint Sparse Principal Component Analysis With F-Norm Minimization For Sparse Modelling: 2D-RJSPCAabstractPrincipal component analysis (PCA) is widely used methods for dimensionality reduction and Lots of variants have been proposed to improve the robustness of algorithm, however, these methods suffer from the fact that PCA is linear combination which makes it difficult to interpret complex nonlinear data, and sensitive to outliers or cannot extract features consistently, i.e., collectively; PCA may still require measuring all input features. 2DPCA based on ℓ1-normhas been recently used for robust dimensionality reduction in the image domain but still sensitive to noise. In this paper, we introduce robust formation of 2DPCA by centering the data using the optimized mean for two-dimensional joint sparse as well as effectively combining the robustness of 2DPCA and the sparsity-inducing lasso regularization. Optimal mean helps to improve the robustness of joint sparse PCA further. The distance in spatial dimension is measure in F-norm and sum of different datapoint uses 1-norm. 2DR-JSPCA imposes joint sparse constraints on its objective function whereas additional plenty term help to deal with outliers efficiently. Both theoretical and empirical results on six publicly available benchmark datasets shows that Optimal mean 2DR-JSPCA provides better performance for dimensionality reduction as compare to non-sparse (2DPCA and 2DPCA-L1) and sparse (SPCA, JSPCA). Muhammad Imran Razzak, Raghib M. Abu-Saris, Michael Blumenstein, Guandong Xu |
IJCNN | 3 |
| 2018 | Person Head Detection in Multiple Scales Using Deep Convolutional Neural NetworksabstractPerson detection is an important problem in computer vision with many real-world applications. The detection of a person is still a challenging task due to variations in pose, occlusions and lighting conditions. The purpose of this study is to detect human heads in natural scenes acquired from a publicly available dataset of Hollywood movies. In this work, we have used state-of-the-art object detectors based on deep convolutional neural networks. These object detectors include region-based convolutional neural networks using region proposals for detections. Also, object detectors that detect objects in the single-shot by looking at the image only once for detections. We have used transfer learning for fine-tuning the network already trained on a massive amount of data. During the fine-tuning process, the models having high mean Average Precision (mAP) are used for evaluation of the test dataset. Experimental results show that Faster R-CNN [18] and SSD MultiBox [13] with VGG16 [21] perform better than YOLO [17] and also demonstrate significant improvements against several baseline approaches. Sultan Daud Khan, Nabin Sharma, Michael Blumenstein |
IJCNN | 4 |
| 2018 | CRISPR/Cas9 cleavage efficiency regression through boosting algorithms and Markov sequence profilingabstractMotivation: CRISPR/Cas9 system is a widely used genome editing tool. A prediction problem of great interests for this system is: how to select optimal single-guide RNAs (sgRNAs), such that its cleavage efficiency is high meanwhile the off-target effect is low. Results: This work proposed a two-step averaging method (TSAM) for the regression of cleavage efficiencies of a set of sgRNAs by averaging the predicted efficiency scores of a boosting algorithm and those by a support vector machine (SVM). We also proposed to use profiled Markov properties as novel features to capture the global characteristics of sgRNAs. These new features are combined with the outstanding features ranked by the boosting algorithm for the training of the SVM regressor. TSAM improved the mean Spearman correlation coefficiencies comparing with the state-of-the-art performance on benchmark datasets containing thousands of human, mouse and zebrafish sgRNAs. Our method can be also converted to make binary distinctions between efficient and inefficient sgRNAs with superior performance to the existing methods. The analysis reveals that highly efficient sgRNAs have lower melting temperature at the middle of the spacer, cut at 5'-end closer parts of the genome and contain more 'A' but less 'G' comparing with inefficient ones. Comprehensive further analysis also demonstrates that our tool can predict an sgRNA's cutting efficiency with consistently good performance no matter it is expressed from an U6 promoter in cells or from a T7 promoter in vitro. Availability and implementation: Online tool is available at http://www.aai-bioinfo.com/CRISPR/. Python and Matlab source codes are freely available at https://github.com/penn-hui/TSAM. Supplementary information: Supplementary data are available at Bioinformatics online. Yi Zheng 0002, Michael Blumenstein, Dacheng Tao, Jinyan Li 0001 |
Bioinform. | 3 |
| 2017 | Drone-vs-Bird detection challenge at IEEE AVSS2017abstractSmall drones are a rising threat due to their possible misuse for illegal activities, in particular smuggling and terrorism. The project SafeShore, funded by the European Commission under the Horizon 2020 program, has launched the “drone-vs-bird detection challenge” to address one of the many technical issues arising in this context. The goal is to detect a drone appearing at some point in a video where birds may be also present: the algorithm should raise an alarm and provide a position estimate only when a drone is present, while not issuing alarms on birds. This paper reports on the challenge proposal, evaluation, and results. Angelo Coluccia, Marian Ghenescu, Tomas Piatrik, Geert De Cubber, Arne Schumann, Lars Wilko Sommer, Johannes Klatte, Tobias Schuchert, Jürgen Beyerer, Mohammad Farhadi, Ruhallah Amandi, Cemal Aker, Sinan Kalkan, Nabin Sharma, Sultan Daud Khan, Khan Makkah, Michael Blumenstein |
AVSS | 18 |
| 2017 | A study on detecting drones using deep convolutional neural networksabstractThe object detection is a challenging problem in computer vision with various potential real-world applications. The objective of this study is to evaluate the deep learning based object detection techniques for detecting drones. In this paper, we have conducted experiments with different Convolutional Neural Network (CNN) based network architectures namely Zeiler and Fergus (ZF), Visual Geometry Group (VGG16) etc. Due to sparse data available for training, networks are trained with pre-trained models using transfer learning. The snapshot of trained models is saved at regular interval during training. The best models having high mean Average Precision (mAP) for each network architecture are used for evaluation on the test dataset. The experimental results show that VGG16 with Faster R-CNN perform better than other architectures on the training dataset. Visual analysis of the test dataset is also presented. Sultan Daud Khan, Nabin Sharma, Michael Blumenstein |
AVSS | 4 |
| 2017 | Can Walking and Measuring Along Chord Bunches Better Describe Leaf Shapes?abstractEffectively describing and recognizing leaf shapes under arbitrary deformations, particularly from a large database, remains an unsolved problem. In this research, we attempted a new strategy of describing shape by walking along a bunch of chords that pass through the shape to measure the regions trespassed. A novel chord bunch walks (CBW) descriptor is developed through the chord walking that effectively integrates the shape image function over the walked chord to reflect the contour features and the inner properties of the shape. For each contour point, the chord bunch groups multiple pairs of chord walks to build a hierarchical framework for a coarse-to-fine description. The proposed CBW descriptor is invariant to rotation, scaling, translation, and mirror transforms. Instead of using the expensive optimal correspondence based matching, an improved Hausdorff distance encoded correspondence information is proposed for efficient yet effective shape matching. In experimental studies, the proposed method obtained substantially higher accuracies with low computational cost over the benchmarks, which indicates the research potential along this direction. Bin Wang 0041, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein, John La Salle |
CVPR | 4 |
| 2017 | A decision-level fusion strategy for multimodal ocular biometric in visible spectrum based on posterior probabilityabstractIn this work, we propose a posterior probability-based decision-level fusion strategy for multimodal ocular biometric in the visible spectrum employing iris, sclera and peri-ocular trait. To best of our knowledge this is the first attempt to design a multimodal ocular biometrics using all three ocular traits. Employing all these traits in combination can help to increase the reliability and universality of the system. For instance in some scenarios, the sclera and iris can be highly occluded or for completely closed eyes scenario, the peri-ocular trait can be relied on for the decision. The proposed system is constituted of three independent traits and their combinations. The classification output of the trait which produces highest posterior probability is to consider as the final decision. An appreciable reliability and universal applicability of ocular trait are achieved in experiments conducted employing the proposed scheme. Abhijit Das 0001, Umapada Pal 0001, Miguel A. Ferrer, Michael Blumenstein |
IJCB | 4 |
| 2017 | SSERBC 2017: Sclera segmentation and eye recognition benchmarking competitionabstractThis paper summarises the results of the Sclera Segmentation and Eye Recognition Benchmarking Competition (SSERBC 2017). It was organised in the context of the International Joint Conference on Biometrics (IJCB 2017). The aim of this competition was to record the recent developments in sclera segmentation and eye recognition in the visible spectrum (using iris, sclera and peri-ocular, and their fusion), and also to gain the attention of researchers on this subject. In this regard, we have used the Multi-Angle Sclera Dataset (MASD version 1). It is comprised of2624 images taken from both the eyes of 82 identities. Therefore, it consists of images of 164 (82×2) eyes. A manual segmentation mask of these images was created to baseline both tasks. Precision and recall based statistical measures were employed to evaluate the effectiveness of the segmentation and the ranks of the segmentation task. Recognition accuracy measure has been employed to measure the recognition task. Manually segmented sclera, iris and peri-ocular regions were used in the recognition task. Sixteen teams registered for the competition, and among them, six teams submitted their algorithms or systems for the segmentation task and two of them submitted their recognition algorithm or systems. The results produced by these algorithms or systems reflect current developments in the literature of sclera segmentation and eye recognition, employing cutting edge techniques. The MASD version 1 dataset with some of the ground truth will be freely available for research purposes. The success of the competition also demonstrates the recent interests of researchers from academia as well as industry on this subject. Abhijit Das 0001, Umapada Pal 0001, Miguel A. Ferrer, Michael Blumenstein, Dejan Stepec, Peter Rot, Ziga Emersic, Peter Peer, Vitomir Struc, S. V. Aruna Kumar, B. S. Harish |
IJCB | 4 |
| 2017 | Linking face images captured from the optical phenomenon in the wild for forensic scienceabstractThis paper discusses the possibility of use of some challenging face images scenario captured from optical phenomenon in the wild for forensic purpose towards individual identification. Occluded and under cover face images in surveillance scenario can be collected from its reflection on a surrounding glass or on a smooth wall that is under the coverage of the surveillance camera and such scenario of face images can be linked for forensic purposes. Another similar scenario that can also be used for forensic is the face images of an individual standing behind a transparent glass wall. To investigate the capability of these images for personal identification this study is conducted. This work investigated different types of features employed in the literature to establish individual identification by such degraded face images. Among them, local region based featured worked best. To achieve higher accuracy and better facial features face image were cropped manually along its close bounding box and noise removal was performed (reflection, etc.). In order to experiment we have developed a database considering the above mentioned scenario, which will be publicly available for academic research. Initial investigation substantiates the possibility of using such face images for forensic purpose. Abhijit Das 0001, Abira Sengupta, Miguel A. Ferrer, Umapada Pal 0001, Michael Blumenstein |
IJCB | 5 |
| 2017 | Legibility and Aesthetic Analysis of HandwritingabstractThis paper deals with computer-based cognitive analysis towards legibility and aesthetics of a handwritten document. The legible text creates a human perception that the writing can be read effortlessly because of its orthographic clarity. The aesthetic property relates to the beautiful appearance of a handwritten document. In this study, we deal with these properties on offline Bengali handwriting. We formulate both legibility and aesthetic analysis tasks as machine learning problems supervised by the human cognitive system. We employ automatically derived feature-based recurrent neural networks to investigate writing legibility. For aesthetics evaluation, we employ hand-crafted feature-based support vector machines (SVMs). We have collected contemporary Bengali handwritings, on which the subjective legibility and aesthetic scores are provided by human readers. On this corpus containing legibility and aesthetic ground-truth information, we executed our experiments. The experimental results obtained on various handwritings are encouraging. Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein |
ICDAR | 3 |
| 2017 | Fourier-Residual for Printer IdentificationabstractPrinter identification is challenging due to advanced software technologies in the field of forgery detection. This paper presents a new idea of using the Fourier transform residual for the identification of documents printed by different printers. The proposed approach first convolves a Laplacian mask with a Fourier transform in the frequency domain to smoothen the edges. Next, we apply an inverse Fourier transform to reconstruct images from smoothed information (RFL). Similarly, the proposed approach reconstructs images using gray information of the input image (RFG). Then the residual is calculated by subtracting RFG from RFL. The set of statistical features, texture and spatial features are extracted from residual images for printer identification. Experimental results with the existing method on our dataset and a standard dataset show that the proposed approach outperforms the existing approach on both the datasets in terms of classification rate, recall, precision and F-measure. Palaiahnakote Shivakumara, Tong Lu 0002, M. Basavanna, Umapada Pal 0001, Michael Blumenstein |
ICDAR | 6 |
| 2017 | Deep Learning Based Face Recognition with Sparse Representation Classification
Eric-Juwei Cheng, Mukesh Prasad, Deepak Puthal, Nabin Sharma, Om Kumar Prasad, Po-Hao Chin, Chin-Teng Lin, Michael Blumenstein |
ICONIP (3) | 8 |
| 2017 | Impact of struck-out text on writer identificationabstractThe presence of struck-out text in handwritten manuscripts may affect the accuracy of automated writer identification. This paper presents a study on such effects of struck-out text. Here we consider offline English and Bengali handwritten document images. At first, the struck-out texts are detected using a hybrid classifier of a CNN (Convolutional Neural Network) and an SVM (Support Vector Machine). Then the writer identification process is activated on normal and struck-out text separately, to ascertain the impact of struck-out texts. For writer identification, we use two methods: (a) a hand-crafted feature-based SVM classifier, and (b) CNN-extracted auto-derived features with a recurrent neural model. For the experimental analysis, we have generated a database from 100 English and 100 Bengali writers. The performance of our system is very encouraging. Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein |
IJCNN | 3 |
| 2017 | Recent advances in video-based human action recognition using deep learning: A reviewabstractVideo-based human action recognition has become one of the most popular research areas in the field of computer vision and pattern recognition in recent years. It has a wide variety of applications such as surveillance, robotics, health care, video searching and human-computer interaction. There are many challenges involved in human action recognition in videos, such as cluttered backgrounds, occlusions, viewpoint variation, execution rate, and camera motion. A large number of techniques have been proposed to address the challenges over the decades. Three different types of datasets namely, single viewpoint, multiple viewpoint and RGB-depth videos, are used for research. This paper presents a review of various state-of-the-art deep learning-based techniques proposed for human action recognition on the three types of datasets. In light of the growing popularity and the recent developments in video-based human action recognition, this review imparts details of current trends and potential directions for future work to assist researchers. Di Wu 0050, Nabin Sharma, Michael Blumenstein |
IJCNN | 3 |
| 2017 | Arbitrarily-oriented multi-lingual text detection in video
Vijeta Khare, Palaiahnakote Shivakumara, Raveendran Paramesran, Michael Blumenstein |
Multim. Tools Appl. | 4 |
| 2017 | Fractals based multi-oriented text detection system for recognition in mobile video images
Palaiahnakote Shivakumara, Liang Wu 0009, Tong Lu 0002, Chew Lim Tan, Michael Blumenstein, Basavaraj S. Anami |
Pattern Recognit. | 5 |
| 2017 | An Efficient Signature Verification Method Based on an Interval Symbolic Representation and a Fuzzy Similarity MeasureabstractIn this paper, an efficient offline signature verification method based on an interval symbolic representation and a fuzzy similarity measure is proposed. In the feature extraction step, a set of local binary pattern-based features is computed from both the signature image and its under-sampled bitmap. Interval-valued symbolic data is then created for each feature in every signature class. As a result, a signature model composed of a set of interval values (corresponding to the number of features) is obtained for each individual's handwritten signature class. A novel fuzzy similarity measure is further proposed to compute the similarity between a test sample signature and the corresponding interval-valued symbolic model for the verification of the test sample. To evaluate the proposed verification approach, a benchmark offline English signature data set (GPDS-300) and a large data set (BHSig260) composed of Bangla and Hindi offline signatures were used. A comparison of our results with some recent signature verification methods available in the literature was provided in terms of average error rate and we noted that the proposed method always outperforms when the number of training samples is eight or more. Alireza Alaei, Srikanta Pal, Umapada Pal 0001, Michael Blumenstein |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2016 | Fast and efficent multimodal eye biometrics using projective dictionary pair learningabstractThis work proposes a projective pairwise dictionary learning-based approach for fast and efficient multimodal eye biometrics. The work uses a faster Projective pairwise Discriminative Dictionary Learning (DL) in contrast to the traditional DL which uses synthesis DL. Projective Pairwise Discriminative Dictionary (PPDD) uses a synthesis dictionary and an analysis dictionary jointly to achieve the goal of pattern representation and discrimination. As the PPDD process of DL is in contrast to the use of l0or l1-norm sparsity constraints on the representation coefficients adopted in most traditional DL, it works faster than other DL. Moreover, the blending of synthesis dictionary and an analysis dictionary also enhance the feature representation of the complex eye patterns. We employed the combination of sclera and iris traits to establish multimodal biometrics. The experimental study and analysis conducted fulfill the hypothesis we considered. In this work we employed a part of the UBIRIS version 1 dataset to conduct the experiments. Abhijit Das 0001, Prabir Mondal, Umapada Pal 0001, Miguel A. Ferrer, Michael Blumenstein |
CEC | 5 |
| 2016 | Named Entity Recognition from Unstructured Handwritten Document ImagesabstractNamed entity recognition is an important topic in the field of natural language processing, whereas in document image processing, such recognition is quite challenging without employing any linguistic knowledge. In this paper we propose an approach to detect named entities (NEs) directly from offline handwritten unstructured document images without explicit character/word recognition, and with very little aid from natural language and script rules. At the preprocessing stage, the document image is binarized, and then the text is segmented into words. The slant/skew/baseline corrections of the words are also performed. After preprocessing, the words are sent for NE recognition. We analyze the structural and positional characteristics of NEs and extract some relevant features from the word image. Then the BLSTM neural network is used for NE recognition. Our system also contains a post-processing stage to reduce the true NE rejection rate. The proposed approach produces encouraging results on both historical and modern document images, including those from an Australian archive, which are reported here for the very first time. Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein |
DAS | 3 |
| 2016 | Document Image Quality Assessment Based on Texture Similarity IndexabstractIn this paper, a full reference document image quality assessment (FR DIQA) method using texture features is proposed. Local binary patterns (LBP) as texture features are extracted at the local and global levels for each image. For each extracted LBP feature set, a similarity measure called the LBP similarity index (LBPSI) is computed. A weighting strategy is further proposed to improve the LBPSI obtained based on local LBP features. The LBPSIs computed for both local and global features are then combined to get the final LBPSI, which also provides the best performance for DIQA. To evaluate the proposed method, two different datasets were used. The first dataset is composed of document images, whereas the second one includes natural scene images. The mean human opinion scores (MHOS) were considered as ground truth for performance evaluation. The results obtained from the proposed LBPSI method indicate a significant improvement in automatically/accurately predicting image quality, especially on the document image-based dataset. Alireza Alaei, Donatello Conte, Michael Blumenstein, Romain Raveaux |
DAS | 3 |
| 2016 | Marginal Noise Reduction in Historical Handwritten Documents - A SurveyabstractThis paper presents a survey on different approaches for removing the marginal noise from document images, and anlaysing the research challenges of those methods relating to handwritten historical datasets. In this survey, historical documents collected from Australian Archives and Libraries are introduced and the associated layout complexities of those document images are also described. Benchmarking other historical databases related to this work is also discussed. This survey discusses the difficulties and suitability of the state-of-the-art methods to remove marginal noise as well as preserving the text content from handwritten historical documents. This survey helps researchers to identify appropriate methods according to the associated marginal noise and also illustrates their drawbacks in order to make suggestions for developing approaches, which are more general and robust for any datasets. Arpita Chakraborty, Michael Blumenstein |
DAS | 2 |
| 2016 | Preserving Text Content from Historical Handwritten DocumentsabstractWe propose a holistic, dynamic method to preserve text content with zero tolerance while removing marginal noise for historical handwritten document images. The key idea is to identify and analyze the region between the sharp peak at the edge and page frame of the text content at each margin. Depending on the proximity of the sharp peak to the text, the text content is then extracted from the document image. This method automatically adapts thresholds for each single document image and is directly applicable to gray-scale images. The proposed method is evaluated on four diverse handwritten historical datasets: Queensland State Archive (QSA), Saint Gall, Parzival and the Prosecution Project. Experimental results show that the proposed method achieves higher accuracy compared with other methods tested on the Saint Gall and Parzival datasets, whilst for the other two Australian datasets, which have been introduced here for the first time, the results are very encouraging. Arpita Chakraborty, Michael Blumenstein |
DAS | 2 |
| 2016 | Performance of an Off-Line Signature Verification Method Based on Texture Features on a Large Indic-Script Signature DatasetabstractIn this paper, a signature verification method based on texture features involving off-line signatures written in two different Indian scripts is proposed. Both Local Binary Patterns (LBP) and Uniform Local Binary Patterns (ULBP), as powerful texture feature extraction techniques, are used for characterizing off-line signatures. The Nearest Neighbour (NN) technique is considered as the similarity metric for signature verification in the proposed method. To evaluate the proposed verification approach, a large Bangla and Hindi off-line signature dataset (BHSig260) comprising 6240 (260×24) genuine signatures and 7800 (260×30) skilled forgeries was introduced and further used for experimentation. We further used the GPDS-100 signature dataset for a comparison. The experiments were conducted, and the verification accuracies were separately computed for the LBP and ULBP texture features. There were no remarkable changes in the results obtained applying the LBP and ULBP features for verification when the BHSig260 and GPDS-100 signature datasets were used for experimentation. Srikanta Pal, Alireza Alaei, Umapada Pal 0001, Michael Blumenstein |
DAS | 4 |
| 2016 | Offline Cursive Bengali Word Recognition Using CNNs with a Recurrent ModelabstractThis paper deals with offline handwritten word recognition of a major Indic script: Bengali. Due to the structure of this script, the characters (mostly ortho-syllables) are frequently overlapping and hard to segment, especially when the writing is cursive. Individual character recognition and the combination of outputs can increase the likelihood of errors. Instead, a better approach can be sending the whole word to a suitable recognizer. Here we use the Convolutional Neural Network (CNN) integrated with a recurrent model for this purpose. Long short-term memory blocks are used as hidden units. Also, the CNN-derived features are employed in a recurrent model with a CTC (Connectionist Temporal Classification) layer to get the output. We have tested our method on three datasets: (a) a publicly available dataset, (b) a new dataset generated by our research group and (c) an unconstrained dataset. The dataset (a) contains 17,091 words, while our dataset (b) contains 107,550 number of words in total. In addition to these, the dataset (c) is comprised of 5,223 words. We have compared our results with those of some earlier work in the area and have found improved performance, which is due to the novel integration of CNNs with the recurrent model. Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein |
ICFHR | 3 |
| 2016 | An Investigation of Novel Combined Features for a Handwritten Short Answer Assessment SystemabstractThis paper proposes an off-line automatic assessment system utilising novel combined feature extraction techniques. The proposed feature extraction techniques are 1) the proposed Water Reservoir, Loop, Modified Direction and Gaussian Grid Feature (WRL_MDGGF), 2) the proposed Gravity, Water Reservoir, Loop, Modified Direction and Gaussian Grid Feature (G_WRL_MDGGF). The proposed feature extraction techniques together with their original features and other combined feature extraction techniques were employed in an investigation of the efficiency of feature extraction techniques on an automatic off-line short answer assessment system. The proposed system utilised two classifiers namely, artificial neural networks and Support Vector Machines (SVMs), two type of datasets and two different thresholds in this investigation. Promising recognition rates of 94.85% and 94.88% were obtained when the proposed WRL_MDGGF and G_WRL_MDGGF were employed, respectively, using SVMs. Hemmaphan Suwanwiwat, Umapada Pal 0001, Michael Blumenstein |
ICFHR | 3 |
| 2016 | Writer identification by training on one script but testing on anotherabstractThis paper deals with identifying a writer from his/her offline handwriting. In a multilingual country where a writer can scribe in multiple scripts, writer identification becomes challenging when we have individual handwriting data in one script while we need to verify/identify a writer from handwriting in another script. In this paper such an issue is addressed with two scripts: English and Bengali. Here we model the task as a classification problem, where training data contains only Bengali handwritten samples and testing is performed on English handwritten texts. This work is based on the understanding that a writer has some inherent stroke characteristics that are independent of the script in which (s)he writes. In this work, some implicit structural and statistical features are extracted, and multiple classifiers are employed for writer identification. Many training sessions are run on a database of 100 writers and the performances are analyzed. We have obtained encouraging results on this database, which show the effectiveness of our method. Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein |
ICPR | 3 |
| 2016 | A quad tree based method for blurred and non-blurred video text frames classification through quality metricsabstractBlur is a common artifact in video, which adds more complexity to text detection and recognition. To achieve good accuracies for text detection and recognition, this paper suggests a new method for classifying blurred and non-blurred frames in video. We explore quality metrics, namely, BRISQUE, NRIQA, GPC and SI, in a new way for classification. We estimate the values of these metrics with the help of predefined samples called reference values. To widen the difference between metric values for better classification, we introduce scaling factors as a non-linear sigmoidal function, which considers the metric of each current frame and its reference and results in templates. Based on the characteristics of metrics, the proposed method finds a relationship between the metrics to derive rules for classification. To classify the frame containing local blur, we explore quad tree division with classification rules which divide non-blurred blocks to identify local blur. We use standard databases, namely, ICDAR 2013, ICDAR 2015 and YVT videos for experimentation, and evaluate the proposed method in terms of text detection and recognition rates given by text detection and binarization methods before and after classification. Vijeta Khare, Palaiahnakote Shivakumara, Ahlad Kumar, Chee Seng Chan, Tong Lu 0002, Michael Blumenstein |
ICPR | 6 |
| 2016 | A brief review of document image retrieval methods: Recent advancesabstractDue to the rapid increase of different digitized documents, the development of a system to automatically retrieve document images from a large collection of structured and unstructured document images is in high demand. Many techniques have been developed to provide an efficient and effective way for retrieving and organizing these document images in the literature. This paper provides an overview of the methods which have been applied for document image retrieval over recent years. It has been found that from a textual perspective, more attention has been paid to the feature extraction methods without using OCR. Fahimeh Alaei, Alireza Alaei, Michael Blumenstein, Umapada Pal 0001 |
IJCNN | 3 |
| 2016 | A blind deconvolution model for scene text detection and recognition in video
Vijeta Khare, Palaiahnakote Shivakumara, Raveendran Paramesran, Michael Blumenstein |
Pattern Recognit. | 4 |
| 2016 | A framework for liveness detection for direct attacks in the visible spectrum for multimodal ocular biometrics
Abhijit Das 0001, Umapada Pal 0001, Miguel A. Ferrer, Michael Blumenstein |
Pattern Recognit. Lett. | 4 |
| 2016 | Contour Restoration of Text Components for Recognition in Video/Scene ImagesabstractText recognition in video/natural scene images has gained significant attention in the field of image processing in many computer vision applications, which is much more challenging than recognition in plain background images. In this paper, we aim to restore complete character contours in video/scene images from gray values, in contrast to the conventional techniques that consider edge images/binary information as inputs for text detection and recognition. We explore and utilize the strengths of zero crossing points given by the Laplacian to identify stroke candidate pixels (SPC). For each SPC pair, we propose new symmetry features based on gradient magnitude and Fourier phase angles to identify probable stroke candidate pairs (PSCP). The same symmetry properties are proposed at the PSCP level to choose seed stroke candidate pairs (SSCP). Finally, an iterative algorithm is proposed for SSCP to restore complete character contours. Experimental results on benchmark databases, namely, the ICDAR family of video and natural scenes, Street View Data, and MSRA data sets, show that the proposed technique outperforms the existing techniques in terms of both quality measures and recognition rate. We also show that character contour restoration is effective for text detection in video and natural scene images. Yirui Wu, Palaiahnakote Shivakumara, Tong Lu 0002, Chew Lim Tan, Michael Blumenstein, G. Hemantha Kumar 0001 |
IEEE Trans. Image Process. | 5 |
| 2015 | ICDAR2015 competition on signature verification and writer identification for on- and off-line skilled forgeries (SigWIcomp2015)abstractThis paper presents the results of the ICDAR 2015 competition on signature verification and writer identification for on- and off-line skilled forgeries jointly organized by PR-researchers and Forensic Handwriting Examiners (FHEs). The aim is to bridge the gap between recent technological developments and forensic casework. Two modalities (signatures and handwritten text) are considered and training and evaluation data are collected and provided by FHEs and PR-researchers. Four tasks are defined for four different languages; Bengali off-line signature verification, Italian off-line signature verification, German on-line signature verification, and English handwritten text based writer identification. In total, 40 systems have participated in this competition. The participants of the signatures modality were motivated to report their results in Likelihood Ratios (LRs). This has made the systems even more interesting for application in forensic casework. For evaluating the performance of the systems, we have used the forensically substantial Cost of Log Likelihood Ratios (Ĉllr) in the case of signatures, and the F-measure in the case of handwritten text. Muhammad Imran Malik, Sheraz Ahmed, Angelo Marcelli, Umapada Pal 0001, Michael Blumenstein, Linda Alewijnse, Marcus Liwicki |
ICDAR | 5 |
| 2015 | Date field extraction from handwritten documents using HMMsabstractAutomatic document interpretation and retrieval is an important task to access handwritten digitized document repositories. In documents, the date is an important field and it has various applications such as date-wise document indexing/retrieval. In this paper a framework has been proposed for automatic date field extraction from handwritten documents. In order to design the system, sliding window-wise Local Gradient Histogram (LGH)-based features and a character-level Hidden Markov Model (HMM)-based approach have been applied for segmentation and recognition. Individual date components such as month-word (month written in word form i.e. January, Jan, etc.), numeral, punctuation and contraction categories are segmented and labelled from a text line. Next, a Histogram of Gradient (HoG)-based features and a Support Vector Machine (SVM)- based classifier have been used to improve the results obtained from the HMM-based recognition system. Subsequently, both numeric and semi-numeric regular expressions of date patterns have been considered for undertaking date pattern extraction in labelled components. The experiments are performed on an English document dataset and the encouraging results obtained from the approach indicate the effectiveness of the proposed system. Ranju Mandal, Partha Pratim Roy 0001, Umapada Pal 0001, Michael Blumenstein |
ICDAR | 4 |
| 2015 | ICDAR2015 Competition on Video Script Identification (CVSI 2015)abstractThis paper presents the final results of the ICDAR 2015 Competition on Video Script Identification. A description and performance of the participating systems in the competition are reported. The general objective of the competition is to evaluate and benchmark the available methods on word-wise video script identification. It also provides a platform for researchers around the globe to particularly address the video script identification problem and video text recognition in general. The competition was organised around four different tasks involving various combinations of scripts comprising tri-script and multi-script scenarios. The dataset used in the competition comprised ten different scripts. In total, six systems were received from five participants over the tasks offered. This report details the competition dataset specifications, evaluation criteria, summary of the participating systems and their performance across different tasks. The systems submitted by Google Inc. were the winner of the competition for all the tasks, whereas the systems received from Huazhong University of Science and Technology (HUST) and Computer Vision Center (CVC) were very close competitors. Nabin Sharma, Ranju Mandal, Rabi Sharma, Umapada Pal 0001, Michael Blumenstein |
ICDAR | 5 |
| 2015 | Multi-lingual text recognition from video framesabstractText recognition from video frames is a challenging task due to low resolution, blur, complex and coloured backgrounds, noise, to mention a few. Consequently, the traditional ways of text recognition from scanned documents having simple backgrounds fails when applied to video text. Although there are various techniques available for text recognition from handwritten and printed documents with simple backgrounds, text recognition from video frames has not been comprehensively investigated, especially for multi-lingual videos. In this paper, we present a technique for multi-lingual video text recognition which involves script identification in the first stage, followed by word and character recognition, and finally the results are refined using a post-processing technique. Considering the inherent problems in videos, a Spatial Pyramid Matching (SPM) based technique, using patch-based SIFT descriptors and SVM classifier, is employed for script identification. In the next stage, a Hidden Markov Model (HMM) based approach is used for word and character recognition, which utilizes the context information. Finally, a lexicon-based post-processing technique is applied to verify and refine the word recognition results. The proposed method was tested on a dataset comprising of 4800 words from three different scripts, namely, Roman (English), Hindi and Bengali. The script identification results obtained are encouraging. The word and character recognition results are also encouraging considering the complexity and problems associated with video text processing. Nabin Sharma, Ranju Mandal, Rabi Sharma, Partha Pratim Roy 0001, Umapada Pal 0001, Michael Blumenstein |
ICDAR | 6 |
| 2015 | A complete automatic short answer assessment system with student identificationabstractThere are only a few studies undertaken in developing automatic assessment systems using handwriting recognition, even though a successful system would undoubtedly benefit the education system as schools and universities in many countries still employ paper-based examinations. To the best of the authors' knowledge, there is no existing work on an automatic off-line short answer assessment system comprising a student identification component. Hence in this paper, the authors propose a system towards this, where a new feature extraction technique called the Enhanced Water Reservoir, Loop and Gaussian Grid Feature, as well as other enhanced feature extraction techniques were utilised. Artificial Neural Networks and Support Vector Machines were employed as the classifiers; they were used for the investigation, and a comparison of the recognition and accuracy rates of the proposed systems, as well as the feature extraction techniques, was undertaken. The proposed assessment system achieved a recognition rate of 87.12% with 91.12% assessment accuracy, and the student identification component obtained a recognition rate of 99.52% with a 100% identification accuracy rate. Hemmaphan Suwanwiwat, Michael Blumenstein, Umapada Pal 0001 |
ICDAR | 2 |
| 2015 | Towards robust flood forecasts using neural networksabstractIn this paper, design of a neural network for a domain-specific problem is described. The problem of concern is forecasting flood events where data is contaminated heavily by noise, training examples have different importance levels and noisy data coincides with the most important ones. To this end, two ideas are explored namely, changing the loss function and integrating a coefficient that reflects on the relative importance of training examples. To this end, backpropagation is re-derived considering implication of having a more general objective function. Independently, inclusion of scores associated with each training examples and its implication of overall loss function and the way weights are optimized is explored. The derived model is implemented in MATLAB and flood data from Talebudgera, Australia is considered for investigations. Compared to the base case being backpropagation, the results suggest that inclusion of scored for training examples corresponds to visible improvement when predicting peaks. Seyyed Adel Alavi Fazel, Hamid Mirfenderesk, Rodger Tomlinson, Michael Blumenstein |
IJCNN | 4 |
| 2015 | Interval-valued symbolic representation based method for off-line signature verificationabstractThe objective of this investigation is to present an interval-symbolic representation based method for offline signature verification. In the feature extraction stage, Connected Components (CC), Enclosed Regions (ER), Basic Features (BF) and Curvelet Feature (CF)-based approaches are used to characterize signatures. Considering the extracted feature vectors, an interval data value is created for each feature extracted from every individual's signatures as an interval-valued symbolic data. This process results in a signature model for each individual that consists of a set of interval values. A similarity measure is proposed as the classifier in this paper. The interval-valued symbolic representation based method has never been used for signature verification considering Indian script signatures. Therefore, to evaluate the proposed method, a Hindi signature database consisting of 2400 (100×24) genuine signatures and 3000 (100×30) skilled forgeries is employed for experimentation. Concerning this large Hindi signature dataset, the highest verification accuracy of 91.83% was obtained on a joint feature set considering all four sets of features, while 2.5%, 13.84% and 8.17% of FAR (False Acceptance Rate), FRR (False Rejection Rate), and AER (Average Error Rate) were achieved, respectively. Srikanta Pal, Alireza Alaei, Umapada Pal 0001, Michael Blumenstein |
IJCNN | 4 |
| 2015 | Bag-of-Visual Words for word-wise video script identification: A studyabstractUse of multiple scripts for information communication through various media is quite common in a multilingual country. Optical character recognition of such document images or videos assists in indexing them for effective information retrieval. Hence, script identification from multi-lingual documents/images is a necessary step for selecting the appropriate OCR, due the absence of a single OCR system capable of handling multiple scripts. Script identification from printed as well as handwritten documents is a well-researched area, but script identification from video frames has not been explored much. Low resolution, blur, noisy background, to mention a few are the major bottle necks when processing video frames, and makes script identification from video images a challenging task. This paper examines the potential of Bag-of-Visual Words based techniques for word-wise script identification from video frames. Two different approaches namely, Bag-Of-Features (BoF) and Spatial Pyramid Matching (SPM), using patch based SIFT descriptors were considered for the current study. SVM Classifier was used for analysing the three popular south Indian scripts, namely Tamil, Telugu and Kannada in combination with English and Hindi. A comparative study of Bag-of-Visual words with traditional script identification techniques involving gradient based features (e.g. HoG) and texture based features (e.g. LBP) is presented. Experimental results shows that patch-based features along with SPM outperformed the traditional techniques and promising accuracies were achieved on 2534 words from the five scripts. The study reveals that patch-based feature can be used for scripts identification in-order to overcome the inherent problems with video frames. Nabin Sharma, Ranju Mandal, Rabi Sharma, Umapada Pal 0001, Michael Blumenstein |
IJCNN | 5 |
| 2015 | Short answer question examination using an automatic off-line handwriting recognition system and a novel combined featureabstractOff-line automatic assessment systems can be an aid for teachers in the marking process. There has been no recent work in the development of off-line automatic assessment systems using handwriting recognition, even though such systems will clearly benefit the education sector. The reason is many schools and universities in many parts of the world still use paper-based examination. This research proposes the use of a newly developed feature extraction technique called the Modified Water Reservoir, Loop and Gaussian Grid Feature, as well as other feature extraction techniques. These techniques were investigated employing artificial neural networks and support vector machines as classifiers to develop an automatic assessment system for marking short answer questions. The system has high assessment accuracy (up to 94.75% for hand printed, 96.09% for cursive handwritten, and 95.71% for hand printed and cursive handwritten combined). The proposed system also includes assessment criteria to augment its accuracy. Hemmaphan Suwanwiwat, Michael Blumenstein, Umapada Pal 0001 |
IJCNN | 2 |
| 2015 | Multi-lingual date field extraction for automatic document retrieval by machine
Ranju Mandal, Partha Pratim Roy 0001, Umapada Pal 0001, Michael Blumenstein |
Inf. Sci. | 4 |
| 2015 | Piece-wise linearity based method for text frame classification in video
Nabin Sharma, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Chew Lim Tan |
Pattern Recognit. | 4 |
| 2014 | Fuzzy logic based selera recognitionabstractIn this paper a sclera recognition and validation system is proposed. Here sclera segmentation was performed by Fuzzy logic-based clustering. Since the selera vessels are not prominent, image enhancement was required. A Fuzzy logic-based Brightness Preserving Dynamic Fuzzy Histogram Equalization and discrete Meyer wavelet was used to enhance the vessel patterns. For feature extraction, the Dense Local Binary Pattern (D-LBP) was used. D-LBP patch descriptors of each training image are used to form a bag of features, which is used to produce the training model. Support Vector Machines (SVMs) are used for classification. The UBIRIS version 1 dataset is used here for experimentation. An encouraging Equal Error Rate (EER) of 4.31% was achieved in our experiments. Abhijit Das 0001, Umapada Pal 0001, Miguel A. Ferrer, Michael Blumenstein |
FUZZ-IEEE | 4 |
| 2014 | Off-Line Handwritten Bilingual Name Recognition for Student Identification in an Automated Assessment SystemabstractThe Student name Identification System (SIS) proposed here was investigated for English and Thai languages combined. The proposed system recognises each name by using an approach for whole word recognition. In the proposed system, the Gaussian Grid Feature (GGF), and Modified Direction Feature (MDF), together with a proposed hybrid feature extraction technique called Water Reservoir, Loop and Gaussian Grid Feature (WRLGGF) were investigated on full word contour images of each name sample. Artificial neural networks and support vector machines were used as classifiers. An encouraging recognition accuracy of 99.25% was achieved employing the proposed technique compared to 98.59% for GGF, and 96.63% using MDF. Hemmaphan Suwanwiwat, Vu Nguyen 0002, Michael Blumenstein, Umapada Pal 0001 |
ICFHR | 3 |
| 2014 | Semi-supervised Online Bayesian Network Learner for Handwritten Characters RecognitionabstractThis work addresses the problem of creating a Bayesian Network based online semi-supervised handwritten character recognisor, which learns continuously over time to make a adaptable recognisor. The proposed method makes learning possible from a continuous inflow of a potentially unlimited amount of data without the requirement for storage. It highlights the use of unlabelled data for boosting the accuracy, especially when labelled data is scarce and expensive unlike unlabelled data. An algorithm is introduced to perform semi-supervised learning based on the combination of novel online ensemble of the Randomized Bayesian network classifiers and a novel online variant of the Expectation Maximization (EM) algorithm. We make use of a novel varying weighting factor to modulate the contribution of unlabelled data. Proposed method was evaluated using online handwritten Tamil characters from the IWFHR 2006 competition dataset. The accuracy obtained was comparable to the state of the art batch learning methods like HMM and SVMs. Rituraj Kunwar, Umapada Pal 0001, Michael Blumenstein |
ICPR | 3 |
| 2014 | Gradient-Angular-Features for Word-wise Video Script IdentificationabstractScript identification at the word level is challenging because of complex backgrounds and low resolution of video. The presence of graphics and scene text in video makes the problem more challenging. In this paper, we employ gradient angle segmentation on words from video text lines. This paper presents new Gradient-Angular-Features (GAF) for video script identification, namely, Arabic, Chinese, English, Japanese, Korean and Tamil. This work enables us to select an appropriate OCR when the frame has words of multi-scripts. We employ gradient directional features for segmenting words from video text lines. For each segmented word, we study the gradient information in effective ways to identify text candidates. The skeleton of the text candidates is analyzed to identify Potential Text Candidates (PTC) by filtering out unwanted text candidates. We propose novel GAF for the PTC to study the structure of the components in the form of cursiveness and softness. The histogram operation on the GAF is performed in different ways to obtain discriminative features. The method is evaluated on 760 words of six scripts having low contrast, complex background, different font sizes, etc. in terms of the classification rate and is compared with an existing method to show the effectiveness of the method. We achieve 88.2% average classification rate. Palaiahnakote Shivakumara, Nabin Sharma, Umapada Pal 0001, Michael Blumenstein, Chew Lim Tan |
ICPR | 4 |
| 2014 | Estuarine flood modelling using artificial neural networksabstractPrediction of water levels at estuaries poses a significant challenge for modelling of floods due to the influence of tidal effects. In this study, a two-stage forecasting system is proposed. In the first stage, the tidal portion of the available records is used to develop a tidal prediction system. The predictions of the first stage are used for flood modelling in the second. Experimental results suggest that the proposed flood modelling approach is advantageous for forecasting flood levels with more than 1 hour lead times. Seyyed Adel Alavi Fazel, Michael Blumenstein, Hamid Mirfenderesk, Rodger Tomlinson |
IJCNN | 2 |
| 2014 | A study on word-level multi-script identification from video framesabstractThe presence of multiple scripts in multi-lingual document images makes Optical Character Recognition (OCR) of such documents a challenging task. Due to the unavailability of a single OCR system which can handle multiple scripts, script identification becomes an essential step for choosing the appropriate OCR. Although, there are various techniques available for script identification from handwritten and printed documents having simple backgrounds, however script identification from video frames has been seldom explored. Video frames are coloured and suffer from low resolution, blur, complex background and noise to mention a few, which makes the script identification process a challenging task. This paper presents a study of various combinations of features and classifiers to explore whether the traditional script identification techniques can be applied to video frames. A texture based feature namely, Local Binary Pattern (LBP), Gradient based features namely, Histogram of Oriented Gradient (HoG) and Gradient Local Auto-Correlation (GLAC) were used in the study. Combination of the features with SVMs and ANNs where used for classification. Three popular scripts, namely English, Bengali and Hindi were considered in the present study. Due to the inherent problems with the video, a super resolution technique was applied as a pre-processing step. Experiments show that the GLAC feature has performed better than the other features, and an accuracy of 94.25% was achieved when testing on 1271 words from three different scripts. The study also reveals that gradient features are more suitable for script identification than the texture features when using traditional script identification techniques on video frames. Nabin Sharma, Umapada Pal 0001, Michael Blumenstein |
IJCNN | 3 |
| 2014 | Off-line handwritten Thai name recognition for student identification in an automated assessment systemabstractIn the field of pattern recognition, off-line handwriting recognition is one of the most intensive areas of study. This paper proposes an automatic off-line Thai language student name identification system which was built as a part of a completed off-line automated assessment system. There is limited work undertaken in developing off-line automatic assessment systems using handwriting recognition. To the authors' knowledge, none of the work on the proposed system has been performed on the Thai language. In addition the proposed system recognises each Thai name by using an approach for whole word recognition, which is different from the work found in the literature as most perform character-based recognition. In this proposed system, the Gaussian Grid Feature (GGF) and the Modified Direction Feature (MDF) extraction techniques are investigated on upper and lower contours, loops from full word contour images of each name sample, and artificial neural networks and support vector machine are used as classifiers. The encouraging recognition rates for both feature extraction techniques were achieved when applied on loop, upper and lower contour images (99.27% accuracy rate was achieved using MDF on artificial neural networks and 99.27% using GGF with a support vector machine classifier). Hemmaphan Suwanwiwat, Vu Nguyen 0002, Michael Blumenstein, Umapada Pal 0001 |
IJCNN | 3 |
| 2013 | ICDAR 2013 Competitions on Signature Verification and Writer Identification for On- and Offline Skilled Forgeries (SigWiComp 2013)abstractThis paper presents the results of the ICDAR2013 competitions on signature verification and writer identification for on- and offline skilled forgeries jointly organized by PR researchers and Forensic Handwriting Examiners (FHEs). The aim is to bridge the gap between recent technological developments and forensic casework. Two modalities (signatures, and handwritten text) are considered where training and evaluation data (in Dutch and Japanese) were collected and provided by FHEs and PR-researchers. Four tasks were defined where the systems had to perform Dutch offline signature verification, Japanese offline signature verification, Japanese online signature verification, and Dutch writer identification. The participants of the signatures modality were motivated to report their results in Likelihood Ratios (LR). This has made the systems even more interesting for application in forensic casework. For evaluation of signatures modality, we used both the traditional Equal Error Rate (EER) and forensically substantial Cost of Log Likelihood Ratios (Ĉllr). The system having the smallest value of the Minimum Cost of Log Likelihood Ratio (Ĉllrmin) is declared winner. For evaluation of the handwritten text modality, we used the precision and accuracy measures and winners are announced on the basis of best F-measure value. Muhammad Imran Malik, Marcus Liwicki, Linda Alewijnse, Wataru Ohyama, Michael Blumenstein, Bryan Found |
ICDAR | 5 |
| 2013 | Word-Wise Script Identification from Video FramesabstractScript identification is an essential step for the efficient use of the appropriate OCR in multilingual document images. There are various techniques available for script identification from printed and handwritten document images, but script identification from video frames has not been explored much. This paper presents a study of some pre-processing techniques and features for word-wise script identification from video frames. Traditional features, namely Zernike moments, Gabor and gradient, have performed well for handwritten and printed documents having simple backgrounds and adequate resolution for OCR. Video frames are mostly coloured and suffer from low resolution, blur, background noise, to mention a few. In this paper, an attempt has been made to explore whether the traditional script identification techniques can be useful in video frames. Three feature extraction techniques, namely Zernike moments, Gabor and gradient features, and SVM classifiers were considered for analyzing three popular scripts, namely English, Bengali and Hindi. Some pre-processing techniques such as super resolution and skeletonization of the original word images were used in order to overcome the inherent problems with video. Experiments show that the super resolution technique with gradient features has performed well, and an accuracy of 87.5% was achieved when testing on 896 words from three different scripts. The study also reveals that the use of proper pre-processing approaches can be helpful in applying traditional script identification techniques to video frames. Nabin Sharma, Sukalpa Chanda, Umapada Pal 0001, Michael Blumenstein |
ICDAR | 4 |
| 2013 | A New Method for Character Segmentation from Multi-oriented Video WordsabstractThis paper presents a two-stage method for multi-oriented video character segmentation. Words segmented from video text lines are considered for character segmentation in the present work. Words can contain isolated or non-touching characters, as well as touching characters. Therefore, the character segmentation problem can be viewed as a two stage problem. In the first stage, text cluster is identified and isolated (non-touching) characters are segmented. The orientation of each word is computed and the segmentation paths are found in the direction perpendicular to the orientation. Candidate segmentation points computed using the top distance profile are used to find the segmentation path between the characters considering the background cluster. In the second stage, the segmentation results are verified and a check is performed to ascertain whether the word component contains touching characters or not. The average width of the components is used to find the touching character components. For segmentation of the touching characters, segmentation points are then found using average stroke width information, along with the top and bottom distance profiles. The proposed method was tested on a large dataset and was evaluated in terms of precision, recall and f-measure. A comparative study with existing methods reveals the superiority of the proposed method. Nabin Sharma, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Chew Lim Tan |
ICDAR | 4 |
| 2013 | Off-line Bangla signature verification: An empirical studyabstractAmong all of the biometric authentication systems, handwritten signatures are considered as the most legally and socially accepted attributes for personal verification. The objective of this paper is to present an empirical contribution towards the understanding of a threshold-based signature verification technique involving off-line Bangla (Bengali) signatures. Experiments on signature verification involving non-English signatures are an important consideration in the signature verification area. Only very few research works employing signatures of Indian script have been considered in the field of non-English signature verification. To fill this gap, a threshold-based scheme for verification considering off-line Bangla signatures is proposed. Some techniques such as under-sampled bitmap, intersection/endpoint and directional chain code are employed for feature extraction. The Nearest Neighbour method is considered for classification. Furthermore, a Bangla signature database, which consists of 2400 (100×24) genuine signatures and 3000 (100×30) forgeries has been created and is employed for experimentation. We obtained a 15.57% Average Error Rate (AER) as the best verification result using directional chain code features employed in this research work. Srikanta Pal, Alireza Alaei, Umapada Pal 0001, Michael Blumenstein |
IJCNN | 4 |
| 2013 | Sclera recognition using dense-SIFTabstractIn this paper we propose a biometric sclera recognition and validation system. Here the sclera segmentation is performed bya time-adaptive active contour-based region growing technique. The sclera vessels are not prominent so image enhancement is required and hence a bank of 2D decomposition. A Haar wavelet multi-resolution filter is used to enhance the vessels pattern for better accuracy. For feature extraction, Dense Scale Invariant Feature Transform (D-SIFT) is used. D-SIFT patch descriptors of each training image are used to form bag of features by using k-means clustering and a spatial pyramid model, which is used to produce the training model. Support Vector Machines (SVMs) are used for classification. The UBIRIS version 1 dataset is used here for experimentation. Anencouraging Equal Error Rate (EER) of 0.66% is attained in the experiments presented. Abhijit Das 0001, Umapada Pal 0001, Miguel A. Ferrer, Michael Blumenstein |
ISDA | 4 |
| 2013 | Signature segmentation and recognition from scanned documentsabstractSignature as a query is important for content-based document image retrieval from a scanned document repository. This paper presents a two-stage approach towards automatic signature segmentation and recognition from scanned document images. In the first stage, signature blocks are segmented from the document using word-wise component extraction and classification. Gradient based features are extracted from each component at the word level to perform the classification task. In the 2nd stage, SIFT (Scale-Invariant Feature Transform) descriptors and Spatial Pyramid Matching (SPM)-based approaches are used for signature recognition. Support Vector Machines (SVMs) are employed as the classifier for both levels in this experiment. The experiments are performed on the publicly available “Tobacco-800” and GPDS [1] datasets and the results obtained from the experiments are promising. Ranju Mandal, Partha Pratim Roy 0001, Umapada Pal 0001, Michael Blumenstein |
ISDA | 4 |
| 2013 | Svm and NN Based Offline Signature VerificationabstractAmong all of the biometric authentication systems, handwritten signatures are considered as the most legally and socially accepted attributes for personal verification. The objective of this paper is to present an empirical contribution toward the understanding of a threshold-based signature verification technique involving off-line Bangla (Bengali) signatures. Experiments on signature verification concerning non-English signatures are an important consideration in the signature verification area. Only very few research works employing signatures of Indian script have been considered in the field of non-English based signature verification. To fill this gap, a threshold-based scheme for the verification of off-line Bangla signatures is proposed. Some techniques such as under-sampled bitmap, intersection/end point and directional chain code are employed for feature extraction. The thresholds are computed based on the similarity measures obtained employing the nearest neighbor classifier. The SVM classifier has also been considered for mainly comparative experimental result generation. Furthermore, a Bangla signature database, which consists of 2400 (100 × 24) genuine signatures and 3000 (100 × 30) forgeries, has been created and is employed for experimentation. An average error rate (AER) of 12.33% was obtained as the best verification result using directional chain code features in this research work. As a comparative study, a different dataset (GPDS-160) has also been considered. Srikanta Pal, Alireza Alaei, Umapada Pal 0001, Michael Blumenstein |
Int. J. Comput. Intell. Appl. | 4 |
| 2012 | A Compact Size Feature Set for the Off-Line Signature Verification ProblemabstractWith increasing computational power, researchers in the area of off-line signature verification have been able to investigate feature extraction techniques that produce large-dimensional feature vectors. However, a large feature vector is not necessarily associated with high performance. This paper investigates the performance of a small feature set consisting of 33 feature values. In the experiments using Support Vector Machines (SVMs), an average error rate (AER) of 16.80% was obtained together with a low false acceptance rate (FAR) for random forgeries of 0.19%. The significant reduction of the error rate was obtained when the proposed global features were employed, which demonstrates their astonishingly high discriminant power. These results suggest that the use of global features for the off-line signature verification problem is worth further investigation. Vu Nguyen 0002, Michael Blumenstein |
Document Analysis Systems | 2 |
| 2012 | Off-Line Bangla Signature VerificationabstractIn the field of information security, biometric systems play an important role. Within biometrics, automatic signature identification and verification has been a strong research area because of the social and legal acceptance and extensive use of the written signature as an individual authentication. Signature verification is a process in which the questioned signature is examined in detail in order to determine whether it belongs to the claimed person or not. Despite substantial research in the field of signature verification involving Western signatures, very few works have been dedicated to non-Western signatures such as Chinese, Japanese, Arabic, or Persian etc. In this paper, the performance of an off-line signature verification system involving Bangla signatures, whose style is distinct from Western scripts, was investigated. The Gaussian Grid feature extraction technique was employed for feature extraction and Support Vector Machines (SVMs) were considered for classification. The Bangla signature database employed in the experiments consisted of 3000 forgeries and 2400 genuine signatures. An encouraging accuracy of 90.4% was obtained from the experiments. Srikanta Pal, Vu Nguyen 0002, Michael Blumenstein, Umapada Pal 0001 |
Document Analysis Systems | 3 |
| 2012 | Recent Advances in Video Based Document Processing: A ReviewabstractExtraction and recognition of text present in video has become a very popular research area in the last decade. Generally, text present in video frames is of different size, orientation, style, etc. with complex backgrounds, noise, low resolution and contrast. These factors make the automatic text extraction and recognition in video frames a challenging task. A large number of techniques have been proposed by various researchers in the recent past to address the problem. This paper presents a review of various state-of-the-art techniques proposed towards different stages (e.g. detection, localization, extraction, etc.) of text information processing in video frames. Looking at the growing popularity and the recent developments in the processing of text in video frames, this review imparts details of current trends and potential directions for further research activities to assist researchers. Nabin Sharma, Umapada Pal 0001, Michael Blumenstein |
Document Analysis Systems | 3 |
| 2012 | A New Method for Arbitrarily-Oriented Text Detection in VideoabstractText detection in video frames plays a vital role in enhancing the performance of information extraction systems because the text in video frames helps in indexing and retrieving video efficiently and accurately. This paper presents a new method for arbitrarily-oriented text detection in video, based on dominant text pixel selection, text representatives and region growing. The method uses gradient pixel direction and magnitude corresponding to Sobel edge pixels of the input frame to obtain dominant text pixels. Edge components in the Sobel edge map corresponding to dominant text pixels are then extracted and we call them text representatives. We eliminate broken segments of each text representatives to get candidate text representatives. Then the perimeter of candidate text representatives grows along the text direction in the Sobel edge map to group the neighboring text components which we call word patches. The word patches are used for finding the direction of text lines and then the word patches are expanded in the same direction in the Sobel edge map to group the neighboring word patches and to restore missing text information. This results in extraction of arbitrarily-oriented text from the video frame. To evaluate the method, we considered arbitrarily-oriented data, non-horizontal data, horizontal data, Hua's data and ICDAR-2003 competition data (Camera images). The experimental results show that the proposed method outperforms the existing method in terms of recall and f-measure. Nabin Sharma, Palaiahnakote Shivakumara, Umapada Pal 0001, Michael Blumenstein, Chew Lim Tan |
Document Analysis Systems | 4 |
| 2012 | Multi-script off-line signature identificationabstractIn this paper, we present an empirical contribution towards the understanding of multi-script signature identification. In the proposed signature identification system, the signatures of Bengali (Bangla), Hindi (Devanagari) and English are considered for the identification process. This system will identify whether a claimed signature belongs to the group of Bengali, Hindi or English signatures. Zernike Moment and histogram of gradient are employed as two different feature extraction techniques. In the proposed system, Support Vector Machines (SVMs) are considered as classifiers for signature identification. A database of 2100 Bangla signatures, 2100 Hindi signatures and 2100 English signatures are used for experimentation. Two different results based on two different feature sets are calculated and analysed. The highest accuracy of 92.14% is obtained based on the gradient features using 4200 (1400 Bangla +1400 Hindi + 1400 English) samples for training and 2100 (700 Bangla +700 Hindi +700 English) samples for testing. Srikanta Pal, Alireza Alaei, Umapada Pal 0001, Michael Blumenstein |
HIS | 4 |
| 2012 | Off-line restricted-set handwritten word recognition for student identification in a short answer question automated assessment systemabstractHandwriting recognition is one of the most intensive areas of study in the field of pattern recognition. Many applications are able to benefit from a robust off-line handwriting recognition technique. An automatic off-line assessment system and a writer identification system are two of those applications. Off-line automatic assessment systems can be an aid for teachers in the marking process; they can reduce the time consumed by the human marker. There has only been limited work undertaken in developing off-line automatic assessment systems using handwriting recognition, and none in developing student identification systems, even though such systems would clearly benefit the education sector. In order to develop a complete off-line automatic assessment system, student identification using full student names is proposed in this paper. The Gaussian Grid and Modified Direction Feature Extraction Techniques are investigated in order to develop the proposed system. The recognition rates achieved using both techniques are encouraging (up to 99.08% for the Modified Direction feature extraction technique, and up to 98.28% for the Gaussian Grid feature extraction technique. Hemmaphan Suwanwiwat, Vu Nguyen 0002, Michael Blumenstein |
HIS | 3 |
| 2012 | Hindi Off-Line Signature VerificationabstractHandwritten Signatures are one of the widely used biometrics for document authentication as well as human authorization. The purpose of this paper is to present an offline signature verification system involving Hindi signatures. Signature verification is a process by which the questioned signature is examined in detail in order to determine whether it belongs to the claimed person or not. Despite of substantial research in the field of signature verification involving Western signatures, very little attention has been dedicated to non-Western signatures such as Chinese, Japanese, Arabic, Persian etc. In this paper, the performance of an off-line signature verification system involving Hindi signatures, whose style is distinct from Western scripts, has been investigated. The gradient and Zernike moment features were employed and Support Vector Machines (SVMs) were considered for verification. To the best of the authors' knowledge, Hindi signatures have never been used for the task of signature verification and this is the first report of using Hindi signatures in this area. The Hindi signature database employed for experimentation consisted of 840 (35x24) genuine signatures and 1050 (35x30) forgeries. An encouraging accuracy of 7.42% FRR and 4.28% FAR were obtained following experimentation when the gradient features were employed. Srikanta Pal, Michael Blumenstein, Umapada Pal 0001 |
ICFHR | 2 |
| 2012 | Off-line English and Chinese signature identification using foreground and background featuresabstractIn the field of information security, the usage of biometrics is growing for user authentication. Automatic signature recognition and verification is one of the biometric techniques, which is only one of several used to verify the identity of individuals. In this paper, a foreground and background based technique is proposed for identification of scripts from bi-lingual (English/Roman and Chinese) off-line signatures. This system will identify whether a claimed signature belongs to the group of English signatures or Chinese signatures. The identification of signatures based on its script is a major contribution for multi-script signature verification. Two background information extraction techniques are used to produce the background components of the signature images. Gradient-based method was used to extract the features of the foreground as well as background components. Zernike Moment feature was also employed on signature samples. Support Vector Machine (SVM) is used as the classifier for signature identification in the proposed system. A database of 1120 (640 English+480 Chinese) signature samples were used for training and 560 (320 English+240 Chinese) signature samples were used for testing the proposed system. An encouraging identification accuracy of 97.70% was obtained using gradient feature from the experiment. Srikanta Pal, Umapada Pal 0001, Michael Blumenstein |
IJCNN | 3 |
| 2012 | Off-line signature verification using G-SURFabstractIn the field of biometric authentication, automatic signature identification and verification has been a strong research area because of the social and legal acceptance and extensive use of the written signature as an easy method for authentication. Signature verification is a process in which the questioned signature is examined in detail in order to determine whether it belongs to the claimed person or not. Signatures provide a secure means for confirmation and authorization in legal documents. So nowadays, signature identification and verification becomes an essential component in automating the rapid processing of documents containing embedded signatures. Sometimes, part-based signature verification can be useful when a questioned signature has lost its original shape due to inferior scanning quality. In order to address the above-mentioned adverse scenario, we propose a new feature encoding technique. This feature encoding is based on the amalgamation of Gabor filter-based features with SURF features (G-SURF). Features generated from a signature are applied to a Support Vector Machine (SVM) classifier. For experimentation, 1500 (50×30) forgeries and 1200 (50×24) genuine signatures from the GPDS signature database were used. A verification accuracy of 97.05% was obtained from the experiments. Srikanta Pal, Sukalpa Chanda, Umapada Pal 0001, Katrin Franke, Michael Blumenstein |
ISDA | 5 |
| 2012 | Concept Learning for $\ensuremath{\ensuremath{\cal E}\ensuremath{\cal L}^{++}}$ by Refinement and Reinforcement
Mahsa Chitsaz, Kewen Wang 0001, Michael Blumenstein, Guilin Qi |
PRICAI | 3 |
| 2011 | Signature Verification Competition for Online and Offline Skilled Forgeries (SigComp2011)abstractThe Netherlands Forensic Institute and the Institute for Forensic Science in Shanghai are in search of a signature verification system that can be implemented in forensic casework and research to objectify results. We want to bridge the gap between recent technological developments and forensic casework. In collaboration with the German Research Center for Artificial Intelligence we have organized a signature verification competition on datasets with two scripts (Dutch and Chinese) in which we asked to compare questioned signatures against a set of reference signatures. We have received 12 systems from 5 institutes and performed experiments on online and offline Dutch and Chinese signatures. For evaluation, we applied methods used by Forensic Handwriting Examiners (FHEs) to assess the value of the evidence, i.e., we took the likelihood ratios more into account than in previous competitions. The data set was quite challenging and the results are very interesting. Marcus Liwicki, Muhammad Imran Malik, C. Elisa van den Heuvel, Xiaohong Chen 0001, Charles Berger 0002, Reinoud Stoel, Michael Blumenstein, Bryan Found |
ICDAR | 7 |
| 2011 | An Application of the 2D Gaussian Filter for Enhancing Feature Extraction in Off-line Signature VerificationabstractSimilar to many other pattern recognition problems, feature extraction contributes significantly to the overall performance of an off-line signature verification system. To be successful, a feature extraction technique must be tolerant to different types of variation whilst preserving essential information of input patterns. In this paper, we describe a grid-based feature extraction technique that utilises directional information extracted from the signature contour, i.e. the chain code histogram. Our experimental results for signature verification indicated that, by applying a suitable 2D Gaussian filter on the matrices containing the chain code histograms, an average error rate (AER) of 13.90% can be obtained whilst maintaining the false acceptance rate (FAR) for random forgeries as low as 0.02%. These figures are comparable or better than those reported by other state of the art feature extraction techniques such as the Modified Direction Feature (MDF) and the Gradient feature. Vu Nguyen 0002, Michael Blumenstein |
ICDAR | 2 |
| 2010 | Segmentation of Inter-neurons in Three Dimensional Brain Imagery
Gervase Tuxworth, Adrian Meedeniya, Michael Blumenstein |
ACIVS (1) | 3 |
| 2010 | Techniques for static handwriting trajectory recovery: a surveyabstractOn-line handwriting recognition systems are usually better than their off-line counterparts thanks to the accessibility of dynamic information such as stroke order, velocity, acceleration, and pressure. Whilst the exact value of velocity as well as acceleration or pressure is unlikely to be recoverable, the temporal order of the strokes or the pen trajectory is shown to be more promising for recovery. The published experimental results suggest that the recovered pen trajectory information actually improves the off-line recognition accuracy. This paper presents an overview and discussion of pen trajectory recovery methods developed to date. Vu Nguyen 0002, Michael Blumenstein |
Document Analysis Systems | 2 |
| 2010 | International Conference on Frontiers in Handwriting Recognition (ICFHR 2010) - Competitions OverviewabstractThe great success and high number of participants in pattern recognition related competitions last years show an important improvement of recognition and classification approaches. This success is unconceivable without the availability of huge datasets of real world data. We have invited for proposals for competitions to be held in the framework of the 12th International Conference on Frontiers in Handwriting Recognition (ICFHR2010). These competitions should aim at evaluating the performance of algorithms and methods for a particular task of Handwriting Recognition. Eight different teams composed of more than one group have submitted their proposals. The subjects of these propositions cover the field of research of handwriting recognition from pre-processing over handwritten document analysis to handwriting text/word recognition. These competitions represent an overview of current research topics and frontiers in handwriting document analysis and recognition. Only 5 competitions have received enough participants (we have defined the threshold to 3 systems) to present their evaluation at the ICFHR 2010. This paper presents the 8 competition proposals with lists of competition organizers and lists of participating systems and approaches. Haikal El Abed, Volker Märgner, Michael Blumenstein |
ICFHR | 3 |
| 2010 | The 4NSigComp2010 Off-line Signature Verification Competition: Scenario 2abstractThe objective of this competition (4NSigComp2010) is to ascertain the performance of automatic off-line signature verifiers to evaluate recent technology developments in the areas of document analysis and machine learning. The current paper focuses on the second scenario, which aims at performance evaluation of off-line signature verification systems on a newly-created large dataset that comprises genuine, simulated signatures produced by unskilled imitators or random signatures (genuine signatures from other writers). Ten systems were evaluated, and some interesting results are presented in terms of accuracy and execution time. The top ranking system attained an overall error of 8.94%. This result interestingly correlates with the top ranking accuracy achieved in a previous signature verification competition at ICDAR 2009. Michael Blumenstein, Miguel A. Ferrer, Jesús Francisco Vargas-Bonilla |
ICFHR | 1 |
| 2010 | Performance Analysis of the Gradient Feature and the Modified Direction Feature for Off-line Signature VerificationabstractFeature extraction is an important process in off-line signature verification. In this work, the performance of two feature extraction techniques, the Modified Direction Feature (MDF) and the gradient feature are compared on the basis of similar experimental settings. In addition, the performance of Support Vector Machines (SVMs) and the squared Mahalanobis distance classifier employing the Gradient Feature are also compared and reported. Without using forgeries for training, experimental results indicated that an average error rate as low as 15.03% could be obtained using the gradient feature and SVMs. Vu Nguyen 0002, Yumiko Kawazoe, Tetsushi Wakabayashi, Umapada Pal 0001, Michael Blumenstein |
ICFHR | 5 |
| 2010 | A new system for breakzone location and the measurement of breaking wave heights and periodsabstractThis paper presents a new system for measuring breakzone locations, breaking wave height and wave periods across the surfzone from a digital video sequence. The system (Wave Pack) aims to provide real-time measurement of breaking and re-breaking wave heights and wave periods using low mounted video camera installations. Following on site data collection and analysis it was found that the Wave Pack system provides a low cost, robust, reliable and accurate system for measuring continuous wave height and period from a low elevation video camera aimed at the target beach under a wide range of wave conditions. These tests have verified the accuracy of Wave Pack in comparison to existing systems. Christopher Lane, Yaniv Gal, Matthew Browne, Andrew Short, Darrell Strauss, Rodger Tomlinson, Kathryn Jackson, Clarence Tan, Michael Blumenstein |
IGARSS | 9 |
| 2009 | Global Features for the Off-Line Signature Verification ProblemabstractGlobal features based on the boundary of a signature and its projections are described for enhancing the process of automated signature verification. The first global feature is derived from the total psilaenergypsila a writer uses to create their signature. The second feature employs information from the vertical and horizontal projections of a signature, focusing on the proportion of the distance between key strokes in the image, and the height/width of the signature. The combination of these features with the Modified Direction Feature (MDF) and the ratio feature showed promising results for the off-line signature verification problem. When being trained using 12 genuine specimens and 400 random forgeries taken from a publicly available database, the Support Vector Machine (SVM) classifier obtained an average error rate (AER) of 17.25%. The false acceptance rate (FAR) for random forgeries was also kept as low as 0.08%. Vu Nguyen 0002, Michael Blumenstein, Graham Leedham |
ICDAR | 2 |
| 2009 | Automated classification of dopaminergic neurons in the rodent brainabstractAccurate morphological characterization of the multiple neuronal classes of the brain would facilitate the elucidation of brain function and the functional changes that underlie neurological disorders such as Parkinson's diseases or Schizophrenia. Manual morphological analysis is very time-consuming and suffers from a lack of accuracy because some cell characteristics are not readily quantified. This paper presents an investigation in automating the classification of dopaminergic neurons located in the brainstem of the rodent, a region critical to the regulation of motor behaviour and is implicated in multiple neurological disorders including Parkinson's disease. Using a Carl Zeiss Axioimager Z1 microscope with Apotome, salient information was obtained from images of dopaminergic neurons using a structural feature extraction technique. A data set of 100 images of neurons was generated and a set of 17 features was used to describe their morphology. In order to identify differences between neurons, 2-dimensional and 3-dimensional image representations were analyzed. This paper compares the performance of three popular classification methods in bioimage classification (Support Vector Machines (SVMs), Back Propagation Neural Networks (BPNNs) and Multinomial Logistic Regression (MLR)), and the results show a significant difference between machine classification (with 97% accuracy) and human expert based classification (72% accuracy). Azadeh Alavi, Brenton Cavanagh, Gervase Tuxworth, Adrian Meedeniya, Alan Mackay-Sim, Michael Blumenstein |
IJCNN | 6 |
| 2009 | Handwritten Shorthand and its Future Potential for Fast Mobile Text entryabstractHandwritten shorthand systems were devised to enable writers to record information on paper at fast speeds, ideally at the speed of speech. While they have been in existence for many years it is only since the 17th Century that widespread usage appeared. Several shorthand systems flourished in the first half of the 20th century until the introduction and widespread use of electronic recording and dictation machines in the 1970's. Since then, shorthand usage has been in rapid decline, but has not yet become a lost skill. Pitman shorthand has been shown to possess unique advantages as a means of fast text entry which is particularly applicable to hand-held devices in mobile environments. This paper presents progress and critical research issues for a Pitman/Renqun Shorthand Online Recognition System. Recognition and transcription experiments are reported which indicate that a correct recognition and transcription rate of around 90% is currently possible. Graham Leedham, Michael Blumenstein |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2007 | Off-line Signature Verification Using Enhanced Modified Direction Features in Conjunction with Neural Classifiers and Support Vector MachinesabstractAs a biometric, signatures have been widely used to identify people. In the context of static image processing, the lack of dynamic information such as velocity, pressure and the direction and sequence of strokes has made the realization of accurate off-line signature verification systems more challenging as compared to their on-line counterparts. In this paper, we propose an effective method to perform off-line signature verification based on intelligent techniques. Structural features are extracted from the signature's contour using the modified direction feature (MDF) and its extended version: the Enhanced MDF (EMDF). Two neural network-based techniques and Support Vector Machines (SVMs) were investigated and compared for the process of signature verification. The classifiers were trained using genuine specimens and other randomly selected signatures taken from a publicly available database of 3840 genuine signatures from 160 volunteers and 4800 targeted forged signatures. A distinguishing error rate (DER) of 17.78% was obtained with the SVM whilst keeping the false acceptance rate for random forgeries (FARR) below 0.16%. Vu Nguyen 0002, Michael Blumenstein, Vallipuram Muthukkumarasamy, Graham Leedham |
ICDAR | 2 |
| 2007 | An investigation of the modified direction feature for cursive character recognition
Michael Blumenstein, Xin Yu Liu, Brijesh K. Verma |
Pattern Recognit. | 1 |
| 2006 | Off-line Signature Verification using the Enhanced Modified Direction Feature and Neural-based ClassificationabstractSignatures continue to be an important biometric for authenticating the identity of human beings. This paper presents an effective method to perform off-line signature verification using unique structural features extracted from the signature's contour. A novel combination of the modified direction feature (MDF) and additional distinguishing features such as the centroid, surface area, length and skew are used for classification. A resilient backpropagation (RBP) neural network and a radial basis function (RBF) network were compared in terms of verification accuracy. Using a publicly available database of 2106 signatures (936 genuine and 1170 forgeries), verification rates of 91.21% and 88.0% were obtained using RBF and RBP respectively. Stephane Armand, Michael Blumenstein, Vallipuram Muthukkumarasamy |
IJCNN | 2 |
| 2006 | Improvement of an Artificial Neural Network Model using Min-Max Preprocessing for the Prediction of Wave-induced Seabed LiquefactionabstractIn the past decade, artificial neural networks (ANNs) have been widely applied to the engineering problems with a complicated system. ANNs are becoming an important alternative option for solving problems in comparison to traditional engineering solutions, which are usually involved in complicated mathematical theories. In this study, we apply an ANN model to the wave-induced seabed liquefaction problem, which is a key issue in the area of coastal and ocean engineering. Furthermore, we adopted an ANN model with preprocessing (MIN-MAX) on difficult training data. This paper demonstrates the capacity of the proposed ANN model using MIN-MAX pre-processing to provide coastal engineers with another effective tool to analyse the stability of seabed sediment. Daeho Cha, Michael Blumenstein, Hong Zhang 0028, Dong-Sheng Jeng |
IJCNN | 2 |
| 2006 | An Exhaustive Search Strategy for Detecting Persons in Beach Scenes using Digital Video Imagery and Neural Network-based ClassificationabstractThis paper presents an investigation of a neural-based technique for detecting and quantifying persons in beach imagery for the purpose of predicting trends of tourist activities at beach sites. The proposed system uses various pre-processing and segmentation techniques to initially isolate potential objects in cluttered scenes. A structural feature extraction technique is then used to represent objects of interest for training a neural classifier. An exhaustive search strategy, incorporating a neural network, is proposed to effectively scan beach images to determine whether objects are "person" or "non-person". Encouraging results are presented for person detection using video imagery collected from a beach site on the coast of Australia. Steve Green, Michael Blumenstein |
IJCNN | 2 |
| 2006 | The Detection of Persons in Cluttered Beach Scenes Using Digital Video Imagery and Neural Network-Based ClassificationabstractThis paper presents an investigation into the detection and quantification of persons in real-world beach scenes for the automated monitoring of public recreation areas. Aside from the obvious use of video and digital imagery for surveillance applications, this research focuses on the analysis of images for the purpose of predicting trends in the intensity of public usage at beach sites in Australia. The proposed system uses image enhancement and segmentation techniques to detect objects in cluttered scenes. Following these steps, a newly proposed feature extraction technique is used to represent salient information in the extracted objects for training of a neural network. The neural classifier is used to distinguish the extracted objects between "person" and "non-person" categories to facilitate analysis of tourist activity. Encouraging results are presented for person classification on a database of real-word beach scene images. Steve Green, Michael Blumenstein, Matthew Browne, Rodger Tomlinson |
Int. J. Comput. Intell. Appl. | 2 |
| 2006 | Empirical Estimation of Nearshore Waves From a Global Deep-Water Wave ModelabstractGlobal wind-wave models such as the National Oceanic and Atmospheric Administration WaveWatch 3 (NWW3) play an important role in monitoring the world's oceans. However, untransformed data at grid points in deep water provide a poor estimate of swell characteristics at nearshore locations, which are often of significant scientific, engineering, and public interest. Explicit wave modeling, such as the Simulating Waves Nearshore (SWAN), is one method for resolving the complex wave transformations affected by bathymetry, winds, and other local factors. However, obtaining accurate bathymetry and determining parameters for such models is often difficult. When target data is available (i.e., from in situ buoys or human observers, empirical alternatives such artificial neural networks (ANNs) and linear regression may be considered for inferring nearshore conditions from offshore model output. Using a sixfold cross-validation scheme, significant wave height$H_s$and period were estimated at one onshore and two nearshore locations. In estimating$H_s$at the shoreline, the validation performance of the best ANN was$r = 0.91$, as compared to those of linear regression (0.82), SWAN (0.78), and the NWW3$H_s$baseline (0.54). Matthew Browne, Darrell Strauss, Bruno Castelle, Michael Blumenstein, Rodger Tomlinson, Christopher Lane |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2006 | Estimation of chemical oxygen demand by ultraviolet spectroscopic profiling and artificial neural networks
Shoshana Fogelman, Michael Blumenstein, Huijun Zhao |
Neural Comput. Appl. | 2 |
| 2006 | Objective Beach-State Classification From Optical Sensing of Cross-Shore Dissipation ProfilesabstractRemote sensing using terrestrial optical charge-coupled device cameras is a useful data collection method for geophysical measurement in the nearshore zone, where in situ measurement is difficult and time consuming. In particular, optical video sensing of the variability in human-visible surface refraction due to the nearshore incident wave field is becoming an established method for distal measurement of nearshore subtidal morphology. We report on the use of a low-mounted shore-normal camera for gathering data on cross-shore dissipative characteristics of a dynamic open beach. Data are analyzed for the purposes of classifying three of Wright and Shorts' intermediate classes of morphological beach state as determined by expert raters. Although these beach states are usually thought of as being distinctive in terms of their longshore bar variability, theory predicts that differences should also be observed in cross-shore dissipative characteristics. Three methods of generating features from statistical features from the archived optical data are described and compared in terms of their ability to discriminate between the beach states. Principal component scores of the percentile distributions were found to provide slightly better classification performance (i.e., 85%, while approximating the data using relatively fewer features), whereas classification using intensity distributions alone resulted in the worst performance, classifying 78% of beach states correctly. Class center moment profiles for each beach state were constructed, and results indicate that cross-shore wave dissipation becomes more disorganized as linear bars devolve into more complex transverse structures Matthew Browne, Darrell Strauss, Rodger Tomlinson, Michael Blumenstein |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2005 | The Neural-based Segmentation of Cursive Words using Enhanced HeuristicsabstractThis paper presents an enhanced heuristic segmenter (EHS) and an improved neural-based segmentation technique for segmenting cursive words and validating prospective segmentation points respectively. The EHS employs two new features, ligature detection and a neural assistant, to locate prospective segmentation points. The improved neural-based segmentation technique can then be used to examine the prospective segmentation points by fusion of confidence values obtained from left and centre character recognition outputs in addition to the segmentation point validation (SPV) output. The improved neural-based segmentation technique uses a recently proposed feature extraction technique (modified direction feature) for representing the segmentation points and characters to enhance the overall segmentation process. The EHS and the neural-based segmentation technique have been implemented and tested on a benchmark database providing encouraging results. Chun Ki Cheng, Michael Blumenstein |
ICDAR | 2 |
| 2004 | A modified direction feature for cursive character recognitionabstractThis paper describes a neural network-based technique for cursive character recognition applicable to segmentation-based word recognition systems. The proposed research builds on a novel feature extraction technique that extracts direction information from the structure of character contours. This principal is extended so that the direction information is integrated with a technique for detecting transitions between background and foreground pixels in the character image. The proposed technique is compared with the standard direction feature extraction technique, providing promising results using segmented characters from the CEDAR benchmark database. Michael Blumenstein, Xin Yu Liu, Brijesh K. Verma |
IJCNN | 1 |
| 2004 | An experimental analysis of GAME: a generic automated marking environmentabstractThis paper describes the Generic Automated Marking Environment (GAME) and provides a detailed analysis of its performance in assessing student programming projects and exercises. GAME has been designed to automatically assess programming assignments written in a variety of languages based on the "structure" of the source code and the correctness of the program's output. Currently, the system is able to mark programs written in Java, C++ and the C language. To use the system, instructors are required to provide a simple "marking schema" for any given assessment item, which includes pertinent information such as the location of files and the model solution. In this research, GAME has been tested on a number of student programming exercises and assignments. The results obtained, have been analysed and compared against a human marker providing encouraging results. Michael Blumenstein, Steve Green, Ann Nguyen, Vallipuram Muthukkumarasamy |
ITiCSE | 1 |
| 2004 | A novel approach for structural feature extraction: Contour vs. direction
Brijesh K. Verma, Michael Blumenstein, Moumita Ghosh |
Pattern Recognit. Lett. | 2 |
| 2003 | A Novel Feature Extraction Technique for the Recognition of Segmented Handwritten CharactersabstractHigh accuracy character recognition techniques can provide useful information for segmentation-based handwritten word recognition systems. This research describes neural network-based techniques for segmented character recognition that may be applied to the segmentation and recognition components of an off-line handwritten word recognition system. Two neural architectures along with two different feature extraction techniques were investigated. A novel technique for character feature extraction is discussed and compared with others in the literature. Recognition results above 80% are reported using characters automatically segmented from the CEDAR benchmark database as well as standard CEDAR alphanumerics. Michael Blumenstein, Brijesh K. Verma, H. Basli |
ICDAR | 1 |
| 2002 | Strategies for Improving a Java-Based, First Year Programming CourseabstractThis paper describes the evolution of a first year Java course at Griffith University-Gold Coast since Semester 1, 2000 to the December 2002. The course was updated to emphasise program design and to implement and evaluate an "objects-as-needed" approach to first year programming. A number of strategies were tested to increase consistency amongst teaching staff, improve delivery of course resources, successfully cater to a wide variety of students and to enhance the learning experience in general. The success of the revised course has been measured by evaluating student feedback and performance. Currently, a focus group-based strategy of evaluation is being adopted to determine students' attitudes to the most recently implemented changes. Michael Blumenstein |
ICCE | 1 |
| 2001 | Analysis of Segmentation Performance on the CEDAR Benchmark DatabaseabstractAnalyses the performance of our improved segmentation algorithm tested on the CEDAR benchmark database of handwritten words. Segmentation is achieved through the extraction of a wide range of information adjacent to or surrounding suspicious segmentation points. Initially, a heuristic technique is employed to search for structural features and to over-segment each word. For each segmentation point that is located, the left character (preceding the segmentation point) and centre character (centred on the segmentation point) are extracted along with other features from the segmentation area. The aforementioned features are presented to trained character and segmentation point validation neural networks to evaluate a number of confidence values. Finally, the confidence values are fused to obtain the final segmentation decision. Based on a detailed analysis, it was observed that the left and centre character networks increased the accuracy of the segmentation algorithm. Michael Blumenstein, Brijesh K. Verma |
ICDAR | 1 |
| 1999 | Neural-based Solutions for the Segmentation and Recognition of Difficult Handwritten Words from a Benchmark DatabaseabstractA new intelligent segmentation technique is proposed that may be used in conjunction with a neural classifier and a simple lexicon for the recognition of difficult handwritten words. A heuristic segmentation algorithm is initially used to over-segment each word. An artificial neural network (ANN) trained with 32,034 segmentation points is then used to verify the validity of the segmentation points found. Following segmentation, character matrices from each word are extracted, normalised and then passed through a global feature extractor, after which a second ANN trained with segmented characters is used for classification. These recognised characters are grouped into words and presented to a variable-length lexicon that utilises a string processing algorithm to compare and retrieve those words with the highest confidences. This research provides promising results for segmentation, character and word recognition. Michael Blumenstein, Brijesh K. Verma |
ICDAR | 1 |
| 1999 | A new segmentation algorithm for handwritten word recognitionabstractAn algorithm for segmenting unconstrained printed and cursive words is proposed. The algorithm initially oversegments handwritten word images (for training and testing) using heuristics and feature detection. An artificial neural network (ANN) is then trained with global features extracted from segmentation points found in words designated for training. Segmentation points located in "test" word images are subsequently extracted and verified using the trained ANN. Two major sets of experiments were conducted, resulting in segmentation accuracies of 75.06% and 76.52%. The handwritten words used for experimentation were taken from the CEDAR CD-ROM. The results obtained for segmentation can easily be used for comparison with other researchers using the same benchmark database. Michael Blumenstein, Brijesh K. Verma |
IJCNN | 1 |