Imran Siddiqi

dblp:09/3932 · DBLP profile ↗
← Back
54ranked-venue papers
5as first author
11since 2021 · last 2024
0000-0002-7203-5195ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 15 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2024 A deep learning framework for historical manuscripts writer identification using data-driven features
Akram Bennour, Merouane Boudraa, Imran Siddiqi, Mohammed Al-Sarem, Mohammad Al-Shabi, Fahad Ghabban
Multim. Tools Appl.3
2024 Correction to: A deep learning framework for historical manuscripts writer identification using data-driven features
Akram Bennour, Merouane Boudraa, Imran Siddiqi, Mohammed Al-Sarem, Mohammad Al-Shabi, Fahad Ghabban
Multim. Tools Appl.3
2023 LSTM-based Siamese neural network for Urdu news story segmentation
Muhammad Nauman Ahmed Bhatti, Imran Siddiqi, Momina Moetesum
Int. J. Document Anal. Recognit.2
2022 Co-clustering based classification of multi-view data
Syed Fawad Hussain, Imran Siddiqi
Appl. Intell.3
2022 Document forgery detection using source printer identification: A comparative study of text-dependent versus text-independent analysis
abstract
Abstract Source printer identification represents an interesting modality for document forgery detection. Establishing the identity of the printer that was employed to print a questioned document allows concluding its authenticity. This paper investigates the effectiveness of deep visual features (learned using convolutional neural networks) in characterization of the source printer. Images of printed documents are divided into small patches as well as characters for extraction of features. An off‐the‐shelf recognition engine is also integrated, allowing experiments in text‐dependent as well as text‐independent modes. Experiments are carried out on a standard data set of documents from 20 different printers and identification rates of 95.52% and 98.06% are reported using patches and characters, respectively. Furthermore, the discriminating power of different characters, as well as their combinations, is also being studied. Unlike many existing techniques, which rely on pre‐segmented characters and report results by comparing same characters only, the proposed technique works on complete images of printed documents and reports high identification rates.
Maryam Bibi, Anmol Hamid, Momina Moetesum, Imran Siddiqi
Expert Syst. J. Knowl. Eng.4
2022 Feature learning and encoding for multi-script writer identification
Abdelillah Semma, Yaâcoub Hannad, Imran Siddiqi, Said Lazrak, Mohamed El Youssfi El Kettani
Int. J. Document Anal. Recognit.3
2022 A survey of visual and procedural handwriting analysis for neuropsychological assessment
abstract
Abstract To date, Artificial Intelligence systems for handwriting and drawing analysis have primarily targeted domains such as writer identification and sketch recognition. Conversely, the automatic characterization of graphomotor patterns asbiomarkersof brain health is a relatively less explored research area. Despite its importance, the work done in this direction is limited and sporadic. This paper aims to provide a survey of related work to provide guidance to novice researchers and highlight relevant study contributions. The literature has been grouped into “visual analysis techniques” and “procedural analysis techniques”. Visual analysis techniques evaluate offline samples of a graphomotor response after completion. On the other hand, procedural analysis techniques focus on the dynamic processes involved in producing a graphomotor reaction. Since the primary goal of both families of strategies is to represent domain knowledge effectively, the paper also outlines the commonly employed handwriting representation and estimation methods presented in the literature and discusses their strengths and weaknesses. It also highlights the limitations of existing processes and the challenges commonly faced when designing such systems. High-level directions for further research conclude the paper.
Momina Moetesum, Moisés Díaz Cabrera, Uzma Masroor, Imran Siddiqi, Gennaro Vessio
Neural Comput. Appl.4
2021 Two-Step Fine-Tuned Convolutional Neural Networks for Multi-label Classification of Children's Drawings
Muhammad Osama Zeeshan, Imran Siddiqi, Momina Moetesum
ICDAR (2)2
2021 Influence of codebook patterns on writer recognition: An experimental study
abstract
Abstract Codebook‐based writer characterization is an effective technique that has been investigated in a number of recent studies on identification and verification of writers. These methods divide a set of writing samples into small units (fragments or graphemes) and cluster these patterns to produce a codebook. Writer of a handwritten sample is then characterized by the probability (distribution) of producing the codebook patterns. In most cases, a small subset of the database under study is employed to produce the codebook while the rest of the database is used in evaluations. This work aims to validate the hypothesis that the codebook simply serves as a representation space to compare different writings and, in most cases, the patterns in the codebook do not significantly influence the identification and verification performance. The hypothesis is validated by generating a number of codebooks using Greek, Arabic and Chinese handwritten samples. Moreover, codebooks using fragments of handwritten music scores, printed text and synthetic data are also investigated. Evaluations on three well‐known handwriting databases (CVL, BFL and IAM) validate the idea that, in general, the codebook patterns do not have a significant impact on characterizing writer from handwriting.
Chawki Djeddi, Imran Siddiqi, Abdeljalil Gattal, Somaya Al-Máadeed, Abdellatif Ennaji
Expert Syst. J. Knowl. Eng.2
2021 Sequence-based dynamic handwriting analysis for Parkinson's disease detection with one-dimensional convolutions and BiGRUs
Moisés Díaz Cabrera, Momina Moetesum, Imran Siddiqi, Gennaro Vessio
Expert Syst. Appl.3
2021 Writer Identification using Deep Learning with FAST Keypoints and Harris corner detector
Abdelillah Semma, Yaâcoub Hannad, Imran Siddiqi, Chawki Djeddi, Mohamed El Youssfi El Kettani
Expert Syst. Appl.3
2020 Dynamic Handwriting Analysis for Parkinson's Disease Identification using C-BiGRU Model
abstract
Parkinson's disease (PD) is commonly characterized by several motor impairments like tremor, muscular rigidity and bradykinesia, that are collectively termed as `Parkinson's disease dysgraphia'. In an attempt to identify these motor-based Parkinsonian symptoms, experts have persistently been evaluating various dynamic attributes of handwriting, like pen pressure/position, stroke speed/trajectory, and on-surface/in-air time taken, captured with the help of online acquisition tools. Such devices not only capture various aspects of handwriting but provide rich sequential information that can be utilized to identify unique patterns from handwriting samples of PD patients. In this paper, we propose a model based on Bidirectional Gated Recurrent Units (BiGRU) to assess the potential of handwriting-based sequential information in the identification of Parkinsonian symptoms. One-dimensional convolution is applied to raw sequences and the resulting feature sequences are employed to train the BiGRU model for prediction. The results of our experiments validate the potential of our proposed technique in comparison to the state-of-the-art.
Momina Moetesum, Imran Siddiqi, Farah Javed, Uzma Masroor
ICFHR2
2020 Recognition of cursive video text using a deep learning framework
abstract
This study focuses on cursive text recognition appearing in videos, using a complete framework of deep neural networks. While mature video optical character recognition systems (V‐OCRs) are available for text in non‐cursive scripts, recognition of cursive scripts is marked by many challenges. These include complex and overlapping ligatures, context‐dependent shape variations and presence of a large number of dots and diacritics. The authors present an analytical technique for recognition of cursive caption text that relies on a combination of convolutional and recurrent neural networks trained in an end‐to‐end framework. Text lines extracted from video frames are preprocessed to segment the background and are fed to a convolutional neural network for feature extraction. The extracted feature sequences are fed to different variants of bi‐directional recurrent neural networks along with the ground truth transcription to learn sequence‐to‐sequence mapping. Finally, a connectionist temporal classification layer is employed to produce the final transcription. Experiments on a data set of more than 40,000 text lines from 11,192 video frames of various News channel videos reported an overall character recognition rate of 97.63%. The proposed work employs Urdu text as a case study but the findings can be generalised to other cursive scripts as well.
Ali Mirza, Imran Siddiqi
IET Image Process.2
2020 Exploring nested ensemble learners using overproduction and choose approach for churn prediction in telecom industry
Mahreen Ahmed, Hammad Afzal, Imran Siddiqi, M. Faisal Amjad, Khawar Khurshid
Neural Comput. Appl.3
2020 Deformation modeling and classification using deep convolutional neural networks for computerized analysis of neuropsychological drawings
Momina Moetesum, Imran Siddiqi, Shoaib Ehsan, Nicole Vincent
Neural Comput. Appl.2
2019 Deep Learning Based Approach for Historical Manuscript Dating
abstract
Digitization of historical manuscripts from premodern eras, has captivated the document analysis and pattern recognition community in recent years. Estimation of the period of production of such documents is a challenging yet favored research problem. In this paper, we present a deep learning based approach to effectively characterize the year of production of sample documents from the Medieval Paleographical Scale (MPS) dataset. By employing transfer learning on a number of popular pre-trained Convolutional Neural Network (CNN) models, we have significantly reduced the Mean Absolute Error (MAE) reported in previous studies.
Anmol Hamid, Maryam Bibi, Momina Moetesum, Imran Siddiqi
ICDAR4
2019 Deformation Classification of Drawings for Assessment of Visual-Motor Perceptual Maturity
abstract
Sketches and drawings are popularly employed in clinical psychology to assess the visual-motor and perceptual development in children and adolescents. Drawn responses by subjects are mostly characterized by high degree of deformations that indicates presence of various visual, perceptual and motor disorders. Classification of deformations is a challenging task due to complex and extensive rule representation. In this study, we propose a novel technique to model clinical manifestations using Deep Convolutional Neural Networks (DCNNs). Drawn responses of nine templates used for assessment of perceptual orientation of individuals are employed as training samples. A number of defined deviations scored in each template are then modeled by applying fine tuning on a pre-trained DCNN architecture. Performance of the proposed technique is evaluated on samples of 106 children. Results of experiments show that pre-trained DCNNs can model and classify a number of deformations across multiple shapes with considerable success. Nevertheless some deformations are represented more reliably than the others. Overall promising classification results are observed that substantiate the effectiveness of our proposed technique.
Momina Moetesum, Imran Siddiqi, Nicole Vincent
ICDAR2
2019 Assessing visual attributes of handwriting for prediction of neurological disorders - A case study on Parkinson's disease
Momina Moetesum, Imran Siddiqi, Nicole Vincent, Florence Cloppet
Pattern Recognit. Lett.2
2018 ICFHR 2018 Competition on Multi-Script Writer Identification
abstract
This paper describes the ICFHR 2018 Competition on Multi-script Writer Identification with details on the competition tasks, databases employed, submitted systems, evaluation protocol and the reported results. The competition was aimed at exploring the traditional writer identification problem in a more challenging scenario of a multi-script environment where training and test samples of writers come from different scripts. Three different databases with handwriting samples in Arabic, French, English, Chinese and Farsi were employed in the six competition tasks. The realized results indicate that while high identification rates are reported in the literature by traditional writer identification systems, identifying writers in a multiscript environment is a much more challenging problem that requires significant investigations to extract effective handwriting representations that are able to characterize the writer across different scripts.
Chawki Djeddi, Somaya Al-Máadeed, Imran Siddiqi, Abdeljalil Gattal, Younes Akbari
ICFHR3
2018 Writer Identification on Historical Documents using Oriented Basic Image Features
abstract
This study addresses the problem of identifying the authorship of historical manuscripts, a challenging task that offers interesting applications for document examiners and paleographers. We exploit handwriting texture as the discriminative attribute characterizing the writer of a given document. The textural information in handwriting is captured using a combination of oriented Basic Image Features (oBIFs) at different scales. Classification is carried out using a number of distance metrics which are combined to arrive at a final decision. A comprehensive series of experiments is carried out using different configurations of the oBIFs and the realized classification rates are compared with the state-of-the-art techniques on this problem
Abdeljalil Gattal, Chawki Djeddi, Imran Siddiqi, Somaya Al-Máadeed
ICFHR3
2018 Data Driven Feature Extraction for Gender Classification using Multi-Script Handwritten Texts
abstract
This paper presents a study on assessing the effectiveness of machine learned features to predict gender of writers from images of handwriting. Pre-trained Convolutional Neural Networks have been employed as feature extractors to discriminate male and female handwriting while classification is carried out using a number of classifiers, Linear Discriminant Analysis (LDA) being the most effective. Feature extraction is carried out by changing the scale of observation using word, patch and page images. Experiments are carried out on English and Arabic handwriting samples of the QUWI database and the realized results demonstrate the effectiveness of machine learned features in predicting gender from handwriting.
Momina Moetesum, Imran Siddiqi, Chawki Djeddi, Yaâcoub Hannad, Somaya Al-Máadeed
ICFHR2
2018 Towards the design of an offline signature verifier based on a small number of genuine samples for training
Walid Bouamra, Chawki Djeddi, Brahim Nini, Moisés Díaz Cabrera, Imran Siddiqi
Expert Syst. Appl.5
2018 Gender classification from offline multi-script handwriting images using oriented Basic Image Features (oBIFs)
Abdeljalil Gattal, Chawki Djeddi, Imran Siddiqi, Youcef Chibani
Expert Syst. Appl.3
2018 Editorial
abstract
It is a matter of immense pleasure to introduce this special issue comprising extended versions of selected papers from the first Mediterranean Conference on Pattern Recognition and Artificial Inte...
Chawki Djeddi, Imran Siddiqi
J. Exp. Theor. Artif. Intell.2
2018 A novel database for automatic processing of Persian handwritten bank checks
Younes Akbari, Mohammad J. Jalili, Javad Sadri, Kazem Nouri, Imran Siddiqi, Chawki Djeddi
Pattern Recognit.5
2017 Classification of Graphomotor Impressions Using Convolutional Neural Networks: An Application to Automated Neuro-Psychological Screening Tests
abstract
Graphomotor impressions are a product of complex cognitive, perceptual and motor skills and are widely used as psychometric tools for the diagnosis of a variety of neuro-psychological disorders. Apparent deformations in these responses are quantified as errors and are used are indicators of various conditions. Contrary to conventional assessment methods where manual analysis of impressions is carried out by trained clinicians, an automated scoring system is marked by several challenges. Prior to analysis, such computerized systems need to extract and recognize individual shapes drawn by subjects on a sheet of paper as an important pre-processing step. The aim of this study is to apply deep learning methods to recognize visual structures of interest produced by subjects. Experiments on figures of Bender Gestalt Test (BGT), a screening test for visuo-spatial and visuo-constructive disorders, produced by 120 subjects, demonstrate that deep feature representation brings significant improvements over classical approaches. The study is intended to be extended to discriminate coherent visual structures between produced figures and expected prototypes.
Haris Bin Nazar, Momina Moetesum, Shoaib Ehsan, Imran Siddiqi, Khurram Khurshid, Nicole Vincent, Klaus D. McDonald-Maier
ICDAR4
2017 Improving handwriting based gender classification using ensemble classifiers
Mahreen Ahmed, Asma Ghulam Rasool, Hammad Afzal, Imran Siddiqi
Expert Syst. Appl.4
2017 Urdu Nastaliq recognition using convolutional-recursive deep learning
Saeeda Naz, Arif Iqbal Umar, Riaz Ahmad 0001, Imran Siddiqi, Saad Bin Ahmed, Muhammad Imran Razzak, Faisal Shafait
Neurocomputing4
2017 Wavelet-based gender detection on off-line handwritten documents using probabilistic finite state automata
Younes Akbari, Kazem Nouri, Javad Sadri, Chawki Djeddi, Imran Siddiqi
Image Vis. Comput.5
2016 Isolated Handwritten Digit Recognition Using oBIFs and Background Features
abstract
This study demonstrates how the combination of oriented Basic Image Features (oBIFs) with the background concavity features can be effectively employed to enhance the performance of isolated digit recognition systems. The features are extracted without any size normalization from the complete image as well as from different regions of the image by applying a uniform grid sampling to the image. Classification is carried out using one-against-all support vector machine (SVM) while the experimental study is conducted on the standard CVL single digit database. A series of evaluations using different feature configurations and combinations realized high recognition rates which are compared with the state-of-the-art methods on this subject.
Abdeljalil Gattal, Chawki Djeddi, Youcef Chibani, Imran Siddiqi
DAS4
2016 ICFHR2016 Competition on Multi-script Writer Demographics Classification Using "QUWI" Database
abstract
This competition is aimed at classification of writer demographics from offline handwritten documents using the QUWI database. QUWI is a bilingual database comprising writing samples of same individuals in Arabic and English. This allows evaluating the performance of different systems in a more challenging multi-script environment. This paper presents the details of the competition tasks, the datasets used in each of the tasks, a brief description of the participating systems, experimental protocol and evaluation criteria and finally the overall rankings of the participants.
Chawki Djeddi, Somaya Al-Máadeed, Abdeljalil Gattal, Imran Siddiqi, Abdellatif Ennaji, Haikal El Abed
ICFHR4
2016 Gender Classification from Offline Handwriting Images Using Textural Features
abstract
Prediction of gender and other demographic attributes of individuals from handwriting samples offers an interesting basic, as well as applied research problem. The correlation between gender and the visual appearance of handwriting has been validated by a number of studies and the present study is based on the same idea. We exploit the textural measurements as the discriminating attribute between male and female writings. The textural information in a writing is captured by applying a bank of Gabor filters to the image of handwriting. The mean and standard deviation values of the filter responses are collected in matrix and the Fourier transform of the matrix is used as a feature. Classification is carried out using a feed forward neural network. The proposed technique evaluated on a subset of the QUWI database realized promising results under different experimental settings.
Ali Mirza, Momina Moetesum, Imran Siddiqi, Chawki Djeddi
ICFHR3
2016 Writer identification using texture descriptors of handwritten fragments
Yaâcoub Hannad, Imran Siddiqi, Mohamed El Youssfi El Kettani
Expert Syst. Appl.2
2016 Offline cursive Urdu-Nastaliq script recognition using multidimensional recurrent neural networks
Saeeda Naz, Arif Iqbal Umar, Riaz Ahmad 0001, Saad Bin Ahmed, Syed Hamad Shirazi, Imran Siddiqi, Muhammad Imran Razzak
Neurocomputing6
2015 ICDAR2015 competition on Multi-script Writer Identification and Gender Classification using 'QUWI' Database
abstract
This competition targets writer identification and gender classification from offline handwritten documents using the QUWI database. The most interesting aspect of the competition is the use of a dataset with writing samples of the same individual in Arabic as well as English. The competition not only allows an objective comparison of different systems but also permits to investigate the performance of traditional script-dependent systems in a multi-script experimental setup. This paper describes the competition details including the competition tasks, the database employed, the methods used by the participating systems, evaluation and ranking criteria and the overall rankings of the participants. The competition received a total of 13 submissions from 8 different institutions. Writer identification tasks received 5 while the gender classification tasks received 8 submissions.
Chawki Djeddi, Somaya Al-Máadeed, Abdeljalil Gattal, Imran Siddiqi, Labiba Souici-Meslati, Haikal El Abed
ICDAR4
2015 Recognition of Urdu ligatures - a holistic approach
abstract
This paper presents an effective segmentation-free and scale-invariant technique for recognition of Urdu ligatures in Nastaliq font. The proposed technique relies on separating the main body of ligatures from the secondary components and training a separate hidden Markov model for each. Features capturing projection, concavity and curvature information of ligatures are extracted using right-to-left sliding windows and are fed to the models for training. The system trained and evaluated on a total of more than 2,000 frequently occurring Urdu ligatures from a standard database realized a recognition rate of 97.93%.
Israr Ud Din, Imran Siddiqi, Shehzad Khalid, Chawki Djeddi
ICDAR2
2015 Automated scoring of Bender Gestalt Test using image analysis techniques
abstract
Drawing tests have been long used by practitioners and researchers for early detection of psychological and neurological impairments. These tests allow subjects to naturally express themselves as opposed to an interview or a written assessment. Bender Gestalt Test (BGT) is a well-known and established neurological test designed to detect signs of perceptual distortions. Subjects are shown a number of geometric patterns for reconstruction and assessments are made by observing properties like rotation, angulations, simplification and closure difficulty. The manual scoring of the test, however, is a time consuming and lengthy procedure especially when a large number of subjects is to be analyzed. This paper proposes the application of image analysis techniques to automatically score a subset of hand drawn images in the BGT test. A comparison of the scores reported by the automated system with those assigned by the psychologists not only reveals the effectiveness of the proposed system but also reflects the huge research potential this area possesses.
Momina Moetesum, Imran Siddiqi, Uzma Masroor, Chawki Djeddi
ICDAR2
2015 Automatic analysis of handwriting for gender classification
Imran Siddiqi, Chawki Djeddi, Ahsen Raza, Labiba Souici-Meslati
Pattern Anal. Appl.1
2014 Evaluation of Texture Features for Offline Arabic Writer Identification
abstract
Biometric identification of persons has mainly been based on fingerprints, face, iris and other similar attributes. We propose a handwriting-based biometric identification system using a large database of Arabic handwritten documents. The system first extracts, from each handwritten sample, a set of features including run lengths, edge-hinge and edge-direction features. These features are used by a Multiclass SVM (Support Vector Machine) classifier. Experiments are conducted on a new large database of Arabic handwritings contributed by 1000 writers. The highest identification rate achieved by the combination of run-length and edge-hinge features stands at 84.10%.
Chawki Djeddi, Labiba Souici-Meslati, Imran Siddiqi, Abdellatif Ennaji, Haikal El Abed, Abdeljalil Gattal
Document Analysis Systems3
2014 LAMIS-MSHD: A Multi-script Offline Handwriting Database
abstract
This paper introduces a new offline handwriting database that was developed to be employed in performance evaluation, result comparison and development of new methods related to handwriting analysis and recognition. The database can particularly be used for signature verification, writer recognition and writer demographics classification. In addition, the database also supports isolated digit recognition, digit/text segmentation and recognition and similar related tasks. The database comprises 600 Arabic and 600 French text samples, 1300 signatures and 21,000 digits. 100 Algerian individuals coming from different age groups and educational backgrounds contributed to the development of database by providing a total of 1300 forms. The database is also accompanied with ground truth data supporting the evaluation of the aforementioned tasks. The main contribution of the database is providing a multi-script platform where same authors contributed samples in French and Arabic. It would be interesting to explore applications like writer recognition and writer demographics classification in a multi-script environment.
Chawki Djeddi, Abdeljalil Gattal, Labiba Souici-Meslati, Imran Siddiqi, Youcef Chibani, Haikal El Abed
ICFHR4
2014 Improving Isolated Digit Recognition Using a Combination of Multiple Features
abstract
This paper investigates the combination of different statistical and structural features for recognition of isolated handwritten digits, a classical pattern recognition problem. The objective of this study is to improve the recognition rates by combining different representations of non-normalized handwritten digits. These features include some global statistics, moments, profile and projection based features and features computed from the contour and skeleton of the digits. Some of these features are extracted from the complete image of digit while others are extracted from different regions of the image by first applying a uniform grid sampling to the image. Classification is carried out using one-against-all SVM. The experiments conducted on the CVL Single Digit Database realized high recognition rates which are comparable to state-of-the-art methods on this subject.
Abdeljalil Gattal, Youcef Chibani, Chawki Djeddi, Imran Siddiqi
ICFHR4
2013 Codebook for Writer Characterization: A Vocabulary of Patterns or a Mere Representation Space?
abstract
Codebook-based representations have been effectively employed for writer identification. Most of the codebook-based methods generate a codebook by clustering a set of patterns extracted from an independent data set. The probability of occurrence of the codebook patterns in a given writing is then used to characterize its author. This study investigates the hypothesis that the codebook is merely a representation space and the codebook patterns themselves do not affect the writer identification performance. The idea is validated by first using codebooks in different scripts from those of writings in question and then by using a synthetically generated codebook. A number of data sets with handwritten samples in Arabic, French, English, German, Urdu and Greek are considered in our series of evaluations. Experiments conducted with different codebooks report interesting results which validate the ideas put forward in this study.
Chawki Djeddi, Imran Siddiqi, Labiba Souici-Meslati, Abdellatif Ennaji
ICDAR2
2013 Multilingual Artificial Text Detection Using a Cascade of Transforms
abstract
This paper presents a method for multilingual artificial text detection and extraction from still images. The proposed detection scheme relies on a cascade of spatial transforms followed by a box counting based fractal dimension approach to exploit the self-similar redundancy of patterns in the shapes of characters in the text. The detected text regions are validated using GLCM based features and are segmented from the background using the proposed binarization scheme. The proposed method is evaluated on five data sets containing textual occurrences in Urdu, English, Chinese, Arabic and Hindi. The experimental results realized show very promising precision and recall rates which are also consistent across different data sets.
Ahsen Raza, Imran Siddiqi, Chawki Djeddi, Abdellatif Ennaji
ICDAR2
2013 Improving Codebook-Based Writer Recognition
abstract
This paper presents an effective method for writer recognition from offline handwritten documents by extending the idea of codebook-based writer recognition. The study is based on generating two codebooks, primary and secondary. The writer of a handwritten document is characterized by the probability of occurrence of codebook patterns in his/her writing. The main contribution of this study is the proposition of a secondary codebook to capture information on connecting strokes in addition to the main strokes. The writing is divided into small windows and, for each window, four small adjacent windows are considered. The patterns in the main and adjacent windows are clustered separately to generate two codebooks. The proposed method evaluated on 650 writers of the IAM database reports an identification rate of 96% and validates the idea that complementing the primary codebook with a secondary codebook serves to enhance the recognition rates.
Muhammed Jehanzeb, Ghazali Sulong, Imran Siddiqi
Int. J. Pattern Recognit. Artif. Intell.3
2013 Text-independent writer recognition using multi-script handwritten texts
Chawki Djeddi, Imran Siddiqi, Labiba Souici-Meslati, Abdellatif Ennaji
Pattern Recognit. Lett.2
2012 Word Spotting Based Retrieval of Urdu Handwritten Documents
abstract
Urdu being one of the most popular languages adopted during different swatches of history has a valuable collection of handwritten scripts in different state libraries of South Asia. Digitizing these collections can serve not only to preserve them but also to make them available to general public. Non existence of an Urdu OCR, however, limits the concept of a digital Urdu library to scanning and manual search of documents only. We present a word spotting based search method for Urdu handwritten text. The text is first segmented into partial words and a set of features is computed from each partial word. The user queries the system using word image. The partial words in the query image are then matched with those in the database and the matched partial words are merged into complete words. The proposed method evaluated on 90 handwritten documents reported encouraging precision and recall rates.
Ali Abidi, Akhtar Jamil, Imran Siddiqi, Khurram Khurshid
ICFHR3
2012 Multi-script Writer Identification Optimized with Retrieval Mechanism
abstract
Identifying the writer of a handwritten document has been an active research area over the last few years with applications in biometrics, forensics, smart meeting rooms and historical document analysis. In this paper, we present a new writer identification system based on a retrieval mechanism. Texture based edge-hinge and run-length features are used to characterize the writing style of an individual. The effectiveness of the proposed system is evaluated on a total of 1583 writing samples in Arabic, German, English, French, and Greek from two different databases. The experimental evaluations reveal that reducing the search space using a writer retrieval mechanism prior to identification improves the identification rates.
Chawki Djeddi, Imran Siddiqi, Labiba Souici-Meslati, Abdellatif Ennaji
ICFHR2
2012 An Unconstrained Benchmark Urdu Handwritten Sentence Database with Automatic Line Segmentation
abstract
In this paper we present and announce a novel off-line sentence database of Urdu handwritten documents along with a few preprocessing and text line segmentation procedures. Despite an increased research interest in Urdu handwritten document analysis over the recent years, a standard benchmark dataset, which could be used in Urdu handwriting recognition tasks, has been missing. Based on our own developed and updated corpus named CENIP-UCCP (Center for Image Processing-Urdu Corpus Construction Project), we have developed an Urdu handwritten database. The corpus is a collection of a variety of Urdu texts that were used to generate forms. These forms were subsequently filled by native writers in their natural handwritings. Six categories of text were used to generate these forms with each category using approximately 66 forms. Up till now, the database comprises 400 digitized forms produced by 200 different writers. The database is completely labeled for content information as well as content detection and supports the evaluation of systems like Urdu handwriting recognition, line segmentation and writer identification. The database was also experimented with the proposed Urdu text line segmentation scheme rendering promising segmentation results.
Ahsen Raza, Imran Siddiqi, Ali Abidi, Fahim Arif
ICFHR2
2011 Towards Searchable Digital Urdu Libraries - A Word Spotting Based Retrieval Approach
abstract
Libraries in South Asia hold huge collections of valuable printed documents in Urdu and it is of interest to digitize these collections to make them more accessible. The unavailability of an OCR for Urdu however limits the concept of a digital Urdu library to scanning of documents only, offering very limited search facility based on manually assigned tags. We address this issue by proposing a word spotting based keyword search method for information retrieval in digitized collections of printed Urdu documents. The proposed method is based on segmentation of Urdu text in to partial words and representing each partial word by a set of features. To search a specific word (or phrase), the user provides a query in the form of an image. Comparing the features of the partial words in the query image with the ones already indexed, the user is provided with a list of documents containing occurrences of the queried word. The system evaluated on 50 Urdu documents exhibited a recall of 95.17% and a precision of 94.3%.
Ali Abidi, Imran Siddiqi, Khurram Khurshid
ICDAR2
2011 Edge-Based Features for Localization of Artificial Urdu Text in Video Images
abstract
Content-based video indexing and retrieval has become an interesting research area with the tremendous growth in the amount of digital media. In addition to the audio-visual content, text appearing in videos can serve as a powerful tool for semantic indexing and retrieval of videos. This paper proposes a method based on edge-features for horizontally aligned artificial Urdu text detection from video images. The system exploits edge based segmentation to extract textual content from videos. We first find the vertical gradients in the input video image and average the gradient magnitude in a fixed neighborhood of each pixel. The resulting image is binarized and the horizontal run length smoothing algorithm (RLSA) is applied to merge possible text regions. An edge density filter is then applied to eliminate noisy non-text regions. Finally, the candidate regions satisfying certain geometrical constraints are accepted as text regions. The proposed approach evaluated on a data set of 150 video images exhibited promising results.
Akhtar Jamil, Imran Siddiqi, Fahim Arif, Ahsen Raza
ICDAR2
2010 Text independent writer recognition using redundant writing patterns with contour-based orientation and curvature features
Imran Siddiqi, Nicole Vincent
Pattern Recognit.1
2009 Combining Contour Based Orientation and Curvature Features for Writer Recognition
Imran Siddiqi, Nicole Vincent
CAIP1
2009 A Set of Chain Code Based Features for Writer Recognition
abstract
This communication presents an effective method for writer recognition in handwritten documents. We have introduced a set of features that are extracted from the contours of handwritten images at different observation levels. At the global level, we extract the histograms of the chain code, the first and second order differential chain codes and, the histogram of the curvature indices at each point of the contour of handwriting. At the local level, the handwritten text is divided into a large number of small adaptive windows and within each window the contribution of each of the eight directions (and their differentials) is counted in the corresponding histograms. Two writings are then compared by computing the distances between their respective histograms. The system trained and tested on two different data sets of 650 and 225 writers respectively, exhibited promising results on writer identification and verification.
Imran Siddiqi, Nicole Vincent
ICDAR1
2007 Writer Identification in Handwritten Documents
abstract
This work presents an effective method for writer identification in handwritten documents. We have developed a local approach, based on the extraction of characteristics that are specific to a writer. To exploit the existence of redundant patterns within a handwriting, the writing is divided into a large number of small sub-images, and the sub-images that are morphologically similar are grouped together in the same classes. The patterns, which occur frequently for a writer are thus extracted. The author of the unknown document is then identified by a Bayesian classifier. The system trained and tested on 50 documents of the same number of authors, reported an identification rate of 94%.
Imran Siddiqi, Nicole Vincent
ICDAR1