Sukalpa Chanda

dblp:06/453 · DBLP profile ↗
← Back
43ranked-venue papers
15as first author
23since 2021 · last 2026
0000-0002-9068-5845ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 10 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 8 first-author · 15 since 2021Databases, data management, data science and information retrieval · 10 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FRETS: Frequency-Enhanced Residual Transformer System for SpO2 Estimation
Surajit Mukherjee, Shahzad Ahmad 0002, Ram Prasad Padhy, Sukalpa Chanda, Umapada Pal 0001
ICPR (8)4
2026 Difficulty-Aware Interleaved Distillation for Robust Cross-Surface Writer Identification
Kumari Priya, Chandranath Adak, Aritra Dey, Soumi Chattopadhyay, Sukalpa Chanda
ICPR (4)5
2026 ACuRE: Accurate Continuity-Regularized SpO2 Estimation Using Liquid Time-Constant Networks
abstract
Blood oxygen saturation (SpO2) is a vital measure of respiratory and circulatory health, essential for detecting hypoxemia in conditions like chronic obstructive pulmonary disease and heart failure. Current non-contact SpO2estimation methods using remote photoplethysmography (rPPG) struggle with motion artifacts, illumination variability, and limited temporal modeling, hindering their practical use. We propose ACuRE, a novel framework that integrates a two-branch 3D-ResNet-18 for AC/DC signal separation, Liquid Time-Constant (LTC) networks for continuous-time dynamics, and a physics-informed partial differential equation (PDE) loss based on mass conservation. ACuRE overcomes these challenges by isolating pulsatile (AC) and baseline (DC) signals for enhanced robustness, using LTC networks to capture nonlinear physiological dynamics, and applying PDE regularization to ensure signal continuity. This achieves a significant reduction in mean absolute error compared to baselines, with strong performance under motion and illumination stress. Evaluated across multiple datasets, ACuRE demonstrates robust accuracy and generalization, offering a scalable solution for video-based health monitoring in telemedicine and low-resource settings. Code available at: https://github.com/Shahzadnit/ACURE_WACV
Shahzad Ahmad 0002, Divya Mishra, Sania Bano, Sukalpa Chanda, Yogesh S. Rawat
WACV4
2026 Exploring the Boundaries of Diffusion Models for Offline Writer Identification with Sparse and Intra-Variable Data
abstract
Offline writer identification poses significant challenges when training data is scarce, and handwriting styles exhibit high intra-writer variability. This scenario is common in practical applications such as forensic analysis and historical document authentication, where only a limited number of handwritten samples are available per writer. In this paper, we explore the viability of using diffusion models to capture writer-specific traits under such challenging conditions. Specifically, we investigate their performance in both text-dependent and text-independent setups, where lexical similarity varies across samples. We propose a novel diffusion-based writer identification framework that integrates a style encoder and handcrafted textural features in a joint training pipeline. Our approach is evaluated on a recent dataset with high intra-writer variability as well as three benchmark datasets (IAM, CERUG-EN, and CVL). Experimental results demonstrate that while diffusion models excel in text-dependent scenarios, their generalization capability diminishes in text-independent settings due to the entanglement of content and style features. This study highlights both the promise and the current limitations of generative diffusion models for fine-grained handwriting style modeling. We identify avenues for improving generalization through disentangled representations, domain adaptation, and hybrid discriminative-generative architectures. The proposed framework contributes to the growing efforts toward scalable, style-aware writer identification in real-world, unconstrained handwriting scenarios.
Aritra Dey, Chandranath Adak, Kumari Priya, Soumi Chattopadhyay, Sukalpa Chanda
WACV5
2025 IndicSideFace: A Dataset for Advancing Deepfake Detection on Side-Face Perspectives of Indian Subjects
abstract
The rapid advancement of generative models and their misuse have made deepfake detection a crucial area of research. However, existing datasets and detection techniques predominantly focus on frontal-face perspectives, leaving sideface views largely underexplored. To bridge this gap, we present IndicSideFace, a novel dataset specifically curated for advancing deepfake detection on side-face perspectives of Indian subjects. This dataset encompasses a diverse range of side-face angles, varying lighting conditions, and demographic attributes, providing a comprehensive benchmark for evaluating detection algorithms. Our experiments using state-of-the-art models highlight the unique challenges posed by side-face deepfakes, such as partial facial feature visibility and uncommon head poses. The findings reveal significant limitations in existing detection approaches when applied to side-face perspectives, underscoring the need for specialized solutions. With IndicSideFace, we aim to strengthen the resilience of deepfake detectors and stimulate further research in this critical yet underexplored domain.
Anurag Deo, Aditya Bangar, Chandranath Adak, Rahul Verma, Deepak Nagar, Zahid Akhtar, Soumya Dutta, Soumi Chattopadhyay, Sukalpa Chanda
FG9
2025 Beyond Memorization: Training-Free Style Mixing for Variability in Handwritten Text Generation Using Writer Embedding Injection in Pretrained Diffusion Models
Aniket Gurav, Sukalpa Chanda, Narayanan Chatapuram Krishnan
ICDAR (4)2
2025 Graph Convolutional Teacher-Student Framework for Writer Inspection from Intra-variable Handwritten Words
Kumari Priya, Aritra Dey, Chandranath Adak, Soumi Chattopadhyay, Sukalpa Chanda, Simone Marinai
ICDAR (3)6
2025 TRUST: Time-Domain Residual Unsupervised Stability Technique for Improved Heart Rate Estimation
abstract
Camera-based estimation of vital signs is a promising method for non-contact health monitoring, which analyzes minute changes in video data. However, the creation of accurate models for this task is challenging due to the scarcity of datasets that possess synchronized vital sign recordings. Our research enhances an existing non-contrastive unsupervised learning technique for extracting rPPG signals, which does not necessitate ground-truth signals during the training process. We have incorporated new time-domain loss functions and added a feature stabilization block to improve the model's stability and accuracy in detecting low-level features. Additionally, we have devised a metric to evaluate the feature instability in the model's final layer. Our experiments on four public datasets demonstrate that our method surpasses the performance of current state-of-the-art methods. These advancements make our approach a significant breakthrough in the development of scalable deep-learning models for camera-based heart-rate estimation.
Shahzad Ahmad 0002, Sania Bano, Sukalpa Chanda, Santosh Kumar Vipparthi, M. Subrahmanyam 0001
WACV3
2025 PULSE: Physiological Understanding with Liquid Signal Extraction
abstract
The non-contact estimation of vital signs, particularly heart rate, from video data is a promising method for remote health monitoring. 3D convolutional layers are widely used for this task due to their ability to capture both spatial and temporal features. However, traditional 3D convolutions, while effective in many cases, lack the capacity to adjust dy-namically to the temporal variability inherent in physiological signals such as remote photoplethysmography (rPPG), which are characterized by subtle frequency changes over time. To address this, we propose PULSE (Physiological Understanding with Liquid Signal Extraction), a frame-work that employs Liquid Time-Constant (LTC) models with 3D convolutional layers to enhance temporal sensitivity and improve the extraction of these fine-grained rPPG signals. In PULSE, traditional 3D-conv layers are deployed for ini-tial feature extraction, while LTC-based 3D-conv layers dy-namically adapt and guide the temporal processing, allowing the model to better track and interpret the subtle variations in heart rate signals under different conditions, such as motion artifacts and lighting changes. We evaluated the effectiveness of PULSE in an unsupervised training setting, demonstrating that our solution performs well even in the absence of labeled datasets a common challenge in rPPG signal extraction. Experimental evaluations on three public datasets confirm that PULSE achieves comparable or supe-rior results to existing methods, proving its robustness and efficacy for real-world, non-contact health monitoring applications.
Shahzad Ahmad 0002, Sania Bano, Sachin Verma, Yogesh S. Rawat, Sukalpa Chanda, Santosh Kumar Vipparthi, M. Subrahmanyam 0001
WACV5
2025 Scene text recognition: an Indic perspective
Vasanthan P. Vijayan, Sukalpa Chanda, David S. Doermann, Narayanan Chatapuram Krishnan
Int. J. Document Anal. Recognit.2
2024 A New Impressive and Expressive Features Based Model for Personality Traits Identification
Kunal Biswas, Palaiahnakote Shivakumara, Umapada Pal 0001, Sukalpa Chanda, Xiaojun Wu 0001
ICPR (8)4
2024 A New StyleGAN Latent Space Based Model for Image Style Transfer
Rakesh Dey, Palaiahnakote Shivakumara, Saumik Bhattacharya, Sukalpa Chanda, Umapada Pal 0001
ICPR (11)4
2024 Word-Diffusion: Diffusion-Based Handwritten Text Word Image Generation
Aniket Gurav, Narayanan Chatapuram Krishnan, Sukalpa Chanda
ICPR (19)3
2024 A New Attention Based UNet and Gated Edge Attention Network for Retinal Vessel Segmentation
Ayush Roy, Palaiahnakote Shivakumara, Umapada Pal 0001, Sukalpa Chanda
ICPR (28)4
2023 Writer Identification from Nordic Historical Manuscripts using Transformer Networks
abstract
Handwriting has been used as a form of authentication for the last 1000 years. Forensic analysis of handwriting using computers has been in practice since the 1970s. With the evolution of deep-learning techniques over the last decade, such automated forensic analysis of handwritten text has become dependent on deep-learning techniques. In this paper, we investigate the prowess of transformer networks in the context of identifying the writer of a handwritten sample. We here propose a deep feature embedding-based transformer network, WiT, for writer identification. Experiments were conducted on a historical Nordic manuscript dataset comprising 9253 handwritten samples scribbled by 50 writers for the very first time, and encouraging results were obtained. Rigorous experiments were also conducted to check the noise / damage resiliency of WiT, and the outcomes were quite promising.
Chandranath Adak, Batturi Jaswanth, Zahid Akhtar, Andre Kåsen, Sukalpa Chanda
IJCB5
2023 Pho(SC)-CTC - a hybrid approach towards zero-shot word image recognition
Ravi Bhatt, Anuj Rai, Sukalpa Chanda, Narayanan Chatapuram Krishnan
Int. J. Document Anal. Recognit.3
2022 DAZeTD: Deep Analysis of Zones in Torn Documents
Chandranath Adak, Priyanshi Sharma, Sukalpa Chanda
ICFHR3
2022 GMSRF-Net: An Improved generalizability with Global Multi-Scale Residual Fusion Network for Polyp Segmentation
abstract
Colonoscopy is a gold standard procedure but is highly operator-dependent. Efforts have been made to automate the detection and segmentation of polyps, a precancerous precursor, to effectively minimize missed rate. Widely used computer-aided polyp segmentation systems actuated by encoder-decoder have achieved high performance in terms of accuracy. However, polyp segmentation datasets collected from varied centers can follow different imaging protocols leading to difference in data distribution. As a result, most methods suffer from performance drop when trained and tested on different distributions and therefore, require re-training for each specific dataset. We address this generalizability issue by proposing a global multi-scale residual fusion network (GMSRF-Net). Our proposed network maintains high-resolution representations by performing multi-scale fusion operations across all resolution scales through dense connections while preserving low-level information. To further leverage scale information, we design cross multi-scale attention (CMSA) module that uses multi-scale features to identify, keep, and propagate informative features. Additionally, we introduce multi-scale feature selection (MSFS) modules to perform channel-wise attention that gates irrelevant features gathered through global multi-scale fusion within the GMSRF-Net. The repeated fusion operations gated by CMSA and MSFS demonstrate improved generalizability of our network.Experiments conducted on two different polyp segmentation datasets show that our proposed GMSRF-Net outperforms the previous top-performing state-of-the-art method by 8.34% and 10.31% on unseen CVC-ClinicDB and on unseen Kvasir-SEG, in terms of dice coefficient. Additionally, when tested on unseen CVC-ColonDB, we surpass the state-of-the-art method by 9.38% and 4.04% in terms of dice coefficient, when source dataset is Kvasir-SEG and CVC-ClinicDB, respectively.
Sukalpa Chanda, Debesh Jha, Umapada Pal 0001, Sharib Ali
ICPR2
2022 AGA-GAN: Attribute Guided Attention Generative Adversarial Network with U-Net for face hallucination
Sukalpa Chanda, Umapada Pal 0001
Image Vis. Comput.2
2022 MSRF-Net: A Multi-Scale Residual Fusion Network for Biomedical Image Segmentation
abstract
Methods based on convolutional neural networks have improved the performance of biomedical image segmentation. However, most of these methods cannot efficiently segment objects of variable sizes and train on small and biased datasets, which are common for biomedical use cases. While methods exist that incorporate multi-scale fusion approaches to address the challenges arising with variable sizes, they usually use complex models that are more suitable for general semantic segmentation problems. In this paper, we propose a novel architecture called Multi-Scale Residual Fusion Network (MSRF-Net), which is specially designed for medical image segmentation. The proposed MSRF-Net is able to exchange multi-scale features of varying receptive fields using a Dual-Scale Dense Fusion (DSDF) block. Our DSDF block can exchange information rigorously across two different resolution scales, and our MSRF sub-network uses multiple DSDF blocks in sequence to perform multi-scale fusion. This allows the preservation of resolution, improved information flow and propagation of both high- and low-level features to obtain accurate segmentation maps. The proposed MSRF-Net allows to capture object variabilities and provides improved results on different biomedical datasets. Extensive experiments on MSRF-Net demonstrate that the proposed method outperforms the cutting-edge medical image segmentation methods on four publicly available datasets. We achieve the Dice Coefficient (DSC) of 0.9217, 0.9420, and 0.9224, 0.8824 on Kvasir-SEG, CVC-ClinicDB, 2018 Data Science Bowl dataset, and ISIC-2018 skin lesion segmentation challenge dataset respectively. We further conducted generalizability tests and achieved DSC of 0.7921 and 0.7575 on CVC-ClinicDB and Kvasir-SEG, respectively.
Debesh Jha, Sukalpa Chanda, Umapada Pal 0001, Håvard D. Johansen, Dag Johansen, Michael Riegler 0001, Sharib Ali, Pål Halvorsen
IEEE J. Biomed. Health Informatics3
2021 LoOp: Looking for Optimal Hard Negative Embeddings for Deep Metric Learning
abstract
Deep metric learning has been effectively used to learn distance metrics for different visual tasks like image retrieval, clustering, etc. In order to aid the training process, existing methods either use a hard mining strategy to extract the most informative samples or seek to generate hard synthetics using an additional network. Such approaches face different challenges and can lead to biased embeddings in the former case, and (i) harder optimization (ii) slower training speed (iii) higher model complexity in the latter case. In order to overcome these challenges, we propose a novel approach that looks for optimal hard negatives (LoOp) in the embedding space, taking full advantage of each tuple by calculating the minimum distance between a pair of positives and a pair of negatives. Unlike mining-based methods, our approach considers the entire space between pairs of embeddings to calculate the optimal hard negatives. Extensive experiments combining our approach and representative metric learning losses reveal a significant boost in performance on three benchmark datasets1.
Bhavya Vasudeva, Puneesh Deora, Saumik Bhattacharya, Umapada Pal 0001, Sukalpa Chanda
ICCV5
2021 DCINN: Deformable Convolution and Inception Based Neural Network for Tattoo Text Detection Through Skin Region
Tamal Chowdhury, Palaiahnakote Shivakumara, Umapada Pal 0001, Tong Lu 0002, Ramachandra Raghavendra, Sukalpa Chanda
ICDAR (2)6
2021 Pho(SC)Net: An Approach Towards Zero-Shot Word Image Recognition in Historical Documents
Anuj Rai, Narayanan Chatapuram Krishnan, Sukalpa Chanda
ICDAR (1)3
2020 Recognizing Bengali Word Images - A Zero-Shot Learning Perspective
abstract
Zero-Shot Learning(ZSL) techniques could classify a completely unseen class, which it has never seen before during training. Thus, making it more apt for any real-life classification problem, where it is not possible to train a system with annotated data for all possible class types. This work investigates recognition of word images written in Bengali Script in a ZSL framework. The proposed approach performs Zero-Shot word recognition by coupling deep learned features procured from various CNN architectures along with 13 basic shapes/stroke primitives commonly observed in Bengali script characters. As per the notion of ZSL framework those 13 basic shapes are termed as “Signature/Semantic Attributes”. The obtained results are promising while evaluation was carried out in a Five-Fold cross-validation setup dealing with samples from 250 word classes.
Sukalpa Chanda, Daniel Haitink, Prashant Kumar Prasad, Jochem Baas, Umapada Pal 0001, Lambert Schomaker
ICPR1
2018 Deep Learning for Classification and as Tapped-Feature Generator in Medieval Word-Image Recognition
abstract
Historical manuscripts are the main source of information about past. In recent years, digitization of large quantities of historical handwritten documents is in vogue. This trend gives access to a plethora of information about our medieval past. Such digital archives can be more useful if automatic indexing and retrieval of document images can be provided to the end users of a digital library. An automatic transcription of the full digital archive using traditional Optical Character Recognition (OCR) is still not possible with sufficient accuracy. If full transcription is not available, the end users are interested in indexing and retrieving of particular document pages of their interest. Hence recognition of certain keywords from within the corpus will be sufficient to meet the end users needs. Recently, deep-learning based methods have shown competence in image classification problems. However, one bottleneck with deep-learning based techniques is that it requires a huge amount of training samples per class. Since the number of samples per word class is scarce for collections that are freshly scanned, this is a serious hindrance for direct usage of the deep-learning technique for the purpose of word image recognition in historical document images. This paper aims to investigate the problem of recognizing words from historical document images using a deep-learning based framework for feature extraction and classification while countering the problem of the low amount of image samples using off-line data augmentation techniques. Encouraging results (highest accuracy of 90.03%) were obtained while dealing with 365 different word classes.
Sukalpa Chanda, Emmanuel Okafor, Sébastien Hamel, Dominique Stutzmann, Lambert Schomaker
DAS1
2018 Zero-Shot Learning Based Approach For Medieval Word Recognition using Deep-Learned Features
abstract
Historical manuscripts reflect our past. Recently digitization of large quantities of historical handwritten documents is taking place in every corner of the world, and are being archived. From those digital repositories, automatic text indexing and retrieval system fetch only those documents to an end user that they are interested in. A regular OCR technology is not capable of rendering this service to an end user in a reliable manner. Instead, a word recognition/spotting algorithm performs the task. Word recognition based systems require enough labelled data per class to train the system. Moreover, all word classes need to be taught beforehand. Though word spotting could evade this drawback of prior training, these systems often need to have additional overheads like a language model to deal with "out of lexicon" words. Zero-shot learning could be a possible alternative to counter such situation. A Zero-shot learning algorithm is capable of handling unseen classes, provided the algorithm has been fortified with rich discriminating features and reliable "attribute description" per class during training. Since deeply learned features have enough discriminating power, a deep learning framework has been used here for feature extraction purpose. To the best of our knowledge, this is probably the first work on "out of lexicon" medieval word recognition using a Zero-Shot Learning framework. We obtained very encouraging results(accuracy ≈57% for "out of lexicon" classes) while dealing with 166 training classes and 50 unseen test classes.
Sukalpa Chanda, Jochem Baas, Daniel Haitink, Sébastien Hamel, Dominique Stutzmann, Lambert Schomaker
ICFHR1
2018 Static and Dynamic Synthesis of Bengali and Devanagari Signatures
abstract
Developing an automatic signature verification system is challenging and demands a large number of training samples. This is why synthetic handwriting generation is an emerging topic in document image analysis. Some handwriting synthesizers use the motor equivalence model, the well-established hypothesis from neuroscience, which analyses how a human being accomplishes movement. Specifically, a motor equivalence model divides human actions into two steps: 1) the effector independent step at cognitive level and 2) the effector dependent step at motor level. In fact, recent work reports the successful application to Western scripts of a handwriting synthesizer, based on this theory. This paper aims to adapt this scheme for the generation of synthetic signatures in two Indic scripts, Bengali (Bangla), and Devanagari (Hindi). For this purpose, we use two different online and offline databases for both Bengali and Devanagari signatures. This paper reports an effective synthesizer for static and dynamic signatures written in Devanagari or Bengali scripts. We obtain promising results with artificially generated signatures in terms of appearance and performance when we compare the results with those for real signatures.
Miguel A. Ferrer, Sukalpa Chanda, Moisés Díaz Cabrera, Chayan Kumar Banerjee, Anirban Majumdar 0002, Cristina Carmona-Duarte, Parikshit Acharya, Umapada Pal 0001
IEEE Trans. Cybern.2
2016 Multiple Generation of Bengali Static Signatures
abstract
Handwritten signature datasets are really necessary for the purpose of developing and training automatic signature verification systems. It is desired that all samples in a signature dataset should exhibit both inter-personal and intra-personal variability. A possibility to model this reality seems to be obtained through the synthesis of signatures. In this paper we propose a method based on motor equivalence model theory to generate static Bengali signatures. This theory divides the human action to write mainly into cognitive and motor levels. Due to difference between scripts, we have redesigned our previous synthesizer [1,2], which generates static Western signatures. The experiments assess whether this method can approach the intra and inter-personal variability of the Bengali-100 Static Signature DB from a performance-based validation. The similarities reported in the experimental results proof the ability of the synthesizer to generate signature images in this script.
Moisés Díaz Cabrera, Sukalpa Chanda, Miguel A. Ferrer, Chayan Kumar Banerjee, Anirban Majumdar 0002, Cristina Carmona-Duarte, Parikshit Acharya, Umapada Pal 0001
ICFHR2
2014 Offline Hand-Written Musical Symbol Recognition
abstract
Recognition of offline musical symbols can aid in automatic retrieval of a particular piece of musical notation from a digital repository. Though some work on on-line Musical symbol notations exists, little work has been done on off-line recognition of the symbols. This article proposes a system for offline isolated musical symbol recognition. Efficacy of a texture analysis based feature extraction method is compared with a structural shape descriptor based feature extraction method coupled with a Support Vector Machine (SVM) classifier. Later three different kinds of feature selection techniques were also analyzed to gauge the contribution of each feature in the overall classification process. We compared our results with an existing method and we noted the proposed system exhibited encouraging results and it is better than existing method. The proposed system even worked better when we used MQDF classifier in place of SVM. In a five-fold cross validation experimental framework, considering 3795 music symbols we achieved 97.50% and 98.05% accuracy from SVM and MQDF classifiers, respectively when chain-code histogram features are applied.
Sukalpa Chanda, Debleena Das, Umapada Pal 0001, Fumitaka Kimura
ICFHR1
2014 ℓp-norm multiple kernel learning with low-rank kernels
Alain Rakotomamonjy, Sukalpa Chanda
Neurocomputing2
2013 Word-Wise Script Identification from Video Frames
abstract
Script identification is an essential step for the efficient use of the appropriate OCR in multilingual document images. There are various techniques available for script identification from printed and handwritten document images, but script identification from video frames has not been explored much. This paper presents a study of some pre-processing techniques and features for word-wise script identification from video frames. Traditional features, namely Zernike moments, Gabor and gradient, have performed well for handwritten and printed documents having simple backgrounds and adequate resolution for OCR. Video frames are mostly coloured and suffer from low resolution, blur, background noise, to mention a few. In this paper, an attempt has been made to explore whether the traditional script identification techniques can be useful in video frames. Three feature extraction techniques, namely Zernike moments, Gabor and gradient features, and SVM classifiers were considered for analyzing three popular scripts, namely English, Bengali and Hindi. Some pre-processing techniques such as super resolution and skeletonization of the original word images were used in order to overcome the inherent problems with video. Experiments show that the super resolution technique with gradient features has performed well, and an accuracy of 87.5% was achieved when testing on 896 words from three different scripts. The study also reveals that the use of proper pre-processing approaches can be helpful in applying traditional script identification techniques to video frames.
Nabin Sharma, Sukalpa Chanda, Umapada Pal 0001, Michael Blumenstein
ICDAR2
2012 Text Independent Writer Identification for Oriya Script
abstract
Automatic identification of an individual based on his/her handwriting characteristics is an important forensic tool. In a computational forensic scenario, presence of huge amount of text/information in a questioned document cannot be ensured. Lack of data threatens system reliability in such cases. We here propose a writer identification system for Oriya script which is capable of performing reasonably well even with small amount of text. Experiments with curvature feature are reported here, using Support Vector Machine (SVM) as classifier. We got promising results of 94.00% writer identification accuracy at first top choice and 99% when considering first three top choices.
Sukalpa Chanda, Katrin Franke, Umapada Pal 0001
Document Analysis Systems1
2012 Similar shaped part-based character recognition using G-SURF
abstract
Classification/misclassification of similar shaped characters largely affects OCR accuracy. Sometimes occlusion/insertion of a part of character (due to inferior scanning quality) also makes it look alike another character type. For such adverse situations a part based character recognition system could be more effective. In order to encounter mentioned adverse scenario we propose a new feature encoding technique. This feature encoding is based on the amalgamation of Gabor filter-based features with SURF features (G-SURF). Features generated from a character are provided to Support Vector Machine (SVM) classifier. We obtained an encouraging accuracy on similar shaped characters from three different scripts.
Sukalpa Chanda, Umapada Pal 0001, Katrin Franke
HIS1
2012 Font identification - In context of an Indic script
Sukalpa Chanda, Umapada Pal 0001, Katrin Franke
ICPR1
2012 Off-line signature verification using G-SURF
abstract
In the field of biometric authentication, automatic signature identification and verification has been a strong research area because of the social and legal acceptance and extensive use of the written signature as an easy method for authentication. Signature verification is a process in which the questioned signature is examined in detail in order to determine whether it belongs to the claimed person or not. Signatures provide a secure means for confirmation and authorization in legal documents. So nowadays, signature identification and verification becomes an essential component in automating the rapid processing of documents containing embedded signatures. Sometimes, part-based signature verification can be useful when a questioned signature has lost its original shape due to inferior scanning quality. In order to address the above-mentioned adverse scenario, we propose a new feature encoding technique. This feature encoding is based on the amalgamation of Gabor filter-based features with SURF features (G-SURF). Features generated from a signature are applied to a Support Vector Machine (SVM) classifier. For experimentation, 1500 (50×30) forgeries and 1200 (50×24) genuine signatures from the GPDS signature database were used. A verification accuracy of 97.05% was obtained from the experiments.
Srikanta Pal, Sukalpa Chanda, Umapada Pal 0001, Katrin Franke, Michael Blumenstein
ISDA2
2011 Identification of Indic Scripts on Torn-Documents
abstract
Questioned Document Examination processes often encompass analysis of torn documents. To aid a forensic expert, automatic classification of content type in torn documents might be useful. This helps a forensic expert to sort out similar document fragments from a pile of torn documents. One parameter of similarity could be the script of the text. In this article we propose a method to identify the script in document fragments. Torn documents are normally characterized by text with arbitrary orientation. We use Zernike moment - based feature that is rotation invariant together with Support Vector Machine (SVM) to classify the script type. Subsequently gradient features are used for comparative analysis of results between rotation dependent and rotation invariant feature type. We achieved an overall script-identification accuracy of 81.39% when dealing with 11 different scripts at character/connected-component level and 94.65% at word level.
Sukalpa Chanda, Katrin Franke, Umapada Pal 0001
ICDAR1
2010 Document-Zone Classification in Torn Documents
abstract
Arbitrary orientation and sparse data content are common characteristics of torn document. To ensure accuracy and reliability in computer-based analysis, content-zone segmentation is required. In our previous work, we studied segmentation of handwritten and printed text. A questioned document-piece in the form of an office note, however, might also contain non-text data like logos, graphics, and pictures. Hence a more precise content-zone classification is required. In this paper we propose a two-tier approach for non-text, handwriting and printed text segmentation. The first tier aims to discriminate text and non-text regions. The second tier classifies handwritten and printed text within all text zones identified during the first tier. Gabor features and chain-code features are used in Tier-1 and Tier-2, respectively. By using SVM classifier we successfully identified 97.65% of 31,227 text regions in our current test data. The proposed approach identified 98.69% of printed and 96.39% of handwritten text amongst all identified text regions.
Sukalpa Chanda, Katrin Franke, Umapada Pal 0001
ICFHR1
2010 Text Independent Writer Identification for Bengali Script
abstract
Automatic identification of an individual based on his/her handwriting characteristics is an important forensic tool. In a computational forensic scenario, presence of huge amount of text/information in a questioned document cannot be always ensured. Also, compromising in terms of systems reliability under such situation is not desirable. We here propose a system to encounter such adverse situation in the context of Bengali script. Experiments with discrete directional feature and gradient feature are reported here, along with Support Vector Machine (SVM) as classifier. We got promising results of 95.19% writer identification accuracy at first top choice and 99.03% when considering first three top choices.
Sukalpa Chanda, Katrin Franke, Umapada Pal 0001, Tetsushi Wakabayashi
ICPR1
2010 Script Identification - A Han and Roman Script Perspective
abstract
All Han-based scripts (Chinese, Japanese, and Korean) possess similar visual characteristics. Hence system development for identification of Chinese, Japanese and Korean scripts from a single document page is quite challenging. It is noted that a Han-based document page might also have Roman script in them. A multi-script OCR system dealing with Chinese, Japanese, Korean, and Roman scripts, demands identification of scripts before execution of respective OCR modules. We propose a system to address this problem using directional features along with a Gaussian Kernel-based Support Vector Machine. We got promising results of 98.39% script identification accuracy at character level and 99.85% at block level, when no rejection was considered.
Sukalpa Chanda, Umapada Pal 0001, Katrin Franke, Fumitaka Kimura
ICPR1
2009 Two-stage Approach for Word-wise Script Identification
abstract
A two-stage approach for word-wise identification of English (Roman), Devnagari and Bengali (Bangla) scripts is proposed. This approach balances the tradeoff between recognition accuracy and processing speed. The 1st stage allows identifying scripts with high speed, yet less accuracy when dealing with noisy data. The advanced 2nd stage processes only those samples that yield low recognition confidence in the first stage. For both stages a rough character segmentation is performed and features are computed on segmented character components. Features used in the 1st stage are a 64-dimensional chain-code-histogram feature, while 400-dimensional gradient features are used in the 2nd stage. Final classification of a word to a particular script is done via majority voting of each recognized character component of the word. Extensive experiments with various confidence scores were conducted and reported here. The overall recognition accuracy and speed is remarkable. Correct classification of 98.51% on 11,123 test words is achieved, even when the recognition-confidence is as high as 95% at both stages.
Sukalpa Chanda, Srikanta Pal, Katrin Franke, Umapada Pal 0001
ICDAR1
2009 Word-Wise Thai and Roman Script Identification
abstract
In some Thai documents, a single text line of a printed document page may contain words of both Thai and Roman scripts. For the Optical Character Recognition (OCR) of such a document page it is better to identify, at first, Thai and Roman script portions and then to use individual OCR systems of the respective scripts on these identified portions. In this article, an SVM-based method is proposed for identification of word-wise printed Roman and Thai scripts from a single line of a document page. Here, at first, the document is segmented into lines and then lines are segmented into character groups (words). In the proposed scheme, we identify the script of a character group combining different character features obtained from structural shape, profile behavior, component overlapping information, topological properties, and water reservoir concept, etc. Based on the experiment on 10,000 data (words) we obtained 99.62% script identification accuracy from the proposed scheme.
Sukalpa Chanda, Umapada Pal 0001, Oriol Ramos Terrades
ACM Trans. Asian Lang. Inf. Process.1
2008 Word-wise Sinhala Tamil and English script identification using Gaussian kernel SVM
abstract
There are many documents in Srilanka where a single document page may contain Sinhala, Tamil and English texts. For OCR development of such a document page, it is better to identify different scripts present in the page and then feed the identified portion to the respective OCR module. In this paper, a SVM based technique is proposed for word-wise identification of Sinhala, Tamil and English scripts from a single document page. Structural features, topological features and water reservoir principle based features are mainly used here for the purpose. From the experiment we obtained encouraging results.
Sukalpa Chanda, Srikanta Pal, Umapada Pal 0001
ICPR1
2007 SVM Based Scheme for Thai and English Script Identification
abstract
In some Thai documents, a single text line of a document page may contain both Thai and English scripts. For the optical character recognition (OCR) of such a document page it is better to identify, at first, Thai and English script portions and then to use individual OCR system of the respective scripts on these identified portions. In this paper, a SVM based method is proposed for identification of word-wise printed English and Thai scripts from a single line of a document page. Here, at first, the document is segmented into lines and then lines are segmented into character groups (words). In the proposed scheme, we identify the script of the individual character group combining different character features obtained from structural shape, profile, component overlapping information, topological properties, water reservoir concept etc. Based on the experiment on 6110 data we obtained 99.36% script identification accuracy from the proposed scheme.
Sukalpa Chanda, Oriol Ramos Terrades, Umapada Pal 0001
ICDAR1