Chandranath Adak

dblp:138/2490 · DBLP profile ↗
← Back
35ranked-venue papers
17as first author
21since 2021 · last 2026
0000-0002-9085-2770ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 10 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 9 · 6 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Explainability-Guided Deepfake Detection for High-Fidelity Facial Edits
Bibek Das, Soumi Chattopadhyay, Chandranath Adak, Astitva Pandey, Ashutosh Parihar, Zahid Akhtar, Soumya Dutta, Abdenour Hadid
ICPR (4)3
2026 Diffusion-Latent Invisible Watermarking for Proactive Deepfake Provenance Verification
Bibek Das, Anurag Deo, Chandranath Adak, Soumi Chattopadhyay, Zahid Akhtar, Soumya Dutta, Abdenour Hadid
ICPR (4)3
2026 Difficulty-Aware Interleaved Distillation for Robust Cross-Surface Writer Identification
Kumari Priya, Chandranath Adak, Aritra Dey, Soumi Chattopadhyay, Sukalpa Chanda
ICPR (4)2
2026 Exploring the Boundaries of Diffusion Models for Offline Writer Identification with Sparse and Intra-Variable Data
abstract
Offline writer identification poses significant challenges when training data is scarce, and handwriting styles exhibit high intra-writer variability. This scenario is common in practical applications such as forensic analysis and historical document authentication, where only a limited number of handwritten samples are available per writer. In this paper, we explore the viability of using diffusion models to capture writer-specific traits under such challenging conditions. Specifically, we investigate their performance in both text-dependent and text-independent setups, where lexical similarity varies across samples. We propose a novel diffusion-based writer identification framework that integrates a style encoder and handcrafted textural features in a joint training pipeline. Our approach is evaluated on a recent dataset with high intra-writer variability as well as three benchmark datasets (IAM, CERUG-EN, and CVL). Experimental results demonstrate that while diffusion models excel in text-dependent scenarios, their generalization capability diminishes in text-independent settings due to the entanglement of content and style features. This study highlights both the promise and the current limitations of generative diffusion models for fine-grained handwriting style modeling. We identify avenues for improving generalization through disentangled representations, domain adaptation, and hybrid discriminative-generative architectures. The proposed framework contributes to the growing efforts toward scalable, style-aware writer identification in real-world, unconstrained handwriting scenarios.
Aritra Dey, Chandranath Adak, Kumari Priya, Soumi Chattopadhyay, Sukalpa Chanda
WACV2
2026 Inspecting Offline Handwritten Signature Intra-Variation Over Time: An Empirical Study
abstract
Handwritten signatures serve as crucial personal identifiers and have been extensively used for authentication purposes for a long time in the human race. Signatures exhibit substantial variations influenced by factors such as mood, time, writing speed, and the writing tool used. Understanding the variations that occur in offline handwritten signatures over time for an individual is paramount in forensic investigations, biometric systems, and legal document authentication, even in this digital era. This article presents an empirical study focused on inspecting the intra-variation of offline handwritten signatures over an extended period. To conduct this study comprehensively, we collected an extensive dataset comprising 6400 signature samples from 100 distinct writers scribbled intermittently over several months, providing a rich and diverse set of signatures for analysis. To analyze the collected data effectively, we ensembled some contemporary convolutional models to train a deep architecture. Our experimental results were quite encouraging and may shed light on the tendencies and patterns in signature changes by providing valuable insights for biometric system design and forensic experts.
Kumari Priya, Shivam Anand, Chandranath Adak
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2026 Anomaly-Resilient Temporal QoS Prediction Using Hypergraph Convoluted Transformer Network
abstract
Quality-of-Service (QoS) prediction is a critical task in the service lifecycle, enabling precise and adaptive service recommendations by anticipating performance variations over time in response to evolving network uncertainties and user preferences. However, contemporary QoS prediction methods frequently encounter data sparsity and cold-start issues, which hinder accurate QoS predictions and limit the ability to capture diverse user preferences. Additionally, these methods often assume QoS data reliability, neglecting potential credibility issues such as outliers and the presence of greysheep users and services with atypical invocation patterns. Furthermore, traditional approaches fail to leverage diverse features, including domain-specific knowledge and complex higher-order patterns, essential for accurate QoS predictions. In this paper, we introduce a real-time, trust-aware framework for temporal QoS prediction to address the aforementioned challenges, featuring an end-to- end deep architecture called the Hypergraph Convoluted Transformer Network (HCTN). HCTN combines a hypergraph structure with graph convolution over hyper-edges to effectively address high-sparsity issues by capturing complex, high-order correlations. Complementing this, the transformer network utilizes multi-head attention along with parallel 1D convolutional layers and fully connected dense blocks to capture both fine-grained and coarse-grained dynamic patterns. Additionally, our approach includes a sparsity-resilient solution for detecting greysheep users and services, incorporating their unique characteristics to improve prediction accuracy. Trained with a robust loss function resistant to outliers, HCTN demonstrated state-of-the-art performance on the large-scale WSDREAM-2 datasets for response time and throughput.
Soumi Chattopadhyay, Chandranath Adak
IEEE Trans. Netw. Serv. Manag.3
2025 IndicSideFace: A Dataset for Advancing Deepfake Detection on Side-Face Perspectives of Indian Subjects
abstract
The rapid advancement of generative models and their misuse have made deepfake detection a crucial area of research. However, existing datasets and detection techniques predominantly focus on frontal-face perspectives, leaving sideface views largely underexplored. To bridge this gap, we present IndicSideFace, a novel dataset specifically curated for advancing deepfake detection on side-face perspectives of Indian subjects. This dataset encompasses a diverse range of side-face angles, varying lighting conditions, and demographic attributes, providing a comprehensive benchmark for evaluating detection algorithms. Our experiments using state-of-the-art models highlight the unique challenges posed by side-face deepfakes, such as partial facial feature visibility and uncommon head poses. The findings reveal significant limitations in existing detection approaches when applied to side-face perspectives, underscoring the need for specialized solutions. With IndicSideFace, we aim to strengthen the resilience of deepfake detectors and stimulate further research in this critical yet underexplored domain.
Anurag Deo, Aditya Bangar, Chandranath Adak, Rahul Verma, Deepak Nagar, Zahid Akhtar, Soumya Dutta, Soumi Chattopadhyay, Sukalpa Chanda
FG3
2025 Graph Convolutional Teacher-Student Framework for Writer Inspection from Intra-variable Handwritten Words
Kumari Priya, Aritra Dey, Chandranath Adak, Soumi Chattopadhyay, Sukalpa Chanda, Simone Marinai
ICDAR (3)4
2025 Impact of Visual Context on Noisy Multimodal NMT: An Empirical Study for English to Indian Languages
abstract
Neural Machine Translation (NMT) has made remarkable progress using large-scale textual data, but the potential of incorporating multimodal inputs, especially visual information, remains underexplored in high-resource settings. While prior research has focused on using multimodal data in low-resource scenarios, this study examines how image features impact translation when added to a large-scale, pre-trained unimodal NMT system. Surprisingly, the study finds that images might be redundant in this context. Additionally, the research introduces synthetic noise to assess whether images help the model handle textual noise. Multimodal models slightly outperform text-only models in noisy settings, even when random images are used. The study’s experiments translate from English to Hindi, Bengali, and Malayalam, significantly outperforming state-of-the-art benchmarks. Interestingly, the effect of visual context varies with the level of source text noise: no visual context works best for non-noisy translations, cropped image features are optimal for low noise, and full image features perform better in high-noise scenarios. This sheds light on the role of visual context, especially in noisy settings, and opens up a new research direction for Noisy Neural Machine Translation in multimodal setups. The research emphasizes the importance of combining visual and textual information to improve translation across various environments. Our code is publicly available at https://github.com/babangain/indicMMT .
Baban Gain, Dibyanayan Bandyopadhyay, Samrat Mukherjee, Chandranath Adak, Asif Ekbal
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2025 Demystifying Visual Features of Movie Posters for Multilabel Genre Identification
abstract
In the film industry, movie posters have been an essential part of advertising and marketing for many decades and continue to play a vital role even today in the form of digital posters through online, social media, and over-the-top (OTT) platforms. Typically, movie posters can effectively promote and communicate the essence of a film, such as its genre, visual style/tone, vibe, and storyline cue/theme, which are essential to attract potential viewers. Identifying the genres of a movie often has significant practical applications in recommending the film to target audiences. Previous studies on genre identification have primarily focused on sources such as plot synopses, subtitles, metadata, movie scenes, and trailer videos; however, posters precede the availability of these sources and provide prerelease implicit information to generate mass interest. In this article, we work for automated multilabel movie genre identification only from poster images, without any aid of additional textual/metadata/video information about movies, which is one of the earliest attempts of its kind. Here, we present a deep transformer network with a probabilistic module to identify the movie genres exclusively from the poster. For experiments, we procured 13882 number of posters of 13 genres from the Internet movie database (IMDb), where our model performances were encouraging and even outperformed some major contemporary architectures.
Utsav Kumar Nareti, Chandranath Adak, Soumi Chattopadhyay
IEEE Trans. Comput. Soc. Syst.2
2024 Mu2STS: A Multitask Multimodal Sarcasm-Humor-Differential Teacher-Student Model for Sarcastic Meme Detection
Gitanjali Kumari, Chandranath Adak, Asif Ekbal
ECIR (3)2
2024 Handwriting Intra-Variability Across Surface Transitions: Implications for Writer Identification
Kumari Priya, Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein
ICPR (20)2
2024 Detecting severity of Diabetic Retinopathy from fundus images: A transformer network-based review
Tejas Karkera, Chandranath Adak, Soumi Chattopadhyay
Neurocomputing2
2024 TPMCF: Temporal QoS Prediction Using Multi-Source Collaborative Features
abstract
The e-commerce industry has seen significant growth in recent years due to the introduction of new web service APIs. Quality-of-Service (QoS) parameters, which are fundamental for assessing service performance, have become crucial in evaluating services in the competitive market. Since QoS parameters can vary among users and change over time, accurate QoS predictions have become essential for users when selecting the most suitable services. Existing methods for predicting temporal QoS have hardly achieved the desired accuracy, beset by challenges like data sparsity, the presence of anomalies, and the inability to capture intricate temporal user-service interactions. Although some recent approaches, particularly those founded on recurrent neural network-based sequential architectures, endeavor to model temporal relationships in QoS data, they grapple with performance degradation due to the omission of other pivotal features, such as collaborative relationships and spatial characteristics of users and services. Furthermore, the uniform attention among features across all time-steps can thwart progress in predictive accuracy. This paper addresses these challenges and proffers a scalable strategy for temporal QoS prediction using multi-source collaborative features that not only furnishes heightened responsiveness but also engenders enhanced prediction accuracy. The method amalgamates collaborative features stemming from both users and services, capitalizing on the user-service relationship. Additionally, it integrates spatio-temporal auto-extracted features through the orchestration of graph convolution and a specialized variant of the transformer encoder equipped with multi-head self-attention. The proposed approach has been validated on the WSDREAM-2 benchmark datasets, and the results of these extensive experiments demonstrate that our framework surpasses major state-of-the-art methods in terms of predictive accuracy, all the while upholding robust scalability and reasonable responsiveness.
Soumi Chattopadhyay, Chandranath Adak
IEEE Trans. Netw. Serv. Manag.3
2023 Writer Identification from Nordic Historical Manuscripts using Transformer Networks
abstract
Handwriting has been used as a form of authentication for the last 1000 years. Forensic analysis of handwriting using computers has been in practice since the 1970s. With the evolution of deep-learning techniques over the last decade, such automated forensic analysis of handwritten text has become dependent on deep-learning techniques. In this paper, we investigate the prowess of transformer networks in the context of identifying the writer of a handwritten sample. We here propose a deep feature embedding-based transformer network, WiT, for writer identification. Experiments were conducted on a historical Nordic manuscript dataset comprising 9253 handwritten samples scribbled by 50 writers for the very first time, and encouraging results were obtained. Rigorous experiments were also conducted to check the noise / damage resiliency of WiT, and the outcomes were quite promising.
Chandranath Adak, Batturi Jaswanth, Zahid Akhtar, Andre Kåsen, Sukalpa Chanda
IJCB1
2022 DAZeTD: Deep Analysis of Zones in Torn Documents
Chandranath Adak, Priyanshi Sharma, Sukalpa Chanda
ICFHR1
2022 OffDQ: An Offline Deep Learning Framework for QoS Prediction
abstract
With the increasing trend of web services over the Internet, developing a robust Quality of Service (QoS) prediction algorithm for recommending services in real-time is becoming a challenge today. Designing an efficient QoS prediction algorithm achieving high accuracy, while supporting faster prediction to enable the algorithm to be integrated into a real-time system, is one of the primary focuses in the domain of Services Computing. The major state-of-the-art QoS prediction methods are yet to efficiently meet both criteria simultaneously, possibly due to the lack of analysis of challenges involved in designing the prediction algorithm. In this paper, we systematically analyze the various challenges associated with the QoS prediction algorithm and propose solution strategies to overcome the challenges, and thereby propose a novel offline framework using deep neural architectures for QoS prediction to achieve our goals. Our framework, on the one hand, handles the sparsity of the dataset, captures the non-linear relationship among data, figures out the correlation between users and services to achieve desirable prediction accuracy. On the other hand, our framework being an offline prediction strategy enables faster responsiveness. We performed extensive experiments on the publicly available WS-DREAM dataset to show the trade-off between prediction performance and prediction time. Furthermore, we observed our framework significantly improved one of the parameters (prediction accuracy or responsiveness) without considerably compromising the other as compared to the state-of-the-art methods.
Soumi Chattopadhyay, Richik Chanda, Chandranath Adak
WWW4
2022 Unsupervised Anomaly Detection for Surface Defects With Dual-Siamese Network
abstract
Unsupervised anomaly detection in real industrial scenarios is challenging since the small amount of defect-free images contain limited discriminative information, and anomaly defects are unpredictable. Although nowadays image reconstruction-based methods are widely being used in various anomaly detection applications, they cannot effectively learn semantic representation, which leads to imperfect reconstruction. In this article, anomaly detection is formulated as a joint problem of feature reconstruction and inpainting in the dual-siamese framework. The proposed approach forces the network to model the feature distribution from the normal area and capture the semantic context for discriminating normal and abnormal areas. It first uses a Siamese architecture to capture discriminative features of defect-free samples and its corresponding defective samples generated by the defect random generation module. A dense feature fusion module is then employed to obtain the dense feature representation of dual input. The second Siamese network is proposed to reconstruct and inpaint the dual-dense features of the previous stage. Compared to the existing methods that mostly employ single image reconstruction, it is beneficial to simultaneously reconstruct and inpaint the information of dense discriminative features. The experimental results on the MVTec AD datasets and some major real industrial datasets demonstrate that our method achieves state-of-the-art inspection accuracy.
Xian Tao, Wenzhi Ma, Zhanxin Hou, Zhenfeng Lu, Chandranath Adak
IEEE Trans. Ind. Informatics6
2022 CAHPHF: Context-Aware Hierarchical QoS Prediction With Hybrid Filtering
abstract
With the proliferation of Internet-of-Things and continuous growth in the number of web services at the Internet-scale, the service recommendation is becoming a challenge nowadays. One of the prime aspects influencing the service recommendation is the Quality-of-Service (QoS) parameter, which depicts the performance of a web service. In general, the service provider furnishes the value of the QoS parameters before service deployment. However, in reality, the QoS values of service vary across different users, time, locations, etc. Therefore, estimating the QoS value of service before its execution is an important task, and thus, the QoS prediction has gained significant research attention. Multiple approaches are available in the literature for predicting service QoS. However, these approaches are yet to reach the desired accuracy level. In this article, we study the QoS prediction problem across different users, and propose a novel solution by taking into account the contextual (more specifically, location) information of both services and users. Our proposal includes two key steps: (a) hybrid filtering, and (b) hierarchical prediction mechanism. On the one hand, the hybrid filtering aims to obtain a set of similar users and services, given a target user and a service. On the other hand, the goal of the hierarchical prediction mechanism is to estimate the QoS value accurately by leveraging hierarchical neural-regression. We evaluated our framework on the publicly available WS-DREAM datasets. The experimental results show the outperformance of our framework over the major state-of-the-art approaches.
Ranjana Roy Chowdhury, Soumi Chattopadhyay, Chandranath Adak
IEEE Trans. Serv. Comput.3
2021 Text-line-up: Don't Worry About the Caret
Chandranath Adak, Bidyut B. Chaudhuri, Chin-Teng Lin, Michael Blumenstein
ICDAR (3)1
2021 BigyaPAn: Deep Analysis of Old Paper Advertisement
abstract
In this paper, we work on analyzing old paper advertisement (Ad). An Ad usually contains various types of textual and non-textual objects, which may also be in different orientations. We attempt to detect such objects from an early Indian-print paper Ad database comprising 1500 Ad images. The past major object detectors did not perform well on this database. We propose a deep reinforcement learning-based orientation-aware object detector. Our system learns by itself where to look and what to look of an Ad image. Therefore, it can bypass the impeding zone due to degraded image quality. To find the looking spot, we come up with a foveal transformation. In reinforcement learning, we present a scheme for shaping an internal reward with a top-up. For oriented object detection, we also propose a generic loss function. Our system obtained encouraging results from the experiments performed on the Ad database.
Chandranath Adak, Xian Tao
IJCNN1
2020 Why Not? Tell us the Reason for Writer Dissimilarity
abstract
Writer verification has drawn significant attention over the past few decades due to its extensive applications in forensics and biometrics. In traditional writer verification, handwriting similarity/dissimilarity analysis is mostly performed by extracting two feature vectors from two respective handwritten samples, followed by comparing them in relation to their similarity. In the state-of-the-art writer verification approaches, a distance metric is usually employed in terms of the similarity between two handwritten samples. If the distance between two handwritten samples is greater than a given threshold, then the samples are assumed to be written by two different writers, otherwise, they are considered to be due to the same writer. In this paper, for the very first time, we propose a model that generates English sentences to explain reasons for writer dissimilarity/similarity. First, our proposed model obtains features from handwritten images by employing a convolutional neural network, verifies the writer using a Siamese architecture, and generates English words using a recurrent neural network. Finally, these two networks are merged using an affine transformation to produce an explanatory sentence in support of writer similarity/dissimilarity. We evaluated our model on a handwritten numeral database of 100 writers and obtained promising results.
Chandranath Adak, Bidyut B. Chaudhuri, Chin-Teng Lin, Michael Blumenstein
IJCNN1
2020 Intra-Variable Handwriting Inspection Reinforced With Idiosyncrasy Analysis
abstract
In this paper, we work on intra-variable handwriting, where the writing samples of an individual can vary significantly. Such within-writer variation throws a challenge for automatic writer inspection, where the state-of-the-art methods do not perform well. To deal with intra-variability, we analyze the idiosyncrasy in individual handwriting. We identify/verify the writer from highly idiosyncratic text-patches. Such patches are detected using a deep recurrent reinforcement learning-based architecture. An idiosyncratic score is assigned to every patch, which is predicted by employing deep regression analysis. For writer identification, we propose a deep neural architecture, which makes the final decision by the idiosyncratic score-induced weighted average of patch-based decisions. For writer verification, we propose two algorithms for patch-fed deep feature aggregation, which assist in authentication using a triplet network. The experiments were performed on two databases, where we obtained encouraging results.
Chandranath Adak, Bidyut B. Chaudhuri, Chin-Teng Lin, Michael Blumenstein
IEEE Trans. Inf. Forensics Secur.1
2019 Detecting Named Entities in Unstructured Bengali Manuscript Images
abstract
In this paper, we undertake a task to find named entities directly from unstructured handwritten document images without any intermediate text/character recognition. Here, we do not receive any assistance from natural language processing. Therefore, it becomes more challenging to detect the named entities. We work on Bengali script which brings some additional hurdles due to its own unique script characteristics. Here, we propose a new deep neural network-based architecture to extract the latent features from a text image. The embedding is then fed to a BLSTM (Bidirectional Long Short-Term Memory) layer. After that, the attention mechanism is adapted to an approach for named entity detection. We perform experimentation on two publicly-available offline handwriting repositories containing 420 Bengali handwritten pages in total. The experimental outcome of our system is quite impressive as it attains 95.43% balanced accuracy on overall named entity detection.
Chandranath Adak, Bidyut B. Chaudhuri, Chin-Teng Lin, Michael Blumenstein
ICDAR1
2018 Offline Bengali Writer Verification by PDF-CNN and Siamese Net
abstract
Automated handwriting analysis is a popular area of research owing to the variation of writing patterns. In this research area, writer verification is one of the most challenging branches, having direct impact on biometrics and forensics. In this paper, we deal with offline writer verification on complex handwriting patterns. Therefore, we choose a relatively complex script, i.e., Indic Abugida script Bengali (or, Bangla) containing more than 250 compound characters. From a handwritten sample, the probability distribution functions (PDFs) of some handcrafted features are obtained and input to a convolutional neural network (CNN). For such a CNN architecture, we coin the term "PDFCNN", where handcrafted feature PDFs are hybridized with auto-derived CNN features. Such hybrid features are then fed into a Siamese neural network for writer verification. The experiments are performed on a Bengali offline handwritten dataset of 100 writers. Our system achieves encouraging results, which sometimes exceed the results of state-of-the-art techniques on writer verification.
Chandranath Adak, Simone Marinai, Bidyut B. Chaudhuri, Michael Blumenstein
DAS1
2018 A Study on Idiosyncratic Handwriting with Impact on Writer Identification
abstract
In this paper, we study handwriting idiosyncrasy in terms of its structural eccentricity. In this study, our approach is to find idiosyncratic handwritten text components and model the idiosyncrasy analysis task as a machine learning problem supervised by human cognition. We employ the Inception network for this purpose. The experiments are performed on two publicly available databases and an in-house database of Bengali offline handwritten samples. On these samples, subjective opinion scores of handwriting idiosyncrasy are collected from handwriting experts. We have analyzed the handwriting idiosyncrasy on this corpus which comprises the perceptive ground-truth opinion. We also investigate the effect of idiosyncratic text on writer identification by using the SqueezeNet. The performance of our system is promising.
Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein
ICFHR1
2018 Cognitive Analysis for Reading and Writing of Bengali Conjuncts
abstract
In this paper, we study the difficulties arising in reading and writing of Bengali conjunct characters by human-beings. Such difficulties appear when the human cognitive system faces certain obstructions in effortlessly reading/writing. In our computer-based investigation, we consider the reading/writing difficulty analysis task as a machine learning problem supervised by human perception. To this end, we employ two distinct models: (a) an auto-derived feature-based Inception network and (b) a hand-crafted feature-based SVM (Support Vector Machine). Two commonly used Bengali printed fonts and three contemporary handwritten databases are used for collecting subjective opinion scores from human readers/writers. On this corpus, which contains the perceptive ground-truth opinion of reading/writing complications, we have undertaken to conduct the experiments. The experimental results obtained on various types of conjunct characters are promising.
Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein
IJCNN1
2017 Legibility and Aesthetic Analysis of Handwriting
abstract
This paper deals with computer-based cognitive analysis towards legibility and aesthetics of a handwritten document. The legible text creates a human perception that the writing can be read effortlessly because of its orthographic clarity. The aesthetic property relates to the beautiful appearance of a handwritten document. In this study, we deal with these properties on offline Bengali handwriting. We formulate both legibility and aesthetic analysis tasks as machine learning problems supervised by the human cognitive system. We employ automatically derived feature-based recurrent neural networks to investigate writing legibility. For aesthetics evaluation, we employ hand-crafted feature-based support vector machines (SVMs). We have collected contemporary Bengali handwritings, on which the subjective legibility and aesthetic scores are provided by human readers. On this corpus containing legibility and aesthetic ground-truth information, we executed our experiments. The experimental results obtained on various handwritings are encouraging.
Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein
ICDAR1
2017 Impact of struck-out text on writer identification
abstract
The presence of struck-out text in handwritten manuscripts may affect the accuracy of automated writer identification. This paper presents a study on such effects of struck-out text. Here we consider offline English and Bengali handwritten document images. At first, the struck-out texts are detected using a hybrid classifier of a CNN (Convolutional Neural Network) and an SVM (Support Vector Machine). Then the writer identification process is activated on normal and struck-out text separately, to ascertain the impact of struck-out texts. For writer identification, we use two methods: (a) a hand-crafted feature-based SVM classifier, and (b) CNN-extracted auto-derived features with a recurrent neural model. For the experimental analysis, we have generated a database from 100 English and 100 Bengali writers. The performance of our system is very encouraging.
Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein
IJCNN1
2017 An approach for detecting and cleaning of struck-out handwritten text
Bidyut B. Chaudhuri, Chandranath Adak
Pattern Recognit.2
2016 Named Entity Recognition from Unstructured Handwritten Document Images
abstract
Named entity recognition is an important topic in the field of natural language processing, whereas in document image processing, such recognition is quite challenging without employing any linguistic knowledge. In this paper we propose an approach to detect named entities (NEs) directly from offline handwritten unstructured document images without explicit character/word recognition, and with very little aid from natural language and script rules. At the preprocessing stage, the document image is binarized, and then the text is segmented into words. The slant/skew/baseline corrections of the words are also performed. After preprocessing, the words are sent for NE recognition. We analyze the structural and positional characteristics of NEs and extract some relevant features from the word image. Then the BLSTM neural network is used for NE recognition. Our system also contains a post-processing stage to reduce the true NE rejection rate. The proposed approach produces encouraging results on both historical and modern document images, including those from an Australian archive, which are reported here for the very first time.
Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein
DAS1
2016 Offline Cursive Bengali Word Recognition Using CNNs with a Recurrent Model
abstract
This paper deals with offline handwritten word recognition of a major Indic script: Bengali. Due to the structure of this script, the characters (mostly ortho-syllables) are frequently overlapping and hard to segment, especially when the writing is cursive. Individual character recognition and the combination of outputs can increase the likelihood of errors. Instead, a better approach can be sending the whole word to a suitable recognizer. Here we use the Convolutional Neural Network (CNN) integrated with a recurrent model for this purpose. Long short-term memory blocks are used as hidden units. Also, the CNN-derived features are employed in a recurrent model with a CTC (Connectionist Temporal Classification) layer to get the output. We have tested our method on three datasets: (a) a publicly available dataset, (b) a new dataset generated by our research group and (c) an unconstrained dataset. The dataset (a) contains 17,091 words, while our dataset (b) contains 107,550 number of words in total. In addition to these, the dataset (c) is comprised of 5,223 words. We have compared our results with those of some earlier work in the area and have found improved performance, which is due to the novel integration of CNNs with the recurrent model.
Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein
ICFHR1
2016 Writer identification by training on one script but testing on another
abstract
This paper deals with identifying a writer from his/her offline handwriting. In a multilingual country where a writer can scribe in multiple scripts, writer identification becomes challenging when we have individual handwriting data in one script while we need to verify/identify a writer from handwriting in another script. In this paper such an issue is addressed with two scripts: English and Bengali. Here we model the task as a classification problem, where training data contains only Bengali handwritten samples and testing is performed on English handwritten texts. This work is based on the understanding that a writer has some inherent stroke characteristics that are independent of the script in which (s)he writes. In this work, some implicit structural and statistical features are extracted, and multiple classifiers are employed for writer identification. Many training sessions are run on a database of 100 writers and the performances are analyzed. We have obtained encouraging results on this database, which show the effectiveness of our method.
Chandranath Adak, Bidyut B. Chaudhuri, Michael Blumenstein
ICPR1
2015 Writer Identification from offline isolated Bangla characters and numerals
abstract
Writer identification is an essential component in computational forensic. In this paper, we attempt to do this job based only on isolated characters and numerals. For that, at first, some points of interest (keypoints) on the image are detected by structural analysis and SIFT based detector. Then we calculate a set of features within a certain neighborhood of the keypoint and employ fusion rule on multiple probabilistic SVM classifiers output for writer identification. For experimental analysis, a database containing 212,300 isolated Bangla orthosyllabic characters and numerals are generated with the help of 100 writers. We obtain fairly good result to identify a writer. We also try to find a small set of highly discriminative characters storing extra information about the writing style of an individual.
Chandranath Adak, Bidyut B. Chaudhuri
ICDAR1
2014 An Approach of Strike-Through Text Identification from Handwritten Documents
abstract
A handwritten document may contain strike-through texts. If such texts are fed into an OCR system, the output will be garbage. In this paper, we propose a scheme to detect such strike-through texts/words. Using a graph based model, we represent a textual connected component as a graph. The start/end and intersection points of the ink-strokes of a component are marked as graph nodes. There exists an edge between two nodes if they are connected by object (ink) pixels. By eliminating parallel edges and self loops we obtain a simple, undirected, edge-weighted graph of the text-component. The edge-weight is found by adding horizontal/vertical moves weighted by 1 and diagonal moves weighted by √2. In this graph, we find the shortest path which is nearly as long as the width of the text component and maintains a reasonable degree of straightness. This path, if exist, is identified as the strike-through line. Here we deal with handwritten documents in English, Bengali and Devanagari script. Our approach delivers fairly good results.
Chandranath Adak, Bidyut B. Chaudhuri
ICFHR1