EDBT 2026 Demo / reviewers in the wild / expert
Partha Pratim Roy 0001
dblp:82/5415
· DBLP profile ↗
151ranked-venue papers
24as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 96 · 14 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 62 · 6 first-author · 19 since 2021Databases, data management, data science and information retrieval · 23 · 7 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EmoXFormer: Human-Cognition-Inspired Multimodal Emotion Recognition from Disjoint Modality Datasets
Qurat Ul Ain Aisha, Ji-Hoon Choi, Se-In Choi, Partha Pratim Roy 0001, Byung-Gyu Kim |
ICPR (8) | 4 |
| 2026 | Modified TSception for analyzing driver drowsiness and mental workload from EEG
Gourav Siddhad, Rajkumar Saini, Partha Pratim Roy 0001 |
Neural Comput. Appl. | 4 |
| 2025 | Combining Spatio-Temporal Networks and Graph Attention Architectures for EEG-Based Workload ClassificationabstractEstimation of cognitive workload from EEG signals is a key challenge in advancing neuroergonomic systems and brain-computer interfaces (BCIs). A hybrid approach is presented, combining EEGNet and Graph Attention Networks (GATs) to effectively capture the intricate spatial and temporal dynamics within EEG data. Leveraging the COG-BCI dataset, which includes cognitively demanding tasks such as N-Back and Multi-Attribute Task Battery-II (MATB-II), the model employs GAT’s multi-head attention mechanism to enhance feature extraction and improve classification performance. A dynamic graph, constructed by integrating spatial distances between electrodes and signal correlations, provides a comprehensive framework for neural activity modeling. The hybrid EEGNet-GAT model demonstrates substantial performance improvements, achieving classification accuracies of 92.13% on the N-Back task and 96.08% on MATB-II, significantly surpassing the results from standalone EEGNet. These findings underscore the potential of the EEGNet-GAT architecture for real-time cognitive workload estimation, offering strong applications in BCIs and neuroergonomic environments. Nikhil Panwar, Vishal Pandey, Rownak Tiwari, Partha Pratim Roy 0001 |
ICASSP | 4 |
| 2025 | Quantum-Behaved Particle Swarm Optimization for the Segmentation of Kidney Stone CT ImagesabstractImage segmentation is a crucial element of image processing that divides an image into distinct regions based on pixel intensity, facilitating detailed analysis and interpretation of various image components. Conventional segmentation techniques often struggle with challenges such as local minima entrapment and premature convergence, especially in complex pixel search spaces, which can also be computationally intensive as the number of threshold levels increases. To overcome these limitations, a Quantum-behaved Particle Swarm Optimization (QPSO) approach is employed for multi-level thresholding. By integrating quantum mechanics principles with classical Particle Swarm Optimization, QPSO introduces quantum behaviors that enhance the algorithm’s exploration capabilities, allowing for a more thorough search of the solution space and reducing the risk of premature convergence to local optima. Meanwhile, Kapur’s entropy method is applied to segment images into distinct regions based on optimal pixel values, thereby enhancing segmentation precision. The algorithm’s performance is rigorously evaluated using kidney stone (Nephrolithiasis) datasets from Kaggle, with segmentation quality assessed through metrics such as Mean Squared Error (MSE), Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), Feature Similarity Index Measure (FSIM), optimal thresholds, and best fitness values. Statistical reliability is ensured through the Wilcoxon signed-rank test and the Friedman ranking test. Experimental results demonstrate that QPSO significantly outperforms other algorithms, achieving an SSIM of 0.96, FSIM of 0.98, and PSNR of 29.80, highlighting its efficacy in addressing complex image segmentation challenges. Sajad Ahmad Rather, Akhilesh Kandwal, Mohammad Khalid Pandit, Partha Pratim Roy 0001 |
ICASSP | 4 |
| 2024 | DCDM: Diffusion-Conditioned-Diffusion Model for Scene Text Image Super-Resolution
Shrey Singh, Prateek Keserwani, Masakazu Iwamura, Partha Pratim Roy 0001 |
ECCV (15) | 4 |
| 2024 | Multi-Level Feature Exploration Using LSTM-Based Variational Autoencoder Network for Fall Detection
Anitha Rani Inturi, Vazhora Malayil Manikandan, Partha Pratim Roy 0001, Byung-Gyu Kim |
ICPR (16) | 3 |
| 2024 | EEG-Based Mental Imagery Task Adaptation via Ensemble of Weight-Decomposed Low-Rank Adapters
Taveena Lotey, Aman Verma, Partha Pratim Roy 0001 |
ICPR (11) | 3 |
| 2024 | Enhanced Cross-Task EEG Classification: Domain Adaptation with EEGNet
Vishal Pandey, Nikhil Panwar, Atharva Kumbhar, Partha Pratim Roy 0001, Masakazu Iwamura |
ICPR (11) | 4 |
| 2024 | Chaos Theory Based Gravitational Search Algorithm For Medical Image Segmentation
Sajad Ahmad Rather, Partha Pratim Roy 0001, Sujit Das |
ICPR (28) | 2 |
| 2024 | DrowzEE-G-Mamba: Leveraging EEG and State Space Models for Driver Drowsiness Detection
Gourav Siddhad, Sayantan Dey, Partha Pratim Roy 0001 |
ICPR (27) | 3 |
| 2024 | Awake at the Wheel: Enhancing Automotive Safety Through EEG-Based Fatigue Detection
Gourav Siddhad, Sayantan Dey, Partha Pratim Roy 0001, Masakazu Iwamura |
ICPR (11) | 3 |
| 2024 | Neural Networks Meet Neural Activity: Utilizing EEG for Mental Workload Estimation
Gourav Siddhad, Partha Pratim Roy 0001, Byung-Gyu Kim |
ICPR (11) | 2 |
| 2024 | Enhancing Document Information Analysis with Multi-Task Pre-training: A Robust Approach for Information Extraction in Visually-Rich DocumentsabstractThis paper introduces a deep learning model tailored for document information analysis, emphasizing document classification, entity relation extraction, and document visual question answering. The proposed model leverages transformer-based models to encode all the information present in a document image, including textual, visual, and layout information. The model is pre-trained and subsequently fine-tuned for various document image analysis tasks. The proposed model incorporates three additional tasks during the pre-training phase, including reading order identification of different layout segments in a document image, layout segments categorization as per PubLayNet, and generation of the text sequence within a given layout segment (text block). The model also incorporates a collective pre-training scheme where losses of all the tasks under consideration, including pre-training and fine-tuning tasks with all datasets, are considered. Additional encoder and decoder blocks are added to the RoBERTa network to generate results for all tasks. The proposed model achieved impressive results across all tasks, with an accuracy of 95.87% on the RVL-CDIP dataset for document classification, F1 scores of 0.9306, 0.9804, 0.9794, and 0.8742 on the FUNSD, CORD, SROIE, and Kleister-NDA datasets respectively for entity relation extraction, and an ANLS score of 0.8468 on the DocVQA dataset for visual question answering. The results highlight the effectiveness of the proposed model in understanding and interpreting complex document layouts and content, making it a promising tool for document analysis tasks. Tofik Ali, Partha Pratim Roy 0001 |
IJCNN | 2 |
| 2023 | Motor Activity Recognition Using Eeg Data and Ensemble of Stacked BLSTM-LSTM Network and Transformer ModelabstractWith the rapid development of brain-computer interfaces, the number of applications based on this technology is increasing rapidly. This work proposes a Stacked BLSTM-LSTM, EEG-Transformer, and their ensemble network to predict real-life motor activities of individuals using EEG (ElectroEncephalo-Gram) data. A 32 electrode gel-based EEG recording device has been used to record brain signals from 20 subjects while performing 17 commonly used day-to-day motor activities. The stacked BLSTM-LSTM and EEG Transformer networks predicted the activities with an accuracy of 97.9%, 96.7%, respectively. The ensemble improved the classification accuracy further to 98.5%, which is a considerable improvement over the existing state-of-the-art methods. This study also reveals that raw and delta band frequencies are better in predicting the activities than other frequency bands of the EEG signals. Motor activity recognition has several applications, including rehabilitation, healthcare, gaming, and preventing loss of lives during mitigation of fires, diffusion of bombs, etc., via imitation robots. Pallavi Kaushik, Ilina Tripathi, Partha Pratim Roy 0001 |
ICASSP | 3 |
| 2023 | Crowd Characterization in Surveillance Videos Using Deep-Graph Convolutional Neural NetworkabstractCrowd behavior is a natural phenomenon that can provide valuable insight into the crowd characterization process. Modeling the visual appearance of a large crowd gathering can reveal meaningful information about its dynamics. Parametric modeling can be used to develop efficient and robust crowd monitoring systems. A crowd can be structured or unstructured based on the organization. In this article, crowd characterization has been mapped to a graph classification problem to classify movements based on order parameter ( ϕ ), active force components, and steadiness (Reynolds number). The graphs are constructed from the motion groups obtained using an active Langevin framework. These graphs are processed using a deep graph convolutional neural network for crowd characterization. For experimentation, we have prepared a dataset comprising of videos from popular publicly available datasets and our own recorded videos. The proposed framework has been compared with the latest deep learning-based frameworks in terms of accuracy and area under the curve (AUC). We have obtained a 4%-5% improvement in accuracy and AUC values over the existing frameworks. The insights obtained from the proposed framework can be used for better crowd monitoring and management. Shreetam Behera, Debi Prosad Dogra, Malay Kumar Bandyopadhyay, Partha Pratim Roy 0001 |
IEEE Trans. Cybern. | 4 |
| 2022 | Multimodal Deep Sparse Subspace Clustering for Multiple Stimuli-based Cognitive taskabstractCognitive state assessment can be effectively performed using Electroencephalogram (EEG). However, due to the curse of dimensionality issues of EEG, most of the clustering methods often lead to poor performance. Deep neural network-based representation learning transforms high-dimensional data into lower-dimensional feature space, increasing the clustering performance. This paper proposes an efficient multimodal (spectral-temporal) deep clustering model to evaluate workload levels from multiple stimuli (visual and auditory)-based n-back task. The proposed model extracts the temporal and spectral EEG features from sequence-wise EEG signal and spectral power. The combined spectral-temporal low-dimensional latent feature is passed to the sparse subspace clustering (SSC) model to estimate different workload levels. The temporal and spectral latent features are learned using the Long short-term memory (LSTM) and Convolutional Neural Network (CNN)-based variational autoencoder (VAE) model. In the SSC method, the collection of data points that lies in the union of low-dimensional subspaces forms a cluster. Here, each cluster overcomes the effect of the outliers from subspace, improving the cluster quality. The proposed model achieves the best clustering accuracy of 98.2% in the subject-independent test and a mean clustering accuracy of 95.2%. The proposed model achieves a significant improvement over the state-of-the-art studies. The effectiveness of the model is also evaluated on two other publicly available n-back datasets. The proposed model enhances the future scope of the deep representation learning-based clustering approach for other cognitive tasks. Debashis Das Chakladar, Debasis Samanta, Partha Pratim Roy 0001 |
ICPR | 3 |
| 2022 | Cross-Session Motor Imagery EEG Classification using Self-Supervised Contrastive LearningabstractAmong various Brain-Computer Interfaces (BCI) categories, Electroencephalography (EEG) signals have advantages over their other counterparts, such as fNIRS and fMRI, for their ease of acquisition and affordable recording devices. Despite the ease of data recording, the labeling of the data needs expert knowledge, which is expensive. Also, the data collected from the same subjects in different sessions are prone to varying distribution, making the cross-session data classification more challenging. Supervised learning has been widely used for EEG signal analysis, but its performance is highly dependent on a vast amount of annotated data. Self-supervised learning (SSL) has been proven to be a highly effective technique that allows deep learning architectures to learn general features from unlabelled data without needing an extensive amount of labeled data. This paper proposes a self-supervised contrastive learning method of cross-session based EEG data for motor imagery classification. The pretext task of signal transformations has been used to train the network during the training phase to reduce the distance between the similar transformation pairs and maximize the distance between the dissimilar transformation pairs generated from the EEG signals. It helps the network learn generalized features of cross-session EEG signals with the help of representation learning and in the absence of labeled data. The diverse nature of cross-session EEG data leads the network toward learning more generalized features. The performance of the proposed model upon pretraining on the BCI-IV 2a dataset using self-supervised contrastive learning and the fine-tuning on the same dataset using supervised learning gives the average accuracy of 50.81% and performed significantly better than the compared supervised learning-based methods. Taveena Lotey, Prateek Keserwani, Gaurav Wasnik, Partha Pratim Roy 0001 |
ICPR | 4 |
| 2022 | Video based exercise recognition and correct pose detection
Tushar Rangari, Sudhanshu Kumar 0002, Partha Pratim Roy 0001, Debi Prosad Dogra, Byung-Gyu Kim |
Multim. Tools Appl. | 3 |
| 2022 | Text Region Conditional Generative Adversarial Network for Text Concealment in the WildabstractTextual information appearing on the captured image may contain personal information. In various circumstances, publishing such images in the public domain may create a threat of privacy leak. To avoid these situations, we propose a text concealment method. To accomplish this task, we have used a conditional generator, which is a text region conditioned concealment network. The text regions predicted by the detector network are used as a conditioning criterion for the concealment process. The text region prediction in the form of a word-level bounding box may contain stroke pixels as well as the background pixels. However, to reduce the number of background pixels for proper conditioning of text concealment network, a character level annotation is used for the generator in place of word-level bounding box annotation. It helps to focus more on strokes as compared to the background pixels. A character-level symmetric line representation of text has been proposed to obtain finer level text region prediction as compared to the character-level bounding box. The proposed model is trainable from end-to-end. The text region conditioned generator is trained from the loss of global and local discriminators. The proposed method is validated on public scene text image datasets such as ICDAR 2015, COCO-Text and Synthesis dataset. The proposed architecture shows competitive results as compared to other state-of-the-art approaches. Prateek Keserwani, Partha Pratim Roy 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Robust Scene Text Detection for Partially Annotated Training DataabstractThis article analyzed the impact of training data containing un-annotated text instances, i.e., partial annotation in scene text detection, and proposed a text region refinement approach to address it. Scene text detection is a problem that has attracted the attention of the research community for decades. Impressive results have been obtained for fully supervised scene text detection with recent deep learning approaches. These approaches, however, need a vast amount of completely labeled datasets, and the creation of such datasets is a challenging and time-consuming task. Research literature lacks the analysis of the partial annotation of training data for scene text detection. We have found that the performance of the generic scene text detection method drops significantly due to the partial annotation of training data. We have proposed a text region refinement method that provides robustness against the partially annotated training data in scene text detection. The proposed method works as a two-tier scheme. Text-probable regions are obtained in the first tier by applying hybrid loss that generates pseudo-labels to refine text regions in the second-tier during training. Extensive experiments have been conducted on a dataset generated from ICDAR 2015 by dropping the annotations with various drop rates and on a publicly available SVT dataset. The proposed method exhibits a significant improvement over the baseline and existing approaches for the partially annotated training data. Prateek Keserwani, Rajkumar Saini, Marcus Liwicki, Partha Pratim Roy 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Vehicular Trajectory Classification and Traffic Anomaly Detection in Videos Using a Hybrid CNN-VAE ArchitectureabstractVisual surveillance has become indispensable in the evolution of Intelligent Transportation Systems (ITS). Video object trajectories are key to many of the visual surveillance applications. Classifying varying length time series data such as video object trajectories using conventional neural networks, can be challenging. In this paper, we propose trajectory classification and anomaly detection using a hybrid Convolutional Neural Network (CNN) and Variational Autoencoder (VAE) architecture. First, we introduce a high level features for varying length object trajectories using color gradient representation. In the next stage, a semi-supervised way to annotate moving object trajectories extracted using Temporally Incremental Gravitational Model (TIGM) is used for class labeling. For training, anomalous trajectories are identified using t-Distributed Stochastic Neighbor Embedding (t-SNE). Finally, a hybrid CNN-VAE architecture has been proposed for trajectory classification and anomaly detection. The results obtained using publicly available surveillance video datasets reveal that the proposed method can successfully identify traffic anomalies such as violations in lane driving, sudden speed variations, abrupt termination of vehicle during movement, and vehicles moving in wrong directions. The accuracy of trajectory classification improves by a margin of 1-6% against popular neural networks-based classifiers across various datasets using the proposed high-level features. The gradient representation also improves the anomaly detection accuracy significantly (30-35%). Code and dataset can be found athttps://github.com/santhoshkelathodi/CNN-VAE. Santhosh Kelathodi Kumaran, Debi Prosad Dogra, Partha Pratim Roy 0001, Adway Mitra |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Logo detection using weakly supervised saliency map
Prateek Keserwani, Partha Pratim Roy 0001, Debi Prosad Dogra |
Multim. Tools Appl. | 3 |
| 2021 | 3D word spotting using leap motion sensor
Partha Pratim Roy 0001, Pradeep Kumar 0002, Shweta Patidar, Rajkumar Saini |
Multim. Tools Appl. | 1 |
| 2021 | Understanding crowd flow patterns using active-Langevin model
Shreetam Behera, Debi Prosad Dogra, Malay Kumar Bandyopadhyay, Partha Pratim Roy 0001 |
Pattern Recognit. | 4 |
| 2021 | A novel spatio-temporal Siamese network for 3D signature recognition
Spandan Ghosh, Pradeep Kumar 0002, Erik J. Scheme, Partha Pratim Roy 0001 |
Pattern Recognit. Lett. | 5 |
| 2021 | Modeling local and global behavior for trajectory classification using graph based algorithm
Rajkumar Saini, Pradeep Kumar 0002, Partha Pratim Roy 0001, Umapada Pal 0001 |
Pattern Recognit. Lett. | 3 |
| 2021 | Trajectory-Based Scene Understanding Using Dirichlet Process Mixture ModelabstractAppropriate modeling of a surveillance scene is essential for the detection of anomalies in road traffic. Learning usual paths can provide valuable insight into road traffic conditions and thus can help in identifying unusual routes taken by commuters/vehicles. If usual traffic paths are learned in a nonparametric way, manual interventions in road marking can be avoided. In this paper, we propose an unsupervised and nonparametric method to learn the frequently used paths from the tracks of moving objects in Θ(kn) time, where k denotes the number of paths and n represents the number of tracks. In the proposed method, temporal dependencies of the moving objects are considered to make the clustering meaningful using temporally incremental gravity model (TIGM). In addition, the distance-based scene learning makes it intuitive to estimate the model parameters. Further, we have extended the TIGM hierarchically as a dynamically evolving model (DEM) to represent notable traffic dynamics of a scene. The experimental validation reveals that the proposed method can learn a scene quickly without prior knowledge about the number of paths ( k ). We have compared the results with various state-of-the-art methods. We have also highlighted the advantages of the proposed method over the existing techniques popularly used for designing traffic monitoring applications. It can be used for administrative decision making to control traffic at junctions or crowded places and generate alarm signals, if necessary. Santhosh Kelathodi Kumaran, Debi Prosad Dogra, Partha Pratim Roy 0001, Bidyut B. Chaudhuri |
IEEE Trans. Cybern. | 3 |
| 2020 | End-to-end Triplet Loss based Emotion Embedding System for Speech Emotion RecognitionabstractIn this paper, an end-to-end neural embedding system based on triplet loss and residual learning has been proposed for speech emotion recognition. The proposed system learns the embeddings from the emotional information of the speech utterances. The learned embeddings are used to recognize the emotions portrayed by given speech samples of various lengths. The proposed system implements Residual Neural Network architecture. It is trained using softmax pretraining and triplet loss function. The weights between the fully connected and embedding layers of the trained network are used to calculate the embedding values. The embedding representations of various emotions are mapped onto a hyperplane, and the angles among them are computed using the cosine similarity. These angles are utilized to classify a new speech sample into its appropriate emotion class. The proposed system has demonstrated 91.67% and 64.44% accuracy while recognizing emotions for RAVDESS and IEMOCAP dataset, respectively. Puneet Kumar 0003, Sidharth Jain, Balasubramanian Raman, Partha Pratim Roy 0001, Masakazu Iwamura |
ICPR | 4 |
| 2020 | EEG-Based Cognitive State Assessment Using Deep Ensemble Model and Filter Bank Common Spatial PatternabstractElectroencephalography (EEG) is the most used physiological measure to evaluate the cognitive state of a user efficiently. As EEG inherently suffers from a poor spatial resolution, features extracted from each EEG channel may not be efficiently used for the cognitive state assessment. In this paper, the EEG-based cognitive state assessment has been performed during the mental arithmetic experiment, which includes two cognitive states (task and rest) of a user. To obtain the temporal as well as the spatial resolution of the EEG signal, we combined the Filter Bank Common Spatial Pattern (FBCSP) method and Long Short-Term Memory (LSTM)-based deep ensemble model for classifying the cognitive state of a user. Subject-wise data distribution has been performed due to the execution of a large volume of data in a low computing environment. In the FBCSP method, the input EEG is decomposed into multiple equal-sized frequency bands, and spatial features of each frequency bands are extracted using the Common Spatial Pattern (CSP) algorithm. Next, a feature selection algorithm has been applied to identify the most informative features for classification. The proposed deep ensemble model consists of multiple similar structured LSTM networks that work in parallel. The output of the ensemble model (i.e., the cognitive state of a user) is computed using the average weighted combination of the individual model prediction. This proposed model achieves 87% classification accuracy, and it can also effectively estimate the cognitive state of a user in a low computing environment. Debashis Das Chakladar, Shubhashis Dey, Partha Pratim Roy 0001, Masakazu Iwamura |
ICPR | 3 |
| 2020 | Modeling Extent-of-Texture Information for Ground Terrain RecognitionabstractGround Terrain Recognition is a difficult task as the context information varies significantly over the regions of a ground terrain image. In this paper, we propose a novel approach towards ground-terrain recognition via modeling the Extent-of-Texture information to establish a balance between the order-less texture component and ordered-spatial information locally. At first, the proposed method uses a CNN backbone feature extractor network to capture meaningful information of ground terrain images, and model the extent of texture and shape information locally. Then, it encodes order-less texture information and ordered shape information in a patch-wise manner, and utilizes an intra-domain message passing mechanism to make every patch aware of each other for rich feature learning. Next, the model combines the extent of texture information with the encoded texture information and the extent of shape information with the encoded shape information patch-wise and then exploit Extent of texture (EoT) Guided Inter-domain Message passing module for sharing knowledge about the opposite domain to balance out the order-less texture information with ordered shape information. Finally, Bilinear model outputs a pairwise correlation between the order-less texture information and ordered shape information, and classifier classifies the ground terrain image efficiently. The experimental results indicate the superior performance of the proposed model11The source code of the proposed system is publicly available at https://github.com/ShuvozitGhose/Ground-Terrain-EoT over existing state-of-the-art techniques on DTD, MINC and GTOS-mobile datasets. Shuvozit Ghose, Pinaki Nath Chowdhury, Partha Pratim Roy 0001, Umapada Pal 0001 |
ICPR | 3 |
| 2020 | UDBNET: Unsupervised Document Binarization Network via Adversarial GameabstractDegraded document image binarization is one of the most challenging tasks in the domain of document image analysis. In this paper, we present a novel approach towards document image binarization by introducing three-player minmax adversarial game. We train the network in an unsupervised setup by assuming that we do not have any paired-training data. In our approach, an Adversarial Texture Augmentation Network (ATANet) first superimposes the texture of a degraded reference image over a clean image. Later, the clean image along with its generated degraded version constitute the pseudo paired-data which is used to train the Unsupervised Document Binarization Network (UDBNet). Following this approach, we have enlarged the document binarization datasets as it generates multiple images having same content feature but different textual feature. These generated noisy images are then fed into the UDBNet to get back the clean version. The joint discriminator which is the third-player of our three-player min-max adversarial game tries to couple both the ATANet and UDBNet. The three-player min-max adversarial game stops, when the distributions modelled by the ATANet and the UDBNet align to the same joint distribution over time. Thus, the joint discriminator enforces the UDBNet to perform better on real degraded image. The experimental results indicate the superior performance of the proposed model over existing state-of-the-art algorithm on widely used DIBCO datasets. The source code of the proposed system is publicly available at https://github.com/VIROBO-15/UDBNET. Amandeep Kumar, Shuvozit Ghose, Pinaki Nath Chowdhury, Partha Pratim Roy 0001, Umapada Pal 0001 |
ICPR | 4 |
| 2020 | Estimation of linear motion in dense crowd videos using Langevin model
Shreetam Behera, Debi Prosad Dogra, Malay Kumar Bandyopadhyay, Partha Pratim Roy 0001 |
Expert Syst. Appl. | 4 |
| 2020 | Efficient extraction of consistent bit locations from binarized iris features
Debanjan Sadhya, Kanjar De, Balasubramanian Raman, Partha Pratim Roy 0001 |
Expert Syst. Appl. | 4 |
| 2020 | Zone-based keyword spotting in Bangla and Devanagari documents
Ayan Kumar Bhunia, Partha Pratim Roy 0001, Aneeshan Sain, Umapada Pal 0001 |
Multim. Tools Appl. | 2 |
| 2020 | Fractional Local Neighborhood Intensity Pattern for Image Retrieval using Genetic Algorithm
Shuvozit Ghose, Abhirup Das, Ayan Kumar Bhunia, Partha Pratim Roy 0001 |
Multim. Tools Appl. | 4 |
| 2020 | A study of EEG for enterprise multimedia security
Barjinder Kaur, Partha Pratim Roy 0001 |
Multim. Tools Appl. | 3 |
| 2020 | Fast Griffin Lim based waveform generation strategy for text-to-speech synthesis
Puneet Kumar 0003, Vikas Maddukuri, Nagasai Madamshettib, Kishore KG, Sahit Sai Sriram Kavurub, Balasubramanian Raman, Partha Pratim Roy 0001 |
Multim. Tools Appl. | 8 |
| 2020 | A hybrid classifier combination for home automation using EEG signals
Partha Pratim Roy 0001, Pradeep Kumar 0002, Victor Chang 0001 |
Neural Comput. Appl. | 1 |
| 2020 | A novel feature descriptor for image retrieval by combining modified color histogram and diagonally symmetric co-occurrence texture pattern
Ayan Kumar Bhunia, Avirup Bhattacharyya, Prithaj Banerjee, Partha Pratim Roy 0001, M. Subrahmanyam 0001 |
Pattern Anal. Appl. | 4 |
| 2020 | Retrieval of colour and texture images using local directional peak valley binary pattern
Partha Pratim Roy 0001, Debi Prosad Dogra, Byung-Gyu Kim |
Pattern Anal. Appl. | 2 |
| 2020 | Can we automate diagrammatic reasoning?abstractDiagrammatic reasoning (DR) problems are well known. However, solving DR problems represented in 4 × 1 Raven’s Progressive Matrix (RPM) form using computer vision and pattern recognition has not yet been tried. Emergence of deep learning techniques aided by advanced computing can be exploited to solve such DR problems. In this paper, we propose a new learning framework by combining LSTM and Convolutional LSTM to solve 4 × 1 DR problems. Initially, the elementary geometrical shapes in such problems are detected using a typical CNN-based detector. Next, relations of various shapes are analyzed and a high-level feature set is produced and processed in the LSTM framework. A new 4 × 1 DR dataset has been prepared and made available to the research community. We believe, it will be helpful in advancing this research further. We have compared our method with some of the existing frameworks that can be used for solving RPM-guided DR problems. We have recorded 18–20% increase in the average prediction accuracy as compared to the prior frameworks when applied to RPM-guided DR problems. We believe the CV research community will be interested to carry out similar research, particularly to investigate the feasibility of solving other types of known DR problems. Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001, Dilip K. Prasad |
Pattern Recognit. | 4 |
| 2020 | Video trajectory analysis using unsupervised clustering and multi-criteria rankingabstractAbstract Surveillance camera usage has increased significantly for visual surveillance. Manual analysis of large video data recorded by cameras may not be feasible on a larger scale. In various applications, deep learning-guided supervised systems are used to track and identify unusual patterns. However, such systems depend on learning which may not be possible. Unsupervised methods relay on suitable features and demand cluster analysis by experts. In this paper, we propose an unsupervised trajectory clustering method referred to as t-Cluster. Our proposed method prepares indexes of object trajectories by fusing high-level interpretable features such as origin, destination, path, and deviation. Next, the clusters are fused using multi-criteria decision making and trajectories are ranked accordingly. The method is able to place abnormal patterns on the top of the list. We have evaluated our algorithm and compared it against competent baseline trajectory clustering methods applied to videos taken from publicly available benchmark datasets. We have obtained higher clustering accuracies on public datasets with significantly lesser computation overhead. Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001 |
Soft Comput. | 4 |
| 2020 | Fusion of Spatio-temporal Information for Indic Word Recognition Combining Online and Offline Text DataabstractWe present a novel Indic handwritten word recognition scheme by fusion of spatio-temporal information extracted from handwritten images. The main challenge in Indic word recognition lies in its complexity because of modifiers, touching characters, and compound characters. Hidden Markov Models (HMMs) are being used to model such data due to their ability to learn sequential data, however, the recognition performance is not satisfactory. We propose here a Long Short-Term Memory (LSTM)-based architecture for offline Indic word recognition. Offline recognition methods usually involve spatial data, whereas it has been observed that online recognition schemes show better performance than the offline methodologies. Online information usually refers to the temporal information obtained from the strokes of the pen tip while writing, which is missing in offline word images. In this article, an effort has been made to extract the online temporal information from offline images using stroke recovery and later it is combined with spatial information in LSTM architecture. During recognition, the character models are trained using both offline and extracted pseudo-online handwritten data separately. Finally, a novel fusion scheme has been used to combine them together. From the experiment, it is noted that recognition performance of handwritten Indic words improves considerably due to the fusion scheme of spatial and temporal data. Subham Mukherjee, Pradeep Kumar 0002, Partha Pratim Roy 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2020 | Movie Recommendation System Using Sentiment Analysis From Microblogging DataabstractRecommendation systems (RSs) have garnered immense interest for applications in e-commerce and digital media. Traditional approaches in RSs include such as collaborative filtering (CF) and content-based filtering (CBF) through these approaches that have certain limitations, such as the necessity of prior user history and habits for performing the task of recommendation. To minimize the effect of such limitation, this article proposes a hybrid RS for the movies that leverage the best of concepts used from CF and CBF along with sentiment analysis of tweets from microblogging sites. The purpose to use movie tweets is to understand the current trends, public sentiment, and user response of the movie. Experiments conducted on the public database have yielded promising results. Sudhanshu Kumar 0002, Kanjar De, Partha Pratim Roy 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2019 | Handwriting Recognition in Low-Resource Scripts Using Adversarial LearningabstractHandwritten Word Recognition and Spotting is a challenging field dealing with handwritten text possessing irregular and complex shapes. The design of deep neural network models makes it necessary to extend training datasets in order to introduce variations and increase the number of samples; word-retrieval is therefore very difficult in low-resource scripts. Much of the existing literature comprises preprocessing strategies which are seldom sufficient to cover all possible variations. We propose an Adversarial Feature Deformation Module (AFDM) that learns ways to elastically warp extracted features in a scalable manner. The AFDM is inserted between intermediate layers and trained alternatively with the original framework, boosting its capability to better learn highly informative features rather than trivial ones. We test our meta-framework, which is built on top of popular word-spotting and word-recognition frameworks and enhanced by AFDM, not only on extensive Latin word datasets but also on sparser Indic scripts. We record results for varying sizes of training data, and observe that our enhanced network generalizes much better in the low-data regime; the overall word-error rates and mAP scores are observed to improve as well. Ayan Kumar Bhunia, Abhirup Das, Ankan Bhunia, Perla Sai Raj Kishore, Partha Pratim Roy 0001 |
CVPR | 5 |
| 2019 | User Constrained Thumbnail Generation Using Adaptive ConvolutionsabstractThumbnails are widely used all over the world as a preview for digital images. In this work we propose a deep neural framework to generate thumbnails of any size and aspect ratio, even for unseen values during training, with high accuracy and precision. We use Global Context Aggregation (GCA) and a modified Region Proposal Network (RPN) with adaptive convolutions to generate thumbnails in real time. GCA is used to selectively attend and aggregate the global context information from the entire image while the RPN is used to generate candidate bounding boxes for the thumbnail image. Adaptive convolution eliminates the difficulty of generating thumbnails of various aspect ratios by using filter weights dynamically generated from the aspect ratio information. The experimental results indicate the superior performance of the proposed model1over existing state-of-the-art techniques. Perla Sai Raj Kishore, Ayan Kumar Bhunia, Shuvozit Ghose, Partha Pratim Roy 0001 |
ICASSP | 4 |
| 2019 | Predicting Video-frames Using Encoder-convlstm CombinationabstractVideo generation is an active field of research. With the rise in the amount of available data and economically available processing power in the form of GPUs, deep Learning has been a go-to solution for many real life problems and similarly it is often attempted to solve the problem of video generation using deep learning. Predicting the next set of frames for a given set of frames in a video has seldom been taken up. Each video is composed of a consecutive closely related frames of images. If we consider these frames, the frame in each time-step seems to be related to the frames in the preceding time-steps. Therefore, we have both spatial and temporal data available from any set of consecutive frames in a video. Learning some sort of representation of the images that encodes the spatial data of the images (frames) can be combined with learning how these representations of a particular time-step is related with the next few time-steps is made possible, then prediction of the next few frames for a given set of frames is made possible. Our aim is to propose a simple yet effective model that can achieve this goal. Subham Mukherjee, Spandan Ghosh, Pradeep Kumar 0002, Partha Pratim Roy 0001 |
ICASSP | 5 |
| 2019 | Facial Micro-expression Spotting and Recognition Using Time Contrasted Feature with Visual MemoryabstractFacial micro-expressions are sudden involuntary minute muscle movements which reveal true emotions that people try to conceal. Spotting a micro-expression and recognizing it is a major challenge owing to its short duration and intensity. Many works pursued traditional and deep learning based approaches to solve this issue but compromised on learning low-level features and higher accuracy due to unavailability of datasets. This motivated us to propose a novel joint architecture of spatial and temporal network which extracts time-contrasted features from the feature maps to contrast out micro-expression from rapid muscle movements. The usage of time contrasted features greatly improved the spotting of micro-expression from inconspicuous facial movements. Also, we include a memory module to predict the class and intensity of the micro-expression across the temporal frames of the micro-expression clip. Our method achieves superior performance in comparison to other conventional approaches on CASMEII dataset. Sauradip Nag, Ayan Kumar Bhunia, Aishik Konwer, Partha Pratim Roy 0001 |
ICASSP | 4 |
| 2019 | Zero Shot Learning Based Script Identification in the WildabstractThe text recognition system for natural images or video frames containing multilingual text needs a method to first identify the written script and then recognize the word in the identified script. However, the occurrence of some scripts is rare as compared to others. Due to the availability of a few samples of the rare script, the supervised learning of the deep neural networks is difficult. To overcome this problem, we have proposed a zero-shot learning based method for script identification. We have also proposed architecture for script identification which fuses the global feature vector and the semantic embedding vector. The semantic embedding of the script is obtained by using the spatial dependency of the stroke's sequence via the recurrent neural network. The proposed architecture shows superior results as compared to the baseline approaches. Prateek Keserwani, Kanjar De, Partha Pratim Roy 0001, Umapada Pal 0001 |
ICDAR | 3 |
| 2019 | Improving Document Binarization Via Adversarial Noise-Texture AugmentationabstractBinarization of degraded document images is an elementary step in most problems involving document image analysis. The paper re-visits the binarization problem by introducing an adversarial learning approach. We construct a Texture Augmentation Network that transfers the texture element of a degraded reference document image to a clean binary image. In this way, the network creates multiple versions of the same textual content with various noisy textures, thus enlarging the available document binarization datasets. Finally, the newly generated images are passed through a Binarization network to get back the clean version. By jointly training the two networks we can increase the adversarial robustness of our system. The most significant contribution of our framework is that it does not require any paired data unlike other Deep Learning-based methods [1], [2], [3]. Such a novel approach has never been implemented earlier thus making it the very first of its kind in Document Image Analysis community. Experimental results suggest that the proposed method achieves superior performance over widely used DIBCO datasets. Ankan Bhunia, Ayan Kumar Bhunia, Aneeshan Sain, Partha Pratim Roy 0001 |
ICIP | 4 |
| 2019 | Cogni-Net: Cognitive Feature Learning Through Deep Visual PerceptionabstractCan we ask computers to recognize what we see from brain signals alone? Our paper seeks to utilize the knowledge learnt in the visual domain by popular pre-trained vision models and use it to teach a recurrent model being trained on brain signals to learn a discriminative manifold of the human brain's cognition of different visual object categories in response to perceived visual cues. For this we make use of brain EEG signals triggered from visual stimuli like images and leverage the natural synchronization between images and their corresponding brain signals to learn a novel representation of the cognitive feature space. The concept of knowledge distillation has been used here for training the deep cognition model, CogniNet1, by employing a student-teacher learning technique in order to bridge the process of inter-modal knowledge transfer. The proposed novel architecture obtains state-of-the-art results, significantly surpassing other existing models. The experiments performed by us also suggest that if visual stimuli information like brain EEG signals can be gathered on a large scale, then that would help to obtain a better understanding of the largely unexplored domain of human brain cognition. Pranay Mukherjee, Abhirup Das, Ayan Kumar Bhunia, Partha Pratim Roy 0001 |
ICIP | 4 |
| 2019 | Texture synthesis guided deep hashing for texture image retrievalabstractWith the large scale explosion of images and videos over the internet, efficient hashing methods have been developed to facilitate memory and time efficient retrieval of similar images. However, none of the existing works use hashing to address texture image retrieval mostly because of the lack of sufficiently large texture image databases. Our work addresses this problem by developing a novel deep learning architecture that generates binary hash codes for input texture images. For this, we first pre-train a Texture Synthesis Network (TSN) which takes a texture patch as input and outputs an enlarged view of the texture by injecting newer texture content. Thus it signifies that the TSN encodes the learnt texture specific information in its intermediate layers. In the next stage, a second network gathers the multi-scale feature representations from the TSN's intermediate layers using channel-wise attention, combines them in a progressive manner to a dense continuous representation which is finally converted into a binary hash code with the help of individual and pairwise label information. The new enlarged texture patches from the TSN also help in data augmentation to alleviate the problem of insufficient texture data and are used to train the second stage of the network. Experiments on three public texture image retrieval datasets indicate the superiority of our texture synthesis guided hashing approach over existing state-of-the-art methods. Ayan Kumar Bhunia, Perla Sai Raj Kishore, Pranay Mukherjee, Abhirup Das, Partha Pratim Roy 0001 |
WACV | 5 |
| 2019 | Queuing theory guided intelligent traffic scheduling through video analysis using Dirichlet process mixture model
Santhosh Kelathodi Kumaran, Debi Prosad Dogra, Partha Pratim Roy 0001 |
Expert Syst. Appl. | 3 |
| 2019 | Computer vision-guided intelligent traffic signaling for isolated intersections
Santhosh Kelathodi Kumaran, Shrohan Mohapatra, Debi Prosad Dogra, Partha Pratim Roy 0001, Byung-Gyu Kim |
Expert Syst. Appl. | 4 |
| 2019 | Fingertip detection and tracking for recognition of air-writing in videos
Sohom Mukherjee, Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001 |
Expert Syst. Appl. | 5 |
| 2019 | Enhancing Text Using Emotion Detected from EEG Signals
Akash Gupta 0005, Harsh Sahu, Nihal Nanecha, Pradeep Kumar 0002, Partha Pratim Roy 0001, Victor Chang 0001 |
J. Grid Comput. | 5 |
| 2019 | Development of a clustering based fusion framework for locating the most consistent IrisCodes bits
Debanjan Sadhya, Kanjar De, Balasubramanian Raman, Partha Pratim Roy 0001 |
Inf. Sci. | 4 |
| 2019 | An intelligent recommendation system using gaze and emotion detection
Saurabh Jaiswal, Shubham Virmani, Vishal Sethi, Kanjar De, Partha Pratim Roy 0001 |
Multim. Tools Appl. | 5 |
| 2019 | Word searching in scene image and video frame in multi-script scenario using dynamic shape coding
Partha Pratim Roy 0001, Ayan Kumar Bhunia, Avirup Bhattacharyya, Umapada Pal 0001 |
Multim. Tools Appl. | 1 |
| 2019 | Signature verification approach using fusion of hybrid texture features
Ankan Bhunia, Alireza Alaei, Partha Pratim Roy 0001 |
Neural Comput. Appl. | 3 |
| 2019 | Age and gender classification using brain-computer interface
Barjinder Kaur, Partha Pratim Roy 0001 |
Neural Comput. Appl. | 3 |
| 2019 | Bag-of-visual-words for signature-based multi-script document retrieval
Ranju Mandal, Partha Pratim Roy 0001, Umapada Pal 0001, Michael Blumenstein |
Neural Comput. Appl. | 2 |
| 2019 | A deep one-shot network for query-based logo retrieval
Ayan Kumar Bhunia, Ankan Bhunia, Shuvozit Ghose, Abhirup Das, Partha Pratim Roy 0001, Umapada Pal 0001 |
Pattern Recognit. | 5 |
| 2019 | Script identification in natural scene image and video frames using an attention based Convolutional-LSTM network
Ankan Bhunia, Aishik Konwer, Ayan Kumar Bhunia, Abir Bhowmick, Partha Pratim Roy 0001, Umapada Pal 0001 |
Pattern Recognit. | 5 |
| 2019 | Likelihood learning in modified Dirichlet Process Mixture Model for video analysis
Santhosh Kelathodi Kumaran, Adyasha Chakravarty, Debi Prosad Dogra, Partha Pratim Roy 0001 |
Pattern Recognit. Lett. | 4 |
| 2019 | Recognizing gender from human facial regions using genetic algorithm
Avirup Bhattacharyya, Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra, Samarjit Kar |
Soft Comput. | 3 |
| 2019 | Sub-Stroke-Wise Relative Feature for Online Indic Handwriting RecognitionabstractThe main problem of Bangla (Bengali) and Devanagari handwriting recognition is the shape similarity of characters. There are only a few pieces of work on writer-independent cursive online Indian text recognition, and the shape similarity problem needs more attention from the researchers. To handle the shape similarity problem of cursive characters of Bangla and Devanagari scripts, in this article, we propose a new category of features called ‘ sub-stroke-wise relative feature ’ (SRF) which are based on relative information of the constituent parts of the handwritten strokes. Relative information among some of the parts within a character can be a distinctive feature as it scales up small dissimilarities and enhances discrimination among similar-looking shapes. Also, contextual anticipatory phenomena are automatically modeled by this type of feature, as it takes into account the influence of previous and forthcoming strokes. We have tested popular state-of-the-art feature sets as well as proposed SRF using various (up to 20,000-word) lexicons and noticed that SRF significantly outperforms the state-of-the-art feature sets for online Bangla and Devanagari cursive word recognition. Nilanjana Bhattacharya 0001, Partha Pratim Roy 0001, Umapada Pal 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2019 | Trajectory-Based Surveillance Analysis: A SurveyabstractDue to the advancement of camera hardware and machine learning techniques, video object tracking for surveillance has received noticeable attention from the computer vision research community. Object tracking and trajectory modeling have important applications in surveillance video analysis. For example, trajectory clustering, summarization or synopsis generation, and detection of anomalous or abnormal events in videos are mainly being exploited by the research community. However, barring one research work (which is almost a decade old), there is no recent review that emphasizes the use of video object trajectories, particularly in the perspective of visual surveillance. This paper presents a survey of trajectory-based surveillance applications with a focus on clustering, anomaly detection, summarization, and synopsis generation. The methods reviewed in this paper broadly summarize the abovementioned applications. The main purpose of this survey is to summarize the state-of-the-art video object trajectory analysis techniques used in the indoor and outdoor surveillance. Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Multimodal Gait Recognition With Inertial Sensor Data and Video Using Evolutionary AlgorithmabstractEvolutionary decision fusion has applications in biometric authentication and verification. Gray wolf optimizer (GWO) is one such evolutionary decision fusion approach that can be used to tune the fusion parameters in a multimodal data acquisition system. Human gait is a proven biometric trait with applications in security and authentication. However, acquiring human-gait data can be erroneous due to various factors and multimodal fusion of such erroneous gait data can be challenging. In this paper, we propose a new decision fusion-based approach to solve the above problem. Gait data is recorded simultaneously using motion sensors and visible-light camera. The signals of the motion sensors are modeled using a long short-term memory neural network and corresponding video recordings are processed using a three-dimensional convolutional neural network. GWO has been used to optimize the parameters during fusion. It has been chosen based on the underlying hunting strategy that leads to better approximation of the solution. Interestingly, in our case it converges quicker than other optimization techniques such as genetic algorithm or particle swarm optimization. To test the model, a dataset involving 23 males and females has been recorded while they perform four different types of walks, including, normal walk, fast walk, walking while listening to music, and walking while watching multimedia content on a mobile. An overall accuracy of 91.3% has been recorded across all test scenarios. Results reveal that the proposed study can further be explored to design robust gait biometric systems. Pradeep Kumar 0002, Subham Mukherjee, Rajkumar Saini, Pallavi Kaushik, Partha Pratim Roy 0001, Debi Prosad Dogra |
IEEE Trans. Fuzzy Syst. | 5 |
| 2019 | Temporal Unknown Incremental Clustering Model for Analysis of Traffic Surveillance VideosabstractOptimized scene representation is an important characteristic of a framework for detecting abnormalities on live videos. One of the challenges for detecting abnormalities in live videos is real-time detection of objects in a non-parametric way. Another challenge is to efficiently represent the state of objects temporally across frames. In this paper, a Gibbs sampling-based heuristic model referred to as temporal unknown incremental clustering has been proposed to cluster pixels with motion. Pixel motion is first detected using optical flow and a Bayesian algorithm has been applied to associate pixels belonging to a similar cluster in subsequent frames. The algorithm is fast and produces accurate results in Θ(kn) time, where k is the number of clusters and n the number of pixels. Our experimental validation with publicly available data sets reveals that the proposed framework has good potential to open up new opportunities for real-time traffic analysis. Santhosh Kelathodi Kumaran, Debi Prosad Dogra, Partha Pratim Roy 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | A novel point-line duality feature for trajectory classification
Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra |
Vis. Comput. | 2 |
| 2018 | Perceptual Conditional Generative Adversarial Networks for End-to-End Image Colourization
Shirsendu Sukanta Halder, Kanjar De, Partha Pratim Roy 0001 |
ACCV (2) | 3 |
| 2018 | Air Signature Recognition Using Deep Convolutional Neural Network-Based Sequential ModelabstractDeep convolutional neural networks are becoming extremely popular in classification, especially when the inputs are non-sequential in nature. Though it seems unrealistic to adopt such networks as sequential classifiers, however, researchers have started to use them for applications that primarily deal with sequential data. It is possible, if the sequential data can be represented in the conventional way the inputs are provided in CNNs. Signature recognition is one of the important tasks for biometric applications. Signatures represent the signer's identity. Air signatures can make traditional biometric systems more secure and robust than conventional pen-paper or stylus guided interfaces. In this paper, we propose a new set of geometrical features to represent 3D air signatures captured using Leap motion sensor. The features are then arranged such that they can be fed to a deep convolutional neural network architecture with application specific tuning of the model parameters. It has been observed that the proposed features in combination with the CNN architecture can act as a good sequential classifier when tested on a moderate size air signature dataset. Experimental results reveal that the proposed biometric system performs better as compared to the state-of-the-art geometrical features with average accuracy improvement of 4%. Santosh Kumar Behera, Ajaya Kumar Dash, Debi Prosad Dogra, Partha Pratim Roy 0001 |
ICPR | 4 |
| 2018 | Handwriting Trajectory Recovery using End-to-End Deep Encoder-Decoder NetworkabstractIn this paper, we introduce a novel technique to recover the pen trajectory of offline characters which is a crucial step for handwritten character recognition. Generally, online acquisition approach has more advantage than its offline counterpart as the online technique keeps track of the pen movement. Hence, pen tip trajectory retrieval from offline text can bridge the gap between online and offline methods. Our proposed framework employs sequence to sequence model which consists of an encoder-decoder LSTM module. The proposed encoder module consists of Convolutional LSTM network, which takes an offline character image as the input and encodes the feature sequence to a hidden representation. The output of the encoder is fed to a decoder LSTM and we get the successive coordinate points from every time step of the decoder LSTM. Although the sequence to sequence model is a popular paradigm in various computer vision and language translation tasks, the main contribution of our work lies in designing an end-to-end network for a decade old popular problem in document image analysis community. Tamil, Telugu and Devanagari characters of LIPI Toolkit dataset are used for our experiments. Our proposed method has achieved superior performance compared to the other conventional approaches. Ayan Kumar Bhunia, Abir Bhowmick, Ankan Bhunia, Aishik Konwer, Prithaj Banerjee, Partha Pratim Roy 0001, Umapada Pal 0001 |
ICPR | 6 |
| 2018 | Word Level Font-to-Font Image Translation using Convolutional Recurrent Generative Adversarial NetworksabstractConversion of one font to another font is very useful in real life applications. In this paper, we propose a Convolutional Recurrent Generative model to solve the word level font transfer problem. Our network is able to convert the font style of any printed text images from its current font to the required font. The network is trained end-to-end for the complete word images. Thus it eliminates the necessary pre-processing steps, like character segmentations. We extend our model to conditional setting that helps to learn one-to-many mapping function. We employ a novel convolutional recurrent model architecture in the Generator that efficiently deals with the word images of arbitrary width. It also helps to maintain the consistency of the final images after concatenating the generated image patches of target font. Besides, the Generator and the Discriminator network, we employ a Classification network to classify the generated word images of converted font style to their subsequent font categories. Most of the earlier works related to image translation are performed on square images. Our proposed architecture is the first of its kind which can handle images of varying widths. Word images generally have varying width depending on the number of characters present. Hence, we test our model on a synthetically generated font dataset. We compare our method with some of the state-of-the-art methods for image translation. The superior performance of our network on the same dataset proves the ability of our model to learn the font distributions. Ankan Bhunia, Ayan Kumar Bhunia, Prithaj Banerjee, Aishik Konwer, Abir Bhowmick, Partha Pratim Roy 0001, Umapada Pal 0001 |
ICPR | 6 |
| 2018 | Staff line Removal using Generative Adversarial NetworksabstractStaff line removal is a crucial pre-processing step in Optical Music Recognition. In this paper we propose a novel approach for staff line removal, based on Generative Adversarial Networks. We convert staff line images into patches and feed them into a U-Net, used as Generator. The Generator intends to produce staff-less images at the output. Then the Discriminator does binary classification and differentiates between the generated fake staff-less image and real ground truth staff less image. For training, we use a Loss function which is a weighted combination of L2 loss and Adversarial loss. L2 loss minimizes the difference between real and fake staff-less image. Adversarial loss helps to retrieve more high quality textures in generated images. Thus our architecture supports solutions which are closer to ground truth and it reflects in our results. For evaluation we consider the ICDAR/GREC 2013 staff removal database. Our method achieves superior performance in comparison to other conventional approaches on the same dataset. Aishik Konwer, Ayan Kumar Bhunia, Abir Bhowmick, Ankan Bhunia, Prithaj Banerjee, Partha Pratim Roy 0001, Umapada Pal 0001 |
ICPR | 6 |
| 2018 | Surveillance scene representation and trajectory abnormality detection using aggregation of multiple concepts
Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001 |
Expert Syst. Appl. | 4 |
| 2018 | Local Neighborhood Intensity Pattern-A new texture feature descriptor for image retrieval
Prithaj Banerjee, Ayan Kumar Bhunia, Avirup Bhattacharyya, Partha Pratim Roy 0001, M. Subrahmanyam 0001 |
Expert Syst. Appl. | 4 |
| 2018 | Fast recognition and verification of 3D air signatures using convex hulls
Santosh Kumar Behera, Debi Prosad Dogra, Partha Pratim Roy 0001 |
Expert Syst. Appl. | 3 |
| 2018 | A segmental HMM based trajectory classification using genetic algorithm
Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra |
Expert Syst. Appl. | 2 |
| 2018 | Date-field retrieval in scene image and video frames using text enhancement and shape coding
Partha Pratim Roy 0001, Ayan Kumar Bhunia, Umapada Pal 0001 |
Neurocomputing | 1 |
| 2018 | Multi-oriented text detection and verification in video frames and scene images
Aneeshan Sain, Ayan Kumar Bhunia, Partha Pratim Roy 0001, Umapada Pal 0001 |
Neurocomputing | 3 |
| 2018 | A novel framework of continuous human-activity recognition using Kinect
Rajkumar Saini, Pradeep Kumar 0002, Partha Pratim Roy 0001, Debi Prosad Dogra |
Neurocomputing | 3 |
| 2018 | Independent Bayesian classifier combination based sign language recognition using facial expression
Pradeep Kumar 0002, Partha Pratim Roy 0001, Debi Prosad Dogra |
Inf. Sci. | 2 |
| 2018 | Don't just sign use brain too: A novel multimodal approach for user identification and verification
Rajkumar Saini, Barjinder Kaur, Priyanka Singh 0001, Pradeep Kumar 0002, Partha Pratim Roy 0001, Balasubramanian Raman |
Inf. Sci. | 5 |
| 2018 | A position and rotation invariant framework for sign language recognition (SLR) using Kinect
Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra |
Multim. Tools Appl. | 3 |
| 2018 | Analysis of 3D signatures recorded using leap motion sensor
Santosh Kumar Behera, Debi Prosad Dogra, Partha Pratim Roy 0001 |
Multim. Tools Appl. | 3 |
| 2018 | Text recognition in scene image and video frame using Color Channel selection
Ayan Kumar Bhunia, Partha Pratim Roy 0001, Balasubramanian Raman, Umapada Pal 0001 |
Multim. Tools Appl. | 3 |
| 2018 | Frame selection for OCR from video stream of book flipping
Dibyayan Chakraborty, Partha Pratim Roy 0001, Rajkumar Saini, José M. Álvarez 0004, Umapada Pal 0001 |
Multim. Tools Appl. | 2 |
| 2018 | Exercise classification and event segmentation in Hammersmith Infant Neurological Examination videos
Abdul Fatir Ansari, Partha Pratim Roy 0001, Debi Prosad Dogra |
Mach. Vis. Appl. | 2 |
| 2018 | Cross-language framework for word recognition and spotting of Indic scripts
Ayan Kumar Bhunia, Partha Pratim Roy 0001, Akash Mohta, Umapada Pal 0001 |
Pattern Recognit. | 2 |
| 2018 | A lexicon-free approach for 3D handwriting recognition using classifier combination
Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Umapada Pal 0001 |
Pattern Recognit. Lett. | 3 |
| 2018 | Envisioned speech recognition using EEG sensors
Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Pawan Kumar Sahu, Debi Prosad Dogra |
Pers. Ubiquitous Comput. | 3 |
| 2018 | Unsupervised classification of erroneous video object trajectories
Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001 |
Soft Comput. | 4 |
| 2017 | Stroke-Order Normalization for Online Bangla Handwriting RecognitionabstractStroke order variation within characters is one of the difficult problems in online Bangla handwriting recognition. Moreover, in Bangla, character parts are written in zone-wise manner. Character parts written in the middle zone are generally cursive, while character parts in the upper and lower zones are written using delayed strokes. As online recognition depends on the order of writing, words written with different stroke-order are treated as different words to the online recognizer. To the best of our knowledge, no work has been reported on stroke-order normalization for any Indic script though it is an important aspect of online recognition. In this paper, we propose a stroke-order normalization method for Bangla online recognition using offline and online information. Here, at first, based on the offline information, sub-strokes in a word are ordered according to their relative positions. This results in similar stroke-order among the different instances of the same word. Next, online information of each ordered sub-stroke is used for feature extraction. This normalization approach has several significant advantages, e.g. (i) characters/words having any stroke order can be recognized, (ii) number of word classes is reduced, etc. We have tested our method on a dataset of 6000 words and obtained 74.65% and 90.53% word recognition accuracies, respectively, before and after stroke-order normalization. Thus, stroke-order normalization has enhanced the recognition result drastically (15.88%). Nilanjana Bhattacharya 0001, Umapada Pal 0001, Partha Pratim Roy 0001 |
ICDAR | 3 |
| 2017 | Keyword spotting in doctor's handwriting on medical prescriptions
Partha Pratim Roy 0001, Ayan Kumar Bhunia, Ayan Das 0001, Prithviraj Dhar, Umapada Pal 0001 |
Expert Syst. Appl. | 1 |
| 2017 | HMM-based writer identification in music score documents without staff-line removal
Partha Pratim Roy 0001, Ayan Kumar Bhunia, Umapada Pal 0001 |
Expert Syst. Appl. | 1 |
| 2017 | A multimodal framework for sensor based sign language recognition
Pradeep Kumar 0002, Himaanshu Gauba, Partha Pratim Roy 0001, Debi Prosad Dogra |
Neurocomputing | 3 |
| 2017 | A bio-signal based framework to secure mobile devices
Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra |
J. Netw. Comput. Appl. | 3 |
| 2017 | Rotation and script independent text detection from video frames using sub pixel mapping
Anshul Mittal, Partha Pratim Roy 0001, Priyanka Singh 0001, Balasubramanian Raman |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Tandem hidden Markov models using deep belief networks for offline handwriting recognitionabstractUnconstrained offline handwriting recognition is a challenging task in the areas of document analysis and pattern recognition. In recent years, to sufficiently exploit the supervisory information hidden in document images, much effort has been made to integrate multi-layer perceptrons (MLPs) in either a hybrid or a tandem fashion into hidden Markov models (HMMs). However, due to the weak learnability of MLPs, the learnt features are not necessarily optimal for subsequent recognition tasks. In this paper, we propose a deep architecture-based tandem approach for unconstrained offline handwriting recognition. In the proposed model, deep belief networks are adopted to learn the compact representations of sequential data, while HMMs are applied for (sub-)word recognition. We evaluate the proposed model on two publicly available datasets, i.e., RIMES and IFN/ENIT, which are based on Latin and Arabic languages respectively, and one dataset collected by ourselves called Devanagari (an Indian script). Extensive experiments show the advantage of the proposed model, especially over the MLP-HMMs tandem approaches. Partha Pratim Roy 0001, Guoqiang Zhong 0001, Mohamed Cheriet |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2017 | A Novel framework of EEG-based user identification by analyzing music-listening behavior
Barjinder Kaur, Partha Pratim Roy 0001 |
Multim. Tools Appl. | 3 |
| 2017 | 3D text segmentation and recognition using leap motion
Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra |
Multim. Tools Appl. | 3 |
| 2017 | A multimodal biometric watermarking system for digital images in redundant discrete wavelet transform
Priyanka Singh 0001, Balasubramanian Raman, Partha Pratim Roy 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Analysis of EEG signals and its application to neuromarketing
Mahendra Yadava, Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra |
Multim. Tools Appl. | 4 |
| 2017 | Prediction of advertisement preference by fusing EEG response and sentiment analysis
Himaanshu Gauba, Pradeep Kumar 0002, Partha Pratim Roy 0001, Priyanka Singh 0001, Debi Prosad Dogra, Balasubramanian Raman |
Neural Networks | 3 |
| 2017 | Combination of context-dependent bidirectional long short-term memory classifiers for robust offline handwriting recognition
Youssouf Chherawala, Partha Pratim Roy 0001, Mohamed Cheriet |
Pattern Recognit. Lett. | 2 |
| 2017 | Coupled HMM-based multi-sensor data fusion for sign language recognition
Pradeep Kumar 0002, Himaanshu Gauba, Partha Pratim Roy 0001, Debi Prosad Dogra |
Pattern Recognit. Lett. | 3 |
| 2017 | Cleaning of Online Bangla Free-form Handwritten TextabstractIn the normal free-form handwritten text, repetition (repeated writing of the same stroke several times in the same place), over-writing, and crossing out are very common. In this article, we call the presence of these three types of writing as “noise.” Cleaning to extract useful text from such types of noisy text is an important task for robust recognition. To the best of our knowledge, no work has been reported on cleaning of such noise from online text in any scripts and hence, in this article, we propose an automatic text-cleaning approach for online handwriting recognition. Here, at first, crossing out noise with straight strike-through lines is detected using the straightness criteria of online strokes. Next, regions containing repetition, over-writing, and other types of crossing out are located using the positional information of the overlapping strokes. Stroke density, self-intersections of strokes etc. are computed from the strokes of located regions to predict the type of noise and this type of information is used as follows for their cleaning. For cleaning of crossing outs, all strokes of the crossing-out region are removed. For cleaning repetition and over-writing, strokes written earlier are removed, keeping the latest strokes. Finally, delayed strokes are properly arranged and word is passed to online recognizer. Though recognition of free-form handwriting is quite difficult, in this attempt, we obtained up to 70.71% improvement in word-recognition accuracy after noise cleaning. Nilanjana Bhattacharya 0001, Umapada Pal 0001, Partha Pratim Roy 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2017 | Moving object detection using modified temporal differencing and local fuzzy thresholding
Nihal Paul, Abhishek Midya, Partha Pratim Roy 0001, Debi Prosad Dogra |
J. Supercomput. | 4 |
| 2016 | Comparison of Zone-Features for Online Bengali and Devanagari Word Recognition Using HMMabstractThis paper presents a comparative study of three feature extraction approaches for online handwritten word recognition of two major Indic scripts-Bengali and Devanagari using Hidden Markov Model (HMM). First approach uses feature extraction from whole stroke without local zone division after segmenting the word into its basic strokes. Whereas, other two approaches consider the segmentation of a word into its basic strokes and a local zone wise analysis of each online stroke. Among these two zone wise local features, one takes into account structural and directional features and other uses dominant points, detected from strokes using slope angles, to find the local features. These features are studied in HMM-based word recognition platform. From the comparative study of the word recognition results, we have noted that dominant point based local feature extraction provides best accuracies for both Bengali and Devanagari scripts. We have obtained 90.23% and 93.82% accuracies for Bengali and Devanagari scripts respectively. Rajib Ghosh, Partha Pratim Roy 0001 |
ICFHR | 2 |
| 2016 | A Simple and Efficient Arrowhead Detection Technique in Biomedical ImagesabstractIn biomedical documents/publications, medical images tend to be complex by nature and often contain several regions that are annotated using arrows. In this context, an automated arrowhead detection is a critical precursor to region-of-interest (ROI) labeling and image content analysis. To detect arrowheads, in this paper, images are first binarized using fuzzy binarization technique to segment a set of candidates based on connected component (CC) principle. To select arrow candidates, we use convexity defect-based filtering, which is followed by template matching via dynamic time warping (DTW). The DTW similarity score confirms the presence of arrows in the image. Our test results on biomedical images from imageCLEF 2010 collection shows the interest of the technique, and can be compared with previously reported state-of-the-art results. KC Santosh, Naved Alam, Partha Pratim Roy 0001, Laurent Wendling, Sameer K. Antani, George R. Thoma |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2016 | HMM-based Indic handwritten word recognition using zone segmentation
Partha Pratim Roy 0001, Ayan Kumar Bhunia, Ayan Das 0001, Umapada Pal 0001 |
Pattern Recognit. | 1 |
| 2016 | Feature Set Evaluation for Offline Handwriting Recognition Systems: Application to the Recurrent Neural Network ModelabstractThe performance of handwriting recognition systems is dependent on the features extracted from the word image. A large body of features exists in the literature, but no method has yet been proposed to identify the most promising of these, other than a straightforward comparison based on the recognition rate. In this paper, we propose a framework for feature set evaluation based on a collaborative setting. We use a weighted vote combination of recurrent neural network (RNN) classifiers, each trained with a particular feature set. This combination is modeled in a probabilistic framework as a mixture model and two methods for weight estimation are described. The main contribution of this paper is to quantify the importance of feature sets through the combination weights, which reflect their strength and complementarity. We chose the RNN classifier because of its state-of-the-art performance. Also, we provide the first feature set benchmark for this classifier. We evaluated several feature sets on the IFN/ENIT and RIMES databases of Arabic and Latin script, respectively. The resulting combination model is competitive with state-of-the-art systems. Youssouf Chherawala, Partha Pratim Roy 0001, Mohamed Cheriet |
IEEE Trans. Cybern. | 2 |
| 2015 | A comparative study of features for handwritten Bangla text recognitionabstractRecognition of Bangla handwritten text is difficult due to its complex nature of having modifiers and headlines features. This paper presents a comparative study of different features namely LGH (Local Gradient of Histogram), PHOG (Pyramid Histogram of Oriented Gradient), GABOR, G-PHOG (Combined GABOR and PHOG) and profile feature by Marti-Bunke when applied in middle zone recognition of Bangla words using Hidden Markov Model (HMM) based framework. For this purpose, a zone segmentation method is applied to extract the busy (middle) zones of handwritten words and features are extracted from the middle zone. The system has been tested on a sufficiently large and variation-rich dataset consisting of 11,253 training and 3,856 testing data. From the experiment, it has been noted that PHOG feature outperforms other features in middle zone recognition. Since PHOG feature outperform others, we use this feature for full word recognition, For this purpose initially upper and lower zone components are recognized by PHOG features and SVM classifier. Finally, the zone-wise results are combined by the context information of the corresponding components in each zone to obtain the word level recognition. Ayan Kumar Bhunia, Ayan Das 0001, Partha Pratim Roy 0001, Umapada Pal 0001 |
ICDAR | 3 |
| 2015 | Generation of synthetic training data for handwritten Indic script recognitionabstractThis paper presents a novel approach to create synthetic dataset for word recognition systems. Our purpose is to improve performance of off-line handwritten text recognizers by providing it with additional synthetic training data. Due to lack of proper data-set for many languages it becomes hard to train recognition systems. To solve such problems synthetic handwriting could be used to expand the existing training dataset. Any available digital data from online newspaper and such sources can be used to generate this synthetic data. The digital data is distorted in such a way that the underlying pattern is conserved for identification of the word by both machine and human user. The images hence produced can be used to train any classification system for handwriting recognition. This data can be used independently to train the system or be combined with natural handwritten data to augment the original dataset and improve the accuracy of the results. We experimented using only synthetic data obtaining high recognition accuracy in both character and word recognition. The data was tested on 3 Indian scripts for numerals- Hindi, Bengali and Telugu, and 1 script-Hindi for words, the results achieved hence are highly promising. Shivansh Gaur, Siddhant Sonkar, Partha Pratim Roy 0001 |
ICDAR | 3 |
| 2015 | Study of two zone-based features for online Bengali and Devanagari character recognitionabstractThis paper presents two zone-based feature extraction approaches for online handwritten character recognition of two major Indic scripts-Bengali and Devanagari. Here, each stroke of an online character is divided into a number of local zones. In the first approach, named Zone wise structural and directional features (ZSD), structural and directional features are extracted for each stroke in each of these local zones. In the second approach, named Zone wise slopes of dominant points (ZSDP), the dominant points are detected first from each stroke and next the slope angles between consecutive dominant points are calculated and features are extracted in these local zones. Next, these features are fed to SVM classifier for stroke recognition. The constituent stroke combinations of characters are matched with training data and characters are recognized accordingly. Using ZSD, the recognition performances for Bengali (9,800 test data) and Devanagari (10,000 test data) scripts are 87.48% and 85.10% and with ZSDP, the accuracies are 92.48% and 90.63% respectively. Rajib Ghosh, Partha Pratim Roy 0001 |
ICDAR | 2 |
| 2015 | Date field extraction from handwritten documents using HMMsabstractAutomatic document interpretation and retrieval is an important task to access handwritten digitized document repositories. In documents, the date is an important field and it has various applications such as date-wise document indexing/retrieval. In this paper a framework has been proposed for automatic date field extraction from handwritten documents. In order to design the system, sliding window-wise Local Gradient Histogram (LGH)-based features and a character-level Hidden Markov Model (HMM)-based approach have been applied for segmentation and recognition. Individual date components such as month-word (month written in word form i.e. January, Jan, etc.), numeral, punctuation and contraction categories are segmented and labelled from a text line. Next, a Histogram of Gradient (HoG)-based features and a Support Vector Machine (SVM)- based classifier have been used to improve the results obtained from the HMM-based recognition system. Subsequently, both numeric and semi-numeric regular expressions of date patterns have been considered for undertaking date pattern extraction in labelled components. The experiments are performed on an English document dataset and the encouraging results obtained from the approach indicate the effectiveness of the proposed system. Ranju Mandal, Partha Pratim Roy 0001, Umapada Pal 0001, Michael Blumenstein |
ICDAR | 2 |
| 2015 | Multi-lingual text recognition from video framesabstractText recognition from video frames is a challenging task due to low resolution, blur, complex and coloured backgrounds, noise, to mention a few. Consequently, the traditional ways of text recognition from scanned documents having simple backgrounds fails when applied to video text. Although there are various techniques available for text recognition from handwritten and printed documents with simple backgrounds, text recognition from video frames has not been comprehensively investigated, especially for multi-lingual videos. In this paper, we present a technique for multi-lingual video text recognition which involves script identification in the first stage, followed by word and character recognition, and finally the results are refined using a post-processing technique. Considering the inherent problems in videos, a Spatial Pyramid Matching (SPM) based technique, using patch-based SIFT descriptors and SVM classifier, is employed for script identification. In the next stage, a Hidden Markov Model (HMM) based approach is used for word and character recognition, which utilizes the context information. Finally, a lexicon-based post-processing technique is applied to verify and refine the word recognition results. The proposed method was tested on a dataset comprising of 4800 words from three different scripts, namely, Roman (English), Hindi and Bengali. The script identification results obtained are encouraging. The word and character recognition results are also encouraging considering the complexity and problems associated with video text processing. Nabin Sharma, Ranju Mandal, Rabi Sharma, Partha Pratim Roy 0001, Umapada Pal 0001, Michael Blumenstein |
ICDAR | 4 |
| 2015 | Bayesian classifier for multi-oriented video text recognition system
Sangheeta Roy, Palaiahnakote Shivakumara, Partha Pratim Roy 0001, Umapada Pal 0001, Chew Lim Tan, Tong Lu 0002 |
Expert Syst. Appl. | 3 |
| 2015 | Multi-lingual date field extraction for automatic document retrieval by machine
Ranju Mandal, Partha Pratim Roy 0001, Umapada Pal 0001, Michael Blumenstein |
Inf. Sci. | 2 |
| 2015 | Word spotting in historical documents using primitive codebook and dynamic programming
Partha Pratim Roy 0001, Frédéric Rayar, Jean-Yves Ramel |
Image Vis. Comput. | 1 |
| 2014 | Multi-oriented Text Recognition in Graphical Documents Using HMMabstractThe text lines in graphical documents (e.g., maps, engineering drawings), artistic documents etc., are often annotated in curve lines to illustrate different locations or symbols. For the optical character recognition of such documents, individual text lines from the documents need to be extracted and recognized. Due to presence of multi-oriented characters in such non-structured layout, word recognition is a challenging task. In this paper, we present an approach towards the recognition of scale and orientation invariant text words in graphical documents using Hidden Markov Models (HMM). First, a line extraction method is applied to segment text lines and the method is based on the foreground and background information of the text components. To effectively utilize the background information, a water reservoir concept is used here. For recognition of curved text lines, a path of sliding window is estimated and features extracted from the sliding window are fed to the HMM system for recognition. Local gradient histogram (LGH) based frame-wise feature is used in HMM. The experimental results are evaluated on a dataset of graphical words and we have obtained encouraging results. Partha Pratim Roy 0001, Sangheeta Roy, Umapada Pal 0001 |
Document Analysis Systems | 1 |
| 2014 | A New Method for Writer Identification Based on Histogram Symbolic RepresentationabstractIn this paper, a new model-based writer identification scheme using histogram symbolic representation approach is proposed. In the proposed scheme, initially, some pre-processing techniques are employed to enhance image quality and extract text-lines from each handwritten document image. For each extracted text-line, a set of 92 features are computed based on analysis of connected component, enclosed region, lower and upper contours, fractal code, and Curve let. Considering the extracted feature vectors, a histogram is created for each feature of every writer as a histogram-valued symbolic data. This process results in a handwriting style model for each individual that consists of a set of histograms. To evaluate the proposed scheme, two different handwritten datasets written in two different scripts (Kannada as an Indian based script and English) were used. The first dataset contains 228 pages written in Kannada by 57 people. The other one is the dataset used in SigWiComp2013 composed of 330 document pages written in English by 55 individuals. The same criteria used in the SigWiComp2013 were followed in our evaluation strategy. Concerning the Kannada dataset, an F-measure of 92.79% was obtained when 114 documents were used in learning stage and the rest (114) were used for testing. For the SigWiComp2013 dataset an F-measure of 26.67% was obtained that is fairly comparable to the best result reported in the literature. Alireza Alaei, Partha Pratim Roy 0001 |
ICFHR | 2 |
| 2014 | Writer Identification in Music Score Documents without Staff-Line RemovalabstractWriter identification from musical scores is a challenging task. A few pieces of work on writer identification in musical sheets have been published in the literature but to the best of our knowledge all these work were performed after removal of staff lines from the musical scores. In this paper we propose a symbol-independent writer identification framework using HMM in music score without removing staff lines. The writing style of each writer is modeled using sliding window based LGH feature. To identify the writer of an input musical sheet, all musical lines are fed to writer specific HMM models and each model return a log-likelihood score for the given input. These log-likelihood scores from each HMM models are compared and the writer corresponding to the maximum score is considered as identified writer of the test sample. Next, a page level log-likelihood score is computed for writer identification in each page sample. We have compared our proposed approach with Gaussian Mixture Models (GMMs) based writer identification system in CVC-MUSCIMA data set. The results obtained from an experiment on 50 writers show that the HMM based approach outperforms GMM based approach. Anirban Jyoti Hati, Partha Pratim Roy 0001, Umapada Pal 0001 |
ICFHR | 2 |
| 2014 | Deep-Belief-Network Based Rescoring Approach for Handwritten Word RecognitionabstractThis paper presents a novel verification approach towards improvement of handwriting recognition systems using a word hypotheses rescoring scheme by Deep Belief Networks (DBNs). A recurrent neural network based sequential text recognition system is used at first to provide the N-best recognition hypotheses of word images. Word hypotheses are aligned with the word image to obtain the character boundaries. Then, a verification approach using a DBN classifier is performed for each character segments. DBNs are recently proved to be very effective for a variety of machine learning problems. The character probabilities obtained from DBNs are next combined with the base recognition system. Finally, the N-best recognition hypotheses list is reranked according to the new score. We have compared our proposed approach with an MLP based rescoring approach on the Rimes dataset. The results obtained show that the verification approach using DBNs outperforms that of MLP systems. Partha Pratim Roy 0001, Youssouf Chherawala, Mohamed Cheriet |
ICFHR | 1 |
| 2014 | A Novel Approach of Bangla Handwritten Text Recognition Using HMMabstractThis paper presents a novel approach for offline Bangla (Bengali) handwritten word recognition by Hidden Markov Model (HMM). Due to the presence of complex features such as headline, vowels, modifiers, etc., character segmentation in Bangla script is not easy. Also, the position of vowels and compound characters make the segmentation task of words into characters very complex. To take care of these problems we propose a novel method considering a zone-wise break up of words and next perform HMM based recognition. In particular, the word image is segmented into 3 zones, upper, middle and lower, respectively. The components in middle zone are modeled using HMM. By this zone segmentation approach we reduce the number of distinct component classes compared to total number of classes in Bangla character set. Once the middle zone portion is recognized, HMM based forced alignment is applied in this zone to mark the boundaries of individual components. The segmentation paths are extended later to other zones. Next, the residue components, if any, in upper and lower zones in their respective boundary are combined to achieve the final word level recognition. We have performed a preliminary experiment on a dataset of 10,120 Bangla handwritten words and found that the proposed approach outperforms the custom way of HMM based recognition. Partha Pratim Roy 0001, Sangheeta Roy, Umapada Pal 0001, Fumitaka Kimura |
ICFHR | 1 |
| 2014 | Context-dependent blstm models. Application to offline handwriting recognitionabstractThe BLSTM model has been recently introduced for sequence labeling tasks and provides state-of-the-art performance for handwriting recognition. Its recurrent connections integrate the context at the feature level over a long range. Nevertheless, this context is not explicitly modeled at the label level. Explicit context-modeling strategies have been applied to HMMs with improvement of the recognition rate. In this paper, we study the effect of context modeling on the performance of the BLSTM model. The baseline BLSTM, with context-independent character label, is compared with two context-dependent BLSTM, one modeling the left context and the other the right context. The results show that context-dependent models provide an improvement of the recognition rate, and demonstrate the ability of the BLSTM model to deal with a large number of models, without clustering. We tested our models on the RIMES database of Latin script documents. Youssouf Chherawala, Partha Pratim Roy 0001, Mohamed Chenet |
ICIP | 2 |
| 2014 | Word searching in unconstrained layout using character pair coding
Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001 |
Int. J. Document Anal. Recognit. | 1 |
| 2013 | Feature Design for Offline Arabic Handwriting Recognition: Handcrafted vs Automated?abstractIn handwriting recognition, design of relevant feature is a very important but daunting task. On one hand, handcraft design of features is difficult, depending on expert knowledge and on heuristics. On the other hand, biologically inspired neural networks are able to learn automatically features from the input image, but requires a good underlying model. The goal of this paper is to evaluate the performance of automatically learned features compared to handcrafted features, as they provide a promising alternative to the difficult task of features handcrafting. In this work, the recognition model is based on the long short-term memory (LSTM) and connectionist temporal classification (CTC) neural networks. This model has been shown to outperform the well-known HMM model for various handwriting tasks, thanks to its reliable probabilistic modeling. In its multidimensional form, called MDLSTM, this network is able to automatically learn features from the input image. For evaluation, we compare the MDLSTM learned features and four state-of-the-art handcrafted features. The IFN/ENIT database has been used as benchmark for Arabic word recognition, where the results are promising. Youssouf Chherawala, Partha Pratim Roy 0001, Mohamed Cheriet |
ICDAR | 2 |
| 2013 | Handwritten Musical Document Retrieval Using Music-Score SpottingabstractIn this paper, we present a novel approach for retrieval of handwritten musical documents using a query sequence/word of musical scores. In our algorithm, the musical score-words are described as sequences of symbols generated from a universal codebook vocabulary of musical scores. Staff lines are removed first from musical documents using structural analysis of staff lines and symbol codebook vocabulary is created in offline. Next, using this symbol codebook the music symbol information in each document image is encoded. Given a query sequence of musical symbols in a musical score-line, the symbols in the query are searched in each of these encoded documents. Finally, a sub-string matching algorithm is applied to find query words. For codebook, two different feature extraction methods namely: Zernike Moments and 400 dimensional gradient features are tested and two unsupervised classifiers using SOM and K-Mean are evaluated. The results are compared with a baseline approach of DTW. The performance is measured on a collection of handwritten musical documents and results are promising. Rakesh Malik, Partha Pratim Roy 0001, Umapada Pal 0001, Fumitaka Kimura |
ICDAR | 2 |
| 2013 | A Two-Stage Approach for Word Spotting in Graphical DocumentsabstractPresence of multi-oriented characters, connected characters with graphical lines, intersection of text and symbols with graphical lines/curves etc. are very common in graphical documents. As a result word spotting in graphical documents is still a challenging task that we try to solve (partially) in this paper. The proposed approach proceeds in two stages. In the first stage, recognition of isolated components is done using rotation invariant features and an SVM classifier. The characters having good recognition score and match in the query string are first selected for initial spotting. Because of structural complexity of graphical documents as well as of touching components, we may miss some of the query characters during initial spotting in some documents. In that case, based on the position, size and orientation of the recognized characters in the input document image, regions where missing characters may be located (candidate regions) are defined. In the second stage, Scale Invariant Feature Transform (SIFT) is used to find those missing characters in the candidate regions for possible spotting. Finally, using the position, size, orientation as well as intercharacter gap information of the recognized components, spotting is validated. Experimental results demonstrate that the method is efficient to locate a query word in multi-oriented and/or touching graphical documents. Arundhati Tarafdar, Umapada Pal 0001, Partha Pratim Roy 0001, Nicolas Ragot, Jean-Yves Ramel |
ICDAR | 3 |
| 2013 | Signature segmentation and recognition from scanned documentsabstractSignature as a query is important for content-based document image retrieval from a scanned document repository. This paper presents a two-stage approach towards automatic signature segmentation and recognition from scanned document images. In the first stage, signature blocks are segmented from the document using word-wise component extraction and classification. Gradient based features are extracted from each component at the word level to perform the classification task. In the 2nd stage, SIFT (Scale-Invariant Feature Transform) descriptors and Spatial Pyramid Matching (SPM)-based approaches are used for signature recognition. Support Vector Machines (SVMs) are employed as the classifier for both levels in this experiment. The experiments are performed on the publicly available “Tobacco-800” and GPDS [1] datasets and the results obtained from the experiments are promising. Ranju Mandal, Partha Pratim Roy 0001, Umapada Pal 0001, Michael Blumenstein |
ISDA | 2 |
| 2012 | An Efficient Coarse-to-Fine Indexing Technique for Fast Text Retrieval in Historical DocumentsabstractIn this paper, we present a fast text retrieval system to index and browse degraded historical documents. The indexing and retrieval strategy is designed in a two level, coarse-to-fine approach, to increase the speed of the retrieval process. During the indexing step, the text parts in the images are encoded into sequences of primitives, obtained from two different codebooks: a coarse one corresponding to connected components and a fine one corresponding to glyph primitives. A glyph consists of a single character or a part of a character according to the shape complexity. During the querying step, the coarse and the fine signature are generated from the query image using both codebooks. Then, a bi-level approximate string matching algorithm is applied to find similar words, using coarse approach first, and then the fine approach if necessary, by exploiting predetermined hypothetical locations. An experimental evaluation on datasets of real life document images, gathered from historical books of different scripts, demonstrated the speed improvement and good accuracy in presence of degradation. Partha Pratim Roy 0001, Frédéric Rayar, Jean-Yves Ramel |
Document Analysis Systems | 1 |
| 2012 | Signature Based Document Retrieval Using GHT of Background InformationabstractThis paper deals with signature based document retrieval from documents with cluttered background. Here, a signature object is characterized by spatial features computed from recognition result of background blobs. The background blobs are computed by analyzing character holes and water reservoir zones in different directions. For the indexation purpose, a codebook of the background blobs is created using a set of training data. Zernike Moment feature is extracted from each blob and a K-Mean clustering algorithm is used to create the codebook of blobs. During retrieval, Generalized Hough Transform (GHT) is used to detect the query signature and a voting is casted to find possible location of the query signature in a document. The spatial features computed from background blobs found in the target document are used for GHT. The peak of votes in GHT accumulator validates the hypothesis of the query signature. The proposed method is tested on a collection of mixed documents (handwritten/printed) of various scripts and we have obtained encouraging results from the experiments. Partha Pratim Roy 0001, Souvik Bhowmick, Umapada Pal 0001, Jean-Yves Ramel |
ICFHR | 1 |
| 2012 | Date field extraction in handwritten documents
Ranju Mandal, Partha Pratim Roy 0001, Umapada Pal 0001 |
ICPR | 2 |
| 2012 | Wavelet-gradient-fusion for video text binarization
Sangheeta Roy, Palaiahnakote Shivakumara, Partha Pratim Roy 0001, Chew Lim Tan |
ICPR | 3 |
| 2012 | Text line extraction in graphical documents using background and foreground information
Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001 |
Int. J. Document Anal. Recognit. | 1 |
| 2012 | Signature Segmentation from Machine Printed Documents using Contextual InformationabstractAutomatic signature segmentation from a printed document is a challenging task due to the nature of handwriting of the signatory, overlapping/touching of signature strokes with printed text, graphics, noise, etc. In this paper, we propose a two-stage approach for signature segmentation from a document page. In the first stage, a document is segmented into blocks and then blocks are classified into two classes: signature block and printed word block. Gradient-based features are used for block feature extraction and support vector machine classifier is used for block-wise classification. In the second stage, printed characters that may be present in isolated form or overlapped/touched with signature part are removed from signature blocks. From each of the detected signature blocks, the isolated printed characters (if exist) are removed using context information. To detect overlapping/touching printed stroke in a signature block, at first some hypothetical zones are detected where possible overlapping/touching may occur. Bounding box information of neighboring printed word block and local linearity of character strings near the signature blocks are used to detect hypothetical zones. Next, to detect the overlapping/touching printed strokes in hypothetical zones of a signature block, the corner points of contours obtained by Douglas and Peucker polygonal approximation algorithm and skeleton junction points are used. Finally, the touching strokes of signature are separated from text characters using the contour smoothness information near skeleton junction points. The experiment is performed in "Tobacco-800" dataset [The legacy tobacco document library (ltdl), available at http://legacy.library.ucsf.edu/ , University of California, San Francisco, 2007.] and the results obtained from the experiment are promising. Ranju Mandal, Partha Pratim Roy 0001, Umapada Pal 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2012 | Multi-oriented touching text character segmentation in graphical documents using dynamic programming
Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001, Mathieu Delalandre |
Pattern Recognit. | 1 |
| 2011 | ICDAR 2011 Robust Reading Competition - Challenge 1: Reading Text in Born-Digital Images (Web and Email)abstractThis paper presents the results of the first Challenge of ICDAR 2011 Robust Reading Competition. Challenge 1 is focused on the extraction of text from born-digital images, specifically from images found in Web pages and emails. The challenge was organized in terms of three tasks that look at different stages of the process: text localization, text segmentation and word recognition. In this paper we present the results of the challenge for all three tasks, and make an open call for continuous participation outside the context of ICDAR 2011. Dimosthenis Karatzas, Sergi Robles, Joan Mas Romeu, Farshad Nourbakhsh, Partha Pratim Roy 0001 |
ICDAR | 5 |
| 2011 | Signature Segmentation from Machine Printed Documents Using Conditional Random FieldabstractAutomatic separation of signatures from a document page involves difficult challenges due to the free-flow nature of handwriting, overlapping/touching of signature parts with printed text, noise, etc. In this paper, we have proposed a novel approach for the segmentation of signatures from machine printed signed documents. The algorithm first locates the signature block in the document using word level feature extraction. Next, the signature strokes that touch or overlap with the printed texts are separated. A stroke level classification is then performed using skeleton analysis to separate the overlapping strokes of printed text from the signature. Gradient based features and Support Vector Machine (SVM) are used in our scheme. Finally, a Conditional Random Field (CRF) model energy minimization concept based on approximated labeling by graph cut is applied to label the strokes as "signature" or "printed text" for accurate segmentation of signatures. Signature segmentation experiment is performed in "tobacco" dataset1 and we have obtained encouraging results. Ranju Mandal, Partha Pratim Roy 0001, Umapada Pal 0001 |
ICDAR | 2 |
| 2011 | Word Retrieval in Historical Document Using Character-PrimitivesabstractWord searching and indexing in historical document collections is a challenging problem because, characters in these documents are often touching or broken due to degradation/ ageing effects. For efficient searching in such historical documents, this paper presents a novel approach towards word spotting using string matching of character primitives. We describe the text string as a sequence of primitives which consists of a single character or a part of a character. Primitive segmentation is performed analyzing text background information that is obtained by water reservoir technique. Next, the primitives are clustered using template matching and a codebook of representative primitives is built. Using this primitive codebook, the text information in the document images are encoded and stored. For a query word, we segment it into primitives and encode the word by a string of representative primitives from codebook. Finally, an approximate string matching is applied to find similar words. The matching similarity is used to rank the retrieved words. The proposed method is tested on historical books of French alphabets and we have obtained encouraging results from the experiment. Partha Pratim Roy 0001, Jean-Yves Ramel, Nicolas Ragot |
ICDAR | 1 |
| 2011 | Document seal detection using GHT and character proximity graphs
Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001 |
Pattern Recognit. | 1 |
| 2010 | Query driven word retrieval in graphical documentsabstractIn this paper, we present an approach towards the retrieval of words from graphical document images. In graphical documents, due to presence of multi-oriented characters in non-structured layout, word indexing is a challenging task. The proposed approach uses recognition results of individual components to form character pairs with the neighboring components. An indexing scheme is designed to store the spatial description of components and to access them efficiently. Given a query text word (ascii/unicode format), the character pairs present in it are searched in the document. Next the retrieved character pairs are linked sequentially to form character string. Dynamic programming is applied to find different instances of query words. A string edit distance is used here to match the query word as the objective function. Recognition of multi-scale and multi-oriented character component is done using Support Vector Machine classifier. To consider multi-oriented character strings the features used in the SVM are invariant to character orientation. Experimental results show that the method is efficient to locate a query word from multi-oriented text in graphical documents. Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001 |
Document Analysis Systems | 1 |
| 2010 | Multi-oriented Bangla and Devnagari text recognition
Umapada Pal 0001, Partha Pratim Roy 0001, Nilamadhaba Tripathy, Josep Lladós 0001 |
Pattern Recognit. | 2 |
| 2009 | Seal Detection and Recognition: An Approach for Document IndexingabstractReliable indexing of documents having seal instances can be achieved by recognizing seal information. This paper presents a novel approach for detecting and classifying such multi-oriented seals in these documents. First, Hough Transform based methods are applied to extract the seal regions in documents. Next, isolated text characters within these regions are detected. Rotation and size invariant features and a Support Vector Machine based classifier have been used to recognize these detected text characters. Next, for each pair of character, we encode their relative spatial organization using their distance and angular position with respect to the centre of the seal, and enter this code into a hash table. Given an input seal, we recognize the individual text characters and compute the code for pair-wise character based on the relative spatial organization. The code obtained from the input seal helps to retrieve model hypothesis from the hash table. The seal model to which we get maximum hypothesis is selected for the recognition of the input seal. The methodology is tested to index seal in rotation and size invariant environment and we obtained encouraging results. Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001 |
ICDAR | 1 |
| 2009 | Multi-Oriented and Multi-Sized Touching Character Segmentation Using Dynamic ProgrammingabstractIn this paper, we present a scheme towards the segmentation of English multi-oriented touching strings into individual characters. When two or more characters touch, they generate a big cavity region at the background portion. Using Convex Hull information, we use these background information to find some initial points to segment a touching string into possible primitive segments (a primitive segment consists of a single character or a part of a character). Next these primitive segments are merged to get optimum segmentation and dynamic programming is applied using total likelihood of characters as the objective function. SVM classifier is used to find the likelihood of a character. To consider multi-oriented touching strings the features used in the SVM are invariant to character orientation. Circular ring and convex hull ring based approach has been used along with angular information of the contour pixels of the character to make the feature rotation invariant. From the experiment, we obtained encouraging results. Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001, Mathieu Delalandre |
ICDAR | 1 |
| 2008 | Multi-Oriented English Text Line Extraction Using Background and Foreground InformationabstractIn graphical documents (map, engineering drawing), artistic documents etc. there exist many printed materials where text lines are not parallel to each other and they are multi-oriented and curve in nature. For the OCR of such documents we need to extract individual text lines from the documents. Extraction of individual text lines from multi-oriented and/or curved text document is a difficult problem. In this paper, we propose a novel method to extract individual text lines from such document pages and the method is based on the foreground and background information of the characters of the text. To take care of background information, water reservoir concept is used here. In the proposed scheme at first, individual components are detected and grouped into 3-character clusters using their inter-component distance, size and positional information. Applying concept of graph, initial 3-character clusters are merged to have larger cluster group. Using inter-character background information, orientations of the extreme characters of a larger cluster are decided and based on these orientation, two candidate regions are formed from the cluster. Finally, with the help of these candidate regions, individual lines are extracted. From the experiment, we obtained encouraging result. Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001, Fumitaka Kimura |
Document Analysis Systems | 1 |
| 2008 | Convex hull based approach for multi-oriented character recognition from graphical documentsabstractIn this paper, we present a scheme towards recognition of English character in multi-scale and multi-oriented environments. Graphical document such as map consists of text lines which appear in different orientation. Sometimes, characters in a single word may follow a curvilinear way to annotate the graphical curve lines. For recognition of such multi-scale and multi-oriented characters a Support Vector Machine (SVM) based scheme is presented in this paper. The feature used here is invariant to character orientation. Circular ring and convex hull have been used along with angular information of the contour pixels of the character to make the feature rotation invariant. We tested our proposed scheme on two different datasets. Combining circular and convex hull feature we have obtained 96.73% and 99.56% accuracy in these two datasets. Partha Pratim Roy 0001, Umapada Pal 0001, Josep Lladós 0001, Fumitaka Kimura |
ICPR | 1 |
| 2004 | Multioriented and curved text lines extraction from Indian documentsabstractThere are printed artistic documents where text lines of a single page may not be parallel to each other. These text lines may have different orientations or the text lines may be curved shapes. For the optical character recognition (OCR) of these documents, we need to extract such lines properly. In this paper, we propose a novel scheme, mainly based on the concept of water reservoir analogy, to extract individual text lines from printed Indian documents containing multioriented and/or curve text lines. A reservoir is a metaphor to illustrate the cavity region of a character where water can be stored. In the proposed scheme, at first, connected components are labeled and identified either as isolated or touching. Next, each touching component is classified either straight type (S-type) or curve type (C-type), depending on the reservoir base-area and envelope points of the component. Based on the type (S-type or C-type) of a component two candidate points are computed from each touching component. Finally, candidate regions (neighborhoods of the candidate points) of the candidate points of each component are detected and after analyzing these candidate regions, components are grouped to get individual text lines. Umapada Pal 0001, Partha Pratim Roy 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 2 |