Himadri Mukherjee

dblp:196/7532 · DBLP profile ↗
← Back
24ranked-venue papers
10as first author
16since 2021 · last 2025
0000-0002-6570-1356ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 12 · 6 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A survey on artificial intelligence-based approaches for personality analysis from handwritten documents
Suparna Saha Biswas, Himadri Mukherjee, Ankita Dhar, Sk Md Obaidullah, Kaushik Roy 0004
Int. J. Document Anal. Recognit.2
2024 EDM10: A Polyphonic Stereo Dataset with Identical BGM for Musical Instrument Identification
Himadri Mukherjee, Matteo Marciano, Ankita Dhar, Kaushik Roy 0004
ICPR (20)1
2024 A bi-stage approach to North Indian raga distinction
Debjyoti Basu, Himadri Mukherjee, Matteo Marciano, Shibaprasad Sen, Sajai Vir Singh, Sk Md Obaidullah, Kaushik Roy 0004
Multim. Tools Appl.2
2024 City name recognition for Indian postal automation: Exploring script dependent and independent approach
Somnath Chatterjee, Himadri Mukherjee, Shibaprasad Sen, Sk Md Obaidullah, Kaushik Roy 0004
Multim. Tools Appl.2
2024 MOPO-HBT: A movie poster dataset for title extraction and recognition
Mridul Ghosh, Sayan Saha Roy, Bivan Banik, Himadri Mukherjee, Sk Md Obaidullah, Kaushik Roy 0004
Multim. Tools Appl.4
2024 LIFA: Language identification from audio with LPCC-G features
Himadri Mukherjee, Ankita Dhar, Sk Md Obaidullah, KC Santosh, Santanu Phadikar, Kaushik Roy 0004, Umapada Pal 0001
Multim. Tools Appl.1
2022 SEN: Stack Ensemble Shallow Convolution Neural Network for Signature-based Writer Identification
abstract
Signature-based writer identification (SWI) is an automated segmentation-free holistic approach where a person is identified based on their handwritten signature. Earlier research attempts mainly featured learning-based approaches where writing patterns were detected and fed to machine learning models for determining the writer. Nowadays, a deep learning-based approach is becoming very popular and several works are reported in the literature using such models. In this paper, we propose a two-stage convolution neural network (CNN) architecture that has two properties: (i) at first, two state-of-the-art CNN models namely VGG-19 and EfficientNet-B0 were truncated making them lightweight; (ii) Secondly, a stack ensemble network (SEN) was proposed where the truncated architectures were stacked along with a shallow base CNN model. The proposed system experimented on a newly built multi-script offline signature dataset where three popular Indic scripts namely: Bangla, Roman and Devanagari were considered. The proposed SEN outperforms individual CNN architectures in terms of recognition rate. In addition, the system converges considerably fast as the SEN architecture is shallower compared to heavier traditional networks. Overall, we obtained the highest writer identification accuracy of 99.44%, 99.04%, and 98.61% for Bangla, Roman, and Devanagari, respectively, by the proposed SEN architecture. Furthermore, the dataset used in this paper will be available freely for research purposes from the link mentioned in Section III.
Sk Md Obaidullah, Mridul Ghosh, Himadri Mukherjee, Kaushik Roy 0004, Umapada Pal 0001
ICPR3
2022 CNN based recognition of handwritten multilingual city names
Ramit Kumar Roy, Himadri Mukherjee, Kaushik Roy 0004, Umapada Pal 0001
Multim. Tools Appl.2
2022 Understanding movie poster: transfer-deep learning approach for graphic-rich text recognition
Mridul Ghosh, Sayan Saha Roy, Himadri Mukherjee, Sk Md Obaidullah, KC Santosh, Kaushik Roy 0004
Vis. Comput.3
2021 Identification of Signs of Depression Relapse using Audio-visual Cues: A Preliminary Study
abstract
Depression is a serious mental disorder that affects many individuals across the globe. Depression (unipolar or bipolar) is characterized by a high rate of relapse or recurrence where a person might experience depressive episodes after non-depressive ones. The symptom patterns for recurrent depressive episodes have not been properly analyzed. Thus, there is a pressing need for systems which can monitor the mental health of individuals at risk to detect initial signs of relapse and recurrence. This points towards an automated system which identifies such signs and facilitates in timely treatment. In this paper, we introduce for the first time a deep learning based prospective monitoring system for the identification of relapse signs using audio-visual cues. The proposed model approximates relapse as the similarity between non-depression and depression samples. Experiments were performed on the DAIC-WOZ dataset and a highest accuracy of 73.21% was obtained using a Siamese network-based approach with one-shot learning regime.
Muhammad Muzammel, Alice Othmani, Himadri Mukherjee, Hanan Salam
CBMS3
2021 Automatic Signature-Based Writer Identification in Mixed-Script Scenarios
Sk Md Obaidullah, Mridul Ghosh, Himadri Mukherjee, Kaushik Roy 0004, Umapada Pal 0001
ICDAR (2)3
2021 Towards Automatic Narrative Coherence Prediction
abstract
Research in Psychology has shown that stories people tell about themselves, and how they recall their experiences, reveal a lot about their individual characteristics and mental well-being. The Narrative Coherence Coding Scheme (NaCCS) is a set of guidelines established in psychology research for annotating the “coherence” of a narrative along three dimensions: context, chronology and theme. A significant correlation was found between a narrative’s coherence score and independently collected mental health markers of the narrator. Currently, all coherence annotations are done manually; a time consuming task which drains vital resources. In this paper, we propose an Artificial Intelligence based approach involving Natural Language Processing (NLP) to predict a narrative’s coherence score (4-class classification problem). We explore a number of techniques, ranging from traditional machine learning models such as Support Vector Machines (SVM) to pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers). BERT produced the best results for all dimensions in terms of accuracy: 53.7% (context), 71.8% (chronology), and 69.6% (theme). The location of information in the narratives (beginning, end, throughout) was helpful in improving predictions.
Filip Bendevski, Jumana Ibrahim, Tina Krulec, Theodore Waters, Nizar Habash, Hanan Salam, Himadri Mukherjee, Christin Camia
ICMI7
2021 Deep neural network to detect COVID-19: one architecture for both CT Scans and Chest X-rays
Himadri Mukherjee, Subhankar Ghosh, Ankita Dhar, Sk Md Obaidullah, KC Santosh, Kaushik Roy 0004
Appl. Intell.1
2021 Lung Health Analysis: Adventitious Respiratory Sound Classification Using Filterbank Energies
abstract
Audio-based healthcare technologies are among the most significant applications of pattern recognition and Artificial Intelligence. Lately, a major chunk of the World population has been infected with serious respiratory diseases such as COVID-19. Early recognition of lung health abnormalities can facilitate early intervention, and decrease the mortality rate of the infected population. Research has shown that it is possible to automatically monitor lung health abnormalities through respiratory sounds. In this paper, we propose an approach that employs filter bank energy-based features and Random Forests to classify lung problem types from respiratory sounds. The adventitious sounds, crackles and wheezes appear distinct to the human ear. Moreover, different sounds are characterized by different frequency ranges that are dominant. The proposed approach attempts to distinguish the adventitious sounds (crackles and wheezes) by modeling the human auditory perception of these sounds. Specifically, we propose a respiratory sounds representation technique capable of modeling the dominant frequency range present in such sounds. On a publicly available dataset (ICBHI) of size 6898 cycles spanning over 5[Formula: see text]h, our results can be compared with the state-of-the-art results, in distinguishing two different types of adventitious sounds: crackles and wheezes.
Himadri Mukherjee, Hanan Salam, KC Santosh
Int. J. Pattern Recognit. Artif. Intell.1
2021 LWSINet: A deep learning-based approach towards video script identification
Mridul Ghosh, Himadri Mukherjee, Sk Md Obaidullah, KC Santosh, Nibaran Das, Kaushik Roy 0004
Multim. Tools Appl.2
2021 Identifying language from songs
Himadri Mukherjee, Ankita Dhar, Sk Md Obaidullah, KC Santosh, Santanu Phadikar, Kaushik Roy 0004
Multim. Tools Appl.1
2020 Periodic Change Detection in Fetal Heart Rate Using Cardiotocograph
abstract
Since 1960s, Cardiotocography (CTG) has been considered the primary tool for monitoring fetal health during antepartum and intra-partum periods. It records both Fetal Heart Rate (FHR) and mother's Uterine Contraction Pressure (UCP) simultaneously. However, due to inter and intra-observer variations, the introduction of CTG in fetal care did little to reduce the fetal mortality and morbidity. To ensure that the signs of hypoxia are recognized at the onset it is needed to have a robust and automated clinical decision support system since the visual analysis (clinicians) can be error-prone. In this work, we proposed methods to identify the periodic changes i.e. acceleration and deceleration. Our method detected 987 accelerations and 1755 decelerations from the 556 CTG data. There were 96.6% and 97.3% agreements with the three clinicians estimate for acceleration and deceleration, respectively. Besides, we also proposed a novel method to detect Sinusoidal Heart Rate (SHR) pattern. With Random Forest classifier, the SHR classification accuracy was 93%. The sensitivity and specificity were 93% and 86%, respectively, while both Positive Predictive Value (PPV) and Negative Predictive Value (NPV) were found to be 100%. We conclude that the proposed method can be used as a "gold standard" for SHR identification.
Sahana Das, Himadri Mukherjee, KC Santosh, Chanchal Kumar Saha, Kaushik Roy 0004
CBMS2
2020 Linear Predictive Coefficients-Based Feature to Identify Top-Seven Spoken Languages
abstract
Speech recognition in multilingual scenario is not trivial in the case when multiple languages are used in one conversation. Language must be identified before we process speech recognition as such tools are language-dependent. We present a language identification system (or AI tool) to distinguish top-seven world languages namely Chinese, Spanish, English, Hindi, Arabic, Bangla and Portuguese [G. F. Simons and C. D. Fennig (eds.), Ethnologue: Laguage of the Americas and the Pacific, Twentieth Edn. (SIL Internatinal, 2017)]. The system uses linear predictive coefficients-based feature, i.e. the line spectral pair–grade ratio (LSP–GR) feature, and ensemble learning for classification. Experiments were performed on more than 200[Formula: see text]h of real-world YouTube data and the highest possible accuracy of 96.95% was received. The results can be compared with other machine learning classifiers.
Himadri Mukherjee, Ankita Dhar, Sk Md Obaidullah, KC Santosh, Santanu Phadikar, Kaushik Roy 0004
Int. J. Pattern Recognit. Artif. Intell.1
2020 Ensemble based technique for the assessment of fetal health using cardiotocograph - a case study with standard feature reduction techniques
Sahana Das, Himadri Mukherjee, Sk Md Obaidullah, Kaushik Roy 0004, Chanchal Kumar Saha
Multim. Tools Appl.2
2020 Image-based features for speech signal classification
Himadri Mukherjee, Ankita Dhar, Sk Md Obaidullah, Santanu Phadikar, Kaushik Roy 0004
Multim. Tools Appl.1
2020 CESS-A System to Categorize Bangla Web Text Documents
abstract
Technology has evolved remarkably, which has led to an exponential increase in the availability of digital text documents of disparate domains over the Internet. This makes the retrieval of the information a very much time- and resource-consuming task. Thus, a system that can categorize such documents based on their domains can truly help the users in obtaining the required information with relative ease and also reduce the workload of the search engines. This article presents a text categorization system (CESS) that categorizes text document using newly proposed hybrid features that combines term frequency-inverse document frequency-inverse class frequency and modified chi-square methods. Experiments were performed on real-world Bangla documents from eight domains comprises of 24,29,857 tokens, and the highest accuracy of 99.91% has been obtained with multilayer perceptron-based classification. Also, the experiments were tested on Reuters-21578 and 20 Newsgroups datasets and obtained accuracies of 97.29% and 94.67%, respectively, to show the language-independent nature of the system.
Ankita Dhar, Himadri Mukherjee, Niladri Sekhar Dash, Kaushik Roy 0004
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2019 Deep learning for spoken language identification: Can we visualize speech signal patterns?
Himadri Mukherjee, Subhankar Ghosh, Shibaprasad Sen, Sk Md Obaidullah, KC Santosh, Santanu Phadikar, Kaushik Roy 0004
Neural Comput. Appl.1
2018 A Dravidian Language Identification System
abstract
Speech recognition has established a strong bond with various technological boons for the day to day life of the rustics across the continents. Such advances have not yet propagated to the grassroot level of India, one of the reasons for it being the multilingual nature of our country. We are habituated in using multiple languages while talking, which makes the task of speech recognition challenging thereby making Language Identification an important task. The technique of automatically identifying language from spoken phrases is termed as Automatic Language Identification. The problem of multilingual speech further elevates for South Indian languages which at times become very difficult to distinguish with negligible prior knowledge. In this paper, an Automatic Language Identification System is proposed to distinguish the 4 Dravidian languages which are also known as South Indian languages due to their pre dominant use in the South Indian subcontinent. Dataset size ranged up to the size of 12224 clips and a highest accuracy of 96.46% was obtained by using a newly proposed Line Spectral Pair-Grade (LSP-G) feature along with FURIA based classification technique.
Himadri Mukherjee, Sk Md Obaidullah, Santanu Phadikar, Kaushik Roy 0004
ICPR1
2018 MISNA - A musical instrument segregation system from noisy audio with LPCC-S features and extreme learning
Himadri Mukherjee, Sk Md Obaidullah, Santanu Phadikar, Kaushik Roy 0004
Multim. Tools Appl.1