Rabul Hussain Laskar

dblp:77/2959 · DBLP profile ↗
← Back
66ranked-venue papers
0as first author
41since 2021 · last 2026
0000-0003-3988-394XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 17 since 2021Artificial intelligence and machine learning · 24 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 EffiSign network: a comprehensive approach for sign language recognition
Bhumika Karsh, Rabul Hussain Laskar, Ram Kumar Karsh, Manas Kamal Bhuyan
Multim. Tools Appl.2
2026 Towards generalized real-world image super-resolution: an adaptive zero-shot and efficient generative approach for handling unknown degradations
Masuma Aktar, Kuldeep Singh Yadav, Rabul Hussain Laskar
Neural Comput. Appl.3
2025 EDSGAN: An Edge-Informed Generative Adversarial Network for Enhanced Perceptual Quality Single Image Super-Resolution
abstract
Single Image Super-Resolution (SISR) faces a persistent challenge in reconstructing high-frequency edge details, which are paramount for human perceptual quality. While Generative Adversarial Networks (GANs) have significantly advanced SISR, they often struggle to generate truly sharp and realistic edges, often due to their inherent loss functions. To address this critical limitation, we propose an Edge-Informed Super-Resolution GAN (EDSGAN). EDSGAN employs a dual-path edge-informed discriminator that simultaneously analyzes image content and edge maps, enabling more effective discrimination of realistic versus artifact-prone edge structures. Concurrently, an innovative edge-aware loss guides the generator towards reconstructing perceptually sharper and more accurate edges. Extensive experiments on benchmark datasets demonstrate EDSGAN's superior perceptual quality. Notably, on the Set14 dataset, EDSGAN achieves a remarkable improvement of 9% in LPIPS (from 0.1329 to 0.121) and 8% in PI (from 2.9261 to 2.689) over the ESRGAN baseline. Our method effectively strikes an optimal balance between visual realism and pixel-level accuracy.
Masuma Aktar, Kuldeep Singh Yadav, Rabul Hussain Laskar
TENCON3
2025 A robust digital image watermarking technique in LWT-DCT domain using particle swarm optimization and statistical distortion correction
Saharul Alom Barlaskar, Anish Monsley K., Rabul Hussain Laskar
Multim. Tools Appl.3
2025 Analysis of multimodal fusion strategies in deep learning for ischemic stroke lesion segmentation on computed tomography perfusion data
Chintha Sri Pothu Raju, Bala Chakravarthy Neelapu, Rabul Hussain Laskar, Muhammad Ghulam
Multim. Tools Appl.3
2025 A General Approach to Fully Linearize the Power Amplifiers in mMIMO With Low Complexity
abstract
A radio frequency (RF) power amplifier (PA) plays an important role to amplify the message signal at higher power to transmit it to a distant receiver. Due to a typical nonlinear behavior of the PA at high power transmission, digital predistortion (DPD), exploiting the preinversion of the nonlinearity, is used to linearize the PA. However, in a massive MIMO (mMIMO) transmitter, a single DPD is not sufficient to fully linearize multiple PAs. Further, for the full linearization, assigning a separate DPD to each PA is complex and not economical. In this work, we address these challenges via the proposed low-complexity DPD (LC-DPD) scheme. Initially, we describe the fully-featured DPD (FF-DPD) scheme to linearize the multiple PAs and examine its complexity. Thereafter, using it, we derive the LC-DPD scheme that can adaptively linearize the PAs as per the requirement. The coefficients in the two schemes are learned using the algorithms that adopt indirect learning architecture based recursive prediction error method (ILA-RPEM) due to its adaptive and free from matrix inversion operations. Furthermore, for the LC-DPD structure, we have proposed three algorithms based on correlation of its common coefficients with the distinct coefficients.
Ganesh Prasad, Håkan Johansson, Rabul Hussain Laskar
IEEE Trans. Commun.3
2025 Attention-based Fusion for Stroke Lesion Segmentation on Computed Tomography Perfusion Data
abstract
In recent times, stroke has emerged as a significant threat to humans, transforming affected brain tissue into core and penumbra regions. As the penumbra becomes irreversible over time, early core region segmentation is crucial. Automatic segmentation systems offer an efficient alternative to manual segmentation and aid radiologists in stroke lesion segmentation using computed tomography and Computed Tomography Perfusion (CTP) maps that comprise four parameter maps. This automatic segmentation is increasingly used in interactive, multimedia-based systems for diagnostic tools and AI-driven health applications. Top-performing models that follow the patch-processing approach suffer from high inference times. To incorporate effective feature extraction at image-level inferences, which reduces the inference time, we present a hybrid fusion technique that combines early and bottleneck fusion, leveraging two separate encoders for effective feature extraction. Moreover, fusing the information from various fusion methods arbitrarily may not yield optimal results. Consequently, we have introduced two attention modules, i.e., cross-modal attention and cross-fusion attention modules, designed for the effective integration of features derived from diverse modalities and multiple fusion strategies, respectively. The findings highlight a considerable reduction in computational time alongside achieving a comparable Dice score. Additionally, the incorporation of hybrid fusion and attention modules in the baseline notably increased the Dice score from 0.482 to 0.521 in the validation dataset and achieved 0.48 in the test dataset of ISLES 2018. It also demonstrates competitive performance compared to existing models while maintaining efficient prediction times.
Chintha Sri Pothu Raju, Rabul Hussain Laskar, Zulfiqar Ali 0001, Muhammad Ghulam
ACM Trans. Multim. Comput. Commun. Appl.2
2024 A Low-Complexity DPD to Fully Linearize the Power Amplifiers in a mMIMO Transmitter
abstract
A radio frequency (RF) power amplifier (PA) is crucial to enhance the signal to transmit via antenna over long distances. High-power transmission often leads to nonlinear behavior in the PA, necessitating the use of digital predistortion (DPD) signal processing to restore linearity by preinverting the nonlinearity. However, when dealing with a massive MIMO (mMIMO) transmitter with numerous PAs, a single DPD is not enough, and allocating a separate DPD for each PA is intricate and cost-inefficient. In this study, we tackle these challenges through our proposed low-complexity DPD (LC-DPD) architecture. The LC-DPD has the flexibility to choose the parameters of its architecture as per the desired tradeoff between the performance and complexity in the linearization. It employs learning of its coefficients through algorithms utilizing an indirect learning architecture based recursive prediction error method (ILA-RPEM), which is adaptive and free from matrix inversions.
Ganesh Prasad, Håkan Johansson, Rabul Hussain Laskar
ICC3
2024 Sign Language Gesture Recognition Using YOLOv9 for Medical Attention of Hard of Hearing Population
abstract
For the hard-of-hearing population, communicating medical issues to doctors who do not understand sign language can be challenging. To address this problem, researchers have focused extensively on sign language gesture recognition. However, existing methods often struggle with conversion time and the accuracy of recognizing similar gestures. In this paper, we propose a sign language gesture recognition system utilizing the Yolov9 model. The system's performance was evaluated using a publicly available Kaggle dataset, achieving an impressive Mean Average Precision ([email protected]) of 99.5%. Experimental results show that the proposed method surpasses state-of-the-art techniques in both efficiency and accuracy.
Pranjal Gogoi, Bhumika Karsh, Ram Kumar Karsh, Rabul Hussain Laskar, Manas Kamal Bhuyan
TENCON4
2024 AttV19Net: Attention Based VGG 19 Network for Hand Gesture Recognition with a Prototype
Bhumika Karsh, Rabul Hussain Laskar, Ram Kumar Karsh, Manas Kamal Bhuyan
TENCON2
2024 A cascaded deep learning framework for iris centre localization in facial image
abstract
Abstract Accurate iris centre localization is crucial in many computer vision and facial biometric applications such as gaze estimation, human–computer interaction, iris recognition, and liveness detection. However, it is challenging in an uncontrolled environment due to variations like pose, scale, rotation, specular reflection, and image quality. Therefore, a cascaded deep learning framework for iris centre localization in facial images is proposed that is robust to the abovementioned variations. The proposed approach consists of (i) YOLOv3 for eye detection, (ii) UNet for iris segmentation, and (iii) statistical modelling for iris centre localization. The eyes are first detected using the YOLOv3, and subsequently, iris segmentation is performed within the detected eyes using the UNet. Following iris segmentation, statistical modelling is employed to enhance the localization accuracy of the iris centre. Experiments were performed on benchmark databases, resulting in a standardized error measure SED of 3.405 pixels for BioID and 3.259 pixels for GI4E databases. In addition, the robustness of the proposed eye detection model was further evaluated on the Yale B for illumination variations and the CAS‐PEAL for pose variations.
Naseem Ahmad, Muhammad Ghulam, Kuldeep Singh Yadav, Rabul Hussain Laskar, Ashraf Hossain, Zulfiqar Ali 0001
Expert Syst. J. Knowl. Eng.4
2024 Detection of tuberculosis using customized MobileNet and transfer learning from chest X-ray image
Nirupam Shome, Richik Kashyap, Rabul Hussain Laskar
Image Vis. Comput.3
2024 Design and development of an integrated approach towards detection and tracking of iris using deep learning
Naseem Ahmad, Kuldeep Singh Yadav, Anish Monsley K., Saharul Alom Barlaskar, Rabul Hussain Laskar, Ashraf Hossain
Multim. Tools Appl.5
2024 mIV3Net: modified inception V3 network for hand gesture recognition
Bhumika Karsh, Rabul Hussain Laskar, Ram Kumar Karsh
Multim. Tools Appl.2
2024 Design of a two-stage ASCII recognizer for the case-sensitive inputs in handwritten and gesticulation mode of the text-entry interface
Anish Monsley K., Kuldeep Singh Yadav, Naragoni Saidulu, Saharul Alom Barlaskar, Rabul Hussain Laskar
Multim. Tools Appl.5
2024 Improved mKLT and low layered HG-CNN based dynamic gesture recognition hardware system
Manoj Kumar Sain, Shweta Saboo, Joyeeta Singha, Rabul Hussain Laskar
Multim. Tools Appl.4
2024 Scale-adaptive gesture computing: detection, tracking and recognition in controlled complex environments
Anish Monsley K., Rabul Hussain Laskar
Mach. Vis. Appl.2
2024 mXception and dynamic image for hand gesture recognition
Bhumika Karsh, Rabul Hussain Laskar, Ram Kumar Karsh
Neural Comput. Appl.2
2024 End-to-end bare-hand localization system for human-computer interaction: a comprehensive analysis and viable solution
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar
Vis. Comput.3
2023 GCR-Net: A deep learning-based bare hand detection and gesticulated character recognition system for human-computer interaction
abstract
Summary Precisely detecting bare hands and recognizing the characters are two major stages in gesticulated character recognition systems. It is very challenging to implement them in an uncontrolled environment. Additional variations, particularly (i) background feature domination (BFD) effect and motion blur in detection, (ii) gesturing style, pattern, and case sensitivity in recognition, make the system more complex. To address these challenges, a gesticulated character recognition (GCR‐Net) model is designed. To detect the bare hand precisely, a pixel‐wise segmentation approach, HandSNet, is presented, which is able to overcome the BFD effect. To handle the motion blur in the frames, a tracking module comprised of a point‐tracker and Kalman filter is applied. To reduce the computational time, a mini‐SqueezeNet network is designed, which is used in HandSNet and recognition models as the backend network. It has 0.39 million parameters only. Four separate deep convolutional neural networks (DCNNs) are connected with the network section module at the recognition end. This network selection module activates one DCNN at a time to recognize the gesticulated characters accurately. The proposed GCR‐Net reduces the complexity between similar characters and provides a high precision rate compared to the existing approaches.
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar
Concurr. Comput. Pract. Exp.3
2023 Exploration of deep learning models for localizing bare-hand in the practical environment
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar, Naseem Ahmad
Eng. Appl. Artif. Intell.3
2023 Deep learning based spatio-temporal hand gesture recognition system in complex environment
abstract
Abstract Gesture recognition nowadays has grabbed the attention of researchers as they represent human behaviour in multiple practical ways. Amongst a variety of gestures available, hand gestures play an essential role in the field of human‐computer interaction when recognised efficiently in complex and dynamic environments. In this paper, we propose a dynamic hand gesture recognition system to recognise hand gestures appearing in different indoor and outdoor environments. Hand detection and tracking uses a two‐level system resulting in the formation of gesture trajectory in challenging conditions in which existing detection and tracking algorithms could not do so. A set of 45 features is provided as input to the various classification techniques. The redundancy problem has been reduced by selecting a set of optimum features using the analysis of variance method, which ranks the list of features. An incremental feature selection technique calculates recognition accuracy by selecting features according to rankings. This system provides an accuracy of 96.32% when used with machine learning and 97.5% when used with deep learning techniques. Recognition accuracy is calculated for various environments, including an extra hand, multiple persons in the video frame, and outdoor environment. All machine‐learning classifiers are combined using classifier combination to calculate the accuracy according to the majority‐voting rule. Based on the experimental results, it has been observed that deep learning provides better results compared to machine learning.
Shweta Saboo, Joyeeta Singha, Rabul Hussain Laskar
Expert Syst. J. Knowl. Eng.3
2023 Efficient hand segmentation for rehabilitation tasks using a convolution neural network with attention
H. Pallab Jyoti Dutta, Manas Kamal Bhuyan, Debanga Raj Neog, Karl F. MacDorman, Rabul Hussain Laskar
Expert Syst. Appl.5
2023 Self co-articulation removal and hybrid classifier-feature combination for dynamic hand gesture recognition
Shweta Saboo, Joyeeta Singha, Rabul Hussain Laskar
Multim. Tools Appl.3
2023 Gesture objects detection and tracking for virtual text entry keyboard interface
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar
Multim. Tools Appl.3
2023 A robust DNN model for text-independent speaker identification using non-speaker embeddings in diverse data conditions
Nirupam Shome, Banala Saritha, Richik Kashyap, Rabul Hussain Laskar
Neural Comput. Appl.4
2023 Detection, tracking, and recognition of isolated multi-stroke gesticulated characters
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar, Manas Kamal Bhuyan
Pattern Anal. Appl.3
2022 Design and development of a vision-based system for detection, tracking and recognition of isolated dynamic bare hand gesticulated characters
abstract
Abstract Detection and tracking are the vital stages to form the gesture trajectory in gesture recognition. It becomes more challenging when the variations in illumination, pose, position, occlusion, scale, speed, blurring effect and complex environment are introduced. Additionally, the background feature domination effect affects the existing deep learning models. A semantic segmentation model is implemented in this work to detect the bare hand to overcome these challenges. A pre‐trained network VGG‐16 is utilized by training with the proposed NITS S‐Net database. Evaluation of the SegNet model is done on EgoHands, Oxford and OUHands databases. To track the bare hand, a SegNet‐based detection and tracking approach is proposed using Kalman filter and point‐tracker. This model achieves 97.01% accuracy (a relative improvement of ~8% from the baseline models) at 0.068 s per frame computational time on NITS hand gesture database VIIIB. The gesticulated characters, that is, alphabets, numbers, operators, special characters, are gesticulated without any constraints on the pattern/strokes. To recognize these 95 multi‐stroke gestures, a deep convolutional neural network (DCNN) is presented using AlexNet. The DCNN model achieves 97.60% (a relative improvement of ~14% from the baseline models) accuracy on the NITS hand gesture database VIIIB merged. Evaluation of the handwritten EMNIST merged (balanced) database resulted in average recognition accuracy of 91.60%.
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar, Manas Kamal Bhuyan
Expert Syst. J. Knowl. Eng.3
2022 Development of an intelligent recognition system for dynamic mid-air gesticulation of isolated alphanumeric keys
Anish Monsley K., Kuldeep Singh Yadav, Rabul Hussain Laskar, Manas Kamal Bhuyan
Expert Syst. Appl.3
2022 Dynamic hand gesture recognition using combination of two-level tracker and trajectory-guided features
Shweta Saboo, Joyeeta Singha, Rabul Hussain Laskar
Multim. Syst.3
2022 A selective region-based detection and tracking approach towards the recognition of dynamic bare hand gesture using deep neural network
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar, Songhita Misra, Manas Kamal Bhuyan
Multim. Syst.3
2022 Eye center localization using gradient and intensity information under uncontrolled environment
Manir Ahmed, Rabul Hussain Laskar
Multim. Tools Appl.2
2022 Genetic algorithm based optimized watermarking technique using hybrid DCNN-SVR and statistical approach for watermark extraction
Saharul Alom Barlaskar, Sajai Vir Singh, Anish Monsley K., Rabul Hussain Laskar
Multim. Tools Appl.4
2021 Optimal New Node Insertion for Strong Minimum Energy Topology in IoT Networks
abstract
To provide seamless services by low powered small devices having stringent energy constraint in Internet of Things (IoT) networks, it is highly sought over the past few decades to design a competent energy-aware network. In that sense, thereof, the strong minimum energy topology (SMET) is investigated in the existing works to minimize the total power consumption while maintaining the strong connectivity between any pair of nodes (small IoT devices) in the network consisting only bidirectional links. Nevertheless, to significantly improve it, in this paper, we further explore the SMET with respect to insertion of a new node among the existing nodes which is defined as node insertion problem (NIP) for SMET (NIP-SMET). It has been proved that the NIP-SMET is NP-complete, therefore, we propose a heuristic based on Prim-incremental power greedy heuristic to solve it in polynomial time. Also, analytically, it has been shown that the NIP-SMET can provide significant energy-aware improvement over SMET. Via obtained numerical results, we find that the proposed heuristic can reduce the total power dissipation by 50% against the existing heuristic for SMET.
Ganesh Prasad, Deepak Mishra 0001, Rabul Hussain Laskar
CCNC3
2021 Precise eye center localization in a practical environment
abstract
Eye center localization plays a crucial role in computer vision and pattern recognition applications such as face recognition, gaze estimation, driver drowsiness detection, etc. Accurate eye center localization is challenging because of face pose orientation, illumination variations, eye occlusion, and specular reflections. This paper proposes a precise eye center localization method to solve the problems stated above. In this paper, multi-iris shape features extract the number of possible eye locations in the upper half of the face region, reducing the non-eye regions and improving computational speed. Then, Convolutional neural networks (CNN) extract high-level features, and the support vector machine (SVM) classifier is used for classification. Experiments were performed on AR, CASPAEL, and BioID databases to test the performance of the proposed model under varying pose, occlusion, and illumination variations, etc. Extensive experiments suggest that the proposed method provides better performance as compared to other existing methods.
Naseem Ahmad, Rabul Hussain Laskar, Ashraf Hossain, Manir Ahmed
TENCON2
2021 Recognition of isolated characters across different input interfaces using 2D DCNN
abstract
Recognition of the characters has gained much attention due to its potential applications like document analysis, license plate detection, house number detection, virtual text entry system, etc., in pattern recognition. However, it is very challenging to recognize the characters under the variations in pattern, style, translation, scale, rotation. This work develops a computationally efficient deep learning model to recognize handwritten, printable, and gesticulated characters. For gesture, the NITS gesticulated database having 60 characters (10 digits, 26 English uppercase alphabets, 4 operations, 18 special symbols) is proposed with the variation in pattern, style, scale in this work. To evaluate the ability and robustness of the proposed model, the handwritten characters (MNIST, EMNIST), printable characters (SVHN, Chars74) databases are considered. This network achieves 94.55%, 89.54%, 87.33, and 93.90% recognition accuracy on NITS gesticulated, EMNIST merge (balanced), SVHN, and Chars74 databases.
Kuldeep Singh Yadav, Anish Monsley K., Saharul Alom Barlaskar, Naseem Ahmad, Rabul Hussain Laskar, Manas Kamal Bhuyan
TENCON5
2021 A comparative analysis between late fusion of features approach and ensemble of multiple classifiers approach for image classification
abstract
Summary In recent times the late fusion of features approach for high‐level features extracted by multiple deep convolutional neural networks (DCNNs) has proven to be very effective in the computer vision field, especially for object classification problems. Pretrained DCNNs DenseNet‐121 and ResNet‐18 are retrained, keeping the number of output nodes equal to the number of classes present in the dataset. The last fully connected layers of these networks thereby get adapted to transform the high‐level features to a low‐dimensional feature map. Then these maps are fused to improve the performance of the model. On the other hand, an ensemble of multiple classifiers reduces the overfitting problem by combining multiple models prediction matrices. In this work, the prediction matrices of two Logsoftmax multiclass classifiers are combined. The feature maps for these two classifiers are extracted using pretrained DenseNet and ResNet. This study compares the late fusion of high‐level features approach and ensemble of multiple classifiers approach for object classification problems. Experimentation has been carried out on two benchmark datasets, such as CIFAR‐10 and CIFAR‐100, and it achieves 96.48% and 83.33% of test accuracy for ensemble of multiple classifiers and the late feature fusion approach. The proposed method has been compared with other deep architectures and datasets.
Khanjan Choudhury, R. Murugan, Mohammad Azharuddin Laskar, Rabul Hussain Laskar
Concurr. Comput. Pract. Exp.4
2021 HiLAM-aligned kernel discriminant analysis for text-dependent speaker verification
Mohammad Azharuddin Laskar, Rabul Hussain Laskar
Expert Syst. Appl.2
2021 Segregation of meaningful strokes, a pre-requisite for self co-articulation removal in isolated dynamic gestures
abstract
Abstract Gesture formation, a pre‐processing step, has its importance when variations in patterns, scale, and speed come into play. Self co‐articulations are intentional movements performed by an individual to complete a gesture, whose presence in the trajectory alters its original meaning. For recognition, most researchers have directly used the trajectory formed along with these self co‐articulated strokes, with a few removing it using visible trait‐like velocity. Usage of velocity has shortcomings as gesturing in air differs from gesturing over a solid surface; hence, we propose a gesture formation model, which incorporates global and local measures to remove these self co‐articulations. The global measure uses Euclidean distance, instantaneous velocity, and polarity calculated from the complete gesture, while the local measure segments the gesture into stroke‐level segments by using the minimum–maximum‐polarity algorithm and applies the selective bypass rules. The proposed model, when experimented on gestures patterns with premeditated speed variation, has a mean error rate of 0.0069 and 7.40% self co‐articulations;individuals’ natural gesticulation has a mean error rate of 0.0371 and 12.07% self co‐articulations. Experimentation on each gesture of NITS hand gesture databases showed a relative improvement of 40% (accuracy 97%) over the existing baseline models.
Anish Monsley K., Kuldeep Singh Yadav, Songhita Misra, Manas Kamal Bhuyan, Rabul Hussain Laskar
IET Image Process.6
2021 Modeling a Virtual Bare-Hand Interface System Using a Robust Hand Detection Approach for HCI
abstract
A practically deployable gesture recognition system is developed using a robust hand detection method implemented using a motion-based image segmentation process and a two-level bare hand classification model, which is integrated with a gesture classification system of 58 gestures using new robust features. Since detection of bare hand is affected by nonideal conditions, multiple color-texture features are analyzed in this study. In the second stage of the system, 18 new ASCII characters are introduced and analyzed along with the existing 40 characters (alphabets, numbers, and arithmetic operators). New 15 dimensional features are introduced along with the existing features to enhance the classification accuracy of the gestures. Significance of features statistically tested using one-way analysis of variance (ANOVA), Kruskal–Wallis and Friedman test, which are sequentially ranked and evaluated using incremental feature selection (IFS) method. Performance of the proposed hand detection system is observed to be 12.5% higher than the existing hand detection system under clean conditions, while 46.4% higher under the nonideal conditions. Performance of 58 gestures classification model has improved by 12.08% (Naïve Bayes), 8.86% (ELM), 10.83% (SVM), 8.02% ([Formula: see text]NN), and 6.61% (ANN) after using the new features. Majority voting-based classifier fusion method further improves the performance of the gesture recognition system by 3.88%, which is validated by Turkey’s HSD test.
Songhita Misra, G. Sridevi, Rabul Hussain Laskar
Int. J. Pattern Recognit. Artif. Intell.3
2021 Evaluation of accurate iris center and eye corner localization method in a facial image for gaze estimation
Manir Ahmed, Rabul Hussain Laskar
Multim. Syst.2
2020 Cascade convolutional neural network-long short-term memory recurrent neural networks for automatic tonal and nontonal preclassification-based Indian language identification
abstract
Abstract This work presents an automatic tonal/nontonal preclassification‐based Indian language identification (LID) system. Languages are firstly classified into tonal and nontonal categories, and then, individual languages are identified from the languages of the respective categories. This work proposes the use of pitch Chroma and formant features for this task, and also investigates how Mel‐frequency Cepstral Coefficients (MFCCs) complement these features. It further explores block processing (BP), pitch synchronous analysis (PSA)‐ and glottal closure regions (GCRs)‐based approaches for feature extraction, using syllables as basic units. Cascade convolutional neural network (CNN)‐long short‐term memory (LSTM) model using syllable‐level features has been developed. National Institute of Technology Silchar language database (NITS‐LD) and OGI‐Multilingual Telephone Speech Corpus (OGI‐MLTS) have been used for experimental validation. The proposed system based on the score combination of Cascade CNN‐LSTM models of Chroma (extracted from BP method), first two formants and MFCCs (both extracted from GCR method) reports the highest accuracies. In the preclassification stage, the observed accuracies are 91%, 87.3%, and 85.1% for NITS‐LD, for 30 s, 10 s, and 3 s test data respectively. For OGI‐MLTS database, the respective accuracies are 86.7%, 83.1%, and 80.6%. That amounts to absolute improvements of 11.6%, 12.3%, and 13.9% for NITS‐LD, and 12.5%, 11.9%, and 12.6% for OGI‐MLTS database with respect to that of the baseline system. The proposed preclassification‐based LID system shows improvements of 7.3%, 6.4%, and 7.4% for NITS‐LD and 6.1%, 6.7%, and 7.2% for OGI‐MLTS database over the baseline system for the three respective test data conditions.
Chuya China Bhanja, Mohammad Azharuddin Laskar, Rabul Hussain Laskar
Expert Syst. J. Knowl. Eng.3
2020 A fuzzy-clustering-based hierarchical i-vector/probabilistic linear discriminant analysis system for text-dependent speaker verification
abstract
Abstract In the i‐vector/probabilistic linear discriminant analysis (PLDA) technique, the PLDA backend classifier is modelled on i‐vectors. PLDA defines an i‐vector subspace that compensates the unwanted variability and helps to discriminate among speaker‐phrase pairs. The channel or session variability manifested in i‐vectors are known to be nonlinear in nature. PLDA training, however, assumes the variability to be linearly separable, thereby causing loss of important discriminating information. Besides, the i‐vector estimation, itself, is known to be poor in case of short utterances. This paper attempts to address these issues using a simple hierarchy‐based system. A modified fuzzy‐clustering technique is employed to divide the feature space into more characteristic feature subspaces using vocal source features. Thereafter, a separate i‐vector/PLDA model is trained for each of the subspaces. The sparser alignment owing to subspace‐specific universal background model and the relatively reduced dimensions of variability in individual subspaces help to train more effective i‐vector/PLDA models. Also, vocal source features are complementary to mel frequency cepstral coefficients, which are transformed into i‐vectors using mixture model technique. As a consequence, vocal source features and i‐vectors tend to have complementary information. Thus using vocal source features for classification in a hierarchy tree may help to differentiate some of the speaker‐phrase classes, which otherwise are not easily discriminable based on i‐vectors. The proposed technique has been validated on Part 1 of RSR2015 database, and it shows a relative equal error rate reduction of up to 37.41% with respect to the baseline i‐vector/PLDA system.
Mohammad Azharuddin Laskar, Rabul Hussain Laskar
Expert Syst. J. Knowl. Eng.2
2020 Removal of 'Salt & Pepper' noise from color images using adaptive fuzzy technique based on histogram estimation
Amarjit Roy, Lalit Manam, Rabul Hussain Laskar
Multim. Tools Appl.3
2020 SVM-based robust image watermarking technique in LWT domain using different sub-bands
Mohiul Islam, Amarjit Roy, Rabul Hussain Laskar
Neural Comput. Appl.3
2020 HiLAM-state discriminative multi-task deep neural network in dynamic time warping framework for text-dependent speaker verification
Mohammad Azharuddin Laskar, Rabul Hussain Laskar
Speech Commun.2
2019 Music Genre Recognition Using Residual Neural Networks
abstract
Genre is an abstract, yet a characteristic feature of music. Existing works for automatic genre classification compute a set of features from the audio and design a classifier on top of it. Such models, in general, compute these features over a relatively long duration of the audio. In this paper, a residual neural network based model is proposed for genre classification which is trained on short clips of just 3 seconds duration. Also, traditional genre classification algorithms will assign a single genre to an audio clip. However, it is well established that different genres have overlapping characteristics. Considering this ambiguous nature of the genre, the model proposed in this work can assign three genre labels to a music clip, with each genre associated with some probability. The proposed model has an error rate of 18%, 9%, and 5.5% while predicting into top-1, top-2 and top-3 genres for a music clip respectively. We demonstrate in this work that the predictions made by the classifier align with the broader understood meaning of genre in a realistic setting.
Dipjyoti Bisharad, Rabul Hussain Laskar
TENCON2
2019 Music genre recognition using convolutional recurrent neural network architecture
abstract
Abstract The genre is an abstract feature, but still, it is considered to be one of the important characteristics of music. Genre recognition forms an essential component for a large number of commercial music applications. Most of the existing music genre recognition algorithms are based on manual feature extraction techniques. These extracted features are used to develop a classifier model to identify the genre. However, in many cases, it has been observed that a set of features giving excellent accuracy fails to explain the underlying typical characteristics of music genres. It has also been observed that some of the features provide a satisfactory level of performance on a particular dataset but fail to provide similar performance on other datasets. Hence, each dataset mostly requires manual selection of appropriate acoustic features to achieve an adequate level of performance on it. In this paper, we propose a genre recognition algorithm that uses almost no handcrafted features. The convolutional recurrent neural network‐based model proposed in this study is trained on melspectrogram extracted from 3‐s duration audio clips taken from GTZAN dataset. The proposed model provides an accuracy of 85.36% on 10‐class genre classification. The same model has been trained and tested on 10 genres of MagnaTagATune dataset having 18,476 clips of 29‐s duration. The model has yielded an accuracy of 86.06%. The experimental results suggest that the proposed architecture with melspectrogram as input feature is capable of providing consistent performances across the different datasets
Dipjyoti Bisharad, Rabul Hussain Laskar
Expert Syst. J. Knowl. Eng.2
2019 Comparative framework for vision-based gesturing modes and implementation of robust colour-marker detector for practical environments
abstract
In this study, the authors have provided a detailed analysis based on experimental analysis, user‐feedback, and the literature survey to choose the appropriate gesturing mode between colour‐markers and bare‐hand for virtual dynamic gesturing system. Salient factors such as lexicon duration, muscle strain, target users, detection/tracking accuracy and so on, indicated that colour‐marker based models can outperform bare‐hand systems in real‐world scenarios. Colour‐markers are more robust to factors such as uneven illumination, affine transformations, and holds a better similarity to the natural way of writing on paper. However, colour‐markers can easily be erroneously detected/tracked due to the presence of imposters in the background. Most of the existing colour‐marker systems are developed with this limitation that reduces the naturalness and confines its application to a narrowed environment. This study addresses multiple types of frequently occurring imposters which are static/dynamic in nature and occlude the genuine colour‐marker often times. The imposters are distinguished based on factors such as speed, start/end position, the area of motion and so on. Some uncertain types of imposters which are partly similar to genuine gestures are distinguished based on the randomness present in the trajectories and classification models. The proposed techniques have achieved an accuracy of 98.26% in distinguishing the given imposter categories in this system.
Songhita Misra, Rabul Hussain Laskar
IET Image Process.2
2019 Eye center localization in a facial image based on geometric shapes of iris and eyelid under natural variability
Manir Ahmed, Rabul Hussain Laskar
Image Vis. Comput.2
2019 Integrated features and GMM Based Hand Detector Applied to Character Recognition System under Practical Conditions
Songhita Misra, Rabul Hussain Laskar
Multim. Tools Appl.2
2019 Fuzzy SVM based fuzzy adaptive filter for denoising impulse noise from color images
Amarjit Roy, Rabul Hussain Laskar
Multim. Tools Appl.2
2018 Malaria infected erythrocyte classification based on a hybrid classifier using microscopic images of thin blood smear
Salam Shuleenda Devi, Amarjit Roy, Joyeeta Singha, Shah Alam Sheikh, Rabul Hussain Laskar
Multim. Tools Appl.5
2018 Erratum to: Malaria infected erythrocyte classification based on a hybrid classifier using microscopic images of thin blood smear
Salam Shuleenda Devi, Amarjit Roy, Joyeeta Singha, Shah Alam Sheikh, Rabul Hussain Laskar
Multim. Tools Appl.5
2018 Geometric distortion correction based robust watermarking scheme in LWT-SVD domain with digital watermark extraction using SVM
Mohiul Islam, Rabul Hussain Laskar
Multim. Tools Appl.2
2018 Image authentication based on robust image hashing with geometric correction
Ram Kumar Karsh, Arunav Saikia, Rabul Hussain Laskar
Multim. Tools Appl.3
2018 Hybrid classifier based life cycle stages analysis for malaria-infected erythrocyte using thin blood smear images
Salam Shuleenda Devi, Rabul Hussain Laskar, Shah Alam Sheikh
Neural Comput. Appl.2
2018 Vision-based hand gesture recognition of alphabets, numbers, arithmetic operators and ASCII characters in order to develop a virtual text-entry interface system
Songhita Misra, Joyeeta Singha, Rabul Hussain Laskar
Neural Comput. Appl.3
2018 Dynamic hand gesture recognition using vision-based approach for human-computer interaction
Joyeeta Singha, Amarjit Roy, Rabul Hussain Laskar
Neural Comput. Appl.3
2017 Combination of adaptive vector median filter and weighted mean filter for removal of high-density impulse noise from colour images
abstract
In this study, a combination of adaptive vector median filter (VMF) and weighted mean filter is proposed for removal of high‐density impulse noise from colour images. In the proposed filtering scheme, the noisy and non‐noisy pixels are classified based on the non‐causal linear prediction error. For a noisy pixel, the adaptive VMF is processed over the pixel where the window size is adapted based on the availability of good pixels. Whereas, a non‐noisy pixel is substituted with the weighted mean of the good pixels of the processing window. The experiments have been carried out on a large database for different classes of images, and the performance is measured in terms of peak signal‐to‐noise ratio, mean squared error, structural similarity and feature similarity index. It is observed from the experiments that the proposed filter outperforms (∼1.5 to 6 dB improvement) some of the existing noise removal techniques not only at low density impulse noise but also at high‐density impulse noise.
Amarjit Roy, Joyeeta Singha, Lalit Manam, Rabul Hussain Laskar
IET Image Process.4
2017 Hand gesture recognition using two-level speed normalization, feature selection and classifier fusion
Joyeeta Singha, Rabul Hussain Laskar
Multim. Syst.2
2017 Analysis and extraction of LP-residual for its application in speaker verification system under uncontrolled noisy environment
Songhita Misra, Rabul Hussain Laskar, Ujwala Baruah, Tushar Kanti Das, Partha Saha, Suman Paul Choudhury
Multim. Tools Appl.2
2016 Self co-articulation detection and trajectory guided recognition for dynamic hand gestures
abstract
Hand gestures are a natural way of communication among humans in everyday life. Presence of spatiotemporal variations and unwanted movements within a gesture called self co‐articulation makes the segmentation a challenging task. The study reveals that the self co‐articulation may be used as one of the feature to enhance the performance of hand gesture recognition system. It was detected from the gesture trajectory by addition of speed information along with the pause in the gesture spotting phase. Moreover, a new set of novel features in the feature extraction stage was used such as position of the hand, self co‐articulated features, ratio and distance features. The ANN and SVM were used to develop two independent models using new set of features as input. The models based on CRF and HCRF was used to develop the baseline system for the present study. The experimental results suggest that the proposed new set of features provides improvement in terms of accuracy using ANN (7.48%) and SVM (9.38%) based models as compared with baseline CRF based model. There are also significant improvements in the performances of both ANN (2.08%) and SVM (3.98%) based models as compared with HCRF based model.
Joyeeta Singha, Rabul Hussain Laskar
IET Comput. Vis.2
2016 Effect of variation in gesticulation pattern in dynamic hand gesture recognition system
Joyeeta Singha, Songhita Misra, Rabul Hussain Laskar
Neurocomputing3
2016 Impulse noise removal using SVM classification based fuzzy filter from gray scale images
Amarjit Roy, Joyeeta Singha, Salam Shuleenda Devi, Rabul Hussain Laskar
Signal Process.4
2016 A pre-processing method for improvement of vowel onset point detection under noisy conditions
Partha Saha, Rabul Hussain Laskar, Mohammad Azharuddin Laskar
Speech Commun.2