Md Baharul Islam

dblp:175/3784 · also Md. Baharul Islam · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
10since 2021 · last 2024
0000-0002-9928-5776ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2024 WNet: A dual-encoded multi-human parsing network
abstract
Abstract In recent years, multi‐human parsing has become a focal point in research, yet prevailing methods often rely on intermediate stages and lacking pixel‐level analysis. Moreover, their high computational demands limit real‐world efficiency. To address these challenges and enable real‐time performance, low‐latency end‐to‐end network is proposed. This approach leverages vision transformer and convolutional neural network in a dual‐encoded network, featuring a lightweight Transformer‐based vision encoder) and a convolution encoder based on Darknet. This combination adeptly captures long‐range dependencies and spatial relationships. Incorporating a fuse block enables the seamless merging of features from the encoders. Residual connections in the decoder design amplify information flow. Experimental validation on crowd instance‐level human parsing and look into person datasets showcases the WNet's effectiveness, achieving high‐speed multi‐human parsing at 26.7 frames per second. Ablation studies further underscore WNet's capabilities, emphasizing its efficiency and accuracy in complex multi‐human parsing tasks.
Md Imran Hosen, Tarkan Aydin, Md Baharul Islam
IET Image Process.3
2023 Denoising Diffusion Probabilistic Model for Retinal Image Generation and Segmentation
abstract
Experts use retinal images and vessel trees to detect and diagnose various eye, blood circulation, and brain-related diseases. However, manual segmentation of retinal images is a time-consuming process that requires high expertise and is difficult due to privacy issues. Many methods have been proposed to segment images, but the need for large retinal image datasets limits the performance of these methods. Several methods synthesize deep learning models based on Generative Adversarial Networks (GAN) to generate limited sample varieties. This paper proposes a novel Denoising Diffusion Probabilistic Model (DDPM) that outperformed GANs in image synthesis. We developed a Retinal Trees (ReTree) dataset consisting of retinal images, corresponding vessel trees, and a segmentation network based on DDPM trained with images from the ReTree dataset. In the first stage, we develop a two-stage DDPM that generates vessel trees from random numbers belonging to a standard normal distribution. Later, the model is guided to generate fundus images from given vessel trees and random distribution. The proposed dataset has been evaluated quantitatively and qualitatively. Quantitative evaluation metrics include Frechet Inception Distance (FID) score, Jaccard similarity coefficient, Cohen’s kappa, Matthew’s Correlation Coefficient (MCC), precision, recall, F1-score, and accuracy. We trained the vessel segmentation model with synthetic data to validate our dataset’s efficiency and tested it on authentic data. Our developed dataset and source code is available at https://github.com/AAleka/retree.
Alnur Alimanov, Md Baharul Islam
ICCP2
2023 Optimized video compression with residual split attention and swin-block artifact contraction
Afsana Ahsan Jeny, Md Baharul Islam
J. Vis. Commun. Image Represent.2
2023 Consistent Video Inpainting Using Axial Attention-Based Style Transformer
abstract
Maintaining spatial and temporal consistency in the inpainted video area of the video is a challenging problem. Recent research focuses on flow information for synthesizing temporally smooth pixels while neglecting semantic structural coherence across the video frames. Thus, it suffers from over-smoothing and shadowy outlines that significantly degrade the inpainted video quality. We propose an end-to-end consistent video inpainting model that will substantially improve the inpainted video region to overcome this problem. The model employs a deep encoder (DE), axial attention block (AAB), style transformer, and decoder to enhance video inpainting with a realistic structure. A deep encoder (DE) encodes features effectively while the axial attention block (AAB) recreates all retrieved attributes by merging recoverable multi-scale characteristics with local spatial structures. Then, a novel-style transformer with the style manipulation block (SMB) fills the missing area with rich visual elements and temporal coherence. We use two publicly available benchmark datasets to assess the model's performance. Experimental results demonstrate that our method performs better than the state-of-the-art methods by a large margin. Besides, an extensive ablation study validates the model's performance.
Masum Shah Junayed, Md Baharul Islam
IEEE Trans. Multim.2
2022 An Efficient End-To-End Image Compression Transformer
abstract
Image and video compression received significant research attention and expanded their applications. Existing entropy estimation-based methods combine with hyperprior and local context, limiting their efficacy. This paper introduces an efficient end-to-end transformer-based image compression model, which generates a global receptive field to tackle the long-range correlation issues. A hyper encoder-decoder-based transformer block employs a multi-head spatial reduction self-attention (MHSRSA) layer to minimize the computational cost of the self-attention layer and enable rapid learning of multi-scale and high-resolution features. A Casual Global Anticipation Module (CGAM) is designed to construct highly informative adjacent contexts utilizing channel-wise linkages and identify global reference points in the latent space for end-to-end rate-distortion optimization (RDO). Experimental results demonstrate the effectiveness and competitive performance of the KODAK dataset.
Afsana Ahsan Jeny, Masum Shah Junayed, Md Baharul Islam
ICIP3
2022 Stereoscopic video quality measurement with fine-tuning 3D ResNets
Hassan Imani, Md Baharul Islam, Masum Shah Junayed, Tarkan Aydin, Nafiz Arica
Multim. Tools Appl.2
2021 Deep Covariance Feature and CNN-based End-to-End Masked Face Recognition
abstract
With the emergence of the global epidemic of COVID-19, face recognition systems have achieved much attention as contactless identity verification methods. However, covering a considerable part of the face by the mask poses severe challenges for conventional face recognition systems. This paper proposes an automated Masked Face Recognition (MFR) system based on the combination of a mask occlusion discarding technique and a deep-learning model. Initially, a pre-processing step is carried out in which the images pass three filters. Then, a Convolutional Neural Network (CNN) model is proposed to extract the features from unoccluded regions of the faces (i.e., eyes and forehead). These feature maps are employed to obtain covariance-based features. Two extra layers, i.e., Bitmap and Eigenvalue, are designed to reduce the dimension and concatenate these covariance feature matrices. The deep covariance features are quantized to codebooks combined based on Bag-of-Features (BoF) paradigm. Finally, a global histogram is created based on these codebooks and utilized for training an SVM classifier. The proposed method is trained and evaluated on Real-World-Masked-Face-Recognition-Dataset (RMFRD) and Simulated-Masked-Face-Recognition-Dataset (SMFRD) achieves an accuracy of 95.07% and 92.32 %, respectively, showing its competitive performance compared to the state-of-the-art. Experimental results prove that our system has high robustness against noisy data and illumination variations.
Masum Shah Junayed, Arezoo Sadeghzadeh, Md Baharul Islam
FG3
2021 An Efficient Video Desnowing and Deraining Method with a Novel Variant Dataset
Arezoo Sadeghzadeh, Md Baharul Islam, Reza Zaker
ICVS2
2021 Machine Vision-based Expert System for Automated Cucumber Diseases Recognition and Classification
abstract
Automated cucumber disease detection may significantly provide agricultural assistance for remote farmers. Due to having the similarity symptoms, it is challenging to differentiate between various forms of cucumber disease. This paper proposes an automated solution to recognize and classify the cucumber disease using different computer vision-based techniques. In light of this circumstance, we design a computerized cucumber disease recognition system that analyzes images collected by mobile phones and can recognize diseases to assist rural farmers in dealing with the situation. In our method, a discriminating feature set is initially extracted from the input images. Then, K-means clustering segmentation separates the disease-affected regions from the remaining image part. Finally, the diseases are classified using five different classification algorithms. Different evaluation metrics, including accuracy, precision, sensitivity, specificity, False-Positive Rate (FPR), False-Negative Rate (FNR), are used to analyze the classifier’s performance. We have carried out several experiments to illustrate the use of the proposed expert system. Our experiments showed that random forest exceeds all other classifiers regarding the total number of metrics used, with an accuracy of 85.84% on our dataset.
Afsana Ahsan Jeny, Masum Shah Junayed, Md Baharul Islam, Hassan Imani, A. F. M. Shahen Shah
INISTA3
2021 Real-Time YOLO-based Heterogeneous Front Vehicles Detection
abstract
The perception of the complex road environment is a critical factor in autonomous driving, which has become the research focus in intelligent vehicles. In this paper, a real-time front vehicle detection system is proposed to ensure safe driving in a complex environment, particularly in congested megacities. This system is based on the YOLO model, which effectively detects and classifies various vehicles from both images and videos. It improves detection accuracy by modifying a feature extraction-based backbone. To the authors’ best knowledge, this is the first time that vehicle detection is implemented on the recently published DhakaAI dataset. Compared to the other available datasets for object detection, such as KITTI, the DhakaAI dataset has a complex environment with numerous vehicles (21 different types). Experimental results demonstrate that the proposed system outperforms the state-of-the-art object detectors. In this method, the mAP (mean average precision) and the FPS (frame per second) is increased by 2.97% and 1.47, 4.64% and 5.57, 4.75% and 3.02, compared to the RetinaNet, SSD, and Faster RCNN on this dataset, respectively.
Masum Shah Junayed, Md Baharul Islam, Arezoo Sadeghzadeh, Tarkan Aydin
INISTA2
2020 EczemaNet: A Deep CNN-based Eczema Diseases Classification
abstract
Eczema is the most common among all types of skin diseases. A solution for this disease is very crucial for patients to have better treatment. Eczema is usually detected manually by doctors or dermatologists. It is tough to distinguish between different types of Eczema because of the similarities in symptoms. In recent years, several attempts have been taken to automate the detection of skin diseases with much accuracy. Many methods such as Image Processing Techniques, Machine Learning algorithms are getting used to execute segmentation and classification of skin diseases. It is found that among all those skin disease detection systems, particularly detection work on eczema disease is rare. There is also insufficiency in eczema disease dataset. In this paper, we propose a novel deep CNN-based approach for classifying five different classes of Eczema with our collected dataset. Data augmentation is used to transform images for better performance. Regularization techniques such as batch normalization and dropout helped to reduce overfitting. Our proposed model achieved an accuracy of 96.2%, which exceeded the performance of the state of the arts.
Masum Shah Junayed, Abu Noman Md Sakib, Nipa Anjum, Md Baharul Islam, Afsana Ahsan Jeny
IPAS4
2019 Warping-Based Stereoscopic 3D Video Retargeting With Depth Remapping
abstract
Due to the recent availability of different stereoscopic display devices and online 3D media resources (e.g. 3D movies), there is a growing demand for stereoscopic video retargeting that can automatically resize a given stereoscopic video to fit the target display device. In this paper, we propose a warping-based approach that can simultaneously resize and remap the depth of a stereoscopic video to produce a better 3D viewing experience. Firstly, our method computes the significance map for each stereo video frame. It then performs volume warping using non-homogeneous scaling optimization to resize the stereoscopic video. A depth remapping constraint is used to remap the depth and a constraint is applied to preserve the significant content during warping process. Experimental results demonstrate the effectiveness of our method in preserving the significant content, ensuring motion consistency, and enhancing the depth perception of the retargeted video sequences within the comfort depth range.
Md Baharul Islam, Lai-Kuan Wong, Kok-Lim Low, Chee-Onn Wong
WACV1
2018 Aesthetics-Driven Stereoscopic 3-D Image Recomposition With Depth Adaptation
abstract
Due to the availability and affordability of the stereoscopic equipment (e.g., stereo camera, lens, and display devices), stereoscopic image manipulation has been receiving considerable research attention in recent years. In this paper, we present a semiautomatic, aesthetic-driven stereoscopic image recomposition approach, which capacitates the change of the spatial position of the foreground object(s) in a given stereoscopic image to enhance human visual aesthetics. Our algorithm recomposes both the left and right stereo images simultaneously using a global optimization algorithm. To maximize image aesthetics, our algorithm minimizes a set of aesthetic quality errors, which is derived from selected photographic composition rules. In addition, depth adaptation is applied to the resized objects and change in vertical disparity of the resulting stereo image pair is minimized to ensure a pleasant three-dimensional (3-D) viewing experience. Our method can be used to perform stereoscopic image retargeting and recomposition simultaneously by providing the target image scale as the input. Empirical evaluations demonstrate the effectiveness of our approach in enhancing the aesthetics of stereoscopic 3-D images. Notably, depth adaptation is shown to play an important role in aesthetics enhancement.
Md Baharul Islam, Lai-Kuan Wong, Kok-Lim Low, Chee-Onn Wong
IEEE Trans. Multim.1
2017 A survey of aesthetics-driven image recomposition
Md Baharul Islam, Lai-Kuan Wong, Chee-Onn Wong
Multim. Tools Appl.1
2015 Semantics-Preserving Warping for Stereoscopic Image Retargeting
Chun-Hau Tan, Md Baharul Islam, Lai-Kuan Wong, Kok-Lim Low
PSIVT2