Masum Shah Junayed

dblp:220/3915 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0003-3592-4601ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2025 Graph Laplacian Transformer with Progressive Sampling for Prostate Cancer Grading
Masum Shah Junayed, John Derek Van Vessem, Gahie Nam, Sheida Nabavi
MICCAI (12)1
2024 A Fused Transformer-based Model for Gene Expression Prediction using Histopathology Images
abstract
Spatial transcriptomics (ST) is a cutting-edge technology that enables the spatial localization and analysis of gene expression within tissue sections. Despite its transformative potential, ST is constrained by high costs and limited spatial resolution, making it less accessible and challenging to implement widely. To overcome these limitations, we propose a fused transformer-based approach designed to predict high-density gene expression profiles directly from Whole Slide Images (WSIs). Our model capitalizes on the multi-scale hierarchical structure of WSIs, integrating several key components: ResNet50 and transformer encoders to extract features at various scales, and a Fused Transformer Block (FTB) that effectively aggregates these features. The FTB incorporates Spatial Positional Embeddings (SPE) to maintain spatial context, cross attention to prioritize relevant features across different scales, and Random Mask Attention (RMA) to concentrate the model’s attention on the most significant patterns. To validate our approach, we conducted extensive experiments using three public spatial transcriptomics datasets (HBC, HER2+, SCC) and three additional Visium datasets from 10X Genomics. Our results demonstrate that the proposed model not only preserves essential spatial relationships for mapping gene expressions to tissue morphology but also outperforms current state-of-the-art methods.
Masum Shah Junayed, Afsana Ahsan Jeny, Sheida Nabavi, Ion Mandoiu
BIBM1
2024 A Scaled Mask Attention-based Hybrid model for Survival Prediction
abstract
Whole Slide Images (WSIs) are widely used in medical practice, but predicting survival from histopathological images is particularly challenging due to the vast scale and intricate structure of WSIs. Traditional Convolutional Neural Network (CNN)-based methods, which divide WSIs into multiple patches for analysis, often struggle with the heterogeneity of WSIs. They often lose crucial feature information during patch sampling and fail to fully capture spatial, contextual, and hierarchical interactions. Moreover, many current methods do not adequately consider correlations between patches and struggle to adapt to the extensive dimensions of WSIs. To overcome these limitations, we introduce a novel hybrid model that integrates CNNs, transformers, and Graph Convolutional Networks (GCNs) for end-to-end survival prediction. Firstly, the model employs ResNet-50 and a Learnable Linear Projection (LLP) to extract the complex and high-dimensional features. Subsequently, the Scaled Masked Attention (SMA) proposes to focus on high-attention regions within informative features. Additionally, cosine similarity (CS) emphasizes characteristics within structurally similar features. Finally, the proposed model utilizes GCNs to effectively aggregate the extracted features, capturing the complex correlations between them to enhance survival prediction. To evaluate the performance of the proposed method, we use five public datasets. Extensive experiments and ablation studies demonstrate that our model outperforms existing methods in terms of C-index.
Masum Shah Junayed, Sheida Nabavi
BIBM1
2024 Multi-modal Spatial Clustering for Spatial Transcriptomics Utilizing High-resolution Histology Images
abstract
Understanding the intricate cellular environment within biological tissues is crucial for uncovering insights into complex biological functions. While single-cell RNA sequencing has significantly enhanced our understanding of cellular states, it lacks the spatial context to fully comprehend the cellular environment. Spatial transcriptomics (ST) addresses this limitation by enabling transcriptome-wide profiling while preserving spatial context. One of the principal challenges in ST data analysis is spatial clustering. Modern ST sequencing procedures typically include a high-resolution histology image, which has been shown in previous studies to be closely connected to gene expression profiles. However, current spatial clustering methods often fail to fully utilize the image information, limiting their ability to capture critical spatial and cellular interactions.In this study, we propose the spatial transcriptomics multimodal clustering (stMMC) model, a novel contrastive learningbased deep learning approach that integrates gene expression data with histology image features through a multi-modal parallel graph autoencoder. We tested stMMC against four state-of-the-art baseline models on two public ST datasets. The experiments demonstrated the superior performance of stMMC in terms of ARI and NMI and an ablation study validated the contributions of key components.
Bingjun Li, Mostafa Karami, Masum Shah Junayed, Sheida Nabavi
BIBM3
2023 Consistent Video Inpainting Using Axial Attention-Based Style Transformer
abstract
Maintaining spatial and temporal consistency in the inpainted video area of the video is a challenging problem. Recent research focuses on flow information for synthesizing temporally smooth pixels while neglecting semantic structural coherence across the video frames. Thus, it suffers from over-smoothing and shadowy outlines that significantly degrade the inpainted video quality. We propose an end-to-end consistent video inpainting model that will substantially improve the inpainted video region to overcome this problem. The model employs a deep encoder (DE), axial attention block (AAB), style transformer, and decoder to enhance video inpainting with a realistic structure. A deep encoder (DE) encodes features effectively while the axial attention block (AAB) recreates all retrieved attributes by merging recoverable multi-scale characteristics with local spatial structures. Then, a novel-style transformer with the style manipulation block (SMB) fills the missing area with rich visual elements and temporal coherence. We use two publicly available benchmark datasets to assess the model's performance. Experimental results demonstrate that our method performs better than the state-of-the-art methods by a large margin. Besides, an extensive ablation study validates the model's performance.
Masum Shah Junayed, Md Baharul Islam
IEEE Trans. Multim.1
2022 An Efficient End-To-End Image Compression Transformer
abstract
Image and video compression received significant research attention and expanded their applications. Existing entropy estimation-based methods combine with hyperprior and local context, limiting their efficacy. This paper introduces an efficient end-to-end transformer-based image compression model, which generates a global receptive field to tackle the long-range correlation issues. A hyper encoder-decoder-based transformer block employs a multi-head spatial reduction self-attention (MHSRSA) layer to minimize the computational cost of the self-attention layer and enable rapid learning of multi-scale and high-resolution features. A Casual Global Anticipation Module (CGAM) is designed to construct highly informative adjacent contexts utilizing channel-wise linkages and identify global reference points in the latent space for end-to-end rate-distortion optimization (RDO). Experimental results demonstrate the effectiveness and competitive performance of the KODAK dataset.
Afsana Ahsan Jeny, Masum Shah Junayed, Md Baharul Islam
ICIP2
2022 Stereoscopic video quality measurement with fine-tuning 3D ResNets
Hassan Imani, Md Baharul Islam, Masum Shah Junayed, Tarkan Aydin, Nafiz Arica
Multim. Tools Appl.3
2021 Deep Covariance Feature and CNN-based End-to-End Masked Face Recognition
abstract
With the emergence of the global epidemic of COVID-19, face recognition systems have achieved much attention as contactless identity verification methods. However, covering a considerable part of the face by the mask poses severe challenges for conventional face recognition systems. This paper proposes an automated Masked Face Recognition (MFR) system based on the combination of a mask occlusion discarding technique and a deep-learning model. Initially, a pre-processing step is carried out in which the images pass three filters. Then, a Convolutional Neural Network (CNN) model is proposed to extract the features from unoccluded regions of the faces (i.e., eyes and forehead). These feature maps are employed to obtain covariance-based features. Two extra layers, i.e., Bitmap and Eigenvalue, are designed to reduce the dimension and concatenate these covariance feature matrices. The deep covariance features are quantized to codebooks combined based on Bag-of-Features (BoF) paradigm. Finally, a global histogram is created based on these codebooks and utilized for training an SVM classifier. The proposed method is trained and evaluated on Real-World-Masked-Face-Recognition-Dataset (RMFRD) and Simulated-Masked-Face-Recognition-Dataset (SMFRD) achieves an accuracy of 95.07% and 92.32 %, respectively, showing its competitive performance compared to the state-of-the-art. Experimental results prove that our system has high robustness against noisy data and illumination variations.
Masum Shah Junayed, Arezoo Sadeghzadeh, Md Baharul Islam
FG1
2021 Machine Vision-based Expert System for Automated Cucumber Diseases Recognition and Classification
abstract
Automated cucumber disease detection may significantly provide agricultural assistance for remote farmers. Due to having the similarity symptoms, it is challenging to differentiate between various forms of cucumber disease. This paper proposes an automated solution to recognize and classify the cucumber disease using different computer vision-based techniques. In light of this circumstance, we design a computerized cucumber disease recognition system that analyzes images collected by mobile phones and can recognize diseases to assist rural farmers in dealing with the situation. In our method, a discriminating feature set is initially extracted from the input images. Then, K-means clustering segmentation separates the disease-affected regions from the remaining image part. Finally, the diseases are classified using five different classification algorithms. Different evaluation metrics, including accuracy, precision, sensitivity, specificity, False-Positive Rate (FPR), False-Negative Rate (FNR), are used to analyze the classifier’s performance. We have carried out several experiments to illustrate the use of the proposed expert system. Our experiments showed that random forest exceeds all other classifiers regarding the total number of metrics used, with an accuracy of 85.84% on our dataset.
Afsana Ahsan Jeny, Masum Shah Junayed, Md Baharul Islam, Hassan Imani, A. F. M. Shahen Shah
INISTA2
2021 Real-Time YOLO-based Heterogeneous Front Vehicles Detection
abstract
The perception of the complex road environment is a critical factor in autonomous driving, which has become the research focus in intelligent vehicles. In this paper, a real-time front vehicle detection system is proposed to ensure safe driving in a complex environment, particularly in congested megacities. This system is based on the YOLO model, which effectively detects and classifies various vehicles from both images and videos. It improves detection accuracy by modifying a feature extraction-based backbone. To the authors’ best knowledge, this is the first time that vehicle detection is implemented on the recently published DhakaAI dataset. Compared to the other available datasets for object detection, such as KITTI, the DhakaAI dataset has a complex environment with numerous vehicles (21 different types). Experimental results demonstrate that the proposed system outperforms the state-of-the-art object detectors. In this method, the mAP (mean average precision) and the FPS (frame per second) is increased by 2.97% and 1.47, 4.64% and 5.57, 4.75% and 3.02, compared to the RetinaNet, SSD, and Faster RCNN on this dataset, respectively.
Masum Shah Junayed, Md Baharul Islam, Arezoo Sadeghzadeh, Tarkan Aydin
INISTA1
2020 EczemaNet: A Deep CNN-based Eczema Diseases Classification
abstract
Eczema is the most common among all types of skin diseases. A solution for this disease is very crucial for patients to have better treatment. Eczema is usually detected manually by doctors or dermatologists. It is tough to distinguish between different types of Eczema because of the similarities in symptoms. In recent years, several attempts have been taken to automate the detection of skin diseases with much accuracy. Many methods such as Image Processing Techniques, Machine Learning algorithms are getting used to execute segmentation and classification of skin diseases. It is found that among all those skin disease detection systems, particularly detection work on eczema disease is rare. There is also insufficiency in eczema disease dataset. In this paper, we propose a novel deep CNN-based approach for classifying five different classes of Eczema with our collected dataset. Data augmentation is used to transform images for better performance. Regularization techniques such as batch normalization and dropout helped to reduce overfitting. Our proposed model achieved an accuracy of 96.2%, which exceeded the performance of the state of the arts.
Masum Shah Junayed, Abu Noman Md Sakib, Nipa Anjum, Md Baharul Islam, Afsana Ahsan Jeny
IPAS1
2018 A Model for Identifying Historical Landmarks of Bangladesh from Image Content Using a Depth-Wise Convolutional Neural Network
Afsana Ahsan Jeny, Masum Shah Junayed, Syeda Tanjila Atik, Sazzad Mahamd
ISDA (1)2