VLDB 2026 Research / reviewers in the wild / expert
Md. Milon Islam
dblp:250/3239
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-4535-5978ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and CaptioningabstractMicroscopic assessment of histopathology images is vital for accurate cancer diagnosis and treatment. Whole Slide Image (WSI) classification and captioning have become crucial tasks in computer-aided pathology. However, microscopic WSIs face challenges such as redundant patches and unknown patch positions due to subjective pathologist captures. Moreover, generating automatic pathology captions remains a significant challenge. To address these challenges, a novel GNN-ViTCap framework is introduced for classification and caption generation from histopathological microscopic images. A visual feature extractor is used to extract feature embeddings. The redundant patches are then removed by dynamically clustering images using deep embedded clustering and extracting representative images through a scalar dot attention mechanism. The graph is formed by constructing edges from the similarity matrix, connecting each node to its nearest neighbors. Therefore, a graph neural network is utilized to extract and represent contextual information from both local and global areas. The aggregated image embeddings are then projected into the language model’s input space using a linear layer and combined with input caption tokens to fine-tune the large language models for caption generation. Our proposed method is validated using the BreakHis and PatchGastric microscopic datasets. The GNN-ViTCap method achieves an F1-Score of 0.934 and AUC of 0.963 for classification, along with BLEU@4 = 0.811 and METEOR = 0.569 for captioning. Experimental analysis demonstrates that the GNN-ViTCap architecture outper-forms state-of-the-art (SOTA) approaches, providing a reliable and efficient approach for patient diagnosis using microscopy images. S. M. Taslim Uddin Raju, Md. Milon Islam, Md. Rezwanul Haque, Hamdi Altaheri, Fakhri Karray |
IJCNN | 2 |
| 2025 | MMFformer: Multimodal Fusion Transformer Network for Depression DetectionabstractDepression is a serious mental health illness that significantly affects an individual’s well-being and quality of life, making early detection crucial for adequate care and treatment. Detecting depression is often difficult, as it is based primarily on subjective evaluations during clinical interviews. Hence, the early diagnosis of depression, thanks to the content of social networks, has become a prominent research area. The extensive and diverse nature of user-generated information poses a significant challenge, limiting the accurate extraction of relevant temporal information and the effective fusion of data across multiple modalities. This paper introduces MMF-former, a multimodal depression detection network designed to retrieve depressive spatio-temporal high-level patterns from multimodal social media information. The transformer network with residual connections captures spatial features from videos, and a transformer encoder is exploited to design important temporal dynamics in audio. Moreover, the fusion architecture fused the extracted features through late and intermediate fusion strategies to find out the most relevant intermodal correlations among them. Finally, the proposed network is assessed on two large-scale depression detection datasets, and the results clearly reveal that it surpasses existing state-of-the-art approaches, improving the F1-Score by 13.92% for D-Vlog dataset and 7.74% for LMVD dataset. The code is made available publicly at https://github.com/rezwanh001/Large-Scale-Multimodal-Depression-Detection. Md. Rezwanul Haque, Md. Milon Islam, S. M. Taslim Uddin Raju, Hamdi Altaheri, Lobna Nassar, Fakhri Karray |
SMC | 2 |
| 2025 | MDD-Net: Multimodal Depression Detection through Mutual TransformerabstractDepression is a major mental health condition that severely impacts the emotional and physical well-being of individuals. The simple nature of data collection from social media platforms has attracted significant interest in properly utilizing this information for mental health research. A Multimodal Depression Detection Network (MDD-Net), utilizing acoustic and visual data obtained from social media networks, is proposed in this work where mutual transformers are exploited to efficiently extract and fuse multimodal features for efficient depression detection. The MDD-Net consists of four core modules: an acoustic feature extraction module for retrieving relevant acoustic attributes, a visual feature extraction module for extracting significant high-level patterns, a mutual transformer for computing the correlations among the generated features and fusing these features from multiple modalities, and a detection layer for detecting depression using the fused feature representations. The extensive experiments are performed using the multimodal D-Vlog dataset, and the findings reveal that the developed multimodal depression detection network surpasses the state-of-the-art by up to 17.37% for F1-Score, demonstrating the greater performance of the proposed system. The source code is accessible at https://github.com/rezwanh001/Multimodal-Depression-Detection. Md. Rezwanul Haque, Md. Milon Islam, S. M. Taslim Uddin Raju, Hamdi Altaheri, Lobna Nassar, Fakhri Karray |
SMC | 2 |
| 2024 | TransUAAE-CapGen: Caption Generation from Histopathological Patches through Transformer and UNet-Based Adversarial AutoencoderabstractCaptioning Whole Slide Images (WSIs) for pathological analysis is an essential but not extensively explored aspect of computer-aided pathological diagnosis. Challenges arise from insufficient datasets and the effectiveness of model training. Generating automatic caption reports for various gastric adenocarcinoma images is another challenge. In this paper, we introduce a hybrid method referred to as TransUAAE-CapGen to generate histopathological captions from WSI patches. The TransUAAE-CapGen architecture consists of a hybrid UNet-based Advereasrial Autoencoder (AAE) for feature extraction and a transformer for caption generation. The hybrid UNet-based AAE extracted complex tissue properties from histopathological patches, transforming them into low-dimensional embeddings. The embeddings are then fed into the transformer to generate concise captions. Our proposed method is validated using the PatchGastricADC22 dataset. The TransUAAE-CapGen model provides the best estimated accuracy of BLEU-4 = 86.8%, METEOR = 59.6%, a ROUGE = 89.3%, and CIDEr = 7.72%. Experimental analysis indicates that the TransUAAE-CapGen architecture outperforms the traditional LSTM-based model for the caption generation task. Our findings reveal that the proposed architecture can effectively generate accurate and precise reports for medical image analysis. S. M. Taslim Uddin Raju, Abdul Raqeeb Mohammad, Md. Milon Islam, Fakhri Karray |
SMC | 3 |
| 2023 | Internet of Things: Device Capabilities, Architectures, Protocols, and Smart Applications in Healthcare DomainabstractNowadays, the Internet has spread to practically every country around the world and is having unprecedented effects on people’s lives. The Internet of Things (IoT) is getting more popular and has a high level of interest in both practitioners and academicians in the age of wireless communication due to its diverse applications. The IoT is a technology that enables everyday things to become savvier, everyday computation toward becoming intellectual, and everyday communication to become a little more insightful. In this article, the most common and popular IoT device capabilities, architectures, and protocols are demonstrated in brief to provide a clear overview of the IoT technology to the researchers in this area. The common IoT device capabilities, including hardware (Raspberry Pi, Arduino, and ESP8266) and software (operating systems (OSs), and built-in tools) platforms are described in detail. The widely used architectures that have recently evolved and used are the three-layer architecture, service-oriented architecture, and middleware-based architecture. The popular protocols for IoT are demonstrated which include constrained application protocol, message queue telemetry transport, extensible messaging and presence protocol, advanced message queuing protocol, data distribution service, low power wireless personal area network, Bluetooth low energy, and ZigBee that are frequently utilized to develop smart IoT applications. Additionally, this research provides an in-depth overview of the potential healthcare applications based on IoT technologies in the context of addressing various healthcare concerns. Finally, this article summarizes state-of-the-art knowledge, highlights open issues and shortcomings, and provides recommendations for further studies which would be quite beneficial to anyone with a desire to work in this field and make breakthroughs to get expertise in this area. Md. Milon Islam, Sheikh Nooruddin, Fakhri Karray, Muhammad Ghulam |
IEEE Internet Things J. | 1 |
| 2022 | Multimodal Human Activity Recognition for Smart Healthcare ApplicationsabstractHuman Activity Recognition (HAR) has emerged as a potential research topic for smart healthcare owing to the fast growth of wearable and smart devices in recent years. The significant applications of HAR in ambient assisted living environments include monitoring the daily activities of elderly and cognitively impaired individuals to assist them by observing their health status. In this research, we present a deep learning-based fusion approach for multimodal HAR that fuses the different modalities of data to obtain robust outcomes. Here, Convolutional Neural Networks (CNNs) retrieve the high-level attributes from the image data, and the Convolutional Long Short Term Memory (ConvLSTM) is utilized to capture significant patterns from the multi-sensory data. Finally, the extracted features from the modalities are fused through self-attention mechanisms that enhance the relevant activity data and inhibit the superfluous and possibly confusing information by measuring their compatibility. Lastly, extensive tests have been performed to measure the efficiency and robustness of the developed fusion approach using the UP-Fall detection dataset. It is evident from the experimental findings that the proposed fusion technique outperforms the existing state-of-the-art and achieves relatively better performance. Md. Milon Islam, Sheikh Nooruddin, Fakhri Karray |
SMC | 1 |