VLDB 2026 Research / reviewers in the wild / expert
Ashraful Islam
dblp:227/5727
· DBLP profile ↗
21ranked-venue papers
9as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 5 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Benchmarking CNN Architectures Using Transfer Learning for Lipomatous Tumor Classification: A Comprehensive Study
Md. Zafor Iqbal, Hasibul Haque, Fahmida Jamil, Redwan Rahman, Siam Tahsin Bhuiyan, Rashedur Rahman, Ashraful Islam, Saadia Binte Alam |
IEA/AIE (2) | 7 |
| 2025 | A Theoretical Framework for Decentralized Intrusion Detection in Smart Networks Using Blockchain and Machine Learning
Moinul Alam, Mostafa Hasan, Arvil Nath Akash, Md. Junayed Hossain, Ashraful Islam, Sanzar Adnan Alam |
AINA (8) | 5 |
| 2025 | Adaptive and Scalable Reinforcement Learning Method for Multi-Site Transportation OptimizationabstractExisting deep reinforcement learning methodologies in logistics optimization tend to confront a tripartite of inherent defects, including suboptimal environmental exploration, slow convergence, and insufficient dynamic adaptability to time-varying transportation, which greatly limits their effectiveness in real-world scenarios. To deal with these limitations, in this paper we propose an adaptive transport optimization algorithm based on reinforcement learning in order to optimize the multi-location product transportation scenario. We take the gym environment, ProductTransportEnv as our experimental simulator by configuring all concerned parameters, including dynamic inventory levels, stochastic demand patterns, delivery delays, variable transport costs, stockout penalties, etc. The method integrates Proximal Policy Optimization and Deep Q-Networks to learn efficient transportation strategies and balance operational costs and service objectives as well. The environment employs a holistic reward structure to balance revenue maximization, transportation cost reduction, and stockout penalties. Experimental results demonstrate the robustness and scalability across diverse demand scenarios, with PPO achieving stable policy convergence and DQN excelling in rapid reward generation. Comparative analyses highlight our method’s ability to reduce inventory depletion rates and improve reward consistency. Visualizations of decision-making processes and economic trade-offs further validate its adaptability in complex logistics networks. This research advances the application of reinforcement learning in supply chain management, offering a scalable solution for real-time optimization of dynamic transportation systems. Ashraful Islam |
IJCNN | 1 |
| 2025 | A Graph-Assisted Digital-Twin-Driven Multiagent Shared Offloading for Internet of VehiclesabstractVehicular edge computing (VEC) allows vehicles to process part of the tasks locally at the network edge while offloading the rest of the tasks to a centralized cloud server for processing. A massive volume of tasks generated by the Internet of Vehicles (IoV) leads to buffer overflow that causes higher latency. Elevating latency, in turn, can increase network energy consumption. Both higher latency and energy consumption lead to a degradation of network performance. Therefore, VEC design requires a balance between latency and energy consumption tradeoff. To reduce overwhelming amount of offloading to edge servers, a cooperative cluster-based shared offloading strategy has been proposed in this work. We use digital twin technology in VEC for managing and adapting to environmental dynamic changes. Then, we leverage Lyapunov (Ly) optimization to transform the stochastic offloading problem into a more manageable deterministic form. Finally, we present a decentralized coordination graph (CG)-driven Ly-based multiagent deep deterministic policy gradient (CG-LyMADDPG) algorithm that trains agents toward energy efficient optimal offloading policy while maintaining queue stability at a maximum delay constraint. The experimental result shows that the proposed learning significantly outperforms the baseline algorithms for energy savings while maintain queue stability. Md. Zahangir Alam, Suryaia Rahman, Md. Asif Bin Khaled, Ashraful Islam, Abbas Jamalipour |
IEEE Internet Things J. | 4 |
| 2024 | Using Transformers for Emotion Recognition in Bangla Text: A Comparative Study of MultiBERT and BanglaBERT with Data AugmentationabstractEmotion recognition is an exacting task due to the complexity and diversity of human emotions, as well as their tendency to be open to multiple interpretations. Despite these challenges, emotion recognition from textual data has achieved promising results in recent years for the English language. Numerous studies have been conducted on this topic, leading to significant advancements in various fields, including business, politics, education, research, and healthcare. However, accurate emotion recognition from Bengali (Bangla) text remains difficult due to limited resources available for the Bangla language. Additionally, the available datasets for Bengali text are often poorly annotated or imbalanced, which significantly hampers the performance of text classifiers and causes overfitting. Data augmentation can help address this issue. By addressing class imbalance through data augmentation via back-translation, we analyzed and evaluated the impact of augmentation on both real and synthetic datasets. We used transformer-based models, Multi-BERT and BanglaBERT. Data augmentation considerably boosts the accuracy by 8% and 6% for MultiBERT and BanglaBERT respectively, along with a 13% increase in F1-score for both the models. An ablation study also evaluated the effects of hyperparameters on our models. Nabarun Halder, Tanjina Piash Proma, Jahanggir Hossain Setu, Arafat Noor, Ashraful Islam, M. Ashraful Amin |
ICMLA | 5 |
| 2024 | ECGInsight: A Web Application-Based Approach to Myocardial Infarction Detection From ECG Image Reports Utilizing ResNetabstractCardiovascular illness is frequently the cause of my-ocardial infarction (MI), also referred to as a heart attack, which happens when the blood supply to a portion of the heart muscle is interrupted or diminished. This event can range from being asymptomatic to causing significant declines in cardiovascular function and premature mortality. An electrocardiogram (ECG) is vital for diagnosing and assessing MI, as it measures the heart's electrical activity and can detect various cardiac conditions. In this study, we developed a Residual Network (ResNet)-based model for MI detection using ECG signals. The model was trained on the PTB-XL dataset, which includes 21,837 annotated ECG signals, and tested on converted ECG signals from the Mendeley Data ECG Images dataset. Our model achieved high precision, recall, and F1-scores, demonstrating robust performance with an accuracy of 98.95 % on the validation set and 98.63 % on the test set. In addition, we developed ECGInsight, a web application that lets users take or upload a picture of an ECG report. The photo is converted to ECG signals, which are then processed by ResNet to produce predictions with confidence levels. This study underscores the potential of Deep Learning (DL) models in automated MI detection and highlights the practical utility of the ECGInsight web application in making MI detection accessible to users, ultimately contributing to advancements in cardiovascular healthcare. Jahanggir Hossain Setu, Syed Tangim Pasha, Nabarun Halder, Sankar Sikder, Ashraful Islam, Md. Zahangir Alam |
ICMLA | 5 |
| 2024 | Optimizing Software Release Management with GPT-Enabled Log Anomaly Detection
Jahanggir Hossain Setu, Md. Shazzad Hossain, Nabarun Halder, Ashraful Islam, M. Ashraful Amin |
ICPR (2) | 4 |
| 2024 | Empowering Multiclass Classification and Data Augmentation of Arabic News Articles Through Transformer ModelabstractArabic Natural Language Processing (NLP) presents unique challenges due to the complexity of the language, including variations in dialects, rich morphology, and context-dependent semantics. These linguistic intricacies make accurate classification a challenging task. This study presents the development and evaluation of an Arabic Bidirectional Encoder Representations from Transformers (AraBERT) model trained on a publicly available dataset named UltimateArabic dataset, which comprises 10 diverse classes. The aim is to perform multiclass classification and to address the imbalance class problem, using AraVec, of Arabic text data in the field of NLP. The model achieved an overall accuracy of 98% on augmented data, showcasing its proficiency in categorizing a wide range of Arabic texts including medical, technology, art, culture, diversity, society, religion, sports, politics, and economy. The AraBERT model exhibits strong classification performance compared to non-augmented data across the majority of categories, with an encouraging raise on macro-average precision of 4%, recall of 4%, and F1-score of 4%. This study underscores the potential of AraBERT along with AraVec for diverse applications in imbalanced Arabic text classification tasks and highlights the need for continued research to address the complexities of multiclass classification in Arabic NLP. Jahanggir Hossain Setu, Nabarun Halder, Sankar Sikder, Ashraful Islam, Md. Zahangir Alam |
IJCNN | 4 |
| 2023 | SceneCalib: Automatic Targetless Calibration of Cameras and Lidars in Autonomous DrivingabstractAccurate camera-to-lidar calibration is a requirement for sensor data fusion in many 3D perception tasks. In this paper, we present SceneCalib, a novel method for simultaneous self-calibration of extrinsic and intrinsic parameters in a system containing multiple cameras and a lidar sensor. Existing methods typically require specially designed calibration targets and human operators, or they only attempt to solve for a subset of calibration parameters. We resolve these issues with a fully automatic method that requires no explicit correspondences between camera images and lidar point clouds, allowing for robustness to many outdoor environments. Furthermore, the full system is jointly calibrated with explicit cross-camera constraints to ensure that camera-to-camera and camera-to-lidar extrinsic parameters are consistent. Ayon Sen, Anton Mitrokhin, Ashraful Islam |
ICRA | 4 |
| 2023 | Designing Healthcare Relational Agents: A Conceptual Framework with User-Centered Design GuidelinesabstractThis paper presents a conceptual framework for designing relational agents (RAs) in healthcare contexts, developed through the findings from multiple user studies on RAs about their acceptance, efficacy, and usability. The framework emphasizes a user-centered design (UCD) approach that takes into account the unique needs and preferences of patients, non-patient users, and healthcare professionals (HCPs). Based on the results of these studies, we analyzed and refined the RA designs and proposed a UCD-based conceptual framework for designing effective and user-friendly healthcare RAs. The paper aims to provide an initial resource for researchers, designers, and developers interested in developing RAs for healthcare contexts by thinking of UCD techniques. Ashraful Islam, Beenish M. Chaudhry, Aminul Islam 0001 |
ISCC | 1 |
| 2023 | SMOTE Oversampling and Near Miss Undersampling Based Diabetes Diagnosis from Imbalanced Dataset with XAI VisualizationabstractThis study investigated the predictive ability of ten different machine learning (ML) models for diabetes using a dataset that was not evenly distributed. Additionally, the study evaluated the effectiveness of two oversampling and undersampling methods, namely the Synthetic Minority Oversampling Technique (SMOTE) and the Near-Miss algorithm. Explainable Artificial Intelligence (XAI) techniques were employed to enhance the interpretability of the model's predictions. The results indicate that the extreme gradient boosting (XGB) model combined with SMOTE oversampling technique exhibited the highest accuracy and an F1-score of 99% and 1.00 respectively. Furthermore, the utilization of XAI methods increased the dependability of the model's decision-making process, rendering it more appropriate for clinical use. These results imply that integrating XAI with ML and oversampling techniques can enhance the early detection and management of diabetes, leading to better diagnosis and intervention. Nasim Mahmud Nayan, Ashraful Islam, Muhammad Usama Islam, Eshtiak Ahmed, Mohammad Mobarak Hossain, Md. Zahangir Alam |
ISCC | 2 |
| 2023 | Self-supervised Learning with Local Contrastive Loss for Detection and Semantic SegmentationabstractWe present a self-supervised learning (SSL) method suitable for semi-global tasks such as object detection and semantic segmentation. We enforce local consistency between self-learned features that represent corresponding image locations of transformed versions of the same image, by minimizing a pixel-level local contrastive (LC) loss during training. LC-loss can be added to existing self-supervised learning methods with minimal overhead. We evaluate our SSL approach on two downstream tasks – object detection and semantic segmentation, using COCO, PASCAL VOC, and CityScapes datasets. Our method outperforms the existing state-of-the-art SSL approaches by 1.9% on COCO object detection, 1.4% on PASCAL VOC detection, and 0.6% on CityScapes segmentation. Ashraful Islam, Ben Lundell, Harpreet Sawhney, Sudipta Sinha, Peter Morales, Richard J. Radke |
WACV | 1 |
| 2022 | ConfLab: A Data Collection Concept, Dataset, and Benchmark for Machine Analysis of Free-Standing Social Interactions in the WildabstractRecording the dynamics of unscripted human interactions in the wild is challenging due to the delicate trade-offs between several factors: participant privacy, ecological validity, data fidelity, and logistical overheads. To address these, following a 'datasets for the community by the community' ethos, we propose the Conference Living Lab (ConfLab): a new concept for multimodal multisensor data collection of in-the-wild free-standing social conversations. For the first instantiation of ConfLab described here, we organized a real-life professional networking event at a major international conference. Involving 48 conference attendees, the dataset captures a diverse mix of status, acquaintance, and networking motivations. Our capture setup improves upon the data fidelity of prior in-the-wild datasets while retaining privacy sensitivity: 8 videos (1920x1080, 60 fps) from a non-invasive overhead view, and custom wearable sensors with onboard recording of body motion (full 9-axis IMU), privacy-preserving low-frequency audio (1250 Hz), and Bluetooth-based proximity. Additionally, we developed custom solutions for distributed hardware synchronization at acquisition, and time-efficient continuous annotation of body keypoints and actions at high sampling rates. Our benchmarks showcase some of the open research tasks related to in-the-wild privacy-preserving social data analysis: keypoints detection from overhead camera views, skeleton-based no-audio speaker detection, and F-formation detection. Chirag Raman, José Vargas Quiros, Stephanie Tan, Ashraful Islam, Ekin Gedik, Hayley Hung |
NeurIPS | 4 |
| 2021 | A Hybrid Attention Mechanism for Weakly-Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization is a challenging vision task due to the absence of ground-truth temporal locations of actions in the training videos. With only video-level supervision during training, most existing methods rely on a Multiple Instance Learning (MIL) framework to predict the start and end frame of each action category in a video. However, the existing MIL-based approach has a major limitation of only capturing the most discriminative frames of an action, ignoring the full extent of an activity. Moreover, these methods cannot model background activity effectively, which plays an important role in localizing foreground activities. In this paper, we present a novel framework named HAM-Net with a hybrid attention mechanism which includes temporal soft, semi-soft and hard attentions to address these issues. Our temporal soft attention module, guided by an auxiliary background class in the classification module, models the background activity by introducing an ``action-ness'' score for each video snippet. Moreover, our temporal semi-soft and hard attention modules, calculating two attention scores for each video snippet, help to focus on the less discriminative frames of an action to capture the full action boundary. Our proposed approach outperforms recent state-of-the-art methods by at least 2.2% mAP at IoU threshold 0.5 on the THUMOS14 dataset, and by at least 1.3% mAP at IoU threshold 0.75 on the ActivityNet1.2 dataset. Ashraful Islam, Chengjiang Long, Richard J. Radke |
AAAI | 1 |
| 2021 | Identifying User Personas for Engagement with a COVID-19 Health Service Delivery Relational AgentabstractUser personas play a crucial role in user-centered design (UCD) methodology-based technology development. In this way, appropriate user personas improve the design of interaction interfaces between a relational agent (RA) and its users, particularly when the RA is about delivering health services to the general population. This article discusses the approach for establishing patient personas to help in designing and developing interaction scenarios and dialogues between a health service RA and its target users during various COVID-19 stages. We employed a mixed-method research that was completed in two phases involving both healthcare professionals (HCPs) and individuals who have encountered COVID-19. Our approach resulted in the identification of three COVID-19 related user personas i.e., experiencing symptoms, COVID-19 positive with mild symptoms, and recovering from COVID-19. Ashraful Islam, Beenish M. Chaudhry |
HAI | 1 |
| 2021 | A Broad Study on the Transferability of Visual Representations with Contrastive LearningabstractTremendous progress has been made in visual representation learning, notably with the recent success of self-supervised contrastive learning methods. Supervised contrastive learning has also been shown to outperform its cross-entropy counterparts by leveraging labels for choosing where to contrast. However, there has been little work to explore the transfer capability of contrastive learning to a different domain. In this paper, we conduct a comprehensive study on the transferability of learned representations of different contrastive approaches for linear evaluation, full-network transfer, and few-shot recognition on 12 downstream datasets from different domains, and object detection tasks on MSCOCO and VOC0712. The results show that the contrastive approaches learn representations that are easily transferable to a different downstream task. We further observe that the joint objective of self-supervised contrastive loss with cross-entropy/supervised-contrastive loss leads to better transferability of these models over their supervised counterparts. Our analysis reveals that the representations learned from the contrastive approaches contain more low/mid-level semantics than cross-entropy models, which enables them to quickly adapt to a new task. Our codes and models will be publicly available to facilitate future research on transferability of visual representations.1 Ashraful Islam, Chun-Fu Chen 0001, Rameswar Panda, Leonid Karlinsky, Richard J. Radke, Rogério Feris |
ICCV | 1 |
| 2021 | Design Validation of a Workplace Stress Management Mobile App for Healthcare Workers During COVID-19 and Beyond
Beenish M. Chaudhry, Ashraful Islam |
MobiQuitous | 2 |
| 2021 | Dynamic Distillation Network for Cross-Domain Few-Shot Recognition with Unlabeled DataabstractMost existing works in few-shot learning rely on meta-learning the network on a large base dataset which is typically from the same domain as the target dataset. We tackle the problem of cross-domain few-shot learning where there is a large shift between the base and target domain. The problem of cross-domain few-shot recognition with unlabeled target data is largely unaddressed in the literature. STARTUP was the first method that tackles this problem using self-training. However, it uses a fixed teacher pretrained on a labeled base dataset to create soft labels for the unlabeled target samples. As the base dataset and unlabeled dataset are from different domains, projecting the target images in the class-domain of the base dataset with a fixed pretrained model might be sub-optimal. We propose a simple dynamic distillation-based approach to facilitate unlabeled images from the novel/base dataset. We impose consistency regularization by calculating predictions from the weakly-augmented versions of the unlabeled images from a teacher network and matching it with the strongly augmented versions of the same images from a student network. The parameters of the teacher network are updated as exponential moving average of the parameters of the student network. We show that the proposed network learns representation that can be easily adapted to the target domain even though it has not been trained with target-specific classes during the pretraining phase. Our model outperforms the current state-of-the art method by 4.4% for 1-shot and 3.6% for 5-shot classification in the BSCD-FSL benchmark, and also shows competitive performance on traditional in-domain few-shot learning task. Ashraful Islam, Chun-Fu Chen 0001, Rameswar Panda, Leonid Karlinsky, Rogério Feris, Richard J. Radke |
NeurIPS | 1 |
| 2020 | DOA-GAN: Dual-Order Attentive Generative Adversarial Network for Image Copy-Move Forgery Detection and LocalizationabstractImages can be manipulated for nefarious purposes to hide content or to duplicate certain objects through copy-move operations. Discovering a well-crafted copy-move forgery in images can be very challenging for both humans and machines; for example, an object on a uniform background can be replaced by an image patch of the same background. In this paper, we propose a Generative Adversarial Network with a dual-order attention model to detect and localize copy-move forgeries. In the generator, the first-order attention is designed to capture copy-move location information, and the second-order attention exploits more discriminative features for the patch co-occurrence. Both attention maps are extracted from the affinity matrix and are used to fuse location-aware and co-occurrence features for the final detection and localization branches of the network. The discriminator network is designed to further ensure more accurate localization results. To the best of our knowledge, we are the first to propose such a network architecture with the 1st-order attention mechanism from the affinity matrix. We have performed extensive experimental validation and our state-of-the-art results strongly demonstrate the efficacy of the proposed approach. Ashraful Islam, Chengjiang Long, Arslan Basharat, Anthony Hoogs |
CVPR | 1 |
| 2020 | Use of Remote Sensing Satellite Images in Rice Area Monitoring System of BangladeshabstractIn this work, we developed a satellite-based Boro rice monitoring system which will be capable of collecting data on rice production area of whole Bangladesh along with a graphical statistics of rice production area every season. In this monitoring system, Bangladesh has been divided into 7 regions and the production area of Boro rice for these 7 regions has been calculated using only the MODIS (Moderate Resolution Imaging Spectroradiometer) satellite data. Our system does not require any ground truth data. The principal approach requires remotely detected information from MODIS at 250m250m spatial resolutions gained more than two distinct time periods: sowing season which is from December 1 to January 10 and growing season which is from January 11 to April 10 throughout the year (Boro Seasons). Total of eight years 2008-2015 MODIS data have been accumulated and an improvised Normalized Difference Vegetation Index (NDVI) threshold of the Boro rice model has been used for rice area extraction. This monitoring system has some features like region-wise and area-wise information, search options, improvement/deterioration analysis, and rice area extraction map view. The rice monitoring system can be used by the Department of Agriculture of Bangladesh or Food ministry for the collection of accurate and up-to-date information on the country's rice production. It will also be able to provide information on Boro rice production readily accessible to decision-makers and government agencies to help respond better during and after natural disasters or in any need. Moreover, we believe that our developments of forecasting the Boro rice yield would be useful for the decision-makers in addressing food security in Bangladesh. Kazi A. Kalpoma, Rumman Ali, Ashiqur Rahman, Ashraful Islam |
IGARSS | 4 |
| 2020 | Weakly Supervised Temporal Action Localization Using Deep Metric LearningabstractTemporal action localization is an important step towards video understanding. Most current action localization methods depend on untrimmed videos with full temporal annotations of action instances. However, it is expensive and time-consuming to annotate both action labels and temporal boundaries of videos. To this end, we propose a weakly supervised temporal action localization method that only requires video-level action instances as supervision during training. We propose a classification module to generate action labels for each segment in the video, and a deep metric learning module to learn the similarity between different action instances. We jointly optimize a balanced binary cross-entropy loss and a metric loss using a standard backpropagation algorithm. Extensive experiments demonstrate the effectiveness of both of these components in temporal localization. We evaluate our algorithm on two challenging untrimmed video datasets: THUMOS14 and ActivityNet1.2. Our approach improves the current state-of-the-art result for THUMOS14 by 6.5% mAP at IoU threshold 0.5, and achieves competitive performance for ActivityNet1.2. Ashraful Islam, Richard J. Radke |
WACV | 1 |