EDBT 2026 Demo / reviewers in the wild / expert
Md Jahidul Islam
dblp:203/5051
· DBLP profile ↗
17ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 7 since 2021Systems, architecture and hardware · 12 · 4 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiRQA: Hierarchical Ranking and Quality Alignment for Opinion-Unaware Image Quality Assessment
Vaishnav Ramesh, Md Jahidul Islam |
Mach. Vis. Appl. | 3 |
| 2025 | Word2Wave: Language Driven Mission Programming for Efficient Subsea Deployments of Marine RobotsabstractThis paper explores the design and development of a language-based interface for dynamic mission programming of autonomous underwater vehicles (AUVs). The proposed 'Word2Wave' (W2W) framework enables interactive programming and parameter configuration of AUVs for remote subsea missions. The W2W framework includes: (i) a set of novel language rules and command structures for efficient language-to-mission mapping; (ii) a GPT-based prompt engineering module for training data generation; (iii) a small language model (SLM)-based sequence-to-sequence learning pipeline for mission command generation from human speech or text; and (iv) a novel user interface for 2D mission map visualization and human-machine interfacing. The proposed learning pipeline adapts an SLM named T5-Small that can learn language-to-mission mapping from processed language data effectively, providing robust and efficient performance. In addition to a benchmark evaluation with state-of-the-art, we conduct a user interaction study to demonstrate the effectiveness of W2W over commercial AUV programming interfaces. Across participants, W2W-based programming required less than 10% time for mission programming compared to traditional interfaces; it is deemed to be a simpler and more natural paradigm for subsea mission programming with a usability score of 76.25. W2W opens up promising future research opportunities on hands-free AUV mission programming for efficient subsea deployments. Ruo Chen, David Blow, Adnan Abdullah, Md Jahidul Islam |
ICRA | 4 |
| 2025 | Optimized Vessel Segmentation: A Structure-Agnostic Approach With Small Vessel Enhancement and Morphological CorrectionabstractAccurate segmentation of blood vessels is essential for various clinical assessments and postoperative analyses. However, the inherent challenges of vascular imaging-such as sparsity, fine granularity, low contrast, data distribution variability, and the critical need for preserving topological integrity-make generalized vessel segmentation particularly complex. While specialized segmentation methods have been developed for specific anatomical regions, their over-reliance on tailored models hinders broader applicability and generalization. General-purpose segmentation models introduced in medical imaging often fail to address critical vascular characteristics, including the connectivity of segmentation results. In this study, we propose OVS-Net, an optimized vessel segmentation framework designed to generalize across diverse vessel structures and imaging modalities. It introduces a dual-branch architecture design for improving small vessel segmentation and a morphology-aware correction module to preserve vascular topology and connectivity. We compiled a comprehensive multi-modality dataset from 17 datasets to train and benchmark the proposed OVS-Net against 6 SAM-based methods and 17 expert models under various conditions. The results demonstrate that our approach achieves superior segmentation accuracy, generalization, and a 34.6% improvement in connectivity, underscoring its potential for clinical applications. The code and dataset information are available at https://github.com/Hk416mod2/OVS-Net. Dongning Song, Weijian Huang, Jiarun Liu, Md Jahidul Islam, Hao Yang 0026, Shuqiang Wang, Hairong Zheng, Shanshan Wang 0002 |
IEEE Trans. Image Process. | 4 |
| 2024 | Edge-Centric Real-Time Segmentation for Autonomous Underwater Cave ExplorationabstractThis paper addresses the challenge of deploying machine learning (ML)-based segmentation models on edge platforms to facilitate real-time scene segmentation for Autonomous Underwater Vehicles (AUVs) in underwater cave exploration and mapping scenarios. We focus on three ML models-U-Net, CaveSeg, and YOLOv8n-deployed on four edge platforms: Raspberry Pi-4, Intel Neural Compute Stick 2 (NCS2), Google Edge TPU, and NVIDIA Jetson Nano. Experimental results reveal that mobile models with modern architectures, such as YOLOv8n, and specialized models for semantic segmentation, like U-Net, offer higher accuracy with lower latency. YOLOv8n emerged as the most accurate model, achieving a 72.5 Intersection Over Union (IoU) score. Meanwhile, the U-Net model deployed on the Coral Dev board delivered the highest speed at 79.24 FPS and the lowest energy consumption at 6.23 mJ. The detailed quantitative analyses and comparative results presented in this paper offer critical insights for deploying cave segmentation systems on underwater robots, ensuring safe and reliable AUV navigation during cave exploration and mapping missions. Mohammadreza Mohammadi, Adnan Abdullah, Aishneet Juneja, Ioannis M. Rekleitis, Md Jahidul Islam, Ramtin Zand |
ICMLA | 5 |
| 2024 | CaveSeg: Deep Semantic Segmentation and Scene Parsing for Autonomous Underwater Cave ExplorationabstractIn this paper, we present CaveSeg - the first visual learning pipeline for semantic segmentation and scene parsing for AUV navigation inside underwater caves. We address the problem of scarce annotated training data by preparing a comprehensive dataset for semantic segmentation of underwater cave scenes. It contains pixel annotations for important navigation markers (e.g. caveline, arrows), obstacles (e.g. ground plain and overhead layers), scuba divers, and open areas for servoing. Through comprehensive benchmark analyses on cave systems in USA, Mexico, and Spain locations, we demonstrate that robust deep visual models can be developed based on CaveSeg for fast semantic scene parsing of underwater cave environments. In particular, we formulate a novel transformer-based model that is computationally light and offers near real-time execution in addition to achieving state-of-the-art performance. Finally, we explore the design choices and implications of semantic segmentation for visual servoing by AUVs inside underwater caves. The proposed model and benchmark dataset open up promising opportunities for future research in autonomous underwater cave exploration and mapping. Adnan Abdullah, Titon Barua, Reagan Tibbetts, Md Jahidul Islam, Ioannis M. Rekleitis |
ICRA | 5 |
| 2024 | AquaSonic: Acoustic Manipulation of Underwater Data Center Operations and Resource ManagementabstractUnderwater data centers (UDCs) hold promise as next-generation data storage due to their energy efficiency and environmental sustainability benefits. While the natural cooling properties of water save power, the isolated aquatic environment and long-range sound propagation characteristics in water create unique vulnerabilities which differ from those of on-land data centers. Our research discovers the unique vulnerabilities of fault-tolerant storage devices, resource allocation software, and distributed file systems to acoustic injection attacks in UDCs. With a realistic testbed approximating UDC server operations, we empirically characterize the capabilities of acoustic injection underwater and find that an attacker can reduce fault-tolerant RAID 5 storage system throughput by 17% up to 100%. Our closed-water analyses reveal that an attacker can (i) cause unresponsiveness and automatic node removal in a distributed filesystem with only 2.4 minutes of sustained acoustic injection, (ii) induce a distributed database’s latency to increase by up to 92.7% to reduce system reliability, and (iii) induce load-balance managers to redirect up to 74% of resources to a target server to cause overload or force resource colocation. Furthermore, we perform open-water experiments in a lake and find that an attacker can cause controlled throughput degradation at the maximum allowable distance of 6.35 m using a commercial speaker. We also investigate and discuss the effectiveness of standard defenses against acoustic injection attacks. Finally, we formulate a novel machine learning-based detection system that reaches 0% False Positive Rate and 98.2% True Positive Rate trained on our dataset of profiled hard disk drives under 30-second FIO benchmark execution. With this work, we aim to help manufacturers proactively protect UDCs against acoustic injection attacks and ensure the security of subsea computing infrastructures. Jennifer Sheldon, Weidong Zhu 0002, Adnan Abdullah, S. Hrushikesh Bhupathiraju, Takeshi Sugawara 0001, Kevin R. B. Butler, Md Jahidul Islam, Sara Rampazzi |
SP | 7 |
| 2023 | Deep Note: Can Acoustic Interference Damage the Availability of Hard Disk Storage in Underwater Data Centers?abstractThe growing worldwide attention toward large-scale subsea data centers has garnered substantial interest from commercial entities which have built and deployed underwater prototypes since 2015. These data centers utilize hard disk drives (HDDs) as a cost-effective method of data storage. However, researchers have demonstrated that acoustic waves can affect the availability and integrity of HDDs and applications that rely on them. These studies are all conducted in air on commercial laptops, hence their applicability and implications in submerged environments remain unexplored. In this position paper, we investigate potential vulnerabilities of storage devices deployed in underwater data centers and subsea storage platforms against targeted acoustic attacks. Based on our initial investigation of a simplified scenario, a victim HDD deployed in an enclosed submerged container is especially vulnerable to those acoustic attacks, which at frequencies ranging from 300 Hz 1300 Hz can result in up to 100% throughput loss and application crashes. Based on these findings, we argue that further study is necessary to assess underwater storage system security and develop effective defenses against overlooked acoustic attacks. Jennifer Sheldon, Weidong Zhu 0002, Adnan Abdullah, Kevin R. B. Butler, Md Jahidul Islam, Sara Rampazzi |
HotStorage | 5 |
| 2023 | Caveline Detection at the Edge for Autonomous Underwater Cave Exploration and MappingabstractThis paper explores the problem of deploying machine learning (ML)-based object detection and segmentation models on edge platforms to enable realtime caveline detection for Autonomous Underwater Vehicles (AUVs) used for under-water cave exploration and mapping. We specifically investigate three ML models, i.e., U-Net, Vision Transformer (ViT), and YOLOv8, deployed on three edge platforms: Raspberry Pi-4, Intel Neural Compute Stick 2 (NCS2), and NVIDIA Jetson Nano. The experimental results unveil clear tradeoffs between model accuracy, processing speed, and energy consumption. The most accurate model has shown to be U-Net with an 85.53 F1-score and 85.38 Intersection Over Union (IoU) value. Meanwhile, the highest inference speed and lowest energy consumption are achieved by the YOLOv8 model deployed on Jetson Nano operating in the high-power and low-power modes, respectively. The comprehensive quantitative analyses and comparative results provided in the paper highlight important nuances that can guide the deployment of caveline detection systems on underwater robots for ensuring safe and reliable AUV navigation during underwater cave exploration and mapping missions. Mohammadreza Mohammadi, Sheng-En Huang, Titon Barua, Ioannis M. Rekleitis, Md Jahidul Islam, Ramtin Zand |
ICMLA | 5 |
| 2023 | UDepth: Fast Monocular Depth Estimation for Visually-guided Underwater RobotsabstractIn this paper, we present a fast monocular depth estimation method for enabling 3D perception capabilities of low-cost underwater robots. We formulate a novel end-to-end deep visual learning pipeline named UDepth, which incorporates domain knowledge of image formation characteristics of natural underwater scenes. First, we adapt a new input space from raw RGB image space by exploiting underwater light attenuation prior, and then devise a least-squared formulation for coarse pixel-wise depth prediction. Subsequently, we extend this into a domain projection loss that guides the end-to-end learning of UDepth on over 9K RGB-D training samples. UDepth is designed with a computationally light MobileNetV2 backbone and a Transformer-based optimizer for ensuring fast inference rates on embedded systems. By domain-aware design choices and through comprehensive experimental analyses, we demonstrate that it is possible to achieve state-of-the-art depth estimation performance while ensuring a small computational footprint. Specifically, with 70 % −80 % less network parameters than existing benchmarks, UDepth achieves comparable and often better depth estimation performance. While the full model offers over 66 FPS (13 FPS) inference rates on a single GPU (CPU core), our domain projection for coarse depth prediction runs at 51.5 FPS rates on single-board Jetson TX2s. The inference pipelines are available at https://github.com/uf-robopi/UDepth. Boxiao Yu, Md Jahidul Islam |
ICRA | 3 |
| 2023 | Weakly Supervised Caveline Detection for AUV Navigation Inside Underwater CavesabstractUnderwater caves are challenging environments that are crucial for water resource management, and for our understanding of hydro-geology and history. Mapping underwater caves is a time-consuming, labor-intensive, and hazardous operation. For autonomous cave mapping by underwater robots, the major challenge lies in vision-based estimation in the complete absence of ambient light, which results in constantly moving shadows due to the motion of the camera-light setup. Thus, detecting and following the caveline as navigation guidance is paramount for robots in autonomous cave mapping missions. In this paper, we present a computationally light caveline detection model based on a novel Vision Transformer (ViT)-based learning pipeline. We address the problem of scarce annotated training data by a weakly supervised formulation where the learning is reinforced through a series of noisy predictions from intermediate sub-optimal models. We validate the utility and effectiveness of such weak supervision for caveline detection and tracking in three different cave locations: USA, Mexico, and Spain. Experimental results demonstrate that our proposed model, CL-ViT, balances the robustness-efficiency trade-off, ensuring good generalization performance while offering 10+ FPS on single-board (Jetson TX2) devices. Boxiao Yu, Reagan Tibbetts, Titon Barua, Ailani Morales, Ioannis M. Rekleitis, Md Jahidul Islam |
IROS | 6 |
| 2020 | Underwater Image Super-Resolution using Deep Residual MultipliersabstractWe present a deep residual network-based generative model for single image super-resolution (SISR) of underwater imagery for use by autonomous underwater robots. We also provide an adversarial training pipeline for learning SISR from paired data. In order to supervise the training, we formulate an objective function that evaluates the perceptual quality of an image based on its global content, color, and local style information. Additionally, we present USR-248, a large-scale dataset of three sets of underwater images of `high' (640×480) and `low' (80 × 60, 160 × 120, and 320×240) resolution. USR-248 contains paired instances for supervised training of 2×, 4×, or 8× SISR models. Furthermore, we validate the effectiveness of our proposed model through qualitative and quantitative experiments and compare the results with several state-of-the-art models' performances. We also analyze its practical feasibility for applications such as scene understanding and attention modeling in noisy visual conditions. Md Jahidul Islam, Sadman Sakib Enan, Peigen Luo, Junaed Sattar |
ICRA | 1 |
| 2020 | Semantic Segmentation of Underwater Imagery: Dataset and BenchmarkabstractIn this paper, we present the first large-scale dataset for semantic Segmentation of Underwater IMagery (SUIM). It contains over 1500 images with pixel annotations for eight object categories: fish (vertebrates), reefs (invertebrates), aquatic plants, wrecks/ruins, human divers, robots, and sea-floor. The images have been rigorously collected during oceanic explorations and human-robot collaborative experiments, and annotated by human participants. We also present a comprehensive benchmark evaluation of several state-of-the-art semantic segmentation approaches based on standard performance metrics. Additionally, we present SUIM-Net, a fully-convolutional deep residual model that balances the trade-off between performance and computational efficiency. It offers competitive performance while ensuring fast end-to-end inference, which is essential for its use in the autonomy pipeline by visually-guided underwater robots. In particular, we demonstrate its usability benefits for visual servoing, saliency prediction, and detailed scene understanding. With a variety of use cases, the proposed model and benchmark dataset open up promising opportunities for future research in underwater robot vision. Md Jahidul Islam, Chelsey Edge, Yuyang Xiao, Peigen Luo, Muntaqim Mehtaz, Christopher Morse, Sadman Sakib Enan, Junaed Sattar |
IROS | 1 |
| 2019 | Robotic Detection of Marine Litter Using Deep Visual Detection ModelsabstractTrash deposits in aquatic environments have a destructive effect on marine ecosystems and pose a long-term economic and environmental threat. Autonomous underwater vehicles (AUVs) could very well contribute to the solution of this problem by finding and eventually removing trash. This paper evaluates a number of deep-learning algorithms performing the task of visually detecting trash in realistic underwater environments, with the eventual goal of exploration, mapping, and extraction of such debris by using AUVs. A large and publicly-available dataset of actual debris in open-water locations is annotated for training a number of convolutional neural network architectures for object detection. The trained networks are then evaluated on a set of images from other portions of that dataset, providing insight into approaches for developing the detection capabilities of an AUV for underwater trash removal. In addition, the evaluation is performed on three different platforms of varying processing power, which serves to assess these algorithms' fitness for real-time applications. Michael Fulton, Jungseok Hong, Md Jahidul Islam, Junaed Sattar |
ICRA | 3 |
| 2018 | Enhancing Underwater Imagery Using Generative Adversarial NetworksabstractAutonomous underwater vehicles (AUVs) rely on a variety of sensors - acoustic, inertial and visual - for intelligent decision making. Due to its non-intrusive, passive nature and high information content, vision is an attractive sensing modality, particularly at shallower depths. However, factors such as light refraction and absorption, suspended particles in the water, and color distortion affect the quality of visual data, resulting in noisy and distorted images. AUVs that rely on visual sensing thus face difficult challenges and consequently exhibit poor performance on vision-driven tasks. This paper proposes a method to improve the quality of visual underwater scenes using Generative Adversarial Networks (GANs), with the goal of improving input to vision-driven behaviors further down the autonomy pipeline. Furthermore, we show how recently proposed methods are able to generate a dataset for the purpose of such underwater image restoration. For any visually-guided underwater robots, this improvement can result in increased safety and reliability through robust visual perception. To that effect, we present quantitative and qualitative data which demonstrates that images corrected through the proposed approach generate more visually appealing images, and also provide increased accuracy for a diver tracking algorithm. Cameron Fabbri, Md Jahidul Islam, Junaed Sattar |
ICRA | 2 |
| 2018 | Dynamic Reconfiguration of Mission Parameters in Underwater Human-Robot CollaborationabstractThis paper presents a real-time programming and parameter reconfiguration method for autonomous underwater robots in human-robot collaborative tasks. Using a set of intuitive and meaningful hand gestures, we develop a syntactically simple framework that is computationally more efficient than a complex, grammar-based approach. In the proposed framework, a convolutional neural network is trained to provide accurate hand gesture recognition; subsequently, a finite-state machine- based deterministic model performs efficient gesture-to-instruction mapping and further improves robustness of the interaction scheme. The key aspect of this framework is that it can be easily adopted by divers for communicating simple instructions to underwater robots without using artificial tags such as fiducial markers or requiring memorization of a potentially complex set of language rules. Extensive experiments are performed both on field-trial data and through simulation, which demonstrate the robustness, efficiency, and portability of this framework in a number of different scenarios. Finally, a user interaction study is presented that illustrates the gain in the ease of use of our proposed interaction framework compared to the existing methods for the underwater domain. Md Jahidul Islam, Marc Ho, Junaed Sattar |
ICRA | 1 |
| 2017 | Mixed-domain biological motion tracking for underwater human-robot interactionabstractWe present a novel algorithm for an autonomous underwater robot to visually detect and follow its companion human diver. Using both spatial-domain and frequency-domain features pertaining to human swimming patterns, we devise an algorithm to visually detect the position and swimming direction of the diver. Our algorithm is unique in the way that it allows detection of arbitrary motion directions, in addition to keeping track of a diver's position through the image sequence over time. A Hidden Markov Model (HMM)-based approach prunes the search-space of all potential trajectories relying on image intensities in the spatial-domain. A diver's motion signature is subsequently detected in a sequence of non-overlapping image subwindows exhibiting human swimming patterns. The pruning step ensures efficient computation by avoiding exponentially large search-spaces, whereas the frequency-domain detection allows us to detect the diver's position and motion direction accurately. Experimental validation of the proposed approach is presented on datasets collected from open-water and closed-water environments. Md Jahidul Islam, Junaed Sattar |
ICRA | 1 |
| 2017 | Underwater multi-robot convoying using visual tracking by detectionabstractWe present a robust multi-robot convoying approach that relies on visual detection of the leading agent, thus enabling target following in unstructured 3-D environments. Our method is based on the idea of tracking-by-detection, which interleaves efficient model-based object detection with temporal filtering of image-based bounding box estimation. This approach has the important advantage of mitigating tracking drift (i.e. drifting away from the target object), which is a common symptom of model-free trackers and is detrimental to sustained convoying in practice. To illustrate our solution, we collected extensive footage of an underwater robot in ocean settings, and hand-annotated its location in each frame. Based on this dataset, we present an empirical comparison of multiple tracker variants, including the use of several convolutional neural networks, both with and without recurrent connections, as well as frequency-based model-free trackers. We also demonstrate the practicality of this tracking-by-detection strategy in real-world scenarios by successfully controlling a legged underwater robot in five degrees of freedom to follow another robot's independent motion. Florian Shkurti, Wei-Di Chang, Peter Henderson 0002, Md Jahidul Islam, Juan Camilo Gamboa Higuera, Jimmy Li 0001, Travis Manderson, Anqi Xu 0003, Gregory Dudek, Junaed Sattar |
IROS | 4 |