Isaac Gerg

dblp:197/8628 · also Isaac D. Gerg · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
10since 2021 · last 2024
0000-0002-3352-6864ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 6 first-author · 10 since 2021
YearPublicationVenuePosition
2024 Progressive Diffusion Autofocus for Synthetic Aperture Sonar Imagery
abstract
Autofocus algorithms for synthetic aperture sonar (SAS) enhance the autonomy of uncrewed underwater vehicles by counteracting environmental factors that often compromise image quality, thus preserving the vehicles’ capacity for accurate perception. Often, these errors are systematic resulting from misestimation of environmental parameters (e.g., sound speed) or sonar position. Traditional autofocusing techniques aimed at correcting these distortions may result in degrading well-focused imagery when applied repeatedly, while single-attempt focusing methods may not fully correct the errors. In this work, we introduce a novel SAS autofocus approach named Progressive Diffusion Autofocus (PDA), which incrementally refines the focus over several iterations, drawing inspiration from recent deep-learning-centric diffusion processes. Our method iteratively improves focus with each iteration, thereby enhancing the image progressively. We evaluate our approach against established methods, such as optimization-based sharpness metrics and a recently introduced deep learning method, Deep Adaptive Phase Learning (DAPL). Our results indicate that our PDA method not only consistently improves image quality, but also minimizes the degradation of images that are already well-focused.
Balkan V. Bingol, Isaac Gerg, Vishal Monga
IGARSS2
2024 Uncovering Bias in Building Damage Assessment from Satellite Imagery
abstract
We identify a bias in a commonly used dataset for building damage detection, evaluate its effects on existing deep learning models, and devise mitigation strategies to overcome it. We find that the data contains significantly more groups of damaged buildings than single ones leading to skewed machine learning evaluations. Consequently, deep learning models heavily rely on surrounding context rather than individual building damage when classifying supporting our claim. Specifically, the dataset includes extraneous damage surrounding buildings such as debris, fallen trees, and other damaged buildings which results in deep neural networks overfitting to these features. We analyze the top-5 solutions of the xView2 challenge, which focuses on building damage classification using satellite imagery as provided by the xBD dataset. Our experiments reveal that these models struggle to accurately identify isolated damaged buildings, potentially causing oversights in critical disaster scenarios and delaying humanitarian aid. Finally, we devise a new augmentation strategy to reduce this bias in disaster datasets and show it improves real-world outcomes.
Dennis Melamed, Cameron Johnson, Isaac Gerg, Russell Blue, Anthony Hoogs, Brian Clipp, Philip Morrone
IGARSS3
2023 A Perceptual Metric Prior on Deep Latent Space Improves Out-Of-Distribution Synthetic Aperture Sonar Image Classification
abstract
Deep learning methods have achieved state-of-the-art performance on various machine learning benchmarks. However, when evaluated on data outside of the training-set distribution, performance can be mixed, even though this data is trivial for humans to discern. One reason for this gap is that deep network training is only tasked with solving a context-limited optimization problem. As long as the loss on the training data is minimized, any suitable set of features and decision boundary geometry are candidate solutions. This can result in the reliance on non-robust features that do not take into account the larger context, leading to poor performance on out-of-distribution (OOD) data that presents no trouble for humans. For example, training on a limited set of 2-dimensional image data makes it difficult to learn relationships that may be obvious in 3 dimensions. To partially mitigate this issue, we propose using a perceptual metric prior (PMP) to influence the latent manifold structure, mimicking the characteristics of human perception and improving OOD performance. Our method is demonstrated on a real-world synthetic aperture sonar (SAS) dataset, showing good performance on OOD imagery, even when only limited training data is available, as is often the case in SAS. Our proposal aims to address the under-specification issue by taking into account how inter-class samples relate to each other, encouraging the latent feature manifold to consider a larger context.
Isaac Gerg, Carl F. Cotner
IGARSS1
2023 Deep Multi-Look Sequence Processing for Synthetic Aperture Sonar Image Segmentation
abstract
Deep learning has enabled significant improvements in semantic image segmentation, especially in underwater imaging domains such as side scan sonar (SSS). In this work, we apply deep learning to synthetic aperture sonar (SAS) imagery, which has an advantage over traditional SSS in which SAS produces coherent high- and constant-resolution imagery. Despite the successes of deep learning, one drawback is the need for abundant labeled training data to enable success. Such abundant labeled data are not always available as in the case of SAS where collections are expensive and obtaining quality ground-truth labels may require diver intervention. To overcome these challenges, we propose a domain-specific deep learning network architecture utilizing a unique property to complex-valued SAS imagery: the ability to resolve angle-of-arrival (AoA) of acoustic returns through$k$-space processing. By sweeping through consecutive incrementally advanced AoA bandpass filters (a process sometimes referred to as multi-look processing), this technique generates a sequence of images emphasizing angle-dependent seafloor scattering and motion from biologics along the seafloor or in the water column. Our proposal, which we call multi-look sequence processing network (MLSP-Net), is a domain-enriched deep neural network architecture that models the multi-look image sequence using a recurrent neural network (RNN) to extract robust features suitable for semantic segmentation of the seafloor without the need for abundant training data. Unlike previous segmentation works in SAS, our model ingests a complex-valued SAS image and affords the ability to learn the AoA filters in$k$-space as part of the training procedure. We show the results on a challenging real-world SAS database, and despite the lack of abundant training data, our proposed method shows superior results over state-of-the-art techniques.
Isaac Gerg, Vishal Monga
IEEE Trans. Geosci. Remote. Sens.1
2023 Resonant Scattering-Inspired Deep Networks for Munition Detection in 3D Sonar Imagery
abstract
Underwater sites affected by unexploded ordnance (UXO) pose significant risk to both human safety and environmental well-being. Sonar imaging is commonly employed to investigate such sites and aid in UXO remediation efforts. However, manually identifying and classifying potential targets in sonar data is challenging and time-consuming. Previous research has explored the use of machine learning models to recognize and categorize targets; however, many of these approaches lack transparency and fail to consider the underlying physical acoustics. Additionally, acquiring sufficient training data for these models can often be problematic. In this study, we present a novel approach by designing neural networks that explicitly account for the unique physics involved in the problem domain. UXOs examined using low-frequency sound frequently exhibit resonant behavior, where the sound is re-radiated after initial geometric scattering, owing to the elastic and compressional properties of the objects. Moreover, such resonant effects are typically absent in clutter objects, making them advantageous in discriminating UXOs from non-UXOs. Consequently, we propose several neural network architectures that leverage these resonant effects, utilizing 3D data obtained from a synthetic aperture sonar (SAS) imaging sonar. Our first proposal incorporates a recurrent neural network to model the physics-based correlation among adjacent time/spatial slices, originating from the resonant phenomena. For our second proposal, we employ intensity imagery of orthogonal projections of the 3D data cube, which capture shape-specific resonant scattering mechanisms unique to specific types of UXOs. To evaluate the effectiveness of our methods, we compare them against recent state-of-the-art algorithms using a real-world 3D SAS dataset. Remarkably, even when confronted with limited training data, our approaches consistently demonstrate superior results. Our findings highlight the significant potential of incorporating physical acoustics into neural network designs for UXO detection and classification, offering improved accuracy and efficiency in underwater remediation operations.
Trung Hoang, Kyle S. Dalton, Isaac Gerg, Thomas E. Blanford, Daniel C. Brown, Vishal Monga
IEEE Trans. Geosci. Remote. Sens.3
2022 Synthetic Aperture Sonar Image Segmentation Using Adaptive, Learned Beam Steering
abstract
Synthetic aperture sonar (SAS) produces high res-olution imagery of the seafloor. Automatic segmentation of this imagery is of interest to a variety of ocean-studying communities. In this work, we propose a novel supervised image segmentation method using adaptive, learned beam steering in the k-space domain in order to exploit the aspect dependent information available in the single-look-complex (SLC) SAS image which is overlooked in current SAS segmentation methods. Our proposal is a deep neural network architecture designed around the domain-specific beam steering geometry conveyed by the k-space domain. Results on a small real-world SAS dataset demonstrate that our proposal gives improved results for a variety of seafloor classes over a state-of-the-art (SOTA) U-Net.
Isaac Gerg, Vishal Monga
IGARSS1
2022 Domain Enriched Deep Networks for Munition Detection in Underwater 3D Sonar Imagery
abstract
Underwater sites impacted by unexploded ordnance (UXO) may pose an unacceptable risk to human and environmen-tal health. Sonar imaging is commonly used to interrogate such sites during UXO remediation, however manually identifying and classifying potential targets is difficult and time intensive. Previous work has explored training machine learning models to recognize and classify targets, however many of these “black-box” approaches fail to model the underlying physical acoustics and require abundant training data which is often hard to obtain. Specifically, UXOs interrogated with low frequency sound often exhibit resonant behavior which re-radiates the sound after the initial scattering due to elastic and compressional properties of the object. Such effects are usually not present in clutter objects, making them advantageous in discriminating UXO from non-UXO. In this work, we propose two neural networks which specifically model resonant scattering effects in order to find UXO from a 3D synthetic aperture sonar (SAS) imaging sonar. We do this by utilizing sequence models which are efficient at modeling the spatially correlated nature of the resonant scattering features. We compare our proposal to two recent state-of-the-art algorithms on a real-world 3D SAS dataset and show superior results even when limited training data is available.
Trung Hoang, Kyle S. Dalton, Isaac Gerg, Thomas E. Blanford, Daniel C. Brown, Vishal Monga
IGARSS3
2022 Structural Prior Driven Regularized Deep Learning for Sonar Image Classification
abstract
Deep learning has been recently shown to improve performance in the domain of synthetic aperture sonar (SAS) image classification. Given the constant resolution with a range of SAS, it is no surprise that deep learning techniques perform so well. Despite deep learning’s recent success, there are still compelling open challenges in reducing the high false alarm rate and enabling success when training imagery is limited, which is a practical challenge that distinguishes the SAS classification problem from standard image classification set-ups where training imagery may be abundant. We address these challenges by exploiting prior knowledge that humans use to grasp the scene. These include unconscious elimination of the image speckle and localization of objects in the scene. We introduce a new deep learning architecture that incorporates these priors with the goal of improving automatic target recognition (ATR) from SAS imagery. Our proposal—called SPDRDL, structural prior driven regularized deep learning—incorporates the previously mentioned priors in a multitask convolutional neural network (CNN) and requires no additional training data when compared to traditional SAS ATR methods. Two structural priors are enforced via regularization terms in the learning of the network: 1) structural similarity prior—enhanced imagery (often through despeckling) aids human interpretation and is semantically similar to the original imagery and 2) structural scene context priors—learned features ideally encapsulate target centering information; hence learning may be enhanced via a regularization that encourages fidelity against known ground truth target shifts (relative target position from scene center). Experiments on a challenging real-world data set reveal that SPDRDL outperforms state-of-the-art deep learning and other competing methods for SAS image classification.
Isaac Gerg, Vishal Monga
IEEE Trans. Geosci. Remote. Sens.1
2022 Iterative, Deep Synthetic Aperture Sonar Image Segmentation
abstract
Synthetic aperture sonar (SAS) systems produce high-resolution images of the seabed environment. Moreover, deep learning has demonstrated superior ability in finding robust features for automating imagery analysis. However, the success of deep learning is conditioned on having lots of labeled training data but obtaining generous pixel-level annotations of SAS imagery is often practically infeasible. This challenge has thus far limited the adoption of deep learning methods for SAS segmentation. Algorithms exist to segment SAS imagery in an unsupervised manner, but they lack the benefit of state-of-the-art learning methods and the results present significant room for improvement. In view of the above, we propose a new iterative algorithm for unsupervised SAS image segmentation combining superpixel formation, deep learning, and traditional clustering methods. We call our method iterative deep unsupervised segmentation (IDUS). IDUS is an unsupervised learning framework that can be divided into four main steps: 1) a deep network estimates class assignments; 2) low-level image features from the deep network are clustered into superpixels; 3) superpixels are clustered into class assignments (which we call pseudo-labels) using$k$-means; and 4) resulting pseudo-labels are used for loss backpropagation of the deep network prediction. These four steps are performed iteratively until convergence. A comparison of IDUS to current state-of-the-art methods on a realistic benchmark dataset for SAS image segmentation demonstrates the benefits of our proposal even as the IDUS incurs a much lower computational burden during inference (actual labeling of a test image). Because our design combines merits of classical superpixel methods with deep learning, practically we demonstrate a very significant benefit in terms of reduced selection bias, i.e., IDUS shows markedly improved robustness against the choice of training images. Finally, we also develop a semi-supervised (SS) extension of IDUS called Iterative Deep SS Segmentation (IDSS) and demonstrate experimentally that it can further enhance performance while outperforming supervised alternatives that exploit the same labeled training imagery.
Yung-Chen Sun, Isaac Gerg, Vishal Monga
IEEE Trans. Geosci. Remote. Sens.2
2021 Real-Time, Deep Synthetic Aperture Sonar (SAS) Autofocus
abstract
Synthetic aperture sonar (SAS) requires precise time-of-flight measurements of the transmitted/received waveform to produce well-focused imagery. It is not uncommon for errors in these measurements to be present resulting in image defocusing. To overcome this, an autofocus algorithm is employed as a post-processing step after image reconstruction to improve image focus. A particular class of these algorithms can be framed as a sharpness/contrast metric-based optimization. To improve convergence, a hand-crafted weighting function to remove “bad” areas of the image is sometimes applied to the image-under-test before the optimization procedure. Additionally, dozens of iterations are necessary for convergence which is a large compute burden for low size, weight, and power (SWaP) systems. We propose a deep learning technique to overcome these limitations and implicitly learn the weighting function in a data-driven manner. Our proposed method, which we call Deep Autofocus, uses features from the single-look-complex (SLC) to estimate the phase correction which is applied in k-space. Furthermore, we train our algorithm on batches of training imagery so that during deployment, only a single iteration of our method is sufficient to autofocus. We show results demonstrating the robustness of our technique by comparing our results to four commonly used image sharpness metrics. Our results demonstrate Deep Autofocus can produce imagery perceptually better than common iterative techniques but at a lower computational cost. We conclude that Deep Autofocus can provide a more favorable cost-quality tradeoff than alternatives with significant potential of future research.
Isaac Gerg, Vishal Monga
IGARSS1
2020 Data Adaptive Image Enhancement and Classification for Synthetic Aperture Sonar
abstract
Deep learning has been recently shown to improve performance in the domain of synthetic aperture sonar (SAS) image classification over existing shallow learning solutions. Given the constant resolution with range of a SAS, it is no surprise that deep learning techniques perform so well; the image of the seafloor produced by a SAS system is almost photographic in quality. Despite the image quality benefits of SAS, there is still room for classification improvement particularly in reducing the number of false alarms. This work addresses this by tackling one facet of the classification pipeline: image enhancement. Specifically, we ask and address the following question: Can we train a deep neural network to simultaneously enhance and classify a SAS image? We will respond in the affirmative as we introduce a new deep learning architecture tackling the problem, Data Adaptive Enhancement and Classification Network (DA-ECNet). DA-ECNet is a deep learning architecture which combines image enhancement as part of the classification procedure eliminating the need for a fixed state-of-the-art despeckling algorithm or enhancement module. Additionally, we train both image enhancement and classification jointly resulting in data adaptive image enhancement. Experiments on a challenging real-world dataset reveal that the proposed DA-ECNet outperforms state of the art deep learning as well as traditional feature based methods for SAS image classification.
Isaac Gerg, David P. Williams, Vishal Monga
IGARSS1
2018 Bridging The GAP: Simultaneous Fine Tuning for Data Re-Balancing
abstract
There are many real-world classification problems wherein the issue of data imbalance (the case when a data set contains substantially more samples for one/many classes than the rest) is unavoidable. While under-sampling the problematic classes is a common solution, this is not a compelling option when the large data class is itself diverse and/or the limited data class is especially small. We suggest a strategy based on recent work concerning limited data problems which utilizes a supplemental set of images with similar properties to the limited data class to aid in the training of a neural network. We show results for our model against other typical methods on a real-world synthetic aperture sonar data set. Code can be found at github.com/JohnMcKay/dataImbalance.
John McKay 0002, Isaac Gerg, Vishal Monga
IGARSS2