Kevin R. Moon

dblp:135/7183 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0002-4457-9988ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Improved Gas Plume Identification Using Nearest Neighbor Methods for Background Estimation
abstract
Longwave infrared (LWIR) hyperspectral imaging (HSI) can be used for many tasks in remote sensing, including detecting and identifying effluent gases by LWIR sensors on airborne platforms. Identification is used after detection to increase confidence in weakly detected plumes, reduce false positives from detection, and distinguish between similar and confounding material signatures. Background estimation is an important step used to reveal the unique spectral characteristics of the detected gas, allowing the identification model to determine what the gas is specifically. The importance of proper background estimation increases when dealing with weak signals, large libraries of gases of interest, and uncommon or heterogeneous backgrounds. In this article, we propose two methods for background estimation: a novel k-nearest segments (KNS) algorithm and the standard k-nearest neighbors (KNN) algorithm. We test our methods and three existing background estimation methods for comparison against global background estimation to determine which performs best at estimating the true background radiance under a plume and for increasing identification confidence using a neural network classification model. We compare the different methods using 640 simulated weak plumes in an urban environment. For identification, our KNS algorithm improves median neural network identification confidence by 53.2%. For background radiance estimation, the KNN algorithm provides a median of 49 times less RMSE than global background estimation. Furthermore, KNN is the easiest method to tune for different plumes, making it an excellent “out of the box” background estimator.
Scout Jarman, Zigfried Hampel-Arias, Adra Carr, Kevin R. Moon
IEEE Trans. Geosci. Remote. Sens.4
2024 Advancing Tabular Data Classification with Graph Neural Networks: A Random Forest Proximity Method
abstract
Graphs are essential for modeling complex relationships, analyzing networks, and offering versatile representations that capture diverse data structures. Graph Neural Networks (GNNs) excel in processing graph-structured data by leveraging the relational information encoded in graph topology. However, not all types of data possess an explicit graph structure, particularly tabular data, which is ubiquitous in the real world. To enable the use of GNNs for tabular data, it is necessary to convert tabular data into graph-structured data. Existing methods for this conversion often lack a generic, straightforward approach with unrestrictive assumptions that can directly apply GNNs to tabular data for downstream tasks. In this paper, we introduce RF-GNN, a novel method that enhances traditional machine learning approaches by transforming tabular data into graph structures and leveraging GNNs. Our approach calculates the similarity between pairs of samples based on Random Forest (RF) proximities, which measure how often the same pair appears in the same terminal nodes of a tree in a Random Forest. This enables the creation of an adjacency matrix for instances of tabular data, allowing the application of GNNs. Extensive experiments on 36 different datasets demonstrate that RF-GNN consistently outperforms traditional machine learning models and recent methods in terms of weighted F1-score. We conduct additional experiments to evaluate the effectiveness of RF-GNN components and settings. The code is available in https://github.com/DSAatUSU/RF-GNN.
Soheila Farokhi, Kevin R. Moon, Hamid Karimi
IEEE Big Data3
2024 Nonparametric Estimation of Non-Smooth Divergences
abstract
Nonparametric estimation of information divergence functionals between two probability densities is an important problem in machine learning. Several estimators exist that guarantee the parametric rate of mean squared error (MSE) of O(1/N) under various assumptions on the smoothness and boundary of the underlying densities, with N being the number of samples. In particular, previous work on ensemble estimation theory derived ensemble estimators of divergence functionals that achieve the parametric rate without requiring knowledge of the densities' support set and are simple to implement. However, these and most other methods all assume some level of differentiability of the divergence functional. This excludes important divergence functionals such as the total variation distance and the Bayes error rate. Here, we show empirically that the ensemble estimation approach for smooth functionals can be applied to less smooth functionals and obtain good convergence rates, suggesting a gap in current theory.
M. Mahbub Hossain, Alan Wisler, Kevin R. Moon
CIKM3
2024 Neural Network Ensembling with Random Features
abstract
This paper presents a novel approach to classification for tabular data. While neural networks excel at capturing nonlinear relationships within features, they often lack the robustness of ensemble methods. To address this limitation, we propose a model that combines the nonlinear learning capabilities of neural networks with the resilience of ensemble methods. Previous research on ensemble neural networks has shown promising results, although the lack of independence among models within the ensemble has been a challenge. Our approach aims to overcome this limitation by introducing a set of independently selected subsets of features to feed into a single neural network and ensembling the results. Experimental results demonstrate the effectiveness of this approach compared to baseline models on diverse datasets.
Jarrod Mau, Kevin R. Moon
ICMLA2
2024 Local Background Estimation for Improved Gas Plume Identification in Hyperspectral Images
abstract
Deep learning identification models have shown promise for identifying gas plumes in Longwave IR hyperspectral images of urban scenes, particularly when a large library of gases are being considered. Because many gases have similar spectral signatures, it is important to properly estimate the signal from a detected plume. Typically, a scene’s global mean spectrum and covariance matrix are estimated to whiten the plume’s signal, which removes the background’s signature from the gas signature. However, urban scenes can have many different background materials that are spatially and spectrally heterogeneous. This can lead to poor identification performance when the global background estimate is not representative of a given local background material. We use image segmentation, along with an iterative background estimation algorithm, to create local estimates for the various background materials that reside underneath a gas plume. Our method outperforms global background estimation on a set of simulated and real gas plumes. This method shows promise in increasing deep learning identification confidence, while being simple and easy to tune when considering diverse plumes.
Scout Jarman, Zigfried Hampel-Arias, Adra Carr, Kevin R. Moon
IGARSS4
2024 Symmetry Discovery Beyond Affine Transformations
abstract
Symmetry detection has been shown to improve various machine learning tasks. In the context of continuous symmetry detection, current state of the art experiments are limited to the detection of affine transformations. Under the manifold assumption, we outline a framework for discovering continuous symmetry in data beyond the affine transformation group. We also provide a similar framework for discovering discrete symmetry. We experimentally compare our method to an existing method known as LieGAN and show that our method is competitive at detecting affine symmetries for large sample sizes and superior than LieGAN for small sample sizes. We also show our method is able to detect continuous symmetries beyond the affine group and is generally more computationally efficient than LieGAN.
Ben Shaw 0003, Abram Magner, Kevin R. Moon
NeurIPS3
2023 Diffusion Transport Alignment
Andrés F. Duque, Guy Wolf, Kevin R. Moon
IDA3
2023 Geometry Regularized Autoencoders
abstract
A fundamental task in data exploration is to extract low dimensional representations that capture intrinsic geometry in data, especially for faithfully visualizing data in two or three dimensions. Common approaches use kernel methods for manifold learning. However, these methods typically only provide an embedding of the input data and cannot extend naturally to new data points. Autoencoders have also become popular for representation learning. While they naturally compute feature extractors that are extendable to new data and invertible (i.e., reconstructing original features from latent representation), they often fail at representing the intrinsic data geometry compared to kernel-based manifold learning. We present a new method for integrating both approaches by incorporating a geometric regularization term in the bottleneck of the autoencoder. This regularization encourages the learned latent representation to follow the intrinsic data geometry, similar to manifold learning algorithms, while still enabling faithful extension to new data and preserving invertibility. We compare our approach to autoencoder models for manifold learning to provide qualitative and quantitative evidence of our advantages in preserving intrinsic structure, out of sample extension, and reconstruction. Our method is easily implemented for big-data applications, whereas other methods are limited in this regard.
Andrés F. Duque, Sacha Morin, Guy Wolf, Kevin R. Moon
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Geometry- and Accuracy-Preserving Random Forest Proximities
abstract
Random forests are considered one of the best out-of-the-box classification and regression algorithms due to their high level of predictive performance with relatively little tuning. Pairwise proximities can be computed from a trained random forest and measure the similarity between data points relative to the supervised task. Random forest proximities have been used in many applications including the identification of variable importance, data imputation, outlier detection, and data visualization. However, existing definitions of random forest proximities do not accurately reflect the data geometry learned by the random forest. In this paper, we introduce a novel definition of random forest proximities called Random Forest-Geometry- and Accuracy-Preserving proximities (RF-GAP). We prove that the proximity-weighted sum (regression) or majority vote (classification) using RF-GAP exactly matches the out-of-bag random forest prediction, thus capturing the data geometry learned by the random forest. We empirically show that this improved geometric representation outperforms traditional random forest proximities in tasks such as data imputation and provides outlier detection and visualization results consistent with the learned data geometry.
Jake S. Rhodes, Adele Cutler, Kevin R. Moon
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Gps-Denied Navigation Using Sar Images And Neural Networks
abstract
Unmanned aerial vehicles (UAV) often rely on GPS for navigation. GPS signals, however, are very low in power and easily jammed or otherwise disrupted. This paper presents a method for determining the navigation errors present at the beginning of a GPS-denied period utilizing data from a synthetic aperture radar (SAR) system. This is accomplished by comparing an online-generated SAR image with a reference image obtained a priori. The distortions relative to the reference image are learned and exploited with a convolutional neural network to recover the initial navigational errors, which can be used to recover the true flight trajectory throughout the synthetic aperture. The proposed neural network approach is able to learn to predict the initial errors on both simulated and real SAR image data.
Teresa White, Jesse Wheeler, Colton Lindstrom, Randall Christensen, Kevin R. Moon
ICASSP5
2021 Ensemble Estimation of Generalized Mutual Information With Applications to Genomics
abstract
Mutual information is a measure of the dependence between random variables that has been used successfully in myriad applications in many fields. Generalized mutual information measures that go beyond classical Shannon mutual information have also received much interest in these applications. We derive the mean squared error convergence rates of kernel density-based plug-in estimators of general mutual information measures between two multidimensional random variablesXandYfor two cases: 1)XandYare continuous; 2)XandYmay have a mixture of discrete and continuous components. Using the derived rates, we propose an ensemble estimator of these information measures called GENIE by taking a weighted sum of the plug-in estimators with varied bandwidths. The resulting ensemble estimators achieve the 1/N parametric mean squared error convergence rate when the conditional densities of the continuous variables are sufficiently smooth. To the best of our knowledge, this is the first nonparametric mutual information estimator known to achieve the parametric convergence rate for the mixture case, which frequently arises in applications (e.g. variable selection in classification). The estimator is simple to implement and it uses the solution to an offline convex optimization problem and simple plug-in estimators. A central limit theorem is also derived for the ensemble estimators and minimax rates are derived for the continuous case. We demonstrate the ensemble estimator for the mixed case on simulated data and apply the proposed estimator to analyze gene relationships in single cell data.
Kevin R. Moon, Kumar Sricharan, Alfred O. Hero III
IEEE Trans. Inf. Theory1
2020 Extendable and invertible manifold learning with geometry regularized autoencoders
abstract
A fundamental task in data exploration is to extract simplified low dimensional representations that capture intrinsic geometry in data, especially for faithfully visualizing data in two or three dimensions. Common approaches to this task use kernel methods for manifold learning. However, these methods typically only provide an embedding of fixed input data and cannot extend to new data points. Autoencoders have also recently become popular for representation learning. But while they naturally compute feature extractors that are both extendable to new data and invertible (i.e., reconstructing original features from latent representation), they have limited capabilities to follow global intrinsic geometry compared to kernel-based manifold learning. We present a new method for integrating both approaches by incorporating a geometric regularization term in the bottleneck of the autoencoder. Our regularization, based on the diffusion potential distances from the recently-proposed PHATE visualization method, encourages the learned latent representation to follow intrinsic data geometry, similar to manifold learning algorithms, while still enabling faithful extension to new data and reconstruction of data in the original feature space from latent coordinates. We compare our approach with leading kernel methods and autoencoder models for manifold learning to provide qualitative and quantitative evidence of our advantages in preserving intrinsic structure, out of sample extension, and reconstruction. Our method is easily implemented for big-data applications, whereas other methods are limited in this regard.
Andrés F. Duque, Sacha Morin, Guy Wolf, Kevin R. Moon
IEEE BigData4
2019 Coarse Graining of Data via Inhomogeneous Diffusion Condensation
abstract
Big data often has emergent structure that exists at multiple levels of abstraction, which are useful for characterizing complex interactions and dynamics of the observations. Here, we consider multiple levels of abstraction via a multiresolution geometry of data points at different granularities. To construct this geometry we define a time-inhomogemeous diffusion process that effectively condenses data points together to uncover nested groupings at larger and larger granularities. This inhomogeneous process creates a deep cascade of intrinsic low pass filters on the data affinity graph that are applied in sequence to gradually eliminate local variability while adjusting the learned data geometry to increasingly coarser resolutions. We provide visualizations to exhibit our method as a "continuously-hierarchical" clustering with directions of eliminated variation highlighted at each step. The utility of our algorithm is demonstrated via neuronal data condensation, where the constructed multiresolution data geometry uncovers the organization, grouping, and connectivity between neurons.
Nathan Brugnone, Smita Krishnaswamy, Alex Gonopolskiy, Mark W. Moyle, Manik Kuchroo, David van Dijk, Kevin R. Moon, Daniel Colón-Ramos, Guy Wolf, Matthew J. Hirn
IEEE BigData7
2018 Direct Ensemble Estimation of Density Functionals
abstract
Estimating density functionals of analog sources is an important problem in statistical signal processing and information theory. Traditionally, estimating these quantities requires either making parametric assumptions about the underlying distributions or using non-parametric density estimation followed by integration. In this paper we introduce a direct nonparametric approach which bypasses the need for density estimation by using the error rates of k-NN classifiers as “data-driven” basis functions that can be combined to estimate a range of density functionals. However, this method is subject to a non-trivial bias that dramatically slows the rate of convergence in higher dimensions. To overcome this limitation, we develop an ensemble method for estimating the value of the basis function which, under some minor constraints on the smoothness of the underlying distributions, achieves the parametric rate of convergence regardless of data dimension.
Alan Wisler, Kevin R. Moon, Visar Berisha
ICASSP2
2017 Information theoretic structure learning with confidence
abstract
Information theoretic measures (e.g. the Kullback Liebler divergence and Shannon mutual information) have been used for exploring possibly nonlinear multivariate dependencies in high dimension. If these dependencies are assumed to follow a Markov factor graph model, this exploration process is called structure discovery. For discrete-valued samples, estimates of the information divergence over the parametric class of multinomial models lead to structure discovery methods whose mean squared error achieves parametric convergence rates as the sample size grows. However, a naive application of this method to continuous nonparametric multivariate models converges much more slowly. In this paper we introduce a new method for nonparametric structure discovery that uses weighted ensemble divergence estimators that achieve parametric convergence rates and obey an asymptotic central limit theorem that facilitates hypothesis testing and other types of statistical validation.
Kevin R. Moon, Morteza Noshad, Salimeh Yasaei Sekeh, Alfred O. Hero III
ICASSP1
2017 Ensemble estimation of mutual information
abstract
We derive the mean squared error convergence rates of kernel density-based plug-in estimators of mutual information measures between two multidimensional random variables X and Y for two cases: 1) X and Y are both continuous; 2) X is continuous and Y is discrete. Using the derived rates, we propose an ensemble estimator of these information measures for the second case by taking a weighted sum of the plug-in estimators with varied bandwidths. The resulting ensemble estimator achieves the 1 /N parametric convergence rate when the conditional densities of the continuous variables are sufficiently smooth. To the best of our knowledge, this is the first nonparametric mutual information estimator known to achieve the parametric convergence rate for this case, which frequently arises in applications (e.g. variable selection in classification). The estimator is simple to implement as it uses the solution to an offline convex optimization problem and simple plug-in estimators. Ensemble estimators that achieve the parametric rate are also derived for the first case (X and Y are both continuous) and another case: 3) X and Y may have any mixture of discrete and continuous components.
Kevin R. Moon, Kumar Sricharan, Alfred O. Hero III
ISIT1
2017 Direct estimation of information divergence using nearest neighbor ratios
abstract
We propose a direct estimation method for Rényi and f-divergence measures based on a new graph theoretical interpretation. Suppose that we are given two sample sets X and Y, respectively with N and M samples, where η := M/N is a constant value. Considering the k-nearest neighbor (k-NN) graph of Y in the joint data set (X, Y), we show that the average powered ratio of the number of X points to the number of Y points among all k-NN points is proportional to Rényi divergence of X and Y densities. A similar method can also be used to estimate f-divergence measures. We derive bias and variance rates, and show that for the class of γ-Hölder smooth functions, the estimator achieves the MSE rate of O(N-2γ/(γ+d)). Furthermore, by using a weighted ensemble estimation technique, for density functions with continuous and bounded derivatives of up to the order d, and some extra conditions at the support set boundary, we derive an ensemble estimator that achieves the parametric MSE rate of O(1/N). Our estimator requires no boundary correction, and remarkably, the boundary issues do not show up. Our approach is also more computationally tractable than other competing estimators, which makes them appealing in many practical applications.
Morteza Noshad, Kevin R. Moon, Salimeh Yasaei Sekeh, Alfred O. Hero III
ISIT2
2016 The intrinsic value of HFO features as a biomarker of epileptic activity
abstract
High frequency oscillations (HFOs) are a promising biomarker of epileptic brain tissue and activity. HFOs additionally serve as a prototypical example of challenges in the analysis of discrete events in high-temporal resolution, intracranial EEG data. Two primary challenges are 1) dimensionality reduction, and 2) assessing feasibility of classification. Dimensionality reduction assumes that the data lie on a manifold with dimension less than that of the features space. However, previous HFO analysis have assumed a linear manifold, global across time, space (i.e. recording electrode/channel), and individual patients. Instead, we assess both a) whether linear methods are appropriate and b) the consistency of the manifold across time, space, and patients. We also estimate bounds on the Bayes classification error to quantify the distinction between two classes of HFOs (those occurring during seizures and those occurring due to other processes). This analysis provides the foundation for future clinical use of HFO features and guides the analysis for other discrete events, such as individual action potentials or multi-unit activity.
Stephen V. Gliske, William C. Stacey, Kevin R. Moon, Alfred O. Hero III
ICASSP3
2016 Improving convergence of divergence functional ensemble estimators
abstract
Recent work has focused on the problem of non-parametric estimation of divergence functionals. Many existing approaches are restrictive in their assumptions on the density support or require difficult calculations at the support boundary which must be known a priori. We derive the MSE convergence rate of a leave-one-out kernel density plug-in divergence functional estimator for general bounded density support sets where knowledge of the support boundary is not required. We generalize the theory of optimally weighted ensemble estimation to derive two estimators that achieve the parametric rate when the densities are sufficiently smooth. The asymptotic distribution of these estimators and tuning parameter selection guidelines are provided. Based on the theory, we propose an empirical estimator of Rényi-α divergence that outperforms the standard kernel density plug-in estimator, especially in higher dimensions.
Kevin R. Moon, Kumar Sricharan, Kristjan Greenewald, Alfred O. Hero III
ISIT1
2014 Image patch analysis and clustering of sunspots: A dimensionality reduction approach
abstract
Sunspots, as seen in white light or continuum images, are associated with regions of high magnetic activity on the Sun, visible on magnetogram images. Their complexity is correlated with explosive solar activity and so classifying these active regions is useful for predicting future solar activity. Current classification of sunspot groups is visually based and suffers from bias. Supervised learning methods can reduce human bias but fail to optimally capitalize on the information present in sunspot images. This paper uses two image modalities (continuum and magnetogram) to characterize the spatial and modal interactions of sunspot and magnetic active region images and presents a new approach to cluster the images. Specifically, in the framework of image patch analysis, we estimate the number of intrinsic parameters required to describe the spatial and modal dependencies, the correlation between the two modalities and the corresponding spatial patterns, and examine the phenomena at different scales within the images. To do this, we use linear and nonlinear intrinsic dimension estimators, canonical correlation analysis, and multiresolution analysis of intrinsic dimension.
Kevin R. Moon, Jimmy J. Li, Véronique Delouille, Fraser Watson, Alfred O. Hero III
ICIP1
2014 Ensemble estimation of multivariate f-divergence
abstract
f-divergence estimation is an important problem in the fields of information theory, machine learning, and statistics. While several divergence estimators exist, relatively few of their convergence rates are known. We derive the MSE convergence rate for a density plug-in estimator of f-divergence. Then by applying the theory of optimally weighted ensemble estimation, we derive a divergence estimator with a convergence rate of O (1 over T) that is simple to implement and performs well in high dimensions. We validate our theoretical results with experiments.
Kevin R. Moon, Alfred O. Hero III
ISIT1
2014 Multivariate f-divergence Estimation With Confidence
Kevin R. Moon, Alfred O. Hero III
NIPS1
2013 Considerations for Ku-Band Scatterometer Calibration Using the Dry-Snow Zone of the Greenland Ice Sheet
abstract
Postlaunch calibration of satellite-borne scatterometers using backscatter data from natural land targets helps to maintain scatterometer accuracy. Due to its temporal stability, the dry-snow zone of the Greenland ice sheet has been proposed in previous studies as a calibration target. Using QuikSCAT data, this letter examines the backscatter properties of the dry-snow zone that are relevant to Ku-band scatterometer calibration including temporal and spatial variabilities, and azimuth-angle and polarization and incidence-angle dependences. The backscatter is found to seasonally vary throughout the dry-snow zone by 0.4 dB on average. Small interannual variations (less than 1.5 dB over a nine year period) in the backscatter are also present in some regions. Azimuth modulation is generally less than 0.4 dB in magnitude and is not significant in some regions of the dry-snow zone. Melting and refreezing appear to cause the quasi-polarization ratio to temporarily decrease. Spatially consistent and relatively temporally stable regions that are well within the interior of the dry-snow zone are best suited as calibration sites.
Kevin R. Moon, David G. Long
IEEE Geosci. Remote. Sens. Lett.1