EDBT 2026 Demo / reviewers in the wild / expert
Nicholas O. Malott
dblp:234/2886
· DBLP profile ↗
8ranked-venue papers in the field
5as first author
5since 2021 · last 2024
0000-0003-3848-4381ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 6 (3 first)Database Systems & Data Management · 1 (1 first)Data Mining & Knowledge Discovery · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Piecewise Computation of Persistent HomologyabstractPersistent Homology (PH) is a widely used tool of Topological Data Analysis (TDA) that measures the persistence of homological features present in data. PH is computed on a sequence of nested sub-complexes that form a filtration ${{\mathcal{K}}_{\mathcal{F}}}$ of data. In general, each nested sub-complex is defined by a scale parameter ϵ that defines the connectivity distances used to for its construction. These sub-complexes are arranged by increasing distances, ϵ0to ϵ∞, in the filtration; PH is then computed from ${{\mathcal{K}}_{\mathcal{F}}}$. However, due to its exponential space and time complexity, computing PH on big data is beyond the capabilities of contemporary machines. This paper explores the Piecewise computation of PH (PwPH) using subset constructions of the filtrations of data. PwPH is a framework for solutions that compute PH of data. In some embodiments, PwPH can be support computations of PH can be assembled into a complete picture of the homologies; in others, PwPH can only compute an approximation of the PH. The general solution of PwPH can be organized to use significantly less memory and provides a partition of computational PH elements that is, in some embodiments, embarrassingly parallel. This paper explores two complementary foundations to organize filtrations for PwPH. Rohit P. Singh, Nicholas O. Malott, Philip A. Wilsey |
IEEE Big Data | 2 |
| 2023 | Scalable Homology Classification through Decomposed Euler Characteristic CurvesabstractTopological Data Analysis (TDA) has demonstrated notable success in data mining by measuring the presence of topological structure in data. Persistent Homology (PH), one popular tool of TDA, examines the nested sequence of graphs formed over a data filtration to characterize unique homology classes. PH suffers from exponential complexity, limiting the approach to relatively small data sets and lower homology class applications. The Euler Characteristic Curve (ECC), a metric closely related to PH, can be computed more efficiently and, in some cases, can be used as a direct replacement to the applications of PH. A recent technique has been introduced to separate the ECC into dimensional components representing the homology classes identified by PH. This study examines the dimensional ECC, provides an improved algorithm, and introduces interpretation of the results as proximity-series representations of topological features; exploration of how to compare and classify ECC curves is detailed and shown in the context of MRA Brain Artery scan classification. Experimental results with ECC demonstrate the effectiveness and scalability to big data sets beyond that of current persistent homology applications. Nicholas O. Malott, Philip A. Wilsey |
IEEE Big Data | 1 |
| 2023 | A Survey on the High-Performance Computation of Persistent HomologyabstractPersistent Homology is a computational method of data mining in the field of Topological Data Analysis. Large-scale data analysis with persistent homology is computationally expensive and memory intensive. The performance of persistent homology has been rigorously studied to optimize data encoding and intermediate data structures for high-performance computation. This paper provides an application-centric survey of the High-Performance Computation of Persistent Homology. Computational topology concepts are reviewed and detailed for a broad data science and engineering audience. Nicholas O. Malott, Shangye Chen, Philip A. Wilsey |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Homology-Separating Triangulated Euler Characteristic CurveabstractTopological Data Analysis (TDA) utilizes concepts from topology to analyze data. In general, TDA considers objects similar based on a topological invariant. Topological invariants are properties of the topological space that are homeomorphic; resilient to deformation in the space. The Euler-Poincaré Characteristic is a classic topological invariant that represents the alternating sum of the vertices, edges, faces, and higherorder cells of a closed surface. Tracking the Euler characteristic over a topological filtration produces an Euler Characteristic Curve (ECC). This study introduces a computational technique to determine the ECC of $\mathbb{R}^{2}$ or $\mathbb{R}^{3}$ data; the technique generalizes to higher dimensions. This technique separates landscapes of lowerorder homologies utilizing triangulations of the space. Nicholas O. Malott, Robert R. Lewis, Philip A. Wilsey |
ICDM | 1 |
| 2021 | Data Reduction and Feature Isolation for Computing Persistent Homology on High Dimensional DataabstractPersistent Homology (PH) is computationally expensive and is thus generally employed with strict limits on the (i) maximum connectivity distance and (ii) dimensions of homology groups to compute (unless working with trivially small data sets). As a result, most studies with PH only work with H0and H1homology groups. This paper examines the identification and isolation of regions of data sets where high dimensional topological features are suspected to be located. These regions are analyzed with PH to characterize the high dimensional homology groups contained in that region. Since only the region around a suspected topological feature is analyzed, it is possible to identify high dimension homologies piecewise and then assemble the results into a scalable characterization of the original data set. Rishi R. Verma, Nicholas O. Malott, Philip A. Wilsey |
IEEE BigData | 2 |
| 2020 | Topology Preserving Data Reduction for Computing Persistent HomologyabstractAn emerging method for data analysis is called Topological Data Analysis (TDA). TDA is based in the mathematical field of topology and examines the properties of spaces under continuous deformation. One of the key tools used for TDA is called persistent homology which considers the connectivity of points in a d-dimensional point cloud at different spatial resolutions to identify topological properties (holes, loops, and voids) in the space. Persistent homology then classifies the topological features by their persistence through the range of spatial connectivity. Unfortunately the memory and run-time complexity of computing persistent homology is exponential and current tools can only process a few thousand points in $\mathbb{R}^{3}$. Fortunately, the use of data reduction techniques enables persistent homology to be applied to much larger point clouds. Techniques to reduce the data range from random sampling of points to clustering the data and using the cluster centroids as the reduced data. While several data reduction approaches appear to preserve the large topological features present in the original point cloud, no systematic study comparing the efficacy of different data clustering techniques in preserving the persistent homology results has been performed. This paper explores the question of topology preserving data reductions and describes formally when and how topological features can be mischaracterized or lost by data reduction techniques. The paper also performs an experimental assessment of data reduction techniques and resilient effects on the persistent homology. In particular, data reduction by random selection is compared to cluster centroids extracted from different data clustering algorithms. Nicholas O. Malott, Aaron M. Sens, Philip A. Wilsey |
IEEE BigData | 1 |
| 2019 | Fast Computation of Persistent Homology with Data Reduction and Data PartitioningabstractPersistent homology is a method of data analysis that is based in the mathematical field of topology. Unfortunately, the run-time and memory complexities associated with computing persistent homology inhibit general use for the analysis of big data. For example, the best tools currently available to compute persistent homology can process only a few thousand data points in ℝ3. Several studies have proposed using sampling or data reduction methods to attack this limit. While these approaches enable the computation of persistent homology on much larger data sets, the methods are approximate. Furthermore, while they largely preserve the results of large topological features, they generally miss reporting information about the small topological features that are present in the data set. While this abstraction is useful in many cases, there are data analysis needs where the smaller features are also significant (e.g., brain artery analysis). This paper explores a combination of data reduction and data partitioning to compute persistent homology on big data that enables the identification of both large and small topological features from the input data set. To reduce the approximation errors that typically accompany data reduction for persistent homology, the described method also includes a mechanism of “upscaling” the data circumscribing the large topological features that are computed from the sampled data. The designed experimental method provides significant results for improving the scale at which persistent homology can be performed. Nicholas O. Malott, Philip A. Wilsey |
IEEE BigData | 1 |
| 2018 | Cluster-based Data Reduction for Persistent HomologyabstractPersistent homology is used for computing topological features of a space at different spatial resolutions. It is one of the main tools from computational topology that is applied to the problems of data analysis. Despite several attempts to reduce its complexity, persistent homology remains expensive in both time and space. These limits are such that the largest data sets to which the method can be applied have the number of points of the order of thousands in ℝ3. This paper explores a technique intended to reduce the number of data points while preserving the salient topological features of the data. The proposed technique enables the computation of persistent homology on a reduced version of the original input data without affecting significant components of the output. Since the run time of persistent homology is exponential in the number of data points, the proposed data reduction method facilitates the computation in a fraction of the time required for the original data. Moreover, the data reduction method can be combined with any existing technique that simplifies the computation of persistent homology. The data reduction is performed by creating small groups of similar data points, called nano-clusters, and then replacing the points within each nano-cluster with its cluster center. The persistence homology of the reduced data differs from that of the original data by an amount bounded by the radius of the nano-clusters. The theoretical analysis is backed by experimental results showing that persistent homology is preserved by the proposed data reduction technique. Anindya Moitra, Nicholas O. Malott, Philip A. Wilsey |
IEEE BigData | 2 |