VLDB 2026 Research / reviewers in the wild / expert
Francis C. Motta
dblp:158/7912
· DBLP profile ↗
7ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-0364-5440ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorSecurity and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Algorithm for Persistent Homology Computation Using Homomorphic EncryptionabstractTopological Data Analysis (TDA) provides a suite of tools that extract shape-based features from high-dimensional data, with applications to modern statistical and machine learning (ML) models. Among these tools, persistent homology (PH) summarizes the topological structure of data in compact representations known as persistence diagrams (PDs). Due to their robustness to noise, interpretability, and compatibility with standard ML architectures, PDs are increasingly used in applications involving sensitive data, such as genomics, cancer research, sensor networks, and finance. Thus, there is a growing need to incorporate TDA methods into secure, end-to-end data analysis pipelines. We present the first adaptation of a fundamental TDA algorithm known as boundary matrix reduction to operate on encrypted data using homomorphic encryption (HE). We provide mathematical guarantees for the correctness of the HE-compatible algorithm under appropriate parameter choices and analyze its computational complexity. We support these theoretical results with two distinct empirical studies: (1) a plaintext simulation that explores the extent to which the theoretically sufficient parameters can be relaxed while still preserving correctness, and (2) a working implementation in the OpenFHE framework that validates correctness on encrypted data. This work lays the foundation for fully encrypted topological computations and opens new directions in privacy-preserving data analysis using TDA. Dominic Gold, Koray Karabina, Francis C. Motta |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2023 | Poster: Computing the Persistent Homology of Encrypted DataabstractTopological Data Analysis (TDA) offers a suite of computational tools that provide quantified shape features of high dimensional data that can be used by modern statistical and predictive machine learning (ML) models. Persistent homology (PH) transforms data (e.g., point clouds, images, time series) into persistence diagrams (PDs)--compact representations of its latent topological structures. Because PDs enjoy inherent noise tolerance, are interpretable, provide a solid basis for data analysis, and can be made compatible with the expansive set of well-established ML model architectures, PH has been widely adopted for model development including on sensitive data. Thus, TDA should be incorporated into secure end-to-end data analysis pipelines. This paper introduces a version of the fundamental algorithm to compute PH on encrypted data using homomorphic encryption (HE). Dominic Gold, Koray Karabina, Francis C. Motta |
CCS | 3 |
| 2022 | Conservation of dynamic characteristics of transcriptional regulatory elements in periodic biological processesabstractBACKGROUND: Cell and circadian cycles control a large fraction of cell and organismal physiology by regulating large periodic transcriptional programs that encompass anywhere from 15 to 80% of the genome despite performing distinct functions. In each case, these large periodic transcriptional programs are controlled by gene regulatory networks (GRNs), and it has been shown through genetics and chromosome mapping approaches in model systems that at the core of these GRNs are small sets of genes that drive the transcript dynamics of the GRNs. However, it is unlikely that we have identified all of these core genes, even in model organisms. Moreover, large periodic transcriptional programs controlling a variety of processes certainly exist in important non-model organisms where genetic approaches to identifying networks are expensive, time-consuming, or intractable. Ideally, the core network components could be identified using data-driven approaches on the transcriptome dynamics data already available. RESULTS: This study shows that a unified set of quantified dynamic features of high-throughput time series gene expression data are more prominent in the core transcriptional regulators of cell and circadian cycles than in their outputs, in multiple organism, even in the presence of external periodic stimuli. Additionally, we observe that the power to discriminate between core and non-core genes is largely insensitive to the particular choice of quantification of these features. CONCLUSIONS: There are practical applications of the approach presented in this study for network inference, since the result is a ranking of genes that is enriched for core regulatory elements driving a periodic phenotype. In this way, the method provides a prioritization of follow-up genetic experiments. Furthermore, these findings reveal something unexpected-that there are shared dynamic features of the transcript abundance of core components of unrelated GRNs that control disparate periodic phenotypes. Francis C. Motta, Robert C. Moseley, Breschine Cummins, Anastasia Deckard, Steven B. Haase |
BMC Bioinform. | 1 |
| 2022 | Experimental guidance for discovering genetic networks through hypothesis reduction on time seriesabstractLarge programs of dynamic gene expression, like cell cyles and circadian rhythms, are controlled by a relatively small "core" network of transcription factors and post-translational modifiers, working in concerted mutual regulation. Recent work suggests that system-independent, quantitative features of the dynamics of gene expression can be used to identify core regulators. We introduce an approach of iterative network hypothesis reduction from time-series data in which increasingly complex features of the dynamic expression of individual, pairs, and entire collections of genes are used to infer functional network models that can produce the observed transcriptional program. The culmination of our work is a computational pipeline, Iterative Network Hypothesis Reduction from Temporal Dynamics (Inherent dynamics pipeline), that provides a priority listing of targets for genetic perturbation to experimentally infer network structure. We demonstrate the capability of this integrated computational pipeline on synthetic and yeast cell-cycle data. Breschine Cummins, Francis C. Motta, Robert C. Moseley, Anastasia Deckard, Sophia Campione, Marcio Gameiro, Tomás Gedeon, Konstantin Mischaikow, Steven B. Haase |
PLoS Comput. Biol. | 2 |
| 2021 | Improved datasets and evaluation methods for the automatic prediction of DNA-binding proteinsabstractMOTIVATION: Accurate automatic annotation of protein function relies on both innovative models and robust datasets. Due to their importance in biological processes, the identification of DNA-binding proteins directly from protein sequence has been the focus of many studies. However, the datasets used to train and evaluate these methods have suffered from substantial flaws. We describe some of the weaknesses of the datasets used in previous DNA-binding protein literature and provide several new datasets addressing these problems. We suggest new evaluative benchmark tasks that more realistically assess real-world performance for protein annotation models. We propose a simple new model for the prediction of DNA-binding proteins and compare its performance on the improved datasets to two previously published models. In addition, we provide extensive tests showing how the best models predict across taxa. RESULTS: Our new gradient boosting model, which uses features derived from a published protein language model, outperforms the earlier models. Perhaps surprisingly, so does a baseline nearest neighbor model using BLAST percent identity. We evaluate the sensitivity of these models to perturbations of DNA-binding regions and control regions of protein sequences. The successful data-driven models learn to focus on DNA-binding regions. When predicting across taxa, the best models are highly accurate across species in the same kingdom and can provide some information when predicting across kingdoms. AVAILABILITY AND IMPLEMENTATION: The data and results for this article can be found at https://doi.org/10.5281/zenodo.5153906. The code for this article can be found at https://doi.org/10.5281/zenodo.5153683. The code, data and results can also be found at https://github.com/AZaitzeff/tools_for_dna_binding_proteins. Alexander Zaitzeff, Nick Leiby, Francis C. Motta, Steven B. Haase, Jed Singer |
Bioinform. | 3 |
| 2019 | Hyperparameter Optimization of Topological Features for Machine Learning ApplicationsabstractThis paper describes a general pipeline for generating optimal vector representations of topological features of data for use with machine learning algorithms. This pipeline can be viewed as a costly black-box function defined over a complex configuration space, each point of which specifies both how features are generated and how predictive models are trained on those features. We propose using state-of-the-art Bayesian optimization algorithms to inform the choice of topological vectorization hyperparameters while simultaneously choosing learning model parameters. We demonstrate the need for and effectiveness of this pipeline using two difficult biological learning problems, and illustrate the nontrivial interactions between topological feature generation and learning model hyperparameters. Francis C. Motta, John Harer, Nick Leiby, Franco Marinozzi, Scott Novotney, Gabe Rocklin, Jed Singer, Devin Strickland, Matthew W. Vaughn, Christopher J. Tralie, Rossella Bedini, Fabiano Bini, Gilberto Bini, Hamed Eramian, Marcio Gameiro, Steven B. Haase, Hugh Haddox |
ICMLA | 1 |
| 2017 | Persistence Images: A Stable Vector Representation of Persistent HomologyabstractMany data sets can be viewed as a noisy sampling of an underlying space, and tools from topological data analysis can characterize this structure for the purpose of knowledge discovery. One such tool is persistent homology, which provides a multiscale description of the homological features within a data set. A useful representation of this homological information is a persistence diagram (PD). Efforts have been made to map PDs into spaces with additional structure valuable to machine learning tasks. We convert a PD to a finite- dimensional vector representation which we call a persistence image (PI), and prove the stability of this transformation with respect to small perturbations in the inputs. The discriminatory power of PIs is compared against existing methods, showing significant performance gains. We explore the use of PIs with vector-based machine learning tools, such as linear sparse support vector machines, which identify features containing discriminating topological information. Finally, high accuracy inference of parameter values from the dynamic output of a discrete dynamical system (the linked twist map) and a partial differential equation (the anisotropic Kuramoto-Sivashinsky equation) provide a novel application of the discriminatory power of PIs. Henry Adams, Tegan Emerson, Michael Kirby, Rachel Neville, Chris Peterson 0001, Patrick D. Shipman, Sofya Chepushtanova, Eric M. Hanson, Francis C. Motta, Lori Ziegelmeier |
J. Mach. Learn. Res. | 9 |