Elana J. Fertig

dblp:47/8711 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0003-3204-342XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Differential cell signaling testing for cell-cell communication inference from single-cell data by dominoSignal
abstract
MOTIVATION: Algorithms for ligand-receptor network inference have emerged as commonly used tools to estimate cell-cell communication from reference single-cell data. Many studies employ these algorithms to compare signaling between conditions and lack methods to statistically identify signals that are significantly different. We previously developed the cell communication inference algorithm Domino, which considers ligand and receptor gene expression in association with downstream transcription factor activity scoring. We developed the dominoSignal software to innovate upon Domino and extend its functionality to test statistically differential cellular signaling. RESULTS: This new functionality includes the compilation of active signals as linkages from multiple subjects in a single-cell data set and testing condition-dependent signaling linkage. The software is applicable for analysis of single-cell data sets with multiple subjects as biological replicates as well as with bootstrapped replicates from data sets with few or pooled subjects. We use simulation studies to benchmark the number of subjects in compared groups and cells within an annotated cell type sufficient to accurately identify differential linkages. We demonstrate the application of the Differential Cell Signaling Test (DCST) in the dominoSignal software to investigate consequences of cancer cell phenotypes and immunotherapy on cell-cell communication in tumor microenvironments. These applications in cancer studies demonstrate the ability of differential cell signaling analysis to infer changes to cell communication networks from therapeutic or experimental perturbations, which is broadly applicable across biological systems. AVAILABILITY: dominoSignal is available through Bioconductor at https://www.bioconductor.org/packages/release/bioc/html/dominoSignal.html.
Jacob T. Mitchell, Orian Stapleton, Kavita Krishnan, Sushma Nagaraj, Dmitrijs Lvovs, Christopher Cherry, Amanda Poissonnier, Wesley Horton, Andrew Adey, Varun Rao, Amanda Huff, Jacquelyn W. Zimmerman, Luciane T. Kagohara, Neeha Zaidi, Lisa M. Coussens, Elizabeth M. Jaffee, Jennifer H. Elisseeff, Elana J. Fertig
Bioinform.18
2025 BIWT: a bioinformatics walkthrough for embedding spatial multiomics in agent-based models for virtual cells
abstract
SUMMARY: Whereas transcriptomic and spatial profiling offer static snapshots of tissue structure, mechanistic models use biological rules to predict how tissues evolve. We present the BioInformatics WalkThrough (BIWT) software to directly initialize spatial agent-based models from single-cell and spatial molecular data. We demonstrate how initialization strategies affect tumor-immune dynamics and spatial clustering, positioning BIWT as a software suite to generate data-driven virtual cells representing both experimental and clinical contexts. AVAILABILITY AND IMPLEMENTATION: The BIWT software is available at https://github.com/PhysiCell-Tools/PhysiCell-Studio. The sample dataset for running the BIWT is available at https://zenodo.org/records/16365625. The code and instructions for reproducing the use case example is available at https://github.com/drbergman/BIWT-Paper.
Daniel R. Bergman, Jeanette A. I. Johnson, Marwa Naji, Max Booth, Heber L. Rocha, Atul Deshpande, Dimitrios N. Sidiropoulos, Tamara Lopez-Vidal, Randy W. Heiland, Luciane T. Kagohara, Robert A. Anders, Elizabeth M. Jaffee, Genevieve L. Stein-O'Brien, Paul Macklin, Elana J. Fertig
Bioinform.16
2024 Leveraging multi-omics data to empower quantitative systems pharmacology in immuno-oncology
abstract
Understanding the intricate interactions of cancer cells with the tumor microenvironment (TME) is a pre-requisite for the optimization of immunotherapy. Mechanistic models such as quantitative systems pharmacology (QSP) provide insights into the TME dynamics and predict the efficacy of immunotherapy in virtual patient populations/digital twins but require vast amounts of multimodal data for parameterization. Large-scale datasets characterizing the TME are available due to recent advances in bioinformatics for multi-omics data. Here, we discuss the perspectives of leveraging omics-derived bioinformatics estimates to inform QSP models and circumvent the challenges of model calibration and validation in immuno-oncology.
Theinmozhi Arulraj, Alberto Ippolito, Elana J. Fertig, Aleksander S. Popel
Briefings Bioinform.5
2024 Computational methods and biomarker discovery strategies for spatial proteomics: a review in immuno-oncology
abstract
Advancements in imaging technologies have revolutionized our ability to deeply profile pathological tissue architectures, generating large volumes of imaging data with unparalleled spatial resolution. This type of data collection, namely, spatial proteomics, offers invaluable insights into various human diseases. Simultaneously, computational algorithms have evolved to manage the increasing dimensionality of spatial proteomics inherent in this progress. Numerous imaging-based computational frameworks, such as computational pathology, have been proposed for research and clinical applications. However, the development of these fields demands diverse domain expertise, creating barriers to their integration and further application. This review seeks to bridge this divide by presenting a comprehensive guideline. We consolidate prevailing computational methods and outline a roadmap from image processing to data-driven, statistics-informed biomarker discovery. Additionally, we explore future perspectives as the field moves toward interfacing with other quantitative domains, holding significant promise for precision care in immuno-oncology.
Haoyang Mi, Shamilene Sivagnanam, Won Jin Ho, Daniel R. Bergman, Atul Deshpande, Alexander S. Baras, Elizabeth M. Jaffee, Lisa M. Coussens, Elana J. Fertig, Aleksander S. Popel
Briefings Bioinform.10
2022 Regression Expression Variation Analysis (REVA): a rank-based multi-dimensional measure of correlation
abstract
Sometimes a simple question arises: how does the distance between two samples in multivariate space compare to another scalar value associated with each sample. Here, inspired by the Kendall rank correlation coefficient, we propose theory for a non-parametric test to statistically test this association based on the neighbors principle implicit in any machine learning algorithm which says that samples with similar labels should be close to one another in feature space as well. Our test, REVA, is independent of the scale of the scalar data, and thus generalizable to any comparison of samples with both high-dimensional data and a scalar. We use U-statistic theory to derive the asymptotic distribution of the new correlation coefficient, developing additional large and finite sample properties along the way. To establish the admissibility of the REVA statistic, and explore the utility and limitations of our model, we compared it to the most widely used distance based correlation coefficient in a range of simulated conditions, demonstrating that REVA does not depend on an assumption of linearity, and is robust to high levels of noise, high dimensions, and the presence of outliers. We apply the resulting statistic to problems in cancer biology motivated by the model that cancer cells with more similar gene expression profiles to one another can be expected to have a more similar response to therapy.
Bahman Afsari, Alexander V. Favorov, Elana J. Fertig, Leslie Cope
ICMLA3
2022 Ten quick tips for deep learning in biology
abstract
Machine learning is a modern approach to problem-solving and task automation.In particular, machine learning is concerned with the development and applications of algorithms that
Benjamin D. Lee, Anthony Gitter, Casey S. Greene, Sebastian Raschka, Finlay Maguire, Alexander J. Titus, Michael D. Kessler, Alexandra Lee, Marc G. Chevrette, Paul Allen Stewart, Thiago Britto-Borges, Evan M. Cofer, Kun-Hsing Yu, Juan Jose Carmona, Elana J. Fertig, Alexandr A. Kalinin, Brandon Signal, Benjamin J. Lengerich, Timothy J. Triche Jr., Simina M. Boca
PLoS Comput. Biol.15
2020 projectR: an R/Bioconductor package for transfer learning via PCA, NMF, correlation and clustering
abstract
MOTIVATION: Dimension reduction techniques are widely used to interpret high-dimensional biological data. Features learned from these methods are used to discover both technical artifacts and novel biological phenomena. Such feature discovery is critically importent in analysis of large single-cell datasets, where lack of a ground truth limits validation and interpretation. Transfer learning (TL) can be used to relate the features learned from one source dataset to a new target dataset to perform biologically driven validation by evaluating their use in or association with additional sample annotations in that independent target dataset. RESULTS: We developed an R/Bioconductor package, projectR, to perform TL for analyses of genomics data via TL of clustering, correlation and factorization methods. We then demonstrate the utility TL for integrated data analysis with an example for spatial single-cell analysis. AVAILABILITY AND IMPLEMENTATION: projectR is available on Bioconductor and at https://github.com/genesofeve/projectR. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Carlo Colantuoni, Loyal A. Goff, Elana J. Fertig, Genevieve L. Stein-O'Brien
Bioinform.4
2020 CoGAPS 3: Bayesian non-negative matrix factorization for single-cell analysis with asynchronous updates and sparse data structures
abstract
BACKGROUND: Bayesian factorization methods, including Coordinated Gene Activity in Pattern Sets (CoGAPS), are emerging as powerful analysis tools for single cell data. However, these methods have greater computational costs than their gradient-based counterparts. These costs are often prohibitive for analysis of large single-cell datasets. Many such methods can be run in parallel which enables this limitation to be overcome by running on more powerful hardware. However, the constraints imposed by the prior distributions in CoGAPS limit the applicability of parallelization methods to enhance computational efficiency for single-cell analysis. RESULTS: We developed a new software framework for parallel matrix factorization in Version 3 of the CoGAPS R/Bioconductor package to overcome the computational limitations of Bayesian matrix factorization for single cell data analysis. This parallelization framework provides asynchronous updates for sequential updating steps of the algorithm to enhance computational efficiency. These algorithmic advances were coupled with new software architecture and sparse data structures to reduce the memory overhead for single-cell data. CONCLUSIONS: Altogether our new software enhance the efficiency of the CoGAPS Bayesian matrix factorization algorithm so that it can analyze 1000 times more cells, enabling factorization of large single-cell data sets.
Thomas Sherman, Tiger Gao, Elana J. Fertig
BMC Bioinform.3
2019 CancerInSilico: An R/Bioconductor package for combining mathematical and statistical modeling to simulate time course bulk and single cell gene expression data in cancer
abstract
Bioinformatics techniques to analyze time course bulk and single cell omics data are advancing. The absence of a known ground truth of the dynamics of molecular changes challenges benchmarking their performance on real data. Realistic simulated time-course datasets are essential to assess the performance of time course bioinformatics algorithms. We develop an R/Bioconductor package, CancerInSilico, to simulate bulk and single cell transcriptional data from a known ground truth obtained from mathematical models of cellular systems. This package contains a general R infrastructure for running cell-based models and simulating gene expression data based on the model states. We show how to use this package to simulate a gene expression data set and consequently benchmark analysis methods on this data set with a known ground truth. The package is freely available via Bioconductor: http://bioconductor.org/packages/CancerInSilico/.
Thomas Sherman, Luciane T. Kagohara, Raymon Cao, Matthew Satriano, Michael Considine, Gabriel Krigsfeld, Ruchira Ranaweera, Sandra A. Jablonski, Genevieve L. Stein-O'Brien, Daria A. Gaykalova, Louis M. Weiner, Christine H. Chung, Elana J. Fertig
PLoS Comput. Biol.15
2018 Splice Expression Variation Analysis (SEVA) for inter-tumor heterogeneity of gene isoform usage in cancer
abstract
Motivation: Current bioinformatics methods to detect changes in gene isoform usage in distinct phenotypes compare the relative expected isoform usage in phenotypes. These statistics model differences in isoform usage in normal tissues, which have stable regulation of gene splicing. Pathological conditions, such as cancer, can have broken regulation of splicing that increases the heterogeneity of the expression of splice variants. Inferring events with such differential heterogeneity in gene isoform usage requires new statistical approaches. Results: We introduce Splice Expression Variability Analysis (SEVA) to model increased heterogeneity of splice variant usage between conditions (e.g. tumor and normal samples). SEVA uses a rank-based multivariate statistic that compares the variability of junction expression profiles within one condition to the variability within another. Simulated data show that SEVA is unique in modeling heterogeneity of gene isoform usage, and benchmark SEVA's performance against EBSeq, DiffSplice and rMATS that model differential isoform usage instead of heterogeneity. We confirm the accuracy of SEVA in identifying known splice variants in head and neck cancer and perform cross-study validation of novel splice variants. A novel comparison of splice variant heterogeneity between subtypes of head and neck cancer demonstrated unanticipated similarity between the heterogeneity of gene isoform usage in HPV-positive and HPV-negative subtypes and anticipated increased heterogeneity among HPV-negative samples with mutations in genes that regulate the splice variant machinery. These results show that SEVA accurately models differential heterogeneity of gene isoform usage from RNA-seq data. Availability and implementation: SEVA is implemented in the R/Bioconductor package GSReg. Contact: [email protected] or [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Bahman Afsari, Theresa Guo, Michael Considine, Liliana Florea, Luciane T. Kagohara, Genevieve L. Stein-O'Brien, Dylan Kelley, Emily Flam, Kristina D. Zambo, Patrick K. Ha, Donald Geman, Michael F. Ochs, Joseph A. Califano, Daria A. Gaykalova, Alexander V. Favorov, Elana J. Fertig
Bioinform.16
2017 StereoGene: rapid estimation of genome-wide correlation of continuous or interval feature data
abstract
MOTIVATION: Genomics features with similar genome-wide distributions are generally hypothesized to be functionally related, for example, colocalization of histones and transcription start sites indicate chromatin regulation of transcription factor activity. Therefore, statistical algorithms to perform spatial, genome-wide correlation among genomic features are required. RESULTS: Here, we propose a method, StereoGene, that rapidly estimates genome-wide correlation among pairs of genomic features. These features may represent high-throughput data mapped to reference genome or sets of genomic annotations in that reference genome. StereoGene enables correlation of continuous data directly, avoiding the data binarization and subsequent data loss. Correlations are computed among neighboring genomic positions using kernel correlation. Representing the correlation as a function of the genome position, StereoGene outputs the local correlation track as part of the analysis. StereoGene also accounts for confounders such as input DNA by partial correlation. We apply our method to numerous comparisons of ChIP-Seq datasets from the Human Epigenome Atlas and FANTOM CAGE to demonstrate its wide applicability. We observe the changes in the correlation between epigenomic features across developmental trajectories of several tissue types consistent with known biology and find a novel spatial correlation of CAGE clusters with donor splice sites and with poly(A) sites. These analyses provide examples for the broad applicability of StereoGene for regulatory genomics. AVAILABILITY AND IMPLEMENTATION: The StereoGene C ++ source code, program documentation, Galaxy integration scripts and examples are available from the project homepage http://stereogene.bioinf.fbb.msu.ru/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Elena D. Stavrovskaya, Tejasvi Niranjan, Elana J. Fertig, Sarah J. Wheelan, Alexander V. Favorov, Andrey A. Mironov
Bioinform.3
2017 PatternMarkers & GWCoGAPS for novel data-driven biomarkers via whole transcriptome NMF
abstract
SUMMARY: Non-negative Matrix Factorization (NMF) algorithms associate gene expression with biological processes (e.g. time-course dynamics or disease subtypes). Compared with univariate associations, the relative weights of NMF solutions can obscure biomarkers. Therefore, we developed a novel patternMarkers statistic to extract genes for biological validation and enhanced visualization of NMF results. Finding novel and unbiased gene markers with patternMarkers requires whole-genome data. Therefore, we also developed Genome-Wide CoGAPS Analysis in Parallel Sets (GWCoGAPS), the first robust whole genome Bayesian NMF using the sparse, MCMC algorithm, CoGAPS. Additionally, a manual version of the GWCoGAPS algorithm contains analytic and visualization tools including patternMatcher, a Shiny web application. The decomposition in the manual pipeline can be replaced with any NMF algorithm, for further generalization of the software. Using these tools, we find granular brain-region and cell-type specific signatures with corresponding biomarkers in GTEx data, illustrating GWCoGAPS and patternMarkers ascertainment of data-driven biomarkers from whole-genome data. AVAILABILITY AND IMPLEMENTATION: PatternMarkers & GWCoGAPS are in the CoGAPS Bioconductor package (3.5) under the GPL license. CONTACT: [email protected] or [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Genevieve L. Stein-O'Brien, Jacob L. Carey, Waishing Lee, Michael Considine, Alexander V. Favorov, Emily Flam, Theresa Guo, Luigi Marchionni, Thomas Sherman, Shawn Sivy, Daria A. Gaykalova, Ronald McKay, Michael F. Ochs, Carlo Colantuoni, Elana J. Fertig
Bioinform.16
2015 The estimation of dimensionality in gene expression data using Nonnegative Matrix Factorization
abstract
Nonnegative matrix factorization and other decomposition methods have proven to offer significant advantages for the interpretation of genome-wide gene expression data. However, unlike analytic methods, they suffer from instability in the inferred factors or patterns as the dimensionality is changed. We present here two statistics, one mathematical and one biological, that estimate the dimensionality. We show that they provide close though not identical estimates, and that they provide strong evidence for elimination of some potential factorizations.
Conor J. Kelton, Waishing Lee, Matthew Rusay, Ondrej Maxian, Elana J. Fertig, Michael F. Ochs
BIBM5
2015 switchBox: an R package for k-Top Scoring Pairs classifier development
abstract
UNLABELLED: k-Top Scoring Pairs (kTSP) is a classification method for prediction from high-throughput data based on a set of the paired measurements. Each of the two possible orderings of a pair of measurements (e.g. a reversal in the expression of two genes) is associated with one of two classes. The kTSP prediction rule is the aggregation of voting among such individual two-feature decision rules based on order switching. kTSP, like its predecessor, Top Scoring Pair (TSP), is a parameter-free classifier relying only on ranking of a small subset of features, rendering it robust to noise and potentially easy to interpret in biological terms. In contrast to TSP, kTSP has comparable accuracy to standard genomics classification techniques, including Support Vector Machines and Prediction Analysis for Microarrays. Here, we describe 'switchBox', an R package for kTSP-based prediction. AVAILABILITY: The 'switchBox' package is freely available from Bioconductor: http://www.bioconductor.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bahman Afsari, Elana J. Fertig, Donald Geman, Luigi Marchionni
Bioinform.2
2014 Preserving biological heterogeneity with a permuted surrogate variable analysis for genomics batch correction
abstract
MOTIVATION: Sample source, procurement process and other technical variations introduce batch effects into genomics data. Algorithms to remove these artifacts enhance differences between known biological covariates, but also carry potential concern of removing intragroup biological heterogeneity and thus any personalized genomic signatures. As a result, accurate identification of novel subtypes from batch-corrected genomics data is challenging using standard algorithms designed to remove batch effects for class comparison analyses. Nor can batch effects be corrected reliably in future applications of genomics-based clinical tests, in which the biological groups are by definition unknown a priori. RESULTS: Therefore, we assess the extent to which various batch correction algorithms remove true biological heterogeneity. We also introduce an algorithm, permuted-SVA (pSVA), using a new statistical model that is blind to biological covariates to correct for technical artifacts while retaining biological heterogeneity in genomic data. This algorithm facilitated accurate subtype identification in head and neck cancer from gene expression data in both formalin-fixed and frozen samples. When applied to predict Human Papillomavirus (HPV) status, pSVA improved cross-study validation even if the sample batches were highly confounded with HPV status in the training set. AVAILABILITY AND IMPLEMENTATION: All analyses were performed using R version 2.15.0. The code and data used to generate the results of this manuscript is available from https://sourceforge.net/projects/psva.
Hilary S. Parker, Jeffrey T. Leek, Alexander V. Favorov, Michael Considine, Xiaoxin Xia, Sameer Chavan, Christine H. Chung, Elana J. Fertig
Bioinform.8
2012 Identifying context-specific transcription factor targets from prior knowledge and gene expression data
abstract
Numerous methodologies, assays, and databases presently provide candidate targets of transcription factors (TFs). However, TFs rarely regulate their targets universally. The context of activation of a TF can change the transcriptional response of targets. Direct multiple regulation typical to mammalian genes complicates direct inference of TF targets from gene expression data. We present a novel statistic that infers context-specific TF regulation based upon the CoGAPS algorithm, which infers overlapping gene expression patterns resulting from coregulation. Numerical experiments with simulated data showed that this statistic correctly inferred targets that are common to multiple TFs, except in cases where the signal from a TF is negligible relative to noise level and signal from other TFs. The statistic is robust to moderate levels of error in the simulated gene sets, identifying fewer false positives than false negatives. Significantly, the regulatory statistic refines the number of transcription factor targets relevant to cell signaling in gastrointestinal stromal tumors (GIST) to genes consistent with the phosphorylation patterns of TFs identified in previous studies. As formulated, the proposed regulatory statistic has wide applicability to inferring set membership in integrated datasets. This statistic could be naturally extended to account for prior probabilities of set membership or to add candidate gene targets.
Elana J. Fertig, Alexander V. Favorov, Michael F. Ochs
BIBM1
2012 Matrix factorization for transcriptional regulatory network inference
abstract
Inference of Transcriptional Regulatory Networks (TRNs) provides insight into the mechanisms driving biological systems, especially mammalian development and disease. Many techniques have been developed for TRN estimation from indirect biochemical measurements. Although successful when initially tested in model organisms, these regulatory models often fail when applied to data from multicellular organisms where multiple regulation and gene reuse increase dramatically. Non-negative matrix factorization techniques were initially introduced to find non-orthogonal patterns in data, making them ideal techniques for inference in cases of multiple regulation. We review these techniques and their application to TRN analysis.
Michael F. Ochs, Elana J. Fertig
CIBCB2
2012 Quantifying the Dynamics of Coupled Networks of Switches and Oscillators
Matthew R. Francis, Elana J. Fertig
RECOMB2
2010 CoGAPS: an R/C++ package to identify patterns and biological process activity in transcriptomic data
abstract
SUMMARY: Coordinated Gene Activity in Pattern Sets (CoGAPS) provides an integrated package for isolating gene expression driven by a biological process, enhancing inference of biological processes from transcriptomic data. CoGAPS improves on other enrichment measurement methods by combining a Markov chain Monte Carlo (MCMC) matrix factorization algorithm (GAPS) with a threshold-independent statistic inferring activity on gene sets. The software is provided as open source C++ code built on top of JAGS software with an R interface. AVAILABILITY: The R package CoGAPS and the C++ package GAPS-JAGS are provided open source under the GNU Lesser Public License (GLPL) with a users manual containing installation and operating instructions. CoGAPS is available through Bioconductor and depends on the rjags package available through CRAN to interface CoGAPS with GAPS-JAGS. URL: http://www.cancerbiostats.onc.jhmi.edu/cogaps.cfm .
Elana J. Fertig, Alexander V. Favorov, Giovanni Parmigiani, Michael F. Ochs
Bioinform.1