Nathan Salomonis

dblp:88/4508 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
4since 2021 · last 2023
0000-0001-9689-2469ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2023 pyInfinityFlow: optimized imputation and analysis of high-dimensional flow cytometry data for millions of cells
abstract
MOTIVATION: While conventional flow cytometry is limited to dozens of markers, new experimental and computational strategies, such as Infinity Flow, allow for the generation and imputation of hundreds of cell surface protein markers in millions of cells. Here, we describe an end-to-end analysis workflow for Infinity Flow data in Python. RESULTS: pyInfinityFlow enables the efficient analysis of millions of cells, without down-sampling, through direct integration with well-established Python packages for single-cell genomics analysis. pyInfinityFlow accurately identifies both common and extremely rare cell populations which are challenging to define from single-cell genomics studies alone. We demonstrate that this workflow can nominate novel markers to design new flow cytometry gating strategies for predicted cell populations. pyInfinityFlow can be extended to diverse cell discovery analyses with flexibility to adapt to diverse Infinity Flow experimental designs. AVAILABILITY AND IMPLEMENTATION: pyInfinityFlow is freely available in GitHub (https://github.com/KyleFerchen/pyInfinityFlow) and on PyPI (https://pypi.org/project/pyInfinityFlow/). Package documentation with tutorials on a test dataset is available by Read the Docs (pyinfinityflow.readthedocs.io). The scripts and data for reproducing the results are available at https://github.com/KyleFerchen/pyInfinityFlow/tree/main/analysis_scripts, along with the raw flow cytometry input data.
Kyle Ferchen, Nathan Salomonis, H. Leighton Grimes
Bioinform.2
2022 CellDrift: inferring perturbation responses in temporally sampled single-cell data
abstract
Cells and tissues respond to perturbations in multiple ways that can be sensitively reflected in the alterations of gene expression. Current approaches to finding and quantifying the effects of perturbations on cell-level responses over time disregard the temporal consistency of identifiable gene programs. To leverage the occurrence of these patterns for perturbation analyses, we developed CellDrift (https://github.com/KANG-BIOINFO/CellDrift), a generalized linear model-based functional data analysis method that is capable of identifying covarying temporal patterns of various cell types in response to perturbations. As compared to several other approaches, CellDrift demonstrated superior performance in the identification of temporally varied perturbation patterns and the ability to impute missing time points. We applied CellDrift to multiple longitudinal datasets, including COVID-19 disease progression and gastrointestinal tract development, and demonstrated its ability to identify specific gene programs associated with sequential biological processes, trajectories and outcomes.
Kang Jin, Daniel J. Schnell, Nathan Salomonis, V. B. Surya Prasath, Rhonda Szczesniak, Bruce J. Aronow
Briefings Bioinform.4
2021 DeepImmuno: deep learning-empowered prediction and generation of immunogenic peptides for T-cell immunity
abstract
Cytolytic T-cells play an essential role in the adaptive immune system by seeking out, binding and killing cells that present foreign antigens on their surface. An improved understanding of T-cell immunity will greatly aid in the development of new cancer immunotherapies and vaccines for life-threatening pathogens. Central to the design of such targeted therapies are computational methods to predict non-native peptides to elicit a T-cell response, however, we currently lack accurate immunogenicity inference methods. Another challenge is the ability to accurately simulate immunogenic peptides for specific human leukocyte antigen alleles, for both synthetic biological applications, and to augment real training datasets. Here, we propose a beta-binomial distribution approach to derive peptide immunogenic potential from sequence alone. We conducted systematic benchmarking of five traditional machine learning (ElasticNet, K-nearest neighbors, support vector machine, Random Forest and AdaBoost) and three deep learning models (convolutional neural network (CNN), Residual Net and graph neural network) using three independent prior validated immunogenic peptide collections (dengue virus, cancer neoantigen and SARS-CoV-2). We chose the CNN as the best prediction model, based on its adaptivity for small and large datasets and performance relative to existing methods. In addition to outperforming two highly used immunogenicity prediction algorithms, DeepImmuno-CNN correctly predicts which residues are most important for T-cell antigen recognition and predicts novel impacts of SARS-CoV-2 variants. Our independent generative adversarial network (GAN) approach, DeepImmuno-GAN, was further able to accurately simulate immunogenic peptides with physicochemical properties and immunogenicity predictions similar to that of real antigens. We provide DeepImmuno-CNN as source code and an easy-to-use web interface.
Balaji Iyer, V. B. Surya Prasath, Yizhao Ni, Nathan Salomonis
Briefings Bioinform.5
2021 Pseudocell Tracer - A method for inferring dynamic trajectories using scRNAseq and its application to B cells undergoing immunoglobulin class switch recombination
abstract
Single cell RNA sequencing (scRNAseq) can be used to infer a temporal ordering of cellular states. Current methods for the inference of cellular trajectories rely on unbiased dimensionality reduction techniques. However, such biologically agnostic ordering can prove difficult for modeling complex developmental or differentiation processes. The cellular heterogeneity of dynamic biological compartments can result in sparse sampling of key intermediate cell states. To overcome these limitations, we develop a supervised machine learning framework, called Pseudocell Tracer, which infers trajectories in pseudospace rather than in pseudotime. The method uses a supervised encoder, trained with adjacent biological information, to project scRNAseq data into a low-dimensional manifold that maps the transcriptional states a cell can occupy. Then a generative adversarial network (GAN) is used to simulate pesudocells at regular intervals along a virtual cell-state axis. We demonstrate the utility of Pseudocell Tracer by modeling B cells undergoing immunoglobulin class switch recombination (CSR) during a prototypic antigen-induced antibody response. Our results revealed an ordering of key transcription factors regulating CSR to the IgG1 isotype, including the concomitant expression of Nfkb1 and Stat6 prior to the upregulation of Bach2 expression. Furthermore, the expression dynamics of genes encoding cytokine receptors suggest a poised IL-4 signaling state that preceeds CSR to the IgG1 isotype.
Derek Reiman, Godhev Kumar Manakkat Vijay, Heping Xu, Andrew Sonin, Dianyu Chen, Nathan Salomonis, Harinder Singh, Aly Azeem Khan
PLoS Comput. Biol.6
2020 Resolving single-cell heterogeneity from hundreds of thousands of cells through sequential hybrid clustering and NMF
abstract
MOTIVATION: The rapid proliferation of single-cell RNA-sequencing (scRNA-Seq) technologies has spurred the development of diverse computational approaches to detect transcriptionally coherent populations. While the complexity of the algorithms for detecting heterogeneity has increased, most require significant user-tuning, are heavily reliant on dimension reduction techniques and are not scalable to ultra-large datasets. We previously described a multi-step algorithm, Iterative Clustering and Guide-gene Selection (ICGS), which applies intra-gene correlation and hybrid clustering to uniquely resolve novel transcriptionally coherent cell populations from an intuitive graphical user interface. RESULTS: We describe a new iteration of ICGS that outperforms state-of-the-art scRNA-Seq detection workflows when applied to well-established benchmarks. This approach combines multiple complementary subtype detection methods (HOPACH, sparse non-negative matrix factorization, cluster 'fitness', support vector machine) to resolve rare and common cell-states, while minimizing differences due to donor or batch effects. Using data from multiple cell atlases, we show that the PageRank algorithm effectively downsamples ultra-large scRNA-Seq datasets, without losing extremely rare or transcriptionally similar yet distinct cell types and while recovering novel transcriptionally distinct cell populations. We believe this new approach holds tremendous promise in reproducibly resolving hidden cell populations in complex datasets. AVAILABILITY AND IMPLEMENTATION: ICGS2 is implemented in Python. The source code and documentation are available at http://altanalyze.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Meenakshi Venkatasubramanian, Kashish Chetal, Daniel J. Schnell, Gowtham Atluri, Nathan Salomonis
Bioinform.5
2014 MAAMD: a workflow to standardize meta-analyses and comparison of affymetrix microarray data
abstract
BACKGROUND: Mandatory deposit of raw microarray data files for public access, prior to study publication, provides significant opportunities to conduct new bioinformatics analyses within and across multiple datasets. Analysis of raw microarray data files (e.g. Affymetrix CEL files) can be time consuming, complex, and requires fundamental computational and bioinformatics skills. The development of analytical workflows to automate these tasks simplifies the processing of, improves the efficiency of, and serves to standardize multiple and sequential analyses. Once installed, workflows facilitate the tedious steps required to run rapid intra- and inter-dataset comparisons. RESULTS: We developed a workflow to facilitate and standardize Meta-Analysis of Affymetrix Microarray Data analysis (MAAMD) in Kepler. Two freely available stand-alone software tools, R and AltAnalyze were embedded in MAAMD. The inputs of MAAMD are user-editable csv files, which contain sample information and parameters describing the locations of input files and required tools. MAAMD was tested by analyzing 4 different GEO datasets from mice and drosophila.MAAMD automates data downloading, data organization, data quality control assesment, differential gene expression analysis, clustering analysis, pathway visualization, gene-set enrichment analysis, and cross-species orthologous-gene comparisons. MAAMD was utilized to identify gene orthologues responding to hypoxia or hyperoxia in both mice and drosophila. The entire set of analyses for 4 datasets (34 total microarrays) finished in ~ one hour. CONCLUSIONS: MAAMD saves time, minimizes the required computer skills, and offers a standardized procedure for users to analyze microarray datasets and make new intra- and inter-dataset comparisons.
Zhuohui Gan, Jianwu Wang 0001, Nathan Salomonis, Jennifer C. Stowe, Gabriel G. Haddad, Andrew D. McCulloch, Ilkay Altintas, Alexander C. Zambon
BMC Bioinform.3
2014 Long Non-Coding RNA and Alternative Splicing Modulations in Parkinson's Leukocytes Identified by RNA Sequencing
abstract
The continuously prolonged human lifespan is accompanied by increase in neurodegenerative diseases incidence, calling for the development of inexpensive blood-based diagnostics. Analyzing blood cell transcripts by RNA-Seq is a robust means to identify novel biomarkers that rapidly becomes a commonplace. However, there is lack of tools to discover novel exons, junctions and splicing events and to precisely and sensitively assess differential splicing through RNA-Seq data analysis and across RNA-Seq platforms. Here, we present a new and comprehensive computational workflow for whole-transcriptome RNA-Seq analysis, using an updated version of the software AltAnalyze, to identify both known and novel high-confidence alternative splicing events, and to integrate them with both protein-domains and microRNA binding annotations. We applied the novel workflow on RNA-Seq data from Parkinson's disease (PD) patients' leukocytes pre- and post- Deep Brain Stimulation (DBS) treatment and compared to healthy controls. Disease-mediated changes included decreased usage of alternative promoters and N-termini, 5'-end variations and mutually-exclusive exons. The PD regulated FUS and HNRNP A/B included prion-like domains regulated regions. We also present here a workflow to identify and analyze long non-coding RNAs (lncRNAs) via RNA-Seq data. We identified reduced lncRNA expression and selective PD-induced changes in 13 of over 6,000 detected leukocyte lncRNAs, four of which were inversely altered post-DBS. These included the U1 spliceosomal lncRNA and RP11-462G22.1, each entailing sequence complementarity to numerous microRNAs. Analysis of RNA-Seq from PD and unaffected controls brains revealed over 7,000 brain-expressed lncRNAs, of which 3,495 were co-expressed in the leukocytes including U1, which showed both leukocyte and brain increases. Furthermore, qRT-PCR validations confirmed these co-increases in PD leukocytes and two brain regions, the amygdala and substantia-nigra, compared to controls. This novel workflow allows deep multi-level inspection of RNA-Seq datasets and provides a comprehensive new resource for understanding disease transcriptome modifications in PD and other neurodegenerative diseases.
Lilach Soreq, Alessandro Guffanti, Nathan Salomonis, Alon Simchovitz, Zvi Israel, Hagai Bergman, Hermona Soreq
PLoS Comput. Biol.3
2013 Whole-Genome rVISTA: a tool to determine enrichment of transcription factor binding sites in gene promoters from transcriptomic data
abstract
SUMMARY: We have developed a web-based query tool, Whole-Genome rVISTA (WGRV), that determines enrichment of transcription factors (TFs) and associated target genes in sets of co-regulated genes. WGRV enables users to query databases containing pre-computed genome coordinates of evolutionarily conserved transcription factor binding sites in the proximal promoters (from 100 bp to 5 kb upstream) of human, mouse and Drosophila genomes. TF binding sites are based on position-weight matrices from the TRANSFAC Professional database. For a given set of co-regulated genes, WGRV returns statistically enriched and evolutionarily conserved binding sites, mapped by the regulatory VISTA (rVISTA) algorithm. Users can then retrieve a list of genes from the query set containing the enriched TF binding sites and their location in the query set promoters. Results are exported in a BED format for rapid visualization in the UCSC genome browser. Flat files of mapped conserved sites and their genomic coordinates are also available for analysis with stand-alone software. AVAILABILITY: http://genome.lbl.gov/cgi-bin/WGRVistaInputCommon.pl.
Inna Dubchak, Matthew Munoz, Alexander Poliakov, Nathan Salomonis, Simon Minovitsky, Rolf Bodmer, Alexander C. Zambon
Bioinform.4
2012 GO-Elite: a flexible solution for pathway and ontology over-representation
abstract
UNLABELLED: We introduce GO-Elite, a flexible and powerful pathway analysis tool for a wide array of species, identifiers (IDs), pathways, ontologies and gene sets. In addition to the Gene Ontology (GO), GO-Elite allows the user to perform over-representation analysis on any structured ontology annotations, pathway database or biological IDs (e.g. gene, protein or metabolite). GO-Elite exploits the structured nature of biological ontologies to report a minimal set of non-overlapping terms. The results can be visualized on WikiPathways or as networks. Built-in support is provided for over 60 species and 50 ID systems, covering gene, disease and phenotype ontologies, multiple pathway databases, biomarkers, and transcription factor and microRNA targets. GO-Elite is available as a web interface, GenMAPP-CS plugin and as a cross-platform application. AVAILABILITY: http://www.genmapp.org/go_elite
Alexander C. Zambon, Stan Gaj, Isaac Ho, Kristina Hanspers, Karen Vranizan, Chris T. A. Evelo, Bruce R. Conklin, Alexander R. Pico, Nathan Salomonis
Bioinform.9
2012 Mosaic: making biological sense of complex networks
abstract
UNLABELLED: We present a Cytoscape plugin called Mosaic to support interactive network annotation, partitioning, layout and coloring based on gene ontology or other relevant annotations. AVAILABILITY: Mosaic is distributed for free under the Apache v2.0 open source license and can be downloaded via the Cytoscape plugin manager. A detailed user manual is available on the Mosaic web site (http://nrnb.org/tools/mosaic).
Chao Zhang 0032, Kristina Hanspers, Allan Kuchinsky, Nathan Salomonis, Dong Xu 0002, Alexander R. Pico
Bioinform.4
2009 Alternative Splicing in the Differentiation of Human Embryonic Stem Cells into Cardiac Precursors
abstract
The role of alternative splicing in self-renewal, pluripotency and tissue lineage specification of human embryonic stem cells (hESCs) is largely unknown. To better define these regulatory cues, we modified the H9 hESC line to allow selection of pluripotent hESCs by neomycin resistance and cardiac progenitors by puromycin resistance. Exon-level microarray expression data from undifferentiated hESCs and cardiac and neural precursors were used to identify splice isoforms with cardiac-restricted or common cardiac/neural differentiation expression patterns. Splice events for these groups corresponded to the pathways of cytoskeletal remodeling, RNA splicing, muscle specification, and cell cycle checkpoint control as well as genes with serine/threonine kinase and helicase activity. Using a new program named AltAnalyze (http://www.AltAnalyze.org), we identified novel changes in protein domain and microRNA binding site architecture that were predicted to affect protein function and expression. These included an enrichment of splice isoforms that oppose cell-cycle arrest in hESCs and that promote calcium signaling and cardiac development in cardiac precursors. By combining genome-wide predictions of alternative splicing with new functional annotations, our data suggest potential mechanisms that may influence lineage commitment and hESC maintenance at the level of specific splice isoforms and microRNA regulation.
Nathan Salomonis, Brandon Nelson, Karen Vranizan, Alexander R. Pico, Kristina Hanspers, Allan Kuchinsky, Linda Ta, Mark Mercola, Bruce R. Conklin
PLoS Comput. Biol.1
2007 GenMAPP 2: new features and resources for pathway analysis
abstract
BACKGROUND: Microarray technologies have evolved rapidly, enabling biologists to quantify genome-wide levels of gene expression, alternative splicing, and sequence variations for a variety of species. Analyzing and displaying these data present a significant challenge. Pathway-based approaches for analyzing microarray data have proven useful for presenting data and for generating testable hypotheses. RESULTS: To address the growing needs of the microarray community we have released version 2 of Gene Map Annotator and Pathway Profiler (GenMAPP), a new GenMAPP database schema, and integrated resources for pathway analysis. We have redesigned the GenMAPP database to support multiple gene annotations and species as well as custom species database creation for a potentially unlimited number of species. We have expanded our pathway resources by utilizing homology information to translate pathway content between species and extending existing pathways with data derived from conserved protein interactions and coexpression. We have implemented a new mode of data visualization to support analysis of complex data, including time-course, single nucleotide polymorphism (SNP), and splicing. GenMAPP version 2 also offers innovative ways to display and share data by incorporating HTML export of analyses for entire sets of pathways as organized web pages. CONCLUSION: GenMAPP version 2 provides a means to rapidly interrogate complex experimental data for pathway-level changes in a diverse range of organisms.
Nathan Salomonis, Kristina Hanspers, Alexander C. Zambon, Karen Vranizan, Steven C. Lawlor, Kam D. Dahlquist, Scott Doniger, Joshua M. Stuart, Bruce R. Conklin, Alexander R. Pico
BMC Bioinform.1