Naftali Kaminski

dblp:88/1597 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
7since 2021 · last 2024
0000-0001-5917-4601ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2024 Predicting lung aging using scRNA-Seq data
abstract
Age prediction based on single cell RNA-Sequencing data (scRNA-Seq) can provide information for patients' susceptibility to various diseases and conditions. In addition, such analysis can be used to identify aging related genes and pathways. To enable age prediction based on scRNA-Seq data, we developed PolyEN, a new regression model which learns continuous representation for expression over time. These representations are then used by PolyEN to integrate genes to predict an age. Existing and new lung aging data we profiled demonstrated PolyEN's improved performance over existing methods for age prediction. Our results identified lung epithelial cells as the most significant predictors for non-smokers while lung endothelial cells led to the best chronological age prediction results for smokers.
Alex Singh, John E. McDonough, Taylor Sterling Adams, Robin Vos, Ruben De Man, Greg Myers, Laurens J. Ceulemans, Bart M. Vanaudenaerde, Wim A. Wuyts, Xiting Yan, Jonas Schupp, James S. Hagood, Naftali Kaminski, Ziv Bar-Joseph
PLoS Comput. Biol.14
2023 A novel Bayesian framework for harmonizing information across tissues and studies to increase cell type deconvolution accuracy
abstract
Computational cell type deconvolution on bulk transcriptomics data can reveal cell type proportion heterogeneity across samples. One critical factor for accurate deconvolution is the reference signature matrix for different cell types. Compared with inferring reference signature matrices from cell lines, rapidly accumulating single-cell RNA-sequencing (scRNA-seq) data provide a richer and less biased resource. However, deriving cell type signature from scRNA-seq data is challenging due to high biological and technical noises. In this article, we introduce a novel Bayesian framework, tranSig, to improve signature matrix inference from scRNA-seq by leveraging shared cell type-specific expression patterns across different tissues and studies. Our simulations show that tranSig is robust to the number of signature genes and tissues specified in the model. Applications of tranSig to bulk RNA sequencing data from peripheral blood, bronchoalveolar lavage and aorta demonstrate its accuracy and power to characterize biological heterogeneity across groups. In summary, tranSig offers an accurate and robust approach to defining gene expression signatures of different cell types, facilitating improved in silico cell type deconvolutions.
Wenxuan Deng, Wei Jiang 0019, Xiting Yan, Ningshan Li, Milica Vukmirovic, Naftali Kaminski, Jing Wang 0003, Hongyu Zhao 0003
Briefings Bioinform.8
2023 Comprehensive visualization of cell-cell interactions in single-cell and spatial transcriptomics with NICHES
abstract
MOTIVATION: Recent years have seen the release of several toolsets that reveal cell-cell interactions from single-cell data. However, all existing approaches leverage mean celltype gene expression values, and do not preserve the single-cell fidelity of the original data. Here, we present NICHES (Niche Interactions and Communication Heterogeneity in Extracellular Signaling), a tool to explore extracellular signaling at the truly single-cell level. RESULTS: NICHES allows embedding of ligand-receptor signal proxies to visualize heterogeneous signaling archetypes within cell clusters, between cell clusters and across experimental conditions. When applied to spatial transcriptomic data, NICHES can be used to reflect local cellular microenvironment. NICHES can operate with any list of ligand-receptor signaling mechanisms, is compatible with existing single-cell packages, and allows rapid, flexible analysis of cell-cell signaling at single-cell resolution. AVAILABILITY AND IMPLEMENTATION: NICHES is an open-source software implemented in R under academic free license v3.0 and it is available at http://github.com/msraredon/NICHES. Use-case vignettes are available at https://msraredon.github.io/NICHES/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Micha Sam Brickman Raredon, Junchen Yang, Neeharika Kothapalli, Wesley Lewis, Naftali Kaminski, Laura E. Niklason, Yuval Kluger
Bioinform.5
2023 iDESC: identifying differential expression in single-cell RNA sequencing data with multiple subjects
abstract
BACKGROUND: Single-cell RNA sequencing (scRNA-seq) technology has enabled assessment of transcriptome-wide changes at single-cell resolution. Due to the heterogeneity in environmental exposure and genetic background across subjects, subject effect contributes to the major source of variation in scRNA-seq data with multiple subjects, which severely confounds cell type specific differential expression (DE) analysis. Moreover, dropout events are prevalent in scRNA-seq data, leading to excessive number of zeroes in the data, which further aggravates the challenge in DE analysis. RESULTS: We developed iDESC to detect cell type specific DE genes between two groups of subjects in scRNA-seq data. iDESC uses a zero-inflated negative binomial mixed model to consider both subject effect and dropouts. The prevalence of dropout events (dropout rate) was demonstrated to be dependent on gene expression level, which is modeled by pooling information across genes. Subject effect is modeled as a random effect in the log-mean of the negative binomial component. We evaluated and compared the performance of iDESC with eleven existing DE analysis methods. Using simulated data, we demonstrated that iDESC had well-controlled type I error and higher power compared to the existing methods. Applications of those methods with well-controlled type I error to three real scRNA-seq datasets from the same tissue and disease showed that the results of iDESC achieved the best consistency between datasets and the best disease relevance. CONCLUSIONS: iDESC was able to achieve more accurate and robust DE analysis results by separating subject effect from disease effect with consideration of dropouts to identify DE genes, suggesting the importance of considering subject effect and dropouts in the DE analysis of scRNA-seq data with multiple subjects.
Taylor Sterling Adams, Ningya Wang, Jonas Schupp, Weimiao Wu, John E. McDonough, Geoffrey Lowell Chupp, Naftali Kaminski, Zuoheng Wang, Xiting Yan
BMC Bioinform.9
2023 Correction: iDESC: identifying differential expression in single-cell RNA sequencing data with multiple subjects
Taylor Sterling Adams, Ningya Wang, Jonas Schupp, Weimiao Wu, John E. McDonough, Geoffrey Lowell Chupp, Naftali Kaminski, Zuoheng Wang, Xiting Yan
BMC Bioinform.9
2022 CINS: Cell Interaction Network inference from Single cell expression data
abstract
Studies comparing single cell RNA-Seq (scRNA-Seq) data between conditions mainly focus on differences in the proportion of cell types or on differentially expressed genes. In many cases these differences are driven by changes in cell interactions which are challenging to infer without spatial information. To determine cell-cell interactions that differ between conditions we developed the Cell Interaction Network Inference (CINS) pipeline. CINS combines Bayesian network analysis with regression-based modeling to identify differential cell type interactions and the proteins that underlie them. We tested CINS on a disease case control and on an aging mouse dataset. In both cases CINS correctly identifies cell type interactions and the ligands involved in these interactions improving on prior methods suggested for cell interaction predictions. We performed additional mouse aging scRNA-Seq experiments which further support the interactions identified by CINS.
Carlos Cosme, Taylor Sterling Adams, Jonas Schupp, Koji Sakamoto, Nikos Xylourgidis, Matthew Ruffalo, Naftali Kaminski, Ziv Bar-Joseph
PLoS Comput. Biol.9
2021 A Markov random field model for network-based differential expression analysis of single-cell RNA-seq data
abstract
BACKGROUND: Recent development of single cell sequencing technologies has made it possible to identify genes with different expression (DE) levels at the cell type level between different groups of samples. In this article, we propose to borrow information through known biological networks to increase statistical power to identify differentially expressed genes (DEGs). RESULTS: We develop MRFscRNAseq, which is based on a Markov random field (MRF) model to appropriately accommodate gene network information as well as dependencies among cell types to identify cell-type specific DEGs. We implement an Expectation-Maximization (EM) algorithm with mean field-like approximation to estimate model parameters and a Gibbs sampler to infer DE status. Simulation study shows that our method has better power to detect cell-type specific DEGs than conventional methods while appropriately controlling type I error rate. The usefulness of our method is demonstrated through its application to study the pathogenesis and biological processes of idiopathic pulmonary fibrosis (IPF) using a single-cell RNA-sequencing (scRNA-seq) data set, which contains 18,150 protein-coding genes across 38 cell types on lung tissues from 32 IPF patients and 28 normal controls. CONCLUSIONS: The proposed MRF model is implemented in the R package MRFscRNAseq available on GitHub. By utilizing gene-gene and cell-cell networks, our method increases statistical power to detect differentially expressed genes from scRNA-seq data.
Biqing Zhu, Taylor Adams, Naftali Kaminski, Hongyu Zhao 0003
BMC Bioinform.5
2018 iDREM: Interactive visualization of dynamic regulatory networks
abstract
The Dynamic Regulatory Events Miner (DREM) software reconstructs dynamic regulatory networks by integrating static protein-DNA interaction data with time series gene expression data. In recent years, several additional types of high-throughput time series data have been profiled when studying biological processes including time series miRNA expression, proteomics, epigenomics and single cell RNA-Seq. Combining all available time series and static datasets in a unified model remains an important challenge and goal. To address this challenge we have developed a new version of DREM termed interactive DREM (iDREM). iDREM provides support for all data types mentioned above and combines them with existing interaction data to reconstruct networks that can lead to novel hypotheses on the function and timing of regulators. Users can interactively visualize and query the resulting model. We showcase the functionality of the new tool by applying it to microglia developmental data from multiple labs.
James S. Hagood, Namasivayam Ambalavanan, Naftali Kaminski, Ziv Bar-Joseph
PLoS Comput. Biol.4
2014 Missing value imputation in high-dimensional phenomic data: imputable or not, and how?
abstract
BACKGROUND: In modern biomedical research of complex diseases, a large number of demographic and clinical variables, herein called phenomic data, are often collected and missing values (MVs) are inevitable in the data collection process. Since many downstream statistical and bioinformatics methods require complete data matrix, imputation is a common and practical solution. In high-throughput experiments such as microarray experiments, continuous intensities are measured and many mature missing value imputation methods have been developed and widely applied. Numerous methods for missing data imputation of microarray data have been developed. Large phenomic data, however, contain continuous, nominal, binary and ordinal data types, which void application of most methods. Though several methods have been developed in the past few years, not a single complete guideline is proposed with respect to phenomic missing data imputation. RESULTS: In this paper, we investigated existing imputation methods for phenomic data, proposed a self-training selection (STS) scheme to select the best imputation method and provide a practical guideline for general applications. We introduced a novel concept of "imputability measure" (IM) to identify missing values that are fundamentally inadequate to impute. In addition, we also developed four variations of K-nearest-neighbor (KNN) methods and compared with two existing methods, multivariate imputation by chained equations (MICE) and missForest. The four variations are imputation by variables (KNN-V), by subjects (KNN-S), their weighted hybrid (KNN-H) and an adaptively weighted hybrid (KNN-A). We performed simulations and applied different imputation methods and the STS scheme to three lung disease phenomic datasets to evaluate the methods. An R package "phenomeImpute" is made publicly available. CONCLUSIONS: Simulations and applications to real datasets showed that MICE often did not perform well; KNN-A, KNN-H and random forest were among the top performers although no method universally performed the best. Imputation of missing values with low imputability measures increased imputation errors greatly and could potentially deteriorate downstream analyses. The STS scheme was accurate in selecting the optimal method by evaluating methods in a second layer of missingness simulation. All source files for the simulation and the real data analyses are available on the author's publication website.
Serena G. Liao, Dongwan D. Kang, Divay Chandra, Jessica Bon, Naftali Kaminski, Frank C. Sciurba, George C. Tseng
BMC Bioinform.6
2012 An R package suite for microarray meta-analysis in quality control, differentially expressed gene analysis and pathway enrichment detection
abstract
SUMMARY: With the rapid advances and prevalence of high-throughput genomic technologies, integrating information of multiple relevant genomic studies has brought new challenges. Microarray meta-analysis has become a frequently used tool in biomedical research. Little effort, however, has been made to develop a systematic pipeline and user-friendly software. In this article, we present MetaOmics, a suite of three R packages MetaQC, MetaDE and MetaPath, for quality control, differentially expressed gene identification and enriched pathway detection for microarray meta-analysis. MetaQC provides a quantitative and objective tool to assist study inclusion/exclusion criteria for meta-analysis. MetaDE and MetaPath were developed for candidate marker and pathway detection, which provide choices of marker detection, meta-analysis and pathway analysis methods. The system allows flexible input of experimental data, clinical outcome (case-control, multi-class, continuous or survival) and pathway databases. It allows missing values in experimental data and utilizes multi-core parallel computing for fast implementation. It generates informative summary output and visualization plots, operates on different operation systems and can be expanded to include new algorithms or combine different types of genomic data. This software suite provides a comprehensive tool to conveniently implement and compare various genomic meta-analysis pipelines. AVAILABILITY: http://www.biostat.pitt.edu/bioinfo/software.htm CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xingbin Wang, Dongwan D. Kang, Kui Shen, Chi Song, Shuya Lu, Lun-Ching Chang, Serena G. Liao, Zhiguang Huo, Shaowu Tang, Naftali Kaminski, Etienne Sibille, George C. Tseng
Bioinform.11
2012 Novel Modeling of Combinatorial miRNA Targeting Identifies SNP with Potential Role in Bone Density
abstract
MicroRNAs (miRNAs) are post-transcriptional regulators that bind to their target mRNAs through base complementarity. Predicting miRNA targets is a challenging task and various studies showed that existing algorithms suffer from high number of false predictions and low to moderate overlap in their predictions. Until recently, very few algorithms considered the dynamic nature of the interactions, including the effect of less specific interactions, the miRNA expression level, and the effect of combinatorial miRNA binding. Addressing these issues can result in a more accurate miRNA:mRNA modeling with many applications, including efficient miRNA-related SNP evaluation. We present a novel thermodynamic model based on the Fermi-Dirac equation that incorporates miRNA expression in the prediction of target occupancy and we show that it improves the performance of two popular single miRNA target finders. Modeling combinatorial miRNA targeting is a natural extension of this model. Two other algorithms show improved prediction efficiency when combinatorial binding models were considered. ComiR (Combinatorial miRNA targeting), a novel algorithm we developed, incorporates the improved predictions of the four target finders into a single probabilistic score using ensemble learning. Combining target scores of multiple miRNAs using ComiR improves predictions over the naïve method for target combination. ComiR scoring scheme can be used for identification of SNPs affecting miRNA binding. As proof of principle, ComiR identified rs17737058 as disruptive to the miR-488-5p:NCOA1 interaction, which we confirmed in vitro. We also found rs17737058 to be significantly associated with decreased bone mineral density (BMD) in two independent cohorts indicating that the miR-488-5p/NCOA1 regulatory axis is likely critical in maintaining BMD in women. With increasing availability of comprehensive high-throughput datasets from patients ComiR is expected to become an essential tool for miRNA-related studies.
Claudia Coronnello, Ryan J. Hartmaier, Arshi Arora, Luai Huleihel, Kusum V. Pandit, Abha S. Bais, Michael Butterworth, Naftali Kaminski, Gary D. Stormo, Steffi Oesterreich, Panayiotis V. Benos
PLoS Comput. Biol.8
2010 Module-based prediction approach for robust inter-study predictions in microarray data
abstract
MOTIVATION: Traditional genomic prediction models based on individual genes suffer from low reproducibility across microarray studies due to the lack of robustness to expression measurement noise and gene missingness when they are matched across platforms. It is common that some of the genes in the prediction model established in a training study cannot be matched to another test study because a different platform is applied. The failure of inter-study predictions has severely hindered the clinical applications of microarray. To overcome the drawbacks of traditional gene-based prediction (GBP) models, we propose a module-based prediction (MBP) strategy via unsupervised gene clustering. RESULTS: K-means clustering is used to group genes sharing similar expression profiles into gene modules, and small modules are merged into their nearest neighbors. Conventional univariate or multivariate feature selection procedure is applied and a representative gene from each selected module is identified to construct the final prediction model. As a result, the prediction model is portable to any test study as long as partial genes in each module exist in the test study. We demonstrate that K-means cluster sizes generally follow a multinomial distribution and the failure probability of inter-study prediction due to missing genes is diminished by merging small clusters into their nearest neighbors. By simulation and applications of real datasets in inter-study predictions, we show that the proposed MBP provides slightly improved accuracy while is considerably more robust than traditional GBP. AVAILABILITY: http://www.biostat.pitt.edu/bioinfo/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhibao Mi, Kui Shen, Nan Song, Chunrong Cheng, Chi Song, Naftali Kaminski, George C. Tseng
Bioinform.6
2008 Alignment and classification of time series gene expression in clinical studies
abstract
MOTIVATION: Classification of tissues using static gene-expression data has received considerable attention. Recently, a growing number of expression datasets are measured as a time series. Methods that are specifically designed for this temporal data can both utilize its unique features (temporal evolution of profiles) and address its unique challenges (different response rates of patients in the same class). RESULTS: We present a method that utilizes hidden Markov models (HMMs) for the classification task. We use HMMs with less states than time points leading to an alignment of the different patient response rates. To focus on the differences between the two classes we develop a discriminative HMM classifier. Unlike the traditional generative HMM, discriminative HMM can use examples from both classes when learning the model for a specific class. We have tested our method on both simulated and real time series expression data. As we show, our method improves upon prior methods and can suggest markers for specific disease and response stages that are not found when using traditional classifiers. AVAILABILITY: Matlab implementation is available from http://www.cs.cmu.edu/~thlin/tram/.
Tien-ho Lin, Naftali Kaminski, Ziv Bar-Joseph
ISMB2
2006 A Patient-Gene Model for Temporal Expression Profiles in Clinical Studies
Naftali Kaminski, Ziv Bar-Joseph
RECOMB1
2005 Comparison of normalization methods for CodeLink Bioarray data
abstract
BACKGROUND: The quality of microarray data can seriously affect the accuracy of downstream analyses. In order to reduce variability and enhance signal reproducibility in these data, many normalization methods have been proposed and evaluated, most of which are for data obtained from cDNA microarrays and Affymetrix GeneChips. CodeLink Bioarrays are a newly emerged, single-color oligonucleotide microarray platform. To date, there are no reported studies that evaluate normalization methods for CodeLink Bioarrays. RESULTS: We compared five existing normalization approaches, in terms of both noise reduction and signal retention: Median (suggested by the manufacturer), CyclicLoess, Quantile, Iset, and Qspline. These methods were applied to two real datasets (a time course dataset and a lung disease-related dataset) generated by CodeLink Bioarrays and were assessed using multiple statistical significance tests. Compared to Median, CyclicLoess and Qspline exhibit a significant and the most consistent improvement in reduction of variability and retention of signal. CyclicLoess appears to retain more signal than Qspline. Quantile reduces more variability than Median in both datasets, yet fails to consistently retain more signal in the time course dataset. Iset does not improve over Median in either noise reduction or signal enhancement in the time course dataset. CONCLUSION: Median is insufficient either to reduce variability or to retain signal effectively for CodeLink Bioarray data. CyclicLoess is a more suitable approach for normalizing these data. CyclicLoess also seems to be the most effective method among the five different normalization strategies examined.
Wei Wu 0023, Nilesh Dave, George C. Tseng, Thomas Richards, Eric P. Xing, Naftali Kaminski
BMC Bioinform.6
2004 Comparative analysis of algorithms for signal quantitation from oligonucleotide microarrays
abstract
MOTIVATION: Recent years' exponential increase in DNA microarrays experiments has motivated the development of many signal quantitation (SQ) algorithms. These algorithms perform various transformations on the actual measurements aimed to enable researchers to compare readings of different genes quantitatively within one experiment and across separate experiments. However, it is relatively unclear whether there is a 'best' algorithm to quantitate microarray data. The ability to compare and assess such algorithms is crucial for any downstream analysis. In this work, we suggest a methodology for comparing different signal quantitation algorithms for gene expression data. Our aim is to enable researchers to compare the effect of different SQ algorithms on the specific dataset they are dealing with. We combine two kinds of tests to assess the effect of an SQ algorithm in terms of signal to noise ratio. To assess noise, we exploit redundancy within the experimental dataset to test the variability of a given SQ algorithm output. For the effect of the SQ on the signal we evaluate the overabundance of differentially expressed genes using various statistical significance tests. RESULTS: We demonstrate our analysis approach with three SQ algorithms for oligonucleotide microarrays. We compare the results of using the dChip software and the RMAExpress software to the ones obtained by using the standard Affymetrix MAS5 on a dataset containing pairs of repeated hybridizations. Our analysis suggests that dChip is more robust and stable than the MAS5 tools for about 60% of the genes while RMAExpress is able to achieve an even greater improvement in terms of signal to noise, for more than 95% of the genes.
Yoseph Barash, Elinor Dehan, Meir Krupsky, Wilbur Franklin, Marc Geraci, Nir Friedman, Naftali Kaminski
Bioinform.7