VLDB 2026 Research / reviewers in the wild / expert
Sheida Nabavi
dblp:35/4705
· DBLP profile ↗
29ranked-venue papers
4as first author
14since 2021 · last 2025
0000-0002-5996-1020ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 25 · 1 first-author · 13 since 2021Computer networks · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Segmentation for Early Tumor Detection in Mammograms Via Temporal Discrepancy Analysis and Dynamic Loss WeightingabstractDetecting discrepancies between prior and current (P&C) mammograms is crucial for identifying subtle temporal changes, such as architectural distortions (AD) and calcifications, which indicate cancer. While radiologists effectively compare P&C images in clinical practice, deep learning models, particularly CNNs, often struggle to capture these changes over time, especially in limited-data scenarios. To address this, we propose a new hybrid deep learning framework specifically designed to analyze P&C discrepancies. The architecture identifies and processes multi-scale tissue changes to enhance segmentation accuracy by including specialized modules, referred to as feature discrepancy blocks (FDBs). Additionally, we develop a composite loss with sigmoid-based dynamic weighting to adaptively balance spatial consistency and boundary precision, ensuring stable learning. In this study, we show that by leveraging temporal context and regularization, our approach outperforms baseline methods, achieving precise binary segmentation masks and superior performance in dice score, mean IoU, sensitivity, and specificity. The code is available at: https://github.com/NabaviLab/MGSeg Afsana Ahsan Jeny, Sahand Hamzehei, Mostafa Karami, Stephen Andrew Baker, Tucker Van Rathe, Clifford Yang, Sheida Nabavi |
ICIP | 7 |
| 2025 | Graph Laplacian Transformer with Progressive Sampling for Prostate Cancer Grading
Masum Shah Junayed, John Derek Van Vessem, Gahie Nam, Sheida Nabavi |
MICCAI (12) | 5 |
| 2025 | Bayesian Transformers and Higher-Order Graph Matching for Cell Tracking in Serial Tissue Sections
Mostafa Karami, Sahand Hamzehei, David Arce, Gianna Raimondi, Linnaea Ostroff, Sheida Nabavi |
MICCAI (12) | 6 |
| 2025 | Advanced Feature Extraction and Outlier Detection for 3D Biological/Biomedical Image Registrationabstract3D image registration is essential in computer vision, medical imaging, and robotics. By aligning images from different perspectives into a single coordinate system, this approach provides a consistent viewpoint for analysis. Using accurate image alignment, we may compare, evaluate, and integrate data from different contexts. This paper describes a new method to register 3D or z-stack microscopy and medical image. It uses a hybrid of traditional and deep learning methods for feature extraction and adaptive likelihood-based methods for finding outliers. The proposed method uses the Scale-invariant Feature Transform (SIFT) and the Residual Network with 50 layers (ResNet50) to extract effective features to obtain precise and accurate representations of image contents. The registration approach also relies on the adaptive Maximum Likelihood Estimation SAmple Consensus (MLESAC) method, which optimizes outlier detection and increases noise and distortion resistance to improve the efficacy of these combined extracted features. This concatenation approach demonstrates robustness, flexibility, and adaptability across a variety of imaging modalities, enabling the registration of complex images with higher precision. Results show that the proposed algorithm outperforms commonly used registration methods, including SIFT, KAZE, Oriented FAST and Rotated BRIEF (ORB), and also registration software tools such as bUnwarpJ, and TurboReg. The algorithm's effectiveness is evaluated in terms of Mutual Information (MI), Phase Congruency-Based (PCB), and Gradient-Based Metrics (GBM). These metrics are applied to two types of datasets, including a brain scan dataset and 3D serial sections of multiplex microscopy image datasets. Sahand Hamzehei, Gianna Raimondi, Rebecca Tripp, Linnaea Ostroff, Sheida Nabavi |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2024 | Benchmarking Distance Functions in Siamese Networks for Current and Prior Mammogram Image AnalysisabstractMammogram image analysis has benefited from advancements in artificial intelligence (AI), particularly through the use of Siamese networks, which, similar to radiologists, compare current and prior mammogram images to enhance diagnostic accuracy. One of the main challenges in employing Siamese networks for this purpose is selecting an effective distance function. Given the complexity of mammogram images and the high correlation between current and prior images, traditional distance functions in Siamese networks often fall short in capturing the subtle, non-linear differences between these correlated features. This study explores the impact of incorporating non-linear and correlation-sensitive distance functions within a Siamese network framework for analyzing paired mammogram images. We benchmarked different distance functions, including Euclidean, Manhattan, Mahalanobis, Radial Basis Function (RBF), and cosine, and introduced a novel combination of RBF with Matern Covariance. Our evaluation revealed that the RBF with Matern Covariance consistently outperformed other functions, emphasizing the importance of addressing non-linearity and correlation in this context. For instance, the ResNet50 model, when paired with this distance function, achieved an accuracy of 0.938, sensitivity of 0.921, precision of 0.955, specificity of 0.958, F1 score of 0.930, and AUC of 0.940. We observed similarly strong performance across other models as well. Furthermore, the robustness of our approach was confirmed through evaluation on a dataset of 30 cross-validation samples, demonstrating its generalizability. These findings underscore the effectiveness of non-linear and correlation-based distance functions in Siamese networks for improving the performance and generalization of mammogram image analysis. All codes used in this paper are available at https://github.com/NabaviLab/Benchmarking_Distance_Functions_in_Siamese_Networks. Sahand Hamzehei, Afsana Ahsan Jeny, Annie Jin, Clifford Yang, Sheida Nabavi |
BIBM | 5 |
| 2024 | A Fused Transformer-based Model for Gene Expression Prediction using Histopathology ImagesabstractSpatial transcriptomics (ST) is a cutting-edge technology that enables the spatial localization and analysis of gene expression within tissue sections. Despite its transformative potential, ST is constrained by high costs and limited spatial resolution, making it less accessible and challenging to implement widely. To overcome these limitations, we propose a fused transformer-based approach designed to predict high-density gene expression profiles directly from Whole Slide Images (WSIs). Our model capitalizes on the multi-scale hierarchical structure of WSIs, integrating several key components: ResNet50 and transformer encoders to extract features at various scales, and a Fused Transformer Block (FTB) that effectively aggregates these features. The FTB incorporates Spatial Positional Embeddings (SPE) to maintain spatial context, cross attention to prioritize relevant features across different scales, and Random Mask Attention (RMA) to concentrate the model’s attention on the most significant patterns. To validate our approach, we conducted extensive experiments using three public spatial transcriptomics datasets (HBC, HER2+, SCC) and three additional Visium datasets from 10X Genomics. Our results demonstrate that the proposed model not only preserves essential spatial relationships for mapping gene expressions to tissue morphology but also outperforms current state-of-the-art methods. Masum Shah Junayed, Afsana Ahsan Jeny, Sheida Nabavi, Ion Mandoiu |
BIBM | 3 |
| 2024 | A Scaled Mask Attention-based Hybrid model for Survival PredictionabstractWhole Slide Images (WSIs) are widely used in medical practice, but predicting survival from histopathological images is particularly challenging due to the vast scale and intricate structure of WSIs. Traditional Convolutional Neural Network (CNN)-based methods, which divide WSIs into multiple patches for analysis, often struggle with the heterogeneity of WSIs. They often lose crucial feature information during patch sampling and fail to fully capture spatial, contextual, and hierarchical interactions. Moreover, many current methods do not adequately consider correlations between patches and struggle to adapt to the extensive dimensions of WSIs. To overcome these limitations, we introduce a novel hybrid model that integrates CNNs, transformers, and Graph Convolutional Networks (GCNs) for end-to-end survival prediction. Firstly, the model employs ResNet-50 and a Learnable Linear Projection (LLP) to extract the complex and high-dimensional features. Subsequently, the Scaled Masked Attention (SMA) proposes to focus on high-attention regions within informative features. Additionally, cosine similarity (CS) emphasizes characteristics within structurally similar features. Finally, the proposed model utilizes GCNs to effectively aggregate the extracted features, capturing the complex correlations between them to enhance survival prediction. To evaluate the performance of the proposed method, we use five public datasets. Extensive experiments and ablation studies demonstrate that our model outperforms existing methods in terms of C-index. Masum Shah Junayed, Sheida Nabavi |
BIBM | 2 |
| 2024 | Multi-modal Spatial Clustering for Spatial Transcriptomics Utilizing High-resolution Histology ImagesabstractUnderstanding the intricate cellular environment within biological tissues is crucial for uncovering insights into complex biological functions. While single-cell RNA sequencing has significantly enhanced our understanding of cellular states, it lacks the spatial context to fully comprehend the cellular environment. Spatial transcriptomics (ST) addresses this limitation by enabling transcriptome-wide profiling while preserving spatial context. One of the principal challenges in ST data analysis is spatial clustering. Modern ST sequencing procedures typically include a high-resolution histology image, which has been shown in previous studies to be closely connected to gene expression profiles. However, current spatial clustering methods often fail to fully utilize the image information, limiting their ability to capture critical spatial and cellular interactions.In this study, we propose the spatial transcriptomics multimodal clustering (stMMC) model, a novel contrastive learningbased deep learning approach that integrates gene expression data with histology image features through a multi-modal parallel graph autoencoder. We tested stMMC against four state-of-the-art baseline models on two public ST datasets. The experiments demonstrated the superior performance of stMMC in terms of ARI and NMI and an ablation study validated the contributions of key components. Bingjun Li, Mostafa Karami, Masum Shah Junayed, Sheida Nabavi |
BIBM | 4 |
| 2024 | Improved allele-specific single-cell copy number estimation in low-coverage DNA-sequencingabstractMOTIVATION: Advances in whole-genome single-cell DNA sequencing (scDNA-seq) have led to the development of numerous methods for detecting copy number aberrations (CNAs), a key driver of genetic heterogeneity in cancer. While most of these methods are limited to the inference of total copy number, some recent approaches now infer allele-specific CNAs using innovative techniques for estimating allele-frequencies in low coverage scDNA-seq data. However, these existing allele-specific methods are limited in their segmentation strategies, a crucial step in the CNA detection pipeline. RESULTS: We present SEACON (Single-cell Estimation of Allele-specific COpy Numbers), an allele-specific copy number profiler for scDNA-seq data. SEACON uses a Gaussian Mixture Model to identify latent copy number states and breakpoints between contiguous segments across cells, filters the segments for high-quality breakpoints using an ensemble technique, and adopts several strategies for tolerating noisy read-depth and allele frequency measurements. Using a wide array of both real and simulated datasets, we show that SEACON derives accurate copy numbers and surpasses existing approaches under numerous experimental conditions, and identify its strengths and weaknesses. AVAILABILITY AND IMPLEMENTATION: SEACON is implemented in Python and is freely available open-source from https://github.com/NabaviLab/SEACON and https://doi.org/10.5281/zenodo.12727008. Samson Weiner, Bingjun Li, Sheida Nabavi |
Bioinform. | 3 |
| 2024 | A multimodal graph neural network framework for cancer molecular subtype classificationabstractBACKGROUND: The recent development of high-throughput sequencing has created a large collection of multi-omics data, which enables researchers to better investigate cancer molecular profiles and cancer taxonomy based on molecular subtypes. Integrating multi-omics data has been proven to be effective for building more precise classification models. Most current multi-omics integrative models use either an early fusion in the form of concatenation or late fusion with a separate feature extractor for each omic, which are mainly based on deep neural networks. Due to the nature of biological systems, graphs are a better structural representation of bio-medical data. Although few graph neural network (GNN) based multi-omics integrative methods have been proposed, they suffer from three common disadvantages. One is most of them use only one type of connection, either inter-omics or intra-omic connection; second, they only consider one kind of GNN layer, either graph convolution network (GCN) or graph attention network (GAT); and third, most of these methods have not been tested on a more complex classification task, such as cancer molecular subtypes. RESULTS: In this study, we propose a novel end-to-end multi-omics GNN framework for accurate and robust cancer subtype classification. The proposed model utilizes multi-omics data in the form of heterogeneous multi-layer graphs, which combine both inter-omics and intra-omic connections from established biological knowledge. The proposed model incorporates learned graph features and global genome features for accurate classification. We tested the proposed model on the Cancer Genome Atlas (TCGA) Pan-cancer dataset and TCGA breast invasive carcinoma (BRCA) dataset for molecular subtype and cancer subtype classification, respectively. The proposed model shows superior performance compared to four current state-of-the-art baseline models in terms of accuracy, F1 score, precision, and recall. The comparative analysis of GAT-based models and GCN-based models reveals that GAT-based models are preferred for smaller graphs with less information and GCN-based models are preferred for larger graphs with extra information. Bingjun Li, Sheida Nabavi |
BMC Bioinform. | 2 |
| 2023 | scGEMOC, A Graph Embedded Contrastive Learning Single-cell Multiomics Clustering ModelabstractRecent advancements in single-cell multiomics sequencing create new research opportunities but also pose challenges, particularly in cell clustering. One major challenge is feature fusion. Early fusion models are robust but ignore the unique distributions of omics and cannot handle various omic dimensions. Most current clustering methods use late fusion, employing independent encoders for each omic. However, the extracted omic features belong to different latent spaces, leading to difficulties in aligning omics. Additionally, current cell clustering methods do not incorporate prior biological knowledge, such as interactions within and across omics, which has been shown plays a key role in defining cell types.To address these shortcomings, we propose a novel, scalable, end-to-end clustering method, called single-cell graph embedding multiomics cluster (scGEMOC). scGEMOC utilizes prior biological knowledge to represent inter- and intra-omics connections as a heterogeneous graph. It applies graph embedding to aggregate omics interaction data as a pseudo omic and employs contrastive learning for effectively aligning omics in the latent space. We evaluated scGEMOC on three public datasets against five state-of-the-art baseline models. scGEMOC achieves superior clustering performance compared to the baseline models on all datasets. An ablation study confirms the significant contribution of each component and identifies the most impactful one. Bingjun Li, Sheida Nabavi |
BIBM | 2 |
| 2021 | Single-cell RNA sequencing data clustering using graph convolutional networksabstractSingle-cell RNA sequencing (scRNAseq) makes it possible to analyze gene expression profiles at the individual cell scale and to discover intrinsic and extrinsic cellular processes in biological research. Cell clustering is one of the most important steps in analyzing scRNAseq data. With rapid developments of single cell sequencing technologies, scRNAseq data grow in size and heterogeneity. However, traditional clustering methods like Kmeans with or without dimension reduction methods, cannot handle high sparse and massive scRNAseq data. Although some deep learning based methods have been proposed to denoise the data and cluster cells simultaneously, learning informative representations of cells for accurate cell clustering is still a challenging problem to be solved. In this work, we propose a deep learning model that combines a deep graph convolutional network (GCN) and a self-supervised mechanism. The GCN considers not only the gene expressions but also the relationship between cells to represent cells. The self-supervised mechanism is employed to provide the clustering assignments of cells. Moreover, we utilize the negative log-likelihood of the negative binomial (NB) function as loss in the data reconstruction due to the assumption that genes expression values can be represented by the NB model. We compared the performance of our proposed method with those of the existing clustering methods for scRNAseq data and conventional clustering methods. Results show that our method achieves better performance in terms of accuracy, adjusted random index (ARI), and normalized mutual information (NMI). Bingjun Li, Sheida Nabavi |
BIBM | 3 |
| 2021 | Single-cell classification using graph convolutional networksabstractBACKGROUND: Analyzing single-cell RNA sequencing (scRNAseq) data plays an important role in understanding the intrinsic and extrinsic cellular processes in biological and biomedical research. One significant effort in this area is the identification of cell types. With the availability of a huge amount of single cell sequencing data and discovering more and more cell types, classifying cells into known cell types has become a priority nowadays. Several methods have been introduced to classify cells utilizing gene expression data. However, incorporating biological gene interaction networks has been proved valuable in cell classification procedures. RESULTS: In this study, we propose a multimodal end-to-end deep learning model, named sigGCN, for cell classification that combines a graph convolutional network (GCN) and a neural network to exploit gene interaction networks. We used standard classification metrics to evaluate the performance of the proposed method on the within-dataset classification and the cross-dataset classification. We compared the performance of the proposed method with those of the existing cell classification tools and traditional machine learning classification methods. CONCLUSIONS: Results indicate that the proposed method outperforms other commonly used methods in terms of classification accuracy and F1 scores. This study shows that the integration of prior knowledge about gene interactions with gene expressions using GCN methodologies can extract effective features improving the performance of cell classification. Sheida Nabavi |
BMC Bioinform. | 3 |
| 2021 | Applying deep learning in digital breast tomosynthesis for automatic breast cancer detection: A reviewabstractThe relatively recent reintroduction of deep learning has been a revolutionary force in the interpretation of diagnostic imaging studies. However, the technology used to acquire those images is undergoing a revolution itself at the very same time. Digital breast tomosynthesis (DBT) is one such technology, which has transformed the field of breast imaging. DBT, a form of three-dimensional mammography, is rapidly replacing the traditional two-dimensional mammograms. These parallel developments in both the acquisition and interpretation of breast images present a unique case study in how modern AI systems can be designed to adapt to new imaging methods. They also present a unique opportunity for co-development of both technologies that can better improve the validity of results and patient outcomes. In this review, we explore the ways in which deep learning can be best integrated into breast cancer screening workflows using DBT. We first explain the principles behind DBT itself and why it has become the gold standard in breast screening. We then survey the foundations of deep learning methods in diagnostic imaging, and review the current state of research into AI-based DBT interpretation. Finally, we present some of the limitations of integrating AI into clinical practice and the opportunities these present in this burgeoning field. Russell Posner, Tianyu Wang 0027, Clifford Yang, Sheida Nabavi |
Medical Image Anal. | 5 |
| 2020 | Convolutional neural network for automated mass segmentation in mammographyabstractBACKGROUND: Automatic segmentation and localization of lesions in mammogram (MG) images are challenging even with employing advanced methods such as deep learning (DL) methods. We developed a new model based on the architecture of the semantic segmentation U-Net model to precisely segment mass lesions in MG images. The proposed end-to-end convolutional neural network (CNN) based model extracts contextual information by combining low-level and high-level features. We trained the proposed model using huge publicly available databases, (CBIS-DDSM, BCDR-01, and INbreast), and a private database from the University of Connecticut Health Center (UCHC). RESULTS: We compared the performance of the proposed model with those of the state-of-the-art DL models including the fully convolutional network (FCN), SegNet, Dilated-Net, original U-Net, and Faster R-CNN models and the conventional region growing (RG) method. The proposed Vanilla U-Net model outperforms the Faster R-CNN model significantly in terms of the runtime and the Intersection over Union metric (IOU). Training with digitized film-based and fully digitized MG images, the proposed Vanilla U-Net model achieves a mean test accuracy of 92.6%. The proposed model achieves a mean Dice coefficient index (DI) of 0.951 and a mean IOU of 0.909 that show how close the output segments are to the corresponding lesions in the ground truth maps. Data augmentation has been very effective in our experiments resulting in an increase in the mean DI and the mean IOU from 0.922 to 0.951 and 0.856 to 0.909, respectively. CONCLUSIONS: The proposed Vanilla U-Net based model can be used for precise segmentation of masses in MG images. This is because the segmentation process incorporates more multi-scale spatial context, and captures more local and global context to predict a precise pixel-wise segmentation map of an input full MG image. These detected maps can help radiologists in differentiating benign and malignant lesions depend on the lesion shapes. We show that using transfer learning, introducing augmentation, and modifying the architecture of the original model results in better performance in terms of the mean accuracy, the mean DI, and the mean IOU in detecting mass lesion compared to the other DL and the conventional models. Dina Abdelhafiz, Jinbo Bi, Reda A. Ammar, Clifford Yang, Sheida Nabavi |
BMC Bioinform. | 5 |
| 2020 | Preprocessing Sequence Coverage Data for More Precise Detection of Copy Number VariationsabstractCopy number variation (CNV) is a type of genomic/genetic variation that plays an important role in phenotypic diversity, evolution, and disease susceptibility. Next generation sequencing (NGS) technologies have created an opportunity for more accurate detection of CNVs with higher resolution. However, efficient and precise detection of CNVs remains challenging due to high levels of noise and biases, data heterogeneity, and the "big data" nature of NGS data. Sequence coverage (readcount) data are mostly used for detecting CNVs, specially for whole exome sequencing data. Readcount data are contaminated with several types of biases and noise that hinder accurate detection of CNVs. In this work, we introduce a novel preprocessing pipeline for reducing noise and biases to improve the detection accuracy of CNVs in heterogeneous NGS data, such as cancer whole exome sequencing data. We have employed several normalization methods to reduce readcount's biases that are due to GC content of reads, read alignment problems, and sample impurity. We have also developed a novel efficient and effective smoothing approach based on Taut String to reduce noise and increase CNV detection power. Using simulated and real data we showed that employing the proposed preprocessing pipeline significantly improves the accuracy of CNV detection. Fatima Zare, Sardar Ansari, Kayvan Najarian, Sheida Nabavi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2019 | Single-cell RNAseq Imputation Based on Matrix Completion with Side InformationabstractDrop-out events in single-cell RNA sequencing cause large numbers of zero values in gene expression matrices. Zero values hinder accurate down-stream analysis of single-cell RNAseq (scRNAseq) data. In this study, to estimate the zero values in scRNAseq data, we proposed a novel method based on the low rank matrix completion approach. The novelty of the proposed method is due to using the gene association information as side information. We observed that incorporating the gene association information can facilitate more accurate recovery of gene expression matrices. To further improve the accuracy of zero imputation, we additionally employed a statistical model to estimate the drop-out probability of each gene to adjust imputed gene expression matrices. We conducted extensive experiments to evaluate the performance of the proposed approach using several datasets. We compared the performance of the proposed method with that of three commonly used zero imputation methods in terms of accuracy in down-stream clustering analysis. Results show that the proposed method has higher or comparable power for accurate zero imputation while has shorter run time compared to the commonly used zero imputation methods. Sheida Nabavi |
BIBM | 2 |
| 2019 | Deep convolutional neural networks for mammography: advances, challenges and applicationsabstractBACKGROUND: The limitations of traditional computer-aided detection (CAD) systems for mammography, the extreme importance of early detection of breast cancer and the high impact of the false diagnosis of patients drive researchers to investigate deep learning (DL) methods for mammograms (MGs). Recent breakthroughs in DL, in particular, convolutional neural networks (CNNs) have achieved remarkable advances in the medical fields. Specifically, CNNs are used in mammography for lesion localization and detection, risk assessment, image retrieval, and classification tasks. CNNs also help radiologists providing more accurate diagnosis by delivering precise quantitative analysis of suspicious lesions. RESULTS: In this survey, we conducted a detailed review of the strengths, limitations, and performance of the most recent CNNs applications in analyzing MG images. It summarizes 83 research studies for applying CNNs on various tasks in mammography. It focuses on finding the best practices used in these research studies to improve the diagnosis accuracy. This survey also provides a deep insight into the architecture of CNNs used for various tasks. Furthermore, it describes the most common publicly available MG repositories and highlights their main features and strengths. CONCLUSIONS: The mammography research community can utilize this survey as a basis for their current and future studies. The given comparison among common publicly available MG repositories guides the community to select the most appropriate database for their application(s). Moreover, this survey lists the best practices that improve the performance of CNNs including the pre-processing of images and the use of multi-view images. In addition, other listed techniques like transfer learning (TL), data augmentation, batch normalization, and dropout are appealing solutions to reduce overfitting and increase the generalization of the CNN models. Finally, this survey identifies the research challenges and directions that require further investigations by the community. Dina Abdelhafiz, Clifford Yang, Reda A. Ammar, Sheida Nabavi |
BMC Bioinform. | 4 |
| 2019 | Comparative analysis of differential gene expression analysis tools for single-cell RNA sequencing dataabstractBACKGROUND: The analysis of single-cell RNA sequencing (scRNAseq) data plays an important role in understanding the intrinsic and extrinsic cellular processes in biological and biomedical research. One significant effort in this area is the detection of differentially expressed (DE) genes. scRNAseq data, however, are highly heterogeneous and have a large number of zero counts, which introduces challenges in detecting DE genes. Addressing these challenges requires employing new approaches beyond the conventional ones, which are based on a nonzero difference in average expression. Several methods have been developed for differential gene expression analysis of scRNAseq data. To provide guidance on choosing an appropriate tool or developing a new one, it is necessary to evaluate and compare the performance of differential gene expression analysis methods for scRNAseq data. RESULTS: In this study, we conducted a comprehensive evaluation of the performance of eleven differential gene expression analysis software tools, which are designed for scRNAseq data or can be applied to them. We used simulated and real data to evaluate the accuracy and precision of detection. Using simulated data, we investigated the effect of sample size on the detection accuracy of the tools. Using real data, we examined the agreement among the tools in identifying DE genes, the run time of the tools, and the biological relevance of the detected DE genes. CONCLUSIONS: In general, agreement among the tools in calling DE genes is not high. There is a trade-off between true-positive rates and the precision of calling DE genes. Methods with higher true positive rates tend to show low precision due to their introducing false positives, whereas methods with high precision show low true positive rates due to identifying few DE genes. We observed that current methods designed for scRNAseq data do not tend to show better performance compared to methods designed for bulk RNAseq data. Data multimodality and abundance of zero read counts are the main characteristics of scRNAseq data, which play important roles in the performance of differential gene expression analysis methods and need to be considered in terms of the development of new methods. Craig E. Nelson, Sheida Nabavi |
BMC Bioinform. | 4 |
| 2018 | Predictive Meta-analysis of Multiple Microarray Datasets: An Application to Classification of Malignant Gliomas
Nurislam Tursynbek, Ghazal Ghahramany, Sheida Nabavi, Amin Zollanvari |
BIBM | 3 |
| 2018 | Copy number variation detection using partial alignment information
Fatima Zare, Sardar Ansari, Kayvan Najarian, Sheida Nabavi |
BIBM | 4 |
| 2018 | Noise cancellation using total variation for copy number variation detectionabstractBACKGROUND: Due to recent advances in sequencing technologies, sequence-based analysis has been widely applied to detecting copy number variations (CNVs). There are several techniques for identifying CNVs using next generation sequencing (NGS) data, however methods employing depth of coverage or read depth (RD) have recently become a main technique to identify CNVs. The main assumption of the RD-based CNV detection methods is that the readcount value at a specific genomic location is correlated with the copy number at that location. However, readcount data's noise and biases distort the association between the readcounts and copy numbers. For more accurate CNV identification, these biases and noise need to be mitigated. In this work, to detect CNVs more precisely and efficiently we propose a novel denoising method based on the total variation approach and the Taut String algorithm. RESULTS: To investigate the performance of the proposed denoising method, we computed sensitivities, false discovery rates and specificities of CNV detection when employing denoising, using both simulated and real data. We also compared the performance of the proposed denoising method, Taut String, with that of the commonly used approaches such as moving average (MA) and discrete wavelet transforms (DWT) in terms of sensitivity of detecting true CNVs and time complexity. The results show that Taut String works better than DWT and MA and has a better power to identify very narrow CNVs. The ability of Taut String denoising in preserving CNV segments' breakpoints and narrow CNVs increases the detection accuracy of segmentation algorithms, resulting in higher sensitivities and lower false discovery rates. CONCLUSIONS: In this study, we proposed a new denoising method for sequence-based CNV detection based on a signal processing technique. Existing CNV detection algorithms identify many false CNV segments and fail in detecting short CNV segments due to noise and biases. Employing an effective and efficient denoising method can significantly enhance the detection accuracy of the CNV segmentation algorithms. Advanced denoising methods from the signal processing field can be employed to implement such algorithms. We showed that non-linear denoising methods that consider sparsity and piecewise constant characteristics of CNV data result in better performance in CNV detection. Fatima Zare, Abdelrahman Hosny, Sheida Nabavi |
BMC Bioinform. | 3 |
| 2017 | Differential gene expression analysis in single-cell RNA sequencing dataabstractDifferential gene expression analysis is one of the significant efforts in single cell RNA sequencing (scRNAseq) analysis to discover the specific changes in expression levels of individual cell types. Since scRNAseq exhibits multimodality, large amounts of zero counts, and sparsity, it is different from the traditional bulk RNA sequencing (RNAseq) data. The new challenges of scRNAseq data promote the development of new methods for identifying differentially expressed (DE) genes. In this study, we proposed a new method, SigEMD, that combines a logistic regression model and a nonparametric method based on Earth Mover's Distance, to precisely and efficiently identify DE genes in scRNAseq data. The regression model is used to reduce the impact of large amounts of zero counts, and the nonparametric method is used to improve the sensitivity of detecting DE genes from multimodal scRNAseq data. By additionally employing gene interaction network information to adjust the final states of DE genes, we further reduce the false positives of calling DE genes. We used simulated data and real data to evaluate the detection accuracy of the proposed method and to compare its performance with those of other differential expression analysis methods. Results indicate that the proposed method has an overall powerful performance in terms of precision in detection, sensitivity, and specificity. Sheida Nabavi |
BIBM | 2 |
| 2017 | Noise cancellation for robust copy number variation detection using next generation sequencing dataabstractHigh-throughput next generation sequencing (NGS) technologies have created an opportunity for detecting copy number variations (CNVs) more accurately. However, efficient and precise detection of CNVs remains challenging due to high levels of noise and biases, data heterogeneity and the “big data” nature of NGS data. In this work, we introduce a novel preprocessing pipeline to improve the detection accuracy of CNVs in heterogeneous NGS data, such as cancer whole exome sequencing data. We employed several normalizations to reduce biases due to GC content, mappability and tumor contamination. We also developed a novel efficient and effective smoothing approach based on the Taut String method to reduce noise and increase the detection power of the CNV detection methods. Fatima Zare, Sardar Ansari, Kayvan Najarian, Sheida Nabavi |
BIBM | 4 |
| 2017 | An evaluation of copy number variation detection tools for cancer using whole exome sequencing dataabstractBACKGROUND: Recently copy number variation (CNV) has gained considerable interest as a type of genomic/genetic variation that plays an important role in disease susceptibility. Advances in sequencing technology have created an opportunity for detecting CNVs more accurately. Recently whole exome sequencing (WES) has become primary strategy for sequencing patient samples and study their genomics aberrations. However, compared to whole genome sequencing, WES introduces more biases and noise that make CNV detection very challenging. Additionally, tumors' complexity makes the detection of cancer specific CNVs even more difficult. Although many CNV detection tools have been developed since introducing NGS data, there are few tools for somatic CNV detection for WES data in cancer. RESULTS: In this study, we evaluated the performance of the most recent and commonly used CNV detection tools for WES data in cancer to address their limitations and provide guidelines for developing new ones. We focused on the tools that have been designed or have the ability to detect cancer somatic aberrations. We compared the performance of the tools in terms of sensitivity and false discovery rate (FDR) using real data and simulated data. Comparative analysis of the results of the tools showed that there is a low consensus among the tools in calling CNVs. Using real data, tools show moderate sensitivity (~50% - ~80%), fair specificity (~70% - ~94%) and poor FDRs (~27% - ~60%). Also, using simulated data we observed that increasing the coverage more than 10× in exonic regions does not improve the detection power of the tools significantly. CONCLUSIONS: The limited performance of the current CNV detection tools for WES data in cancer indicates the need for developing more efficient and precise CNV detection methods. Due to the complexity of tumors and high level of noise and biases in WES data, employing advanced novel segmentation, normalization and de-noising techniques that are designed specifically for cancer data is necessary. Also, CNV detection development suffers from the lack of a gold standard for performance evaluation. Finally, developing tools with user-friendly user interfaces and visualization features can enhance CNV studies for a broader range of users. Fatima Zare, Michelle Dow, Nicholas Monteleone, Abdelrahman Hosny, Sheida Nabavi |
BMC Bioinform. | 5 |
| 2016 | EMDomics: a robust and powerful method for the identification of genes differentially expressed between heterogeneous classesabstractMOTIVATION: A major goal of biomedical research is to identify molecular features associated with a biological or clinical class of interest. Differential expression analysis has long been used for this purpose; however, conventional methods perform poorly when applied to data with high within class heterogeneity. RESULTS: To address this challenge, we developed EMDomics, a new method that uses the Earth mover's distance to measure the overall difference between the distributions of a gene's expression in two classes of samples and uses permutations to obtain q-values for each gene. We applied EMDomics to the challenging problem of identifying genes associated with drug resistance in ovarian cancer. We also used simulated data to evaluate the performance of EMDomics, in terms of sensitivity and specificity for identifying differentially expressed gene in classes with high within class heterogeneity. In both the simulated and real biological data, EMDomics outperformed competing approaches for the identification of differentially expressed genes, and EMDomics was significantly more powerful than conventional methods for the identification of drug resistance-associated gene sets. EMDomics represents a new approach for the identification of genes differentially expressed between heterogeneous classes and has utility in a wide range of complex biomedical conditions in which sample classes show within class heterogeneity. AVAILABILITY AND IMPLEMENTATION: The R package is available at http://www.bioconductor.org/packages/release/bioc/html/EMDomics.html. Sheida Nabavi, Daniel Schmolze, Mayinuer Maitituoheti, Sadhika Malladi, Andrew H. Beck |
Bioinform. | 1 |
| 2010 | An Analytical Approach for Performance Evaluation of Bit-Patterned Media ChannelsabstractIn this work, a new analytical approach is used to evaluate the error performance of bit-patterned media (BPM) magnetic recording channels that employ one-dimensional (1D) and two-dimensional (2D) generalized partial response (GPR) equalizers to combat the significant inter-track interference (ITI) expected in BPM magnetic recording systems. The probability density function of ITI is obtained analytically and is used to estimate the bit error rate (BER) from the Viterbi detector. The proposed method takes into account most of the important factors affecting the BER such as ITI, un-equalized intersymbol interference (ISI), colored noise and the distance and the multiplicity of error events. In this work, it is shown that for 1D channels, modeling ITI and un-equalized ISI by Gaussian PDFs leads to inaccurate BERs and that the non-Gaussian distribution of the ITI and un-equalized ISI must be taken into account for more accurate BER estimates. This method provides fast and accurate estimates of BERs for moderate to high signal-to-noise ratios (SNRs). By using this analytical method, time-consuming numerical simulations for error performance evaluation can be avoided. Sheida Nabavi, Seungjune Jeon, B. V. K. Vijaya Kumar |
IEEE J. Sel. Areas Commun. | 1 |
| 2008 | Mitigating the Effects of Track Mis-Registration in Bit-Patterned MediaabstractIn bit-patterned media (BPM) aimed at magnetic recording densities of 1 Tbit/in2and higher, adjacent tracks may become very close leading to significant inter-track interference (ITI). Read-head offset or track mis-registration (TMR) can further degrade the performance of the channel. To investigate the effects of ITI and TMR on the performance of the channel, we have developed a two-dimensional (2D) pulse response simulator for BPM. Simulation results using these pulse responses suggest that equalizers and detectors optimized for zero TMR may not perform well in the presence of TMR. To mitigate the effects of TMR, we propose a modified trellis for the Viterbi algorithm (VA). The modified VA (MVA) takes into account the ITI while computing branch metrics. Simulation results show that the MVA can improve the bit error rate in the presence of TMR. Sheida Nabavi, B. V. K. Vijaya Kumar, James A. Bain |
ICC | 1 |
| 2007 | Two-Dimensional Generalized Partial Response Equalizer for Bit-Patterned MediaabstractThe use of bit-patterned media is one of the approaches being investigated to extend magnetic recording densities to 1 Tbit/in2and beyond. In patterned media, track pitch may be small causing adjacent tracks to have significant interference on the replay waveform from the main data track. To mitigate the effect of such inter-track interference (ITI), we propose the use of a two-dimensional (2D) generalized partial response (GPR) equalizer. We select both the equalizer and the partial response target using the minimum mean squared error (MMSE) criterion. However, we avoid the need for a 2D Viterbi algorithm by imposing a constraint on the 2D target that forces the adjacent track contributions (in the ideal case) to zero. Simulation results show that this 2D equalizer significantly improves the bit error rate (BER). In this work, the effect of a 2D GPR equalizer on the performance of a patterned media system in the presence of track misregistration (TMR) is also investigated. Based on the simulation results, the 2D equalization method appears to be more tolerant to TMR than the conventional GPR The main drawback of the proposed method is the need for simultaneously acquiring the signals from three adjacent tracks. Sheida Nabavi, B. V. K. Vijaya Kumar |
ICC | 1 |