Zhaolei Zhang

dblp:32/3360 · DBLP profile ↗
← Back
31ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0002-7303-6472ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Databases, data management, data science and information retrieval · 3Systems, architecture and hardware · 2Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 scTrans: Sparse attention powers fast and accurate cell type annotation in single-cell RNA-seq data
abstract
Cell type annotation is crucial in single-cell RNA sequencing data analysis because it enables significant biological discoveries and deepens our understanding of tissue biology. Given the high-dimensional and highly sparse nature of single-cell RNA sequencing data, most existing annotation tools focus on highly variable genes to reduce dimensionality and computational load. However, this approach inevitably results in information loss, potentially weakening the model's generalization performance and adaptability to novel datasets. To mitigate this issue, we developed scTrans, a single cell Transformer-based model, which employs sparse attention to utilize all non-zero genes, thereby effectively reducing the input data dimensionality while minimizing information loss. We validated the speed and accuracy of scTrans by performing cell type annotation on 31 different tissues within the Mouse Cell Atlas. Remarkably, even with datasets nearing a million cells, scTrans efficiently perform cell type annotation in limited computational resources. Furthermore, scTrans demonstrates strong generalization capabilities, accurately annotating cells in novel datasets and generating high-quality latent representations, which are essential for precise clustering and trajectory analysis.
Zhiyi Zou, Ying Liu 0027, Yuting Bai, Jiawei Luo 0001, Zhaolei Zhang
PLoS Comput. Biol.5
2024 PmmNDD: Predicting the Pathogenicity of Missense Mutations in Neurodegenerative Diseases via Ensemble Learning
Xijian Li, Runxuan Tang, Guangcheng Xiao, Xiaochuan Chen, Ruilin He, Zhaolei Zhang, Jiana Luo, Yanjie Wei, Yijun Mao
ISBRA (3)7
2024 Hardware Implementation of a Novel UAV Multi-Trajectory Dual-Mobility Channel Emulator
abstract
As an indispensable part of future sixth generation (6G) wireless communication, unmanned aerial vehicle (UAV) communication has become an inevitable trend and developed rapidly. The channel emulator is an effective and resource-saving tool to evaluate the designed UAV communication systems. In this paper, a novel UAV multi-trajectory dual-mobility channel model is proposed. Then, based on the Coordinate Rotation Digital Computer (CORDIC) method, the hardware implementation of the proposed model is developed in the Field Programmable Gate Array (FPGA). As the hardware output data of channel emulator, the channel impulse responses (CIRs) are acquired. Furthermore, the corresponding power delay profile (PDP) and root-mean-square (RMS) delay spread calculated by the CIRs are analyzed. Finally, by comparing with the simulation results of proposed channel model, the availability and accuracy of the channel emulator is validated.
Jingquan Li, Yu Liu 0020, Jingfan Zhang, Zhaolei Zhang, Hengtai Chang, Jie Huang 0004
WCNC4
2024 Modeling single cell trajectory using forward-backward stochastic differential equations
abstract
Recent advances in single-cell sequencing technology have provided opportunities for mathematical modeling of dynamic developmental processes at the single-cell level, such as inferring developmental trajectories. Optimal transport has emerged as a promising theoretical framework for this task by computing pairings between cells from different time points. However, optimal transport methods have limitations in capturing nonlinear trajectories, as they are static and can only infer linear paths between endpoints. In contrast, stochastic differential equations (SDEs) offer a dynamic and flexible approach that can model non-linear trajectories, including the shape of the path. Nevertheless, existing SDE methods often rely on numerical approximations that can lead to inaccurate inferences, deviating from true trajectories. To address this challenge, we propose a novel approach combining forward-backward stochastic differential equations (FBSDE) with a refined approximation procedure. Our FBSDE model integrates the forward and backward movements of two SDEs in time, aiming to capture the underlying dynamics of single-cell developmental trajectories. Through comprehensive benchmarking on multiple scRNA-seq datasets, we demonstrate the superior performance of FBSDE compared to other methods, highlighting its efficacy in accurately inferring developmental trajectories.
Dehan Kong, Zhaolei Zhang
PLoS Comput. Biol.4
2022 A configurable deep learning framework for medical image analysis
Jianguo Chen 0001, Mimi Zhou, Zhaolei Zhang, Xulei Yang
Neural Comput. Appl.4
2022 Inferring RNA-binding protein target preferences using adversarial domain adaptation
abstract
Precise identification of target sites of RNA-binding proteins (RBP) is important to understand their biochemical and cellular functions. A large amount of experimental data is generated by in vivo and in vitro approaches. The binding preferences determined from these platforms share similar patterns but there are discernable differences between these datasets. Computational methods trained on one dataset do not always work well on another dataset. To address this problem which resembles the classic "domain shift" in deep learning, we adopted the adversarial domain adaptation (ADDA) technique and developed a framework (RBP-ADDA) that can extract RBP binding preferences from an integration of in vivo and vitro datasets. Compared with conventional methods, ADDA has the advantage of working with two input datasets, as it trains the initial neural network for each dataset individually, projects the two datasets onto a feature space, and uses an adversarial framework to derive an optimal network that achieves an optimal discriminative predictive power. In the first step, for each RBP, we include only the in vitro data to pre-train a source network and a task predictor. Next, for the same RBP, we initiate the target network by using the source network and use adversarial domain adaptation to update the target network using both in vitro and in vivo data. These two steps help leverage the in vitro data to improve the prediction on in vivo data, which is typically challenging with a lower signal-to-noise ratio. Finally, to further take the advantage of the fused source and target data, we fine-tune the task predictor using both data. We showed that RBP-ADDA achieved better performance in modeling in vivo RBP binding data than other existing methods as judged by Pearson correlations. It also improved predictive performance on in vitro datasets. We further applied augmentation operations on RBPs with less in vivo data to expand the input data and showed that it can improve prediction performances. Lastly, we explored the predictive interpretability of RBP-ADDA, where we quantified the contribution of the input features by Integrated Gradients and identified nucleotide positions that are important for RBP recognition.
Ying Liu 0027, Ruihui Li, Jiawei Luo 0001, Zhaolei Zhang
PLoS Comput. Biol.4
2022 RNANetMotif: Identifying sequence-structure RNA network motifs in RNA-protein binding sites
abstract
RNA molecules can adopt stable secondary and tertiary structures, which are essential in mediating physical interactions with other partners such as RNA binding proteins (RBPs) and in carrying out their cellular functions. In vivo and in vitro experiments such as RNAcompete and eCLIP have revealed in vitro binding preferences of RBPs to RNA oligomers and in vivo binding sites in cells. Analysis of these binding data showed that the structure properties of the RNAs in these binding sites are important determinants of the binding events; however, it has been a challenge to incorporate the structure information into an interpretable model. Here we describe a new approach, RNANetMotif, which takes predicted secondary structure of thousands of RNA sequences bound by an RBP as input and uses a graph theory approach to recognize enriched subgraphs. These enriched subgraphs are in essence shared sequence-structure elements that are important in RBP-RNA binding. To validate our approach, we performed RNA structure modeling via coarse-grained molecular dynamics folding simulations for selected 4 RBPs, and RNA-protein docking for LIN28B. The simulation results, e.g., solvent accessibility and energetics, further support the biological relevance of the discovered network subgraphs.
Hongli Ma, Zhiyuan Xue, Zhaolei Zhang
PLoS Comput. Biol.5
2022 Development of a Safety Prediction Method for Arterial Roads Based on Big-Data Technology and Stacked AutoEncoder-Gated Recurrent Unit
abstract
Modern complexities associated with an arterial traffic makes existing safety prediction methods insufficient to meet desired standards required by recent developmental needs. This paper proposes an enhanced active safety prediction method based on big-data approach and Stacked AutoEncoder-Gated Recurrent Unit. Firstly, the big-data technology is used to construct a dynamic identification model to recognize real-time operation state and risk state. Secondly, the Stacked AutoEncoder-Gated Recurrent Unit is used to predict a level of safety based on associated recognition results. This paper uses data from working days of Sunset Boulevard, California, from January$1^{\mathrm{st}}$, 2020, to February$28^{\mathrm{th}}$, 2020. The results of analysis show that the accuracy of the proposed dynamic recognition model reaches 98.92%, which is better than existing models such as random forest, K-nearest neighbor, and naïve Bayes models. In addition, it is found that the Stacked AutoEncoder-Gated Recurrent Unit can achieve a prediction accuracy of 95.157% and has significant advantages in terms of efficiency. The proposed methods will provide feasible solutions for actively monitoring safety levels.
Wei Hao 0002, Donglei Rong, Zhaolei Zhang, Qiyu Wu 0003, Young-Ji Byon, Kefu Yi, Jinjun Tang, Nengchao Lyu
IEEE Trans. Intell. Transp. Syst.3
2021 Application and Challenges of Blockchain in Heterogeneous Identity Trust
Zhaolei Zhang, Guishan Dong, Junyan Lin
BlockSys1
2021 ProTICS reveals prognostic impact of tumor infiltrating immune cells in different molecular subtypes
abstract
Different subtypes of the same cancer often show distinct genomic signatures and require targeted treatments. The differences at the cellular and molecular levels of tumor microenvironment in different cancer subtypes have significant effects on tumor pathogenesis and prognostic outcomes. Although there have been significant researches on the prognostic association of tumor infiltrating lymphocytes in selected histological subtypes, few investigations have systemically reported the prognostic impacts of immune cells in molecular subtypes, as quantified by machine learning approaches on multi-omics datasets. This paper describes a new computational framework, ProTICS, to quantify the differences in the proportion of immune cells in tumor microenvironment and estimate their prognostic effects in different subtypes. First, we stratified patients into molecular subtypes based on gene expression and methylation profiles by applying nonnegative tensor factorization technique. Then we quantified the proportion of cell types in each specimen using an mRNA-based deconvolution method. For tumors in each subtype, we estimated the prognostic effects of immune cell types by applying Cox proportional hazard regression. At the molecular level, we also predicted the prognosis of signature genes for each subtype. Finally, we benchmarked the performance of ProTICS on three TCGA datasets and another independent METABRIC dataset. ProTICS successfully stratified tumors into different molecular subtypes manifested by distinct overall survival. Furthermore, the different immune cell types showed distinct prognostic patterns with respect to molecular subtypes. This study provides new insights into the prognostic association between immune cells and molecular subtypes, showing the utility of immune cells as potential prognostic markers. Availability: R code is available at https://github.com/liu-shuhui/ProTICS.
Shuhui Liu, Xuequn Shang 0001, Zhaolei Zhang
Briefings Bioinform.4
2021 Improving domain adaptation in de-identification of electronic health records through self-training
abstract
OBJECTIVE: De-identification is a fundamental task in electronic health records to remove protected health information entities. Deep learning models have proven to be promising tools to automate de-identification processes. However, when the target domain (where the model is applied) is different from the source domain (where the model is trained), the model often suffers a significant performance drop, commonly referred to as domain adaptation issue. In de-identification, domain adaptation issues can make the model vulnerable for deployment. In this work, we aim to close the domain gap by leveraging unlabeled data from the target domain. MATERIALS AND METHODS: We introduce a self-training framework to address the domain adaptation issue by leveraging unlabeled data from the target domain. We validate the effectiveness on 4 standard de-identification datasets. In each experiment, we use a pair of datasets: labeled data from the source domain and unlabeled data from the target domain. We compare the proposed self-training framework with supervised learning that directly deploys the model trained on the source domain. RESULTS: In summary, our proposed framework improves the F1-score by 5.38 (on average) when compared with direct deployment. For example, using i2b2-2014 as the training dataset and i2b2-2006 as the test, the proposed framework increases the F1-score from 76.61 to 85.41 (+8.8). The method also increases the F1-score by 10.86 for mimic-radiology and mimic-discharge. CONCLUSION: Our work demonstrates an effective self-training framework to boost the domain adaptation performance for the de-identification task for electronic health records.
Shun Liao, Jamie Kiros, Jiyang Chen, Zhaolei Zhang
J. Am. Medical Informatics Assoc.4
2021 Firmware code instrumentation technology for internet of things-based services
abstract
Abstract With the rapid development of electronic and information technology, Internet of Things (IoT) devices have become extensively utilised in various fields. Increasing attention has been paid to the performance and security analysis of IoT-based services. Dynamic instrumentation is a common process in software analysis for acquiring runtime information. However, due to the limited software and hardware resources in IoT devices, most dynamic instrumentation tools do not support IoT-based services. In this paper, we provide an analysis tool, IoTDIT, to solve the current problem of runtime detection in IoT-based services. IoTDIT employs static analysis andptracesystem calls to obtain dynamic firmware information, which can aid in firmware performance analysis and security detection. We perform experiments to verify the performance and effectiveness of the proposed instrumentation tool.
Chen Chen 0061, Jinxin Ma, Baojiang Cui, Weikong Qi, Zhaolei Zhang
World Wide Web6
2019 An Open Identity Authentication Scheme Based on Blockchain
Guishan Dong, Yao Hao, Zhaolei Zhang, Haiyang Peng, Shui Yu 0001
ICA3PP (1)4
2019 Tiger Tally: Cross-Domain Scheme for Different Authentication Mechanism
Guishan Dong, Yao Hao, Zhaolei Zhang, Shui Yu 0001
ICA3PP (1)4
2015 A novel motif-discovery algorithm to identify co-regulatory motifs in large transcription factor and microRNA co-regulatory networks in human
abstract
MOTIVATION: Interplays between transcription factors (TFs) and microRNAs (miRNAs) in gene regulation are implicated in various physiological processes. It is thus important to identify biologically meaningful network motifs involving both types of regulators to understand the key co-regulatory mechanisms underlying the cellular identity and function. However, existing motif finders do not scale well for large networks and are not designed specifically for co-regulatory networks. RESULTS: In this study, we propose a novel algorithm CoMoFinder to accurately and efficiently identify composite network motifs in genome-scale co-regulatory networks. We define composite network motifs as network patterns involving at least one TF, one miRNA and one target gene that are statistically significant than expected. Using two published disease-related co-regulatory networks, we show that CoMoFinder outperforms existing methods in both accuracy and robustness. We then applied CoMoFinder to human TF-miRNA co-regulatory network derived from The Encyclopedia of DNA Elements project and identified 44 recurring composite network motifs of size 4. The functional analysis revealed that genes involved in the 44 motifs are enriched for significantly higher number of biological processes or pathways comparing with non-motifs. We further analyzed the identified composite bi-fan motif and showed that gene pairs involved in this motif structure tend to physically interact and are functionally more similar to each other than expected. AVAILABILITY AND IMPLEMENTATION: CoMoFinder is implemented in Java and available for download at http://www.cs.utoronto.ca/∼yueli/como.html.
Cheng Liang 0001, Yue Li 0017, Jiawei Luo 0001, Zhaolei Zhang
Bioinform.4
2015 SignalSpider: probabilistic pattern discovery on multiple normalized ChIP-Seq signal profiles
abstract
MOTIVATION: Chromatin immunoprecipitation (ChIP) followed by high-throughput sequencing (ChIP-Seq) measures the genome-wide occupancy of transcription factors in vivo. Different combinations of DNA-binding protein occupancies may result in a gene being expressed in different tissues or at different developmental stages. To fully understand the functions of genes, it is essential to develop probabilistic models on multiple ChIP-Seq profiles to decipher the combinatorial regulatory mechanisms by multiple transcription factors. RESULTS: In this work, we describe a probabilistic model (SignalSpider) to decipher the combinatorial binding events of multiple transcription factors. Comparing with similar existing methods, we found SignalSpider performs better in clustering promoter and enhancer regions. Notably, SignalSpider can learn higher-order combinatorial patterns from multiple ChIP-Seq profiles. We have applied SignalSpider on the normalized ChIP-Seq profiles from the ENCODE consortium and learned model instances. We observed different higher-order enrichment and depletion patterns across sets of proteins. Those clustering patterns are supported by Gene Ontology (GO) enrichment, evolutionary conservation and chromatin interaction enrichment, offering biological insights for further focused studies. We also proposed a specific enrichment map visualization method to reveal the genome-wide transcription factor combinatorial patterns from the models built, which extend our existing fine-scale knowledge on gene regulation to a genome-wide level. AVAILABILITY AND IMPLEMENTATION: The matrix-algebra-optimized executables and source codes are available at the authors' websites: http://www.cs.toronto.edu/∼wkc/SignalSpider.
Ka-Chun Wong, Yue Li 0017, Chengbin Peng 0001, Zhaolei Zhang
Bioinform.4
2014 A probabilistic approach to explore human miRNA targetome by integrating miRNA-overexpression data and sequence information
abstract
MOTIVATION: Systematic identification of microRNA (miRNA) targets remains a challenge. The miRNA overexpression coupled with genome-wide expression profiling is a promising new approach and calls for a new method that integrates expression and sequence information. RESULTS: We developed a probabilistic scoring method called targetScore. TargetScore infers miRNA targets as the transformed fold-changes weighted by the Bayesian posteriors given observed target features. To this end, we compiled 84 datasets from Gene Expression Omnibus corresponding to 77 human tissue or cells and 113 distinct transfected miRNAs. Comparing with other methods, targetScore achieves significantly higher accuracy in identifying known targets in most tests. Moreover, the confidence targets from targetScore exhibit comparable protein downregulation and are more significantly enriched for Gene Ontology terms. Using targetScore, we explored oncomir-oncogenes network and predicted several potential cancer-related miRNA-messenger RNA interactions. AVAILABILITY AND IMPLEMENTATION: TargetScore is available at Bioconductor: http://www.bioconductor.org/packages/devel/bioc/html/TargetScore.html.
Yue Li 0017, Anna Goldenberg, Ka-Chun Wong, Zhaolei Zhang
Bioinform.4
2014 Mirsynergy: detecting synergistic miRNA regulatory modules by overlapping neighbourhood expansion
abstract
MOTIVATION: Identification of microRNA regulatory modules (MiRMs) will aid deciphering aberrant transcriptional regulatory network in cancer but is computationally challenging. Existing methods are stochastic or require a fixed number of regulatory modules. RESULTS: We propose Mirsynergy, an efficient deterministic overlapping clustering algorithm adapted from a recently developed framework. Mirsynergy operates in two stages: it first forms MiRMs based on co-occurring microRNA (miRNA) targets and then expands each MiRM by greedily including (excluding) mRNAs into (from) the MiRM to maximize the synergy score, which is a function of miRNA-mRNA and gene-gene interactions. Using expression data for ovarian, breast and thyroid cancer from The Cancer Genome Atlas, we compared Mirsynergy with internal controls and existing methods. Mirsynergy-MiRMs exhibit significantly higher functional enrichment and more coherent miRNA-mRNA expression anti-correlation. Based on Kaplan-Meier survival analysis, we proposed several prognostically promising MiRMs and envisioned their utility in cancer research. AVAILABILITY AND IMPLEMENTATION: Mirsynergy is implemented/available as an R/Bioconductor package at www.cs.utoronto.ca/∼yueli/Mirsynergy.html.
Yue Li 0017, Cheng Liang 0001, Ka-Chun Wong, Jiawei Luo 0001, Zhaolei Zhang
Bioinform.5
2014 SNPdryad: predicting deleterious non-synonymous human SNPs using only orthologous protein sequences
abstract
MOTIVATION: The recent advances in genome sequencing have revealed an abundance of non-synonymous polymorphisms among human individuals; subsequently, it is of immense interest and importance to predict whether such substitutions are functional neutral or have deleterious effects. The accuracy of such prediction algorithms depends on the quality of the multiple-sequence alignment, which is used to infer how an amino acid substitution is tolerated at a given position. Because of the scarcity of orthologous protein sequences in the past, the existing prediction algorithms all include sequences of protein paralogs in the alignment, which can dilute the conservation signal and affect prediction accuracy. However, we believe that, with the sequencing of a large number of mammalian genomes, it is now feasible to include only protein orthologs in the alignment and improve the prediction performance. RESULTS: We have developed a novel prediction algorithm, named SNPdryad, which only includes protein orthologs in building a multiple sequence alignment. Among many other innovations, SNPdryad uses different conservation scoring schemes and uses Random Forest as a classifier. We have tested SNPdryad on several datasets. We found that SNPdryad consistently outperformed other methods in several performance metrics, which is attributed to the exclusion of paralogous sequence. We have run SNPdryad on the complete human proteome, generating prediction scores for all the possible amino acid substitutions. AVAILABILITY AND IMPLEMENTATION: The algorithm and the prediction results can be accessed from the Web site: http://snps.ccbr.utoronto.ca:8080/SNPdryad/ CONTACT: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Ka-Chun Wong, Zhaolei Zhang
Bioinform.2
2014 Regression Analysis of Combined Gene Expression Regulation in Acute Myeloid Leukemia
abstract
Gene expression is a combinatorial function of genetic/epigenetic factors such as copy number variation (CNV), DNA methylation (DM), transcription factors (TF) occupancy, and microRNA (miRNA) post-transcriptional regulation. At the maturity of microarray/sequencing technologies, large amounts of data measuring the genome-wide signals of those factors became available from Encyclopedia of DNA Elements (ENCODE) and The Cancer Genome Atlas (TCGA). However, there is a lack of an integrative model to take full advantage of these rich yet heterogeneous data. To this end, we developed RACER (Regression Analysis of Combined Expression Regulation), which fits the mRNA expression as response using as explanatory variables, the TF data from ENCODE, and CNV, DM, miRNA expression signals from TCGA. Briefly, RACER first infers the sample-specific regulatory activities by TFs and miRNAs, which are then used as inputs to infer specific TF/miRNA-gene interactions. Such a two-stage regression framework circumvents a common difficulty in integrating ENCODE data measured in generic cell-line with the sample-specific TCGA measurements. As a case study, we integrated Acute Myeloid Leukemia (AML) data from TCGA and the related TF binding data measured in K562 from ENCODE. As a proof-of-concept, we first verified our model formalism by 10-fold cross-validation on predicting gene expression. We next evaluated RACER on recovering known regulatory interactions, and demonstrated its superior statistical power over existing methods in detecting known miRNA/TF targets. Additionally, we developed a feature selection procedure, which identified 18 regulators, whose activities clustered consistently with cytogenetic risk groups. One of the selected regulators is miR-548p, whose inferred targets were significantly enriched for leukemia-related pathway, implicating its novel role in AML pathogenesis. Moreover, survival analysis using the inferred activities identified C-Fos as a potential AML prognostic marker. Together, we provided a novel framework that successfully integrated the TCGA and ENCODE data in revealing AML-specific regulatory program at global level.
Yue Li 0017, Minggao Liang, Zhaolei Zhang
PLoS Comput. Biol.3
2012 Evolutionary multimodal optimization using the principle of locality
Ka-Chun Wong, Chun-Ho Wu, Ricky K. P. Mok, Chengbin Peng 0001, Zhaolei Zhang
Inf. Sci.5
2011 Differential effects of chromatin regulators and transcription factors on gene regulation: a nucleosomal perspective
abstract
MOTIVATION: Chromatin regulators (CR) and transcription factors (TF) are important trans-acting factors regulating transcription process, and many efforts have been devoted to understand their underlying mechanisms in gene regulation. However, the influences of CR and TF regulation effects on nucleosomes during transcription are still minimally understood, and it remains to be determined the extent to which CR and TF regulatory effect shape the organization of nucleosomes in the genome. In this article we attempted to address this problem and examine the patterns of CR and TF regulation effects from the nucleosome perspective. RESULTS: Our results show that the CR and TF regulatory effects exhibit different paradigms of transcriptional control in Saccharomyces cerevisiae. We grouped yeast genes into two categories, 'CR-sensitive' genes and 'TF-sensitive' genes, based on how their expression profiles change upon deletion of CRs or TFs. We found that genes in these two groups have very different patterns of nucleosome organization. The promoters of CR-sensitive genes tend to have higher nucleosome occupancy, whereas the promoters of TF-sensitive genes are depleted of nucleosomes. Furthermore, the nucleosome profiles of CR-sensitive genes tend to show more dynamic characteristics than TF-sensitive genes. These results reveal that the nucleosome organizations of yeast genes have a strong impact on their mode of regulation, and there are differential regulation effects on nucleosomes between CRs and TFs. AVAILABILITY: http://www.utoronto.ca/zhanglab/papers/bioinfo_2010/.
Xiaojian Shao, Zhaolei Zhang
Bioinform.3
2010 Deep Supervised t-Distributed Embedding
Martin Renqiang Min, Laurens van der Maaten, Zineng Yuan, Anthony J. Bonner, Zhaolei Zhang
ICML5
2010 Gene Expression Variability within and between Human Populations and Implications toward Disease Susceptibility
abstract
Variations in gene expression level might lead to phenotypic diversity across individuals or populations. Although many human genes are found to have differential mRNA levels between populations, the extent of gene expression that could vary within and between populations largely remains elusive. To investigate the dynamic range of gene expression, we analyzed the expression variability of ∼18, 000 human genes across individuals within HapMap populations. Although ∼20% of human genes show differentiated mRNA levels between populations, our results show that expression variability of most human genes in one population is not significantly deviant from another population, except for a small fraction that do show substantially higher expression variability in a particular population. By associating expression variability with sequence polymorphism, intriguingly, we found SNPs in the untranslated regions (5' and 3'UTRs) of these variable genes show consistently elevated population heterozygosity. We performed differential expression analysis on a genome-wide scale, and found substantially reduced expression variability for a large number of genes, prohibiting them from being differentially expressed between populations. Functional analysis revealed that genes with the greatest within-population expression variability are significantly enriched for chemokine signaling in HIV-1 infection, and for HIV-interacting proteins that control viral entry, replication, and propagation. This observation combined with the finding that known human HIV host factors show substantially elevated expression variability, collectively suggest that gene expression variability might explain differential HIV susceptibility across individuals.
Yu Liu 0093, TaeHyung Kim, Martin Renqiang Min, Zhaolei Zhang
PLoS Comput. Biol.5
2009 A Deep Non-linear Feature Mapping for Large-Margin kNN Classification
abstract
KNN is one of the most popular data mining methods for classification, but it often fails to work well with inappropriate choice of distance metric or due to the presence of numerous class-irrelevant features. Linear feature transformation methods have been widely applied to extract class-relevant information to improve kNN classification, which is very limited in many applications. Kernels have also been used to learn powerful non-linear feature transformations, but these methods fail to scale to large datasets. In this paper, we present a scalable non-linear feature mapping method based on a deep neural network pretrained with Restricted Boltzmann Machines for improving kNN classification in a large-margin framework, which we call DNet-kNN. DNet-kNN can be used for both classification and for supervised dimensionality reduction. The experimental results on two benchmark handwritten digit datasets and one newsgroup text dataset show that DNet-kNN has much better performance than large-margin kNN using a linear mapping and kNN based on a deep autoencoder pretrained with Restricted Boltzmann Machines.
Martin Renqiang Min, David A. Stanley, Zineng Yuan, Anthony J. Bonner, Zhaolei Zhang
ICDM5
2009 Learning Random-Walk Kernels for Protein Remote Homology Identification and Motif Discovery
abstract
Random-walk based algorithms are good choices for solving many classification problems with limited labeled data and a large amount of unlabeled data. However, it is difficult to choose the optimal number of random steps, and the results are very sensitive to the parameter chosen. In this paper, we will discuss how to better identify protein remote homology than any other algorithm using a learned random-walk kernel based on a positive linear combination of random-walk kernels with different random steps, which leads to a convex combination of kernels. The resulting kernel has much better prediction performance than the state-of-the-art profile kernel for protein remote homology identification. On the SCOP benchmark dataset, the overall mean ROC50 score on 54 protein families we obtained using the new kernel is above 0.90, which has almost perfect prediction performance on most of the 54 families and has significant improvement over the best published result; moreover, our approach based on learned random-walk kernels can effectively identify meaningful protein sequence motifs that are responsible for discriminating the memberships of protein sequences' remote homology in SCOP.
Martin Renqiang Min, Rui Kuang, Anthony J. Bonner, Zhaolei Zhang
SDM4
2009 Small RNAs Originated from Pseudogenes: cis- or trans-Acting?
abstract
Pseudogenes are significant components of eukaryotic genomes, and some have acquired novel regulatory roles. To date, no study has characterized rice pseudogenes systematically or addressed their impact on the structure and function of the rice genome. In this genome-wide study, we have identified 11,956 non-transposon-related rice pseudogenes, most of which are from gene duplications. About 12% of the rice protein-coding genes, half of which are in singleton families, have a pseudogene paralog. Interestingly, we found that 145 of these pseudogenes potentially gave rise to antisense small RNAs after examining approximately 1.5 million small RNAs from developing rice grains. The majority (>50%) of these antisense RNAs are 24-nucleotides long, a feature often seen in plant repeat-associated small interfering RNAs (siRNAs) produced by RNA-dependent RNA polymerase (RDR2) and Dicer-like protein 3 (DCL3), suggesting that some pseudogene-derived siRNAs may be implicated in repressing pseudogene transcription (i.e., cis-acting). Multiple lines of evidence, however, indicate that small RNAs from rice pseudogenes might also function as natural antisense siRNAs either by interacting with the complementary sense RNAs from functional parental genes (38 cases) or by forming double-strand RNAs with transcripts of adjacent paralogous pseudogenes (2 cases) (i.e., trans-acting). Further examinations of five additional small RNA libraries revealed that pseudogene-derived antisense siRNAs could be produced in specific rice developmental stages or physiological growth conditions, suggesting their potentially important roles in normal rice development. In summary, our results show that pseudogenes derived from protein-coding genes are prevalent in the rice genome, and a subset of them are strong candidates for producing small RNAs with novel regulatory roles. Our findings suggest that pseudogenes of exapted functions may be a phenomenon ubiquitous in eukaryotic organisms.
Xingyi Guo, Zhaolei Zhang, Mark Gerstein, Deyou Zheng
PLoS Comput. Biol.2
2008 A hybrid model for robust detection of transcription factor binding sites
abstract
MOTIVATION: The short and degenerate nature of transcription factor (TF) binding sites contributes towards a low signal to noise ratio making it very difficult to separate them from their background. In order to tackle this problem one needs to look at ways of capturing the underlying biophysical properties that best discriminates TF binding sites from their background DNA. One such discriminatory property lies in the observed compositional differences in the nucleotide levels of TF binding sites and background DNA which are a result of processes such as purifying selection and selective preferences of TF binding sites for particular nucleotides or a combination of nucleotides over others. RESULTS: In this article, we present a hybrid model, referred to as a MonoDi-nucleotide model for robustly detecting TF binding sites. It incorporates both mono- and dinucleotide statistics to optimally partition the base positions of an aligned set of TF binding sites (motif) into a non-redundant sequence of mono and/or dinucleotide segments that maximizes the odds ratio of the binding sites relative to their background DNA. We tested the MonoDi-nucleotide model on the benchmark dataset compiled by Tompa et al. (2005) for assessing computational tools that predict TF binding sites. The performance of the MonoDi-nucleotide model on this data set compares well to, and in many cases exceeds, the performance of existing tools. This is in part attributed to the significant role played by dinucleotides in discriminating TF binding sites from background DNA. AVAILABILITY: A Matlab implementation of the MonoDi-nucleotide model can be found at http://www.utoronto.ca/zhanglab/MonoDi/.
Sumedha Gunewardena, Zhaolei Zhang
Bioinform.2
2008 MAID : An effect size based model for microarray data integration across laboratories and platforms
abstract
BACKGROUND: Gene expression profiling has the potential to unravel molecular mechanisms behind gene regulation and identify gene targets for therapeutic interventions. As microarray technology matures, the number of microarray studies has increased, resulting in many different datasets available for any given disease. The increase in sensitivity and reliability of measurements of gene expression changes can be improved through a systematic integration of different microarray datasets that address the same or similar biological questions. RESULTS: Traditional effect size models can not be used to integrate array data that directly compare treatment to control samples expressed as log ratios of gene expressions. Here we extend the traditional effect size model to integrate as many array datasets as possible. The extended effect size model (MAID) can integrate any array datatype generated with either single or two channel arrays using either direct or indirect designs across different laboratories and platforms. The model uses two standardized indices, the standard effect size score for experiments with two groups of data, and a new standardized index that measures the difference in gene expression between treatment and control groups for one sample data with replicate arrays. The statistical significance of treatment effect across studies for each gene is determined by appropriate permutation methods depending on the type of data integrated. We apply our method to three different expression datasets from two different laboratories generated using three different array platforms and two different experimental designs. Our results indicate that the proposed integration model produces an increase in statistical power for identifying differentially expressed genes when integrating data across experiments and when compared to other integration models. We also show that genes found to be significant using our data integration method are of direct biological relevance to the three experiments integrated. CONCLUSION: High-throughput genomics data provide a rich and complex source of information that could play a key role in deciphering intricate molecular networks behind disease. Here we propose an extension of the traditional effect size model to allow the integration of as many array experiments as possible with the aim of increasing the statistical power for identifying differentially expressed genes.
Ivan Borozan, Bryan Paeper, Jenny E. Heathcote, Aled M. Edwards, Michael G. Katze, Zhaolei Zhang, Ian D. McGilvray
BMC Bioinform.7
2007 Modifying kernels using label information improves SVM classification performance
abstract
Kernel learning methods based on kernel alignment with semidefinite programming (SDP) are often memory intensive and computationally expensive, thus often impractical for problems with large-size dataset. We propose a method using label information to modify kernels based on SVD and a linear mapping. As a result, the new kernel matrix reflects the label-dependent separability of the data in a better way than the original kernel matrix. In addition, our experimental results on USPS handwritten digits and the SCOP dataset, show that the SVM classifier based on the improved kernels has better performance than the SVM classifier based on the original kernels; moreover, SVM based on the improved profile kernel with pull-in homologs (see experiment section for explanations) produced the best results for remote homology detection on the SCOP dataset compared to the published results.
Martin Renqiang Min, Anthony J. Bonner, Zhaolei Zhang
ICMLA3
2006 PseudoPipe: an automated pseudogene identification pipeline
abstract
MOTIVATION: Mammalian genomes contain many 'genomic fossils' i.e. pseudogenes. These are disabled copies of functional genes that have been retained in the genome by gene duplication or retrotransposition events. Pseudogenes are important resources in understanding the evolutionary history of genes and genomes. RESULTS: We have developed a homology-based computational pipeline ('PseudoPipe') that can search a mammalian genome and identify pseudogene sequences in a comprehensive and consistent manner. The key steps in the pipeline involve using BLAST to rapidly cross-reference potential "parent" proteins against the intergenic regions of the genome and then processing the resulting "raw hits" -- i.e. eliminating redundant ones, clustering together neighbors, and associating and aligning clusters with a unique parent. Finally, pseudogenes are classified based on a combination of criteria including homology, intron-exon structure, and existence of stop codons and frameshifts.
Zhaolei Zhang, Nicholas Carriero, Deyou Zheng, John E. Karro, Paul M. Harrison, Mark Gerstein
Bioinform.1