Michael Q. Zhang

dblp:00/1997 · DBLP profile ↗
← Back
39ranked-venue papers
6as first author
1since 2021 · last 2021
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 36 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
19 papers
Bioinformatics and computational biology · 100%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 77% Image and video processing · 23%

Topics — the 30 heaviest of 35, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
genomics
0.632018
CGmapTools improves the precision of heterozygous SNV calls and supports allele-specific methylation detection and visualization in bisulfite-sequencing data · Bioinform. 2018
MICC: an R package for identifying chromatin interactions from ChIA-PET data · Bioinform. 2015
Analysis and visualization of DNA spectrograms: open possibilities for the genome research · ACM Multimedia 2006
Bioinformatics and computational biology › epigenomics › DNA methylation
DNA methylation prediction
0.322017
DIRECTION: a machine learning framework for predicting and characterizing DNA methylation and hydroxymethylation in mammalian genomes · Bioinform. 2017
Predicting methylation status of CpG islands in the human brain · Bioinform. 2006
Bioinformatics and computational biology › epigenomics › DNA methylation
allele-specific methylation
0.312018
CGmapTools improves the precision of heterozygous SNV calls and supports allele-specific methylation detection and visualization in bisulfite-sequencing data · Bioinform. 2018
Bioinformatics and computational biology › epigenomics › differential methylation analysis
differentially methylated region detection
0.312018
CGmapTools improves the precision of heterozygous SNV calls and supports allele-specific methylation detection and visualization in bisulfite-sequencing data · Bioinform. 2018
Bioinformatics and computational biology › epigenomics › DNA methylation
DNA methylation analysis
0.312018
CGmapTools improves the precision of heterozygous SNV calls and supports allele-specific methylation detection and visualization in bisulfite-sequencing data · Bioinform. 2018
Bioinformatics and computational biology › genomics
variant calling
0.312018
CGmapTools improves the precision of heterozygous SNV calls and supports allele-specific methylation detection and visualization in bisulfite-sequencing data · Bioinform. 2018
Bioinformatics and computational biology
epigenomics
0.312017
DIRECTION: a machine learning framework for predicting and characterizing DNA methylation and hydroxymethylation in mammalian genomes · Bioinform. 2017
Bioinformatics and computational biology › epigenomics
chromatin interaction detection
0.212015
MICC: an R package for identifying chromatin interactions from ChIA-PET data · Bioinform. 2015
Bioinformatics and computational biology
sequence analysis
0.232009
Updates to the RMAP short-read mapping software · Bioinform. 2009
ZOOM! Zillions of oligos mapped · Bioinform. 2008
A weight array method for splicing signal analysis · Comput. Appl. Biosci. 1993
Bioinformatics and computational biology › gene regulation › transcription factor binding site prediction
transcription factor binding site analysis
0.222011
Correlated evolution of transcription factors and their binding sites · Bioinform. 2011
Similarity of position frequency matrices for transcription factor binding sites · Bioinform. 2005
Bioinformatics and computational biology › sequence analysis › read mapping
short read alignment
0.222009
Updates to the RMAP short-read mapping software · Bioinform. 2009
ZOOM! Zillions of oligos mapped · Bioinform. 2008
Bioinformatics and computational biology › gene regulation
regulatory genomics
0.222008
Identification of phylogenetically conserved microRNA cis-regulatory elements across 12 Drosophila species · Bioinform. 2008
OSCAR: One-class SVM for accurate recognition of cis-elements · Bioinform. 2007
Bioinformatics and computational biology › gene regulation
transcription factor binding site prediction
0.122007
OSCAR: One-class SVM for accurate recognition of cis-elements · Bioinform. 2007
DWE: Discriminating Word Enumerator · Bioinform. 2005
Bioinformatics and computational biology › transcriptomics › alternative splicing analysis
alternative splicing quantification
0.112011
SpliceTrap: a method to quantify alternative splicing under single cellular conditions · Bioinform. 2011
Bioinformatics and computational biology
gene regulation
0.112011
Correlated evolution of transcription factors and their binding sites · Bioinform. 2011
Bioinformatics and computational biology
transcriptomics
0.112011
SpliceTrap: a method to quantify alternative splicing under single cellular conditions · Bioinform. 2011
Bioinformatics and computational biology › sequence analysis › read mapping
bisulfite-treated read mapping
0.112009
Updates to the RMAP short-read mapping software · Bioinform. 2009
Bioinformatics and computational biology › gene regulation › regulatory element discovery
cis-regulatory element identification
0.112008
Identification of phylogenetically conserved microRNA cis-regulatory elements across 12 Drosophila species · Bioinform. 2008
Bioinformatics and computational biology › phylogenetics
phylogenetic footprinting
0.112008
Identification of phylogenetically conserved microRNA cis-regulatory elements across 12 Drosophila species · Bioinform. 2008
Bioinformatics and computational biology › sequence analysis
read mapping
0.112008
ZOOM! Zillions of oligos mapped · Bioinform. 2008
Bioinformatics and computational biology › gene regulation › gene regulatory network
gene regulatory network analysis
0.112007
Computational prediction of novel components of lung transcriptional networks · Bioinform. 2007
Bioinformatics and computational biology › sequence analysis › motif discovery
motif analysis
0.112007
Computing exact P-values for DNA motifs · Bioinform. 2007
Bioinformatics and computational biology › sequence analysis
motif discovery
0.112007
Computational prediction of novel components of lung transcriptional networks · Bioinform. 2007
Bioinformatics and computational biology › gene regulation › transcription factor analysis
transcription factor prediction
0.112007
Computational prediction of novel components of lung transcriptional networks · Bioinform. 2007
Bioinformatics and computational biology › sequence analysis
DNA sequence analysis
0.112006
Analysis and visualization of DNA spectrograms: open possibilities for the genome research · ACM Multimedia 2006
Visualization and visual analytics
scientific visualization
0.112006
Analysis and visualization of DNA spectrograms: open possibilities for the genome research · ACM Multimedia 2006
Bioinformatics and computational biology › genome annotation
gene prediction
0.012001
Identifying the 3'-terminal exon in human DNA · Bioinform. 2001
Bioinformatics and computational biology › sequence analysis
genomic sequence analysis
0.012001
Promoter Extraction from GenBank (PEG): automatic extraction of eukaryotic promoter sequences in large sets of genes · Bioinform. 2001
Bioinformatics and computational biology › biological database
promoter database
0.011999
SCPD: a promoter database of the yeast Saccharomyces cerevisiae · Bioinform. 1999
Image and video processing
spectral analysis
0.012006
Analysis and visualization of DNA spectrograms: open possibilities for the genome research · ACM Multimedia 2006

Methods — techniques the papers use, named apart from their topics

bayesian inference · 0.5binomial test · 0.3supervised learning · 0.3feature selection · 0.3beam search · 0.3bayesian mixture model · 0.2paired-end RNA-seq · 0.1mutual information · 0.1quality score integration · 0.1paired-end mapping · 0.1dynamic programming · 0.1spectral analysis · 0.1pattern identification · 0.1image processing · 0.1
YearPublicationVenuePosition
2021 Deciphering hierarchical organization of topologically associated domains through change-point testing
abstract
BACKGROUND: The nucleus of eukaryotic cells spatially packages chromosomes into a hierarchical and distinct segregation that plays critical roles in maintaining transcription regulation. High-throughput methods of chromosome conformation capture, such as Hi-C, have revealed topologically associating domains (TADs) that are defined by biased chromatin interactions within them. RESULTS: We introduce a novel method, HiCKey, to decipher hierarchical TAD structures in Hi-C data and compare them across samples. We first derive a generalized likelihood-ratio (GLR) test for detecting change-points in an interaction matrix that follows a negative binomial distribution or general mixture distribution. We then employ several optimal search strategies to decipher hierarchical TADs with p values calculated by the GLR test. Large-scale validations of simulation data show that HiCKey has good precision in recalling known TADs and is robust against random collisions of chromatin interactions. By applying HiCKey to Hi-C data of seven human cell lines, we identified multiple layers of TAD organization among them, but the vast majority had no more than four layers. In particular, we found that TAD boundaries are significantly enriched in active chromosomal regions compared to repressed regions. CONCLUSIONS: HiCKey is optimized for processing large matrices constructed from high-resolution Hi-C experiments. The method and theoretical result of the GLR test provide a general framework for significance testing of similar experimental chromatin interaction data that may not fully follow negative binomial distributions but rather more general mixture distributions.
Haipeng Xing, Yingru Wu, Michael Q. Zhang, Yong Chen 0031
BMC Bioinform.3
2020 2SigFinder: the combined use of small-scale and large-scale statistical testing for genomic island detection from a single genome
abstract
BACKGROUND: Genomic islands are associated with microbial adaptations, carrying genomic signatures different from the host. Some methods perform an overall test to identify genomic islands based on their local features. However, regions of different scales will display different genomic features. RESULTS: We proposed here a novel method "2SigFinder ", the first combined use of small-scale and large-scale statistical testing for genomic island detection. The proposed method was tested by genomic island boundary detection and identification of genomic islands or functional features of real biological data. We also compared the proposed method with the comparative genomics and composition-based approaches. The results indicate that the proposed 2SigFinder is more efficient in identifying genomic islands. CONCLUSIONS: From real biological data, 2SigFinder identified genomic islands from a single genome and reported robust results across different experiments, without annotated information of genomes or prior knowledge from other datasets. 2SigHunter identified 25 Pathogenicity, 1 tRNA, 2 Virulence and 2 Repeats from 27 Pathogenicity, 1 tRNA, 2 Virulence and 2 Repeats, and detected 101 Phage and 28 HEG out of 130 Phage and 36 HEGs in S. enterica Typhi CT18, which shows that it is more efficient in detecting functional features associated with GIs.
Xinnan Xu, Ping-An He 0001, Michael Q. Zhang
BMC Bioinform.5
2018 DE MERVLs are Enriched Around Two-Cell-Specific Genes During Zygotic Genome Activation in Mouse
abstract
ZGA(zygotic genome activation) is not only important for early embryonic development, but also sheds light on understanding totipotency acquisition of mammalian cells. Recent studies has revealed that DUX severs as an important pioneering factor for mouse ZGA initiation, and MERVL(mouse ERV-like element) plays an important role to activate the downstream genes of DUX. However, how MERVLs functions during ZGA remains mysterious. Here we investigated two important aspects of MERVLs regulation, the characteristics of MERVLs regulation, and why DUX4 is needed to activate downstream genes via MERVLs. Then we proposed a ZGA-initiating model supported by Zscan4 case study, to understand and appreciate the significance of MERVLs in ZGA process.
Yisi Li, Michael Q. Zhang, Juntao Gao
SMC2
2018 MTGIpick allows robust identification of genomic islands from a single genome
abstract
Genomic islands (GIs) that are associated with microbial adaptations and carry sequence patterns different from that of the host are sporadically distributed among closely related species. This bias can dominate the signal of interest in GI detection. However, variations still exist among the segments of the host, although no uniform standard exists regarding the best methods of discriminating GIs from the rest of the genome in terms of compositional bias. In the present work, we proposed a robust software, MTGIpick, which used regions with pattern bias showing multiscale difference levels to identify GIs from the host. MTGIpick can identify GIs from a single genome without annotated information of genomes or prior knowledge from other data sets. When real biological data were used, MTGIpick demonstrated better performance than existing methods, as well as revealed potential GIs with accurate sizes missed by existing methods because of a uniform standard. Software and supplementary are freely available at http://bioinfo.zstu.edu.cn/MTGI or https://github.com/bioinfo0706/MTGIpick.
Chaohui Bao, Yabing Hai Hai, Sheng Ma, Wenwen Huo, Yuhua Yao, Zhenyu Xuan, Min Chen 0014, Michael Q. Zhang
Briefings Bioinform.13
2018 CGmapTools improves the precision of heterozygous SNV calls and supports allele-specific methylation detection and visualization in bisulfite-sequencing data
abstract
Motivation: DNA methylation is important for gene silencing and imprinting in both plants and animals. Recent advances in bisulfite sequencing allow detection of single nucleotide variations (SNVs) achieving high sensitivity, but accurately identifying heterozygous SNVs from partially C-to-T converted sequences remains challenging. Results: We designed two methods, BayesWC and BinomWC, that substantially improved the precision of heterozygous SNV calls from ∼80% to 99% while retaining comparable recalls. With these SNV calls, we provided functions for allele-specific DNA methylation (ASM) analysis and visualizing the methylation status on reads. Applying ASM analysis to a previous dataset, we found that an average of 1.5% of investigated regions showed allelic methylation, which were significantly enriched in transposon elements and likely to be shared by the same cell-type. A dynamic fragment strategy was utilized for DMR analysis in low-coverage data and was able to find differentially methylated regions (DMRs) related to key genes involved in tumorigenesis using a public cancer dataset. Finally, we integrated 40 applications into the software package CGmapTools to analyze DNA methylomes. This package uses CGmap as the format interface, and designs binary formats to reduce the file size and support fast data retrieval, and can be applied for context-wise, gene-wise, bin-wise, region-wise and sample-wise analyses and visualizations. Availability and implementation: The CGmapTools software is freely available at https://cgmaptools.github.io/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Weilong Guo, Ping Zhu 0002, Matteo Pellegrini, Michael Q. Zhang, Xiangfeng Wang 0003, Zhongfu Ni
Bioinform.4
2017 DIRECTION: a machine learning framework for predicting and characterizing DNA methylation and hydroxymethylation in mammalian genomes
abstract
MOTIVATION: 5-Methylcytosine and 5-Hydroxymethylcytosine in DNA are major epigenetic modifications known to significantly alter mammalian gene expression. High-throughput assays to detect these modifications are expensive, labor-intensive, unfeasible in some contexts and leave a portion of the genome unqueried. Hence, we devised a novel, supervised, integrative learning framework to perform whole-genome methylation and hydroxymethylation predictions in CpG dinucleotides. Our framework can also perform imputation of missing or low quality data in existing sequencing datasets. Additionally, we developed infrastructure to perform in silico, high-throughput hypotheses testing on such predicted methylation or hydroxymethylation maps. RESULTS: We test our approach on H1 human embryonic stem cells and H1-derived neural progenitor cells. Our predictive model is comparable in accuracy to other state-of-the-art DNA methylation prediction algorithms. We are the first to predict hydroxymethylation in silico with high whole-genome accuracy, paving the way for large-scale reconstruction of hydroxymethylation maps in mammalian model systems. We designed a novel, beam-search driven feature selection algorithm to identify the most discriminative predictor variables, and developed a platform for performing integrative analysis and reconstruction of the epigenome. Our toolkit DIRECTION provides predictions at single nucleotide resolution and identifies relevant features based on resource availability. This offers enhanced biological interpretability of results potentially leading to a better understanding of epigenetic gene regulation. AVAILABILITY AND IMPLEMENTATION: http://www.pradiptaray.com/direction, under CC-by-SA license. CONTACTS: [email protected] or [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Milos Pavlovic, Pradipta Ray, Kristina Pavlovic, Aaron Kotamarti, Min Chen 0014, Michael Q. Zhang
Bioinform.6
2015 MICC: an R package for identifying chromatin interactions from ChIA-PET data
abstract
UNLABELLED: ChIA-PET is rapidly emerging as an important experimental approach to detect chromatin long-range interactions at high resolution. Here, we present Model based Interaction Calling from ChIA-PET data (MICC), an easy-to-use R package to detect chromatin interactions from ChIA-PET sequencing data. By applying a Bayesian mixture model to systematically remove random ligation and random collision noise, MICC could identify chromatin interactions with a significantly higher sensitivity than existing methods at the same false discovery rate. AVAILABILITY AND IMPLEMENTATION: http://bioinfo.au.tsinghua.edu.cn/member/xwwang/MICCusage CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Michael Q. Zhang, Xiaowo Wang
Bioinform.2
2012 Computational Modeling of Mammalian Promoters - (Invited Keynote Talk)
Michael Q. Zhang
ISBRA1
2012 Genome-Wide Localization of Protein-DNA Binding and Histone Modification by a Bayesian Change-Point Method with ChIP-seq Data
abstract
Next-generation sequencing (NGS) technologies have matured considerably since their introduction and a focus has been placed on developing sophisticated analytical tools to deal with the amassing volumes of data. Chromatin immunoprecipitation sequencing (ChIP-seq), a major application of NGS, is a widely adopted technique for examining protein-DNA interactions and is commonly used to investigate epigenetic signatures of diffuse histone marks. These datasets have notoriously high variance and subtle levels of enrichment across large expanses, making them exceedingly difficult to define. Windows-based, heuristic models and finite-state hidden Markov models (HMMs) have been used with some success in analyzing ChIP-seq data but with lingering limitations. To improve the ability to detect broad regions of enrichment, we developed a stochastic Bayesian Change-Point (BCP) method, which addresses some of these unresolved issues. BCP makes use of recent advances in infinite-state HMMs by obtaining explicit formulas for posterior means of read densities. These posterior means can be used to categorize the genome into enriched and unenriched segments, as is customarily done, or examined for more detailed relationships since the underlying subpeaks are preserved rather than simplified into a binary classification. BCP performs a near exhaustive search of all possible change points between different posterior means at high-resolution to minimize the subjectivity of window sizes and is computationally efficient, due to a speed-up algorithm and the explicit formulas it employs. In the absence of a well-established "gold standard" for diffuse histone mark enrichment, we corroborated BCP's island detection accuracy and reproducibility using various forms of empirical evidence. We show that BCP is especially suited for analysis of diffuse histone ChIP-seq data but also effective in analyzing punctate transcription factor ChIP datasets, making it widely applicable for numerous experiment types.
Haipeng Xing, Yifan Mo, Will Liao, Michael Q. Zhang
PLoS Comput. Biol.4
2011 SpliceTrap: a method to quantify alternative splicing under single cellular conditions
abstract
MOTIVATION: Alternative splicing (AS) is a pre-mRNA maturation process leading to the expression of multiple mRNA variants from the same primary transcript. More than 90% of human genes are expressed via AS. Therefore, quantifying the inclusion level of every exon is crucial for generating accurate transcriptomic maps and studying the regulation of AS. RESULTS: Here we introduce SpliceTrap, a method to quantify exon inclusion levels using paired-end RNA-seq data. Unlike other tools, which focus on full-length transcript isoforms, SpliceTrap approaches the expression-level estimation of each exon as an independent Bayesian inference problem. In addition, SpliceTrap can identify major classes of alternative splicing events under a single cellular condition, without requiring a background set of reads to estimate relative splicing changes. We tested SpliceTrap both by simulation and real data analysis, and compared it to state-of-the-art tools for transcript quantification. SpliceTrap demonstrated improved accuracy, robustness and reliability in quantifying exon-inclusion ratios. CONCLUSIONS: SpliceTrap is a useful tool to study alternative splicing regulation, especially for accurate quantification of local exon-inclusion ratios from RNA-seq data. AVAILABILITY AND IMPLEMENTATION: SpliceTrap can be implemented online through the CSH Galaxy server http://cancan.cshl.edu/splicetrap and is also available for download and installation at http://rulai.cshl.edu/splicetrap/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Martin Akerman, Shuying Sun, W. Richard McCombie, Adrian R. Krainer, Michael Q. Zhang
Bioinform.6
2011 Correlated evolution of transcription factors and their binding sites
abstract
MOTIVATION: The interaction between transcription factor (TF) and transcription factor binding site (TFBS) is essential for gene regulation. Mutation in either the TF or the TFBS may weaken their interaction and thus result in abnormalities. To maintain such vital interaction, a mutation in one of the interacting partners might be compensated by a corresponding mutation in its binding partner during the course of evolution. Confirming this co-evolutionary relationship will guide us in designing protein sequences to target a specific DNA sequence or in predicting TFBS for poorly studied proteins, or even correcting and rescuing disease mutations in clinical applications. RESULTS: Based on six, publicly available, experimentally validated TF-TFBS binding datasets for the basic Helix-Loop-Helix (bHLH) family, Homeo family, High-Mobility Group (HMG) family and Transient Receptor Potential channels (TRP) family, we showed that the evolutions of the TFs and their TFBSs are significantly correlated across eukaryotes. We further developed a mutual information-based method to identify co-evolved protein residues and DNA bases. This research sheds light on the dynamic relationship between TF and TFBS during their evolution. The same principle and strategy can be applied to co-evolutionary studies on protein-DNA interactions in other protein families. AVAILABILITY: All the datasets, scripts and other related files have been made freely available at: http://jjwanglab.org/co-evo. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shu Yang 0009, Hari Krishna Yalamanchili, Kwok-Ming Yao, Pak Chung Sham, Michael Q. Zhang, Junwen Wang
Bioinform.6
2011 Histone modification profiles are predictive for tissue/cell-type specific expression of both protein-coding and microRNA genes
abstract
BACKGROUND: Gene expression is regulated at both the DNA sequence level and through modification of chromatin. However, the effect of chromatin on tissue/cell-type specific gene regulation (TCSR) is largely unknown. In this paper, we present a method to elucidate the relationship between histone modification/variation (HMV) and TCSR. RESULTS: A classifier for differentiating CD4+ T cell-specific genes from housekeeping genes using HMV data was built. We found HMV in both promoter and gene body regions to be predictive of genes which are targets of TCSR. For example, the histone modification types H3K4me3 and H3K27ac were identified as the most predictive for CpG-related promoters, whereas H3K4me3 and H3K79me3 were the most predictive for nonCpG-related promoters. However, genes targeted by TCSR can be predicted using other type of HMVs as well. Such redundancy implies that multiple type of underlying regulatory elements, such as enhancers or intragenic alternative promoters, which can regulate gene expression in a tissue/cell-type specific fashion, may be marked by the HMVs. Finally, we show that the predictive power of HMV for TCSR is not limited to protein-coding genes in CD4+ T cells, as we successfully predicted TCSR targeted genes in muscle cells, as well as microRNA genes with expression specific to CD4+ T cells, by the same classifier which was trained on HMV data of protein-coding genes in CD4+ T cells. CONCLUSION: We have begun to understand the HMV patterns that guide gene expression in both tissue/cell-type specific and ubiquitous manner.
Michael Q. Zhang
BMC Bioinform.2
2010 SFSSClass: an integrated approach for miRNA based tumor classification
abstract
BACKGROUND: MicroRNA (miRNA) expression profiling data has recently been found to be particularly important in cancer research and can be used as a diagnostic and prognostic tool. Current approaches of tumor classification using miRNA expression data do not integrate the experimental knowledge available in the literature. A judicious integration of such knowledge with effective miRNA and sample selection through a biclustering approach could be an important step in improving the accuracy of tumor classification. RESULTS: In this article, a novel classification technique called SFSSClass is developed that judiciously integrates a biclustering technique SAMBA for simultaneous feature (miRNA) and sample (tissue) selection (SFSS), a cancer-miRNA network that we have developed by mining the literature of experimentally verified cancer-miRNA relationships and a classifier uncorrelated shrunken centroid (USC). SFSSClass is used for classifying multiple classes of tumors and cancer cell lines. In a part of the investigation, poorly differentiated tumors (PDT) having non diagnostic histological appearance are classified while training on more differentiated tumor (MDT) samples. The proposed method is found to outperform the best known accuracy in the literature on the experimental data sets. For example, while the best accuracy reported in the literature for classifying PDT samples is approximately 76.5%, the accuracy of SFSSClass is found to be approximately 82.3%. The advantage of incorporating biclustering integrated with the cancer-miRNA network is evident from the consistently better performance of SFSSClass (integration of SAMBA, cancer-miRNA network and USC) over USC (eg., approximately 70.5% for SFSSClass versus approximately 58.8% in classifying a set of 17 MDT samples from 9 tumor types, approximately 91.7% for SFSSClass versus approximately 75% in classifying 12 cell lines from 6 tumor types and approximately 82.3% for SFSSClass versus approximately 41.2% in classifying 17 PDT samples from 11 tumor types). CONCLUSION: In this article, we develop the SFSSClass algorithm which judiciously integrates a biclustering technique for simultaneous feature (miRNA) and sample (tissue) selection, the cancer-miRNA network and a classifier. The novel integration of experimental knowledge with computational tools efficiently selects relevant features that have high intra-class and low inter-class similarity. The performance of the SFSSClass is found to be significantly improved with respect to the other existing approaches.
Ramkrishna Mitra, Sanghamitra Bandyopadhyay, Ujjwal Maulik, Michael Q. Zhang
BMC Bioinform.4
2009 Updates to the RMAP short-read mapping software
abstract
Abstract Summary: We report on a major new version of the RMAP software for mapping reads from short-read sequencing technology. General improvements to accuracy and space requirements are included, along with novel functionality. Included in the RMAP software package are tools for mapping paired-end reads, mapping using more sophisticated use of quality scores, collecting ambiguous mapping locations and mapping bisulfite-treated reads. Availability: The applications described in this note are available for download at http://www.cmb.usc.edu/people/andrewds/rmap and are distributed as Open Source software under the GPLv3.0. The software has been tested on Linux and OS X platforms. Contact: [email protected]; [email protected] The RMAP algorithm was introduced by (Smith et al., 2008) as one of the earliest available programs for mapping reads from the Illumina second-generation sequencing technology. One important contribution of RMAP was to incorporate the use of quality scores directly into the mapping process: read positions with too low a quality score were not considered while mapping, and that quality score cutoff could be adjusted by the user. Subsequently, numerous mapping algorithm have appeared (Langmead et al., 2009; Li,H. et al., 2008; Li,R. et al., 2008; Lin et al., 2008; Schatz, 2009; Yanovsky et al., 2008), with improvements in both efficiency and breadth of functionality (e.g. ability to map paired-end reads; integrated SNP calling). Investigators requiring solutions to mapping problems now have many options. As new applications of short-read sequencing emerge, many variations on the analysis task of read mapping emerge. Diversity in performance characteristics of existing mapping tools becomes potentially valuable. We report the first major update to RMAP. The basic algorithmic framework in RMAP is still to preprocess reads and scan the genome, but several modifications have been made and much additional functionality has been included. Importantly, RMAP has a memory footprint that depends on the number of reads being mapped. This feature allows RMAP to be used effectively in cluster environments with commodity nodes, because partitioning the reads allows natural parallelizations with linear reduction in memory requirements per processor core used. Included in this release of the RMAP software package is functionality for mapping paired-end reads, making more sophisticated use of quality scores, collecting mapping locations for ambiguously mapping reads and mapping bisulfite-treated reads.
Andrew D. Smith, Wen-Yu Chung, Emily Hodges, Jude Kendall, Gregory J. Hannon, James Hicks, Zhenyu Xuan, Michael Q. Zhang
Bioinform.8
2009 Gene set-based module discovery in the breast cancer transcriptome
abstract
BACKGROUND: Although microarray-based studies have revealed global view of gene expression in cancer cells, we still have little knowledge about regulatory mechanisms underlying the transcriptome. Several computational methods applied to yeast data have recently succeeded in identifying expression modules, which is defined as co-expressed gene sets under common regulatory mechanisms. However, such module discovery methods are not applied cancer transcriptome data. RESULTS: In order to decode oncogenic regulatory programs in cancer cells, we developed a novel module discovery method termed EEM by extending a previously reported module discovery method, and applied it to breast cancer expression data. Starting from seed gene sets prepared based on cis-regulatory elements, ChIP-chip data, and gene locus information, EEM identified 10 principal expression modules in breast cancer based on their expression coherence. Moreover, EEM depicted their activity profiles, which predict regulatory programs in each subtypes of breast tumors. For example, our analysis revealed that the expression module regulated by the Polycomb repressive complex 2 (PRC2) is downregulated in triple negative breast cancers, suggesting similarity of transcriptional programs between stem cells and aggressive breast cancer cells. We also found that the activity of the PRC2 expression module is negatively correlated to the expression of EZH2, a component of PRC2 which belongs to the E2F expression module. E2F-driven EZH2 overexpression may be responsible for the repression of the PRC2 expression modules in triple negative tumors. Furthermore, our network analysis predicts regulatory circuits in breast cancer cells. CONCLUSION: These results demonstrate that the gene set-based module discovery approach is a powerful tool to decode regulatory programs in cancer cells.
Atsushi Niida, Andrew D. Smith, Seiya Imoto, Hiroyuki Aburatani, Michael Q. Zhang, Tetsu Akiyama
BMC Bioinform.5
2009 The Seventh Asia Pacific Bioinformatics Conference (APBC2009)
abstract
The Asia Pacific Bioinformatics Conference (APBC) series, founded in 2003, is an annual international forum for exploring research, development and applications of Bioinformatics and Computational Biology.The Seventh Asia Pacific Bioinformatics Conference (APBC2009) was held at Tsinghua University,
Michael Q. Zhang, Michael S. Waterman, Xuegong Zhang
BMC Bioinform.1
2009 Challenges in Understanding Genome-Wide DNA Methylation
Michael Q. Zhang, Andrew D. Smith
J. Comput. Sci. Technol.1
2008 Multiobjective fuzzy biclustering in microarray data: Method and a new performance measure
abstract
Objective of any biclustering algorithm in microarray data is to discover a subset of genes that are expressed similarly in a subset of conditions. The boundaries of biclusters usually overlap as genes and conditions may belong to different biclusters with different membership degrees. Hence the notion of fuzzy sets is useful for discovering such overlapping biclusters. In this article an attempt has been made to develop a multiobjective genetic algorithm based approach for probabilistic fuzzy biclustering that minimizes the residual and maximizes cluster size and expression profile variance. A novel variable string length encoding has been proposed in this regard that encodes multiple biclusters in a single string. Also a new performance measure that reflects how a bicluster is statistically distinguished from the background is proposed. Performance of the proposed algorithm has been compared with some well known biclustering algorithms.
Ujjwal Maulik, Anirban Mukhopadhyay 0001, Sanghamitra Bandyopadhyay, Michael Q. Zhang, Xuegong Zhang
IEEE Congress on Evolutionary Computation4
2008 ZOOM! Zillions of oligos mapped
abstract
MOTIVATION: The next generation sequencing technologies are generating billions of short reads daily. Resequencing and personalized medicine need much faster software to map these deep sequencing reads to a reference genome, to identify SNPs or rare transcripts. RESULTS: We present a framework for how full sensitivity mapping can be done in the most efficient way, via spaced seeds. Using the framework, we have developed software called ZOOM, which is able to map the Illumina/Solexa reads of 15x coverage of a human genome to the reference human genome in one CPU-day, allowing two mismatches, at full sensitivity. AVAILABILITY: ZOOM is freely available to non-commercial users at http://www.bioinfor.com/zoom
Michael Q. Zhang, Bin Ma 0002, Ming Li 0001
Bioinform.3
2008 Identification of phylogenetically conserved microRNA cis-regulatory elements across 12 Drosophila species
abstract
MOTIVATION: MicroRNAs are a class of endogenous small RNAs that play regulatory roles. Intergenic miRNAs are believed to be transcribed independently, but the transcriptional control of these crucial regulators is still poorly understood. RESULTS: In this work, phylogenetic footprinting is used to identify conserved cis-regulatory elements (CCEs) surrounding intergenic miRNAs in Drosophila. With a two-step strategy that takes advantage of both alignment-based and motif-based methods, we identified CCEs that are conserved across the 12 fly species. When compared with TRANSFAC database, these CCEs are significantly enriched in known transcription factor binding sites (TFBSs). Moreover, several TFs that play essential roles in Drosophila development (e.g. Adf-1, Abd-B, Sd, Prd, Ubx, Zen and En) are found to be preferentially regulating the miRNA genes. Further analysis revealed many over-represented cis-regulatory modules (CRMs) composed of multiple known TFBSs, motif pairs with significant distance constraints and a number of novel motifs, many of which preferentially occur near the transcription start site of protein-coding genes. Additionally, a number of putative miRNA-TF regulatory feedback loops were also detected. AVAILABILITY: Supplementary Material and the Perl scripts performing two-step phylogenetic footprinting are available at http://bioinfo.au.tsinghua.edu.cn/member/xwwang/mircisreg
Xiaowo Wang, Jin Gu, Michael Q. Zhang, Yanda Li
Bioinform.3
2008 Integrative bioinformatics analysis of transcriptional regulatory programs in breast cancer cells
abstract
BACKGROUND: Microarray technology has unveiled transcriptomic differences among tumors of various phenotypes, and, especially, brought great progress in molecular understanding of phenotypic diversity of breast tumors. However, compared with the massive knowledge about the transcriptome, we have surprisingly little knowledge about regulatory mechanisms underling transcriptomic diversity. RESULTS: To gain insights into the transcriptional programs that drive tumor progression, we integrated regulatory sequence data and expression profiles of breast cancer into a Bayesian Network, and searched for cis-regulatory motifs statistically associated with given histological grades and prognosis. Our analysis found that motifs bound by ELK1, E2F, NRF1 and NFY are potential regulatory motifs that positively correlate with malignant progression of breast cancer. CONCLUSION: The results suggest that these 4 motifs are principal regulatory motifs driving malignant progression of breast cancer. Our method offers a more concise description about transcriptome diversity among breast tumors with different clinical phenotypes.
Atsushi Niida, Andrew D. Smith, Seiya Imoto, Shuichi Tsutsumi, Hiroyuki Aburatani, Michael Q. Zhang, Tetsu Akiyama
BMC Bioinform.6
2008 Using quality scores and longer reads improves accuracy of Solexa read mapping
abstract
BACKGROUND: Second-generation sequencing has the potential to revolutionize genomics and impact all areas of biomedical science. New technologies will make re-sequencing widely available for such applications as identifying genome variations or interrogating the oligonucleotide content of a large sample (e.g. ChIP-sequencing). The increase in speed, sensitivity and availability of sequencing technology brings demand for advances in computational technology to perform associated analysis tasks. The Solexa/Illumina 1G sequencer can produce tens of millions of reads, ranging in length from approximately 25-50 nt, in a single experiment. Accurately mapping the reads back to a reference genome is a critical task in almost all applications. Two sources of information that are often ignored when mapping reads from the Solexa technology are the 3' ends of longer reads, which contain a much higher frequency of sequencing errors, and the base-call quality scores. RESULTS: To investigate whether these sources of information can be used to improve accuracy when mapping reads, we developed the RMAP tool, which can map reads having a wide range of lengths and allows base-call quality scores to determine which positions in each read are more important when mapping. We applied RMAP to analyze data re-sequenced from two human BAC regions for varying read lengths, and varying criteria for use of quality scores. RMAP is freely available for downloading at http://rulai.cshl.edu/rmap/. CONCLUSION: Our results indicate that significant gains in Solexa read mapping performance can be achieved by considering the information in 3' ends of longer reads, and appropriately using the base-call quality scores. The RMAP tool we have developed will enable researchers to effectively exploit this information in targeted re-sequencing projects.
Andrew D. Smith, Zhenyu Xuan, Michael Q. Zhang
BMC Bioinform.3
2008 Identification of Synaptic Targets of Drosophila Pumilio
abstract
Drosophila Pumilio (Pum) protein is a translational regulator involved in embryonic patterning and germline development. Recent findings demonstrate that Pum also plays an important role in the nervous system, both at the neuromuscular junction (NMJ) and in long-term memory formation. In neurons, Pum appears to play a role in homeostatic control of excitability via down regulation of para, a voltage gated sodium channel, and may more generally modulate local protein synthesis in neurons via translational repression of eIF-4E. Aside from these, the biologically relevant targets of Pum in the nervous system remain largely unknown. We hypothesized that Pum might play a role in regulating the local translation underlying synapse-specific modifications during memory formation. To identify relevant translational targets, we used an informatics approach to predict Pum targets among mRNAs whose products have synaptic localization. We then used both in vitro binding and two in vivo assays to functionally confirm the fidelity of this informatics screening method. We find that Pum strongly and specifically binds to RNA sequences in the 3'UTR of four of the predicted target genes, demonstrating the validity of our method. We then demonstrate that one of these predicted target sequences, in the 3'UTR of discs large (dlg1), the Drosophila PSD95 ortholog, can functionally substitute for a canonical NRE (Nanos response element) in vivo in a heterologous functional assay. Finally, we show that the endogenous dlg1 mRNA can be regulated by Pumilio in a neuronal context, the adult mushroom bodies (MB), which is an anatomical site of memory storage.
Gengxin Chen, Wanhe Li, Qing-Shuo Zhang, Michael Regulski, Nishi Sinha, Jody Barditch, Tim Tully, Adrian R. Krainer, Michael Q. Zhang, Josh Dubnau
PLoS Comput. Biol.9
2007 OSCAR: One-class SVM for accurate recognition of cis-elements
abstract
MOTIVATION: Traditional methods to identify potential binding sites of known transcription factors still suffer from large number of false predictions. They mostly use sequence information in a position-specific manner and neglect other types of information hidden in the proximal promoter regions. Recent biological and computational researches, however, suggest that there exist not only locational preferences of binding, but also correlations between transcription factors. RESULTS: In this article, we propose a novel approach, OSCAR, which utilizes one-class SVM algorithms, and incorporates multiple factors to aid the recognition of transcription factor binding sites. Using both synthetic and real data, we find that our method outperforms existing algorithms, especially in the high sensitivity region. The performance of our method can be further improved by taking into account locational preference of binding events. By testing on experimentally-verified binding sites of GATA and HNF transcription factor families, we show that our algorithm can infer the true co-occurring motif pairs accurately, and by considering the co-occurrences of correlated motifs, we not only filter out false predictions, but also increase the sensitivity. AVAILABILITY: An online server based on OSCAR is available at http://bioinfo.au.tsinghua.edu.cn/oscar.
Bo Jiang 0005, Michael Q. Zhang, Xuegong Zhang
Bioinform.2
2007 Computational prediction of novel components of lung transcriptional networks
abstract
MOTIVATION: Little is known regarding the transcriptional mechanisms involved in forming and maintaining epithelial cell lineages of the mammalian respiratory tract. RESULTS: Herein, a motif discovery approach was used to identify novel transcriptional regulators in the lung using genes previously found to be regulated by Foxa2 or Wnt signaling pathways. A human-mouse comparison of both novel and known motifs was also performed. Some of the factors and families identified here were previously shown to be involved epithelial cell differentiation (ETS family, HES-1 and MEIS-1), and ciliogenesis (RFX family), but have never been characterized in lung epithelia. Other unidentified over-represented motifs suggest the existence of novel mammalian lung transcription factors. Of the fraction of motifs examined we describe 25 transcription factor family predictions for lung. Fifteen novel factors were shown here to be expressed in mouse lung, and/or human bronchial or distal lung epithelial tissues or lung epithelial cell lineages. AVAILABILITY: DME: http://rulai.cshl.edu/dme. MATCOMPARE: http://rulai.cshl.edu/MatCompare. MOTIFCLASS is available from the authors.
M. Juanita Martinez, Andrew D. Smith, Bilan Li, Michael Q. Zhang, Kevin S. Harrod
Bioinform.4
2007 Computing exact P-values for DNA motifs
abstract
MOTIVATION: Many heuristic algorithms have been designed to approximate P-values of DNA motifs described by position weight matrices, for evaluating their statistical significance. They often significantly deviate from the true P-value by orders of magnitude. Exact P-value computation is needed for ranking the motifs. Furthermore, surprisingly, the complexity of the problem is unknown. RESULTS: We show the problem to be NP-hard, and present MotifRank, software based on dynamic programming, to calculate exact P-values of motifs. We define the exact P-value on a general and more precise model. Asymptotically, MotifRank is faster than the best exact P-value computing algorithm, and is in fact practical. Our experiments clearly demonstrate that MotifRank significantly improves the accuracy of existing approximation algorithms. AVAILABILITY: MotifRank is available from http://bio.dlg.cn. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jing Zhang 0012, Bo Jiang 0005, Ming Li 0001, John Tromp, Xuegong Zhang, Michael Q. Zhang
Bioinform.6
2007 Statistical significance of cis-regulatory modules
abstract
BACKGROUND: It is becoming increasingly important for researchers to be able to scan through large genomic regions for transcription factor binding sites or clusters of binding sites forming cis-regulatory modules. Correspondingly, there has been a push to develop algorithms for the rapid detection and assessment of cis-regulatory modules. While various algorithms for this purpose have been introduced, most are not well suited for rapid, genome scale scanning. RESULTS: We introduce methods designed for the detection and statistical evaluation of cis-regulatory modules, modeled as either clusters of individual binding sites or as combinations of sites with constrained organization. In order to determine the statistical significance of module sites, we first need a method to determine the statistical significance of single transcription factor binding site matches. We introduce a straightforward method of estimating the statistical significance of single site matches using a database of known promoters to produce data structures that can be used to estimate p-values for binding site matches. We next introduce a technique to calculate the statistical significance of the arrangement of binding sites within a module using a max-gap model. If the module scanned for has defined organizational parameters, the probability of the module is corrected to account for organizational constraints. The statistical significance of single site matches and the architecture of sites within the module can be combined to provide an overall estimation of statistical significance of cis-regulatory module sites. CONCLUSION: The methods introduced in this paper allow for the detection and statistical evaluation of single transcription factor binding sites and cis-regulatory modules. The features described are implemented in the Search Tool for Occurrences of Regulatory Motifs (STORM) and MODSTORM software.
Dustin E. Schones, Andrew D. Smith, Michael Q. Zhang
BMC Bioinform.3
2007 Computational analyses of eukaryotic promoters
abstract
Computational analysis of eukaryotic promoters is one of the most difficult problems in computational genomics and is essential for understanding gene expression profiles and reverse-engineering gene regulation network circuits. Here I give a basic introduction of the problem and recent update on both experimental and computational approaches. More details may be found in the extended references. This review is based on a summer lecture given at Max Planck Institute at Berlin in 2005.
Michael Q. Zhang
BMC Bioinform.1
2007 Neighbor number, valley seeking and clustering
Xuegong Zhang, Michael Q. Zhang, Yanda Li
Pattern Recognit. Lett.3
2006 Analysis and visualization of DNA spectrograms: open possibilities for the genome research
abstract
The demand for technology that can process biological information is becoming more and more obvious and urgent. Existing research in bioinformatics has been focusing on various types of analysis of DNA sequences and various measurements taken at the protein, RNA transcript and DNA level. In this paper we will show the application of spectral analysis and image processing in analyzing DNA sequences of specific structure. In addition, we extend the framework to visualize long DNA sequences and help in identifying patterns that are visible at high resolution of DNA spectral images.
Nevenka Dimitrova, Yee Him Cheung, Michael Q. Zhang
ACM Multimedia3
2006 Predicting methylation status of CpG islands in the human brain
abstract
MOTIVATION: Over 50% of human genes contain CpG islands in their 5'-regions. Methylation patterns of CpG islands are involved in tissue-specific gene expression and regulation. Mis-epigenetic silencing associated with aberrant CpG island methylation is one mechanism leading to the loss of tumor suppressor functions in cancer cells. Large-scale experimental detection of DNA methylation is still both labor-intensive and time-consuming. Therefore, it is necessary to develop in silico approaches for predicting methylation status of CpG islands. RESULTS: Based on a recent genome-scale dataset of DNA methylation in human brain tissues, we developed a classifier called MethCGI for predicting methylation status of CpG islands using a support vector machine (SVM). Nucleotide sequence contents as well as transcription factor binding sites (TFBSs) are used as features for the classification. The method achieves specificity of 84.65% and sensitivity of 84.32% on the brain data, and can also correctly predict about two-third of the data from other tissues reported in the MethDB database. AVAILABILITY: An online predictor based on MethCGI is available at http://166.111.201.7/MethCGI.html CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data available at Bioinformatics online and http://166.111.201.7/help.html.
Shicai Fan, Xuegong Zhang, Michael Q. Zhang
Bioinform.4
2006 Profiling alternatively spliced mRNA isoforms for prostate cancer classification
abstract
BACKGROUND: Prostate cancer is one of the leading causes of cancer illness and death among men in the United States and world wide. There is an urgent need to discover good biomarkers for early clinical diagnosis and treatment. Previously, we developed an exon-junction microarray-based assay and profiled 1532 mRNA splice isoforms from 364 potential prostate cancer related genes in 38 prostate tissues. Here, we investigate the advantage of using splice isoforms, which couple transcriptional and splicing regulation, for cancer classification. RESULTS: As many as 464 splice isoforms from more than 200 genes are differentially regulated in tumors at a false discovery rate (FDR) of 0.05. Remarkably, about 30% of genes have isoforms that are called significant but do not exhibit differential expression at the overall mRNA level. A support vector machine (SVM) classifier trained on 128 signature isoforms can correctly predict 92% of the cases, which outperforms the classifier using overall mRNA abundance by about 5%. It is also observed that the classification performance can be improved using multivariate variable selection methods, which take correlation among variables into account. CONCLUSION: These results demonstrate that profiling of splice isoforms is able to provide unique and important information which cannot be detected by conventional microarrays.
Hai-Ri Li, Jian-Bing Fan, Jessica Wang-Rodriguez, Tracy Downs, Xiang-Dong Fu, Michael Q. Zhang
BMC Bioinform.7
2005 Similarity of position frequency matrices for transcription factor binding sites
abstract
MOTIVATION: Transcription-factor binding sites (TFBS) in promoter sequences of higher eukaryotes are commonly modeled using position frequency matrices (PFM). The ability to compare PFMs representing binding sites is especially important for de novo sequence motif discovery, where it is desirable to compare putative matrices to one another and to known matrices. RESULTS: We describe a PFM similarity quantification method based on product multinomial distributions, demonstrate its ability to identify PFM similarity and show that it has a better false positive to false negative ratio compared to existing methods. We grouped TFBS frequency matrices from two libraries into matrix families and identified the matrices that are common and unique to these libraries. We identified similarities and differences between the skeletal-muscle-specific and non-muscle-specific frequency matrices for the binding sites of Mef-2, Myf, Sp-1, SRF and TEF of Wasserman and Fickett. We further identified known frequency matrices and matrix families that were strongly similar to the matrices given by Wasserman and Fickett. We provide methodology and tools to compare and query libraries of frequency matrices for TFBSs. AVAILABILITY: Software is available to use over the Web at http://rulai.cshl.edu/MatCompare SUPPLEMENTARY INFORMATION: Database and clustering statistics, matrix families and representatives are available at http://rulai.cshl.edu/MatCompare/Supplementary.
Dustin E. Schones, Pavel Sumazin, Michael Q. Zhang
Bioinform.3
2005 DWE: Discriminating Word Enumerator
abstract
MOTIVATION: Tissue-specific transcription factor binding sites give insight into tissue-specific transcription regulation. RESULTS: We describe a word-counting-based tool for de novo tissue-specific transcription factor binding site discovery using expression information in addition to sequence information. We incorporate tissue-specific gene expression through gene classification to positive expression and repressed expression. We present a direct statistical approach to find overrepresented transcription factor binding sites in a foreground promoter sequence set against a background promoter sequence set. Our approach naturally extends to synergistic transcription factor binding site search. We find putative transcription factor binding sites that are overrepresented in the proximal promoters of liver-specific genes relative to proximal promoters of liver-independent genes. Our results indicate that binding sites for hepatocyte nuclear factors (especially HNF-1 and HNF-4) and CCAAT/enhancer-binding protein (C/EBPbeta) are the most overrepresented in proximal promoters of liver-specific genes. Our results suggest that HNF-4 has strong synergistic relationships with HNF-1, HNF-4 and HNF-3beta and with C/EBPbeta. AVAILABILITY: Programs are available for use over the Web at http://rulai.cshl.edu/tools/dwe.
Pavel Sumazin, Gengxin Chen, Naoya Hata, Andrew D. Smith, Theresa Zhang, Michael Q. Zhang
Bioinform.6
2001 Identifying the 3'-terminal exon in human DNA
abstract
MOTIVATION: We present JTEF, a new program for finding 3' terminal exons in human DNA sequences. This program is based on quadratic discriminant analysis, a standard non-linear statistical pattern recognition method. The quadratic discriminant functions used for building the algorithm were trained on a set of 3' terminal exons of type 3tuexon (those containing the true STOP codon). RESULTS: We showed that the average predictive accuracy of JTEF is higher than the presently available best programs (GenScan and Genemark.hmm) based on a test set of 65 human DNA sequences with 121 genes. In particular JTEF performs well on larger genomic contigs containing multiple genes and significant amounts of intergenic DNA. It will become a valuable tool for genome annotation and gene functional studies. AVAILABILITY: JTEF is available free for academic users on request from ftp://cshl.org/pub/science/mzhanglab/JTEF and will be made available through the World Wide Web (http://argon.cshl.org/).
Jack E. Tabaska, Ramana V. Davuluri, Michael Q. Zhang
Bioinform.3
2001 Promoter Extraction from GenBank (PEG): automatic extraction of eukaryotic promoter sequences in large sets of genes
abstract
UNLABELLED: Promoter Extraction from GenBank (PEG) extracts promoter sequences for large sets of genes using information present in GenBank. For a gene whose promoter sequence is not found, PEG will attempt to extract promoter sequences of the orthologous genes instead. AVAILABILITY: It is freely available to academic users at ftp://cshl.org/pub/science/mzhanglab/theresa/. CONTACT: [email protected]; [email protected]
Theresa Zhang, Michael Q. Zhang
Bioinform.2
2000 Discriminant Analysis and Its Application in DNA Sequence Motif Recognition
abstract
Identification of functional motifs in a DNA sequence is fundamentally a statistical pattern recognition problem. Discriminant analysis is widely used for solving such problems. This paper will review two basic parametric methods: LDA (linear discriminant analysis) and QDA (quadratic discriminant analysis). Their usage in recognition of splice sites and exons in the human genome will be demonstrated.
Michael Q. Zhang
Briefings Bioinform.1
1999 SCPD: a promoter database of the yeast Saccharomyces cerevisiae
abstract
MOTIVATION: In order to facilitate a systematic study of the promoters and transcriptionally regulatory cis-elements of the yeast Saccharomyces cerevisiae on a genomic scale, we have developed a comprehensive yeast-specific promoter database, SCPD. RESULTS: Currently SCPD contains 580 experimentally mapped transcription factor (TF) binding sites and 425 transcriptional start sites (TSS) as its primary data entries. It also contains relevant binding affinity and expression data where available. In addition to mechanisms for promoter information (including sequence) retrieval and a data submission form, SCPD also provides some simple but useful tools for promoter sequence analysis. AVAILABILITY: SCPD can be accessed from the URL http://cgsigma.cshl.org/jian. The database is continually updated.
Michael Q. Zhang
Bioinform.2
1993 A weight array method for splicing signal analysis
abstract
A new method of sequence analysis, using a weight array method (WAM), which generalizes the traditional Staden weight matrix method (WMM), is proposed. With the help of a statistical mechanical model, the discriminant function is identified with the energy function describing macromolecular interactions. The method is applied to the study of 5'-splice signals in Schizosaccharomyces pombe pre-mRNA sequences. The results show that there may exist weak pairwise correlations within the signals and that our method can help to better discriminate these signals. Experiments are proposed to test the predictions of the theory.
Michael Q. Zhang, T. G. Marr
Comput. Appl. Biosci.1