EDBT 2026 Demo / reviewers in the wild / expert
Xiao Sun 0006
dblp:30/202-6
· DBLP profile ↗
18ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0003-1048-7775ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Decoding cancer heterogeneity through super-enhancer landscapes: from subtype discovery to therapeutic opportunityabstractSuper-enhancers (SEs) are clusters of enhancers with potent regulatory capabilities. They play a crucial role in shaping cellular identity and driving the progression of various diseases, including cancer. SEs exhibit significant heterogeneity across different cell types and cancer subtypes. Analysis of SE landscapes can recapitulate existing classification systems and unveil novel SE-driven epigenetic subtypes. In this review, we summarized the latest advancements in cancer subtype identification based on SE-related characteristics, outlined the typical analytical workflows adopted in such studies, and explored the biological and clinical significance of SE-driven subtypes. Furthermore, we discussed the field's key challenges and emerging technologies to highlight future research directions. SE analysis provides a robust framework for dissecting cancer heterogeneity. This approach offers novel epigenetic perspectives and support for the realization of personalized medicine. Shuyang Cai, Zhenchang Wang, Xiao Sun 0006 |
Briefings Bioinform. | 3 |
| 2025 | G4SNVHunter: An R/Bioconductor Package for Evaluating SNV-Induced Disruption of G-Quadruplex Structures Leveraging the G4Hunter AlgorithmabstractG-quadruplexes (G4s) are nucleic acid secondary structures with important regulatory functions. Single-nucleotide variants (SNVs), one of the most common forms of genetic variation, can potentially impact the formation of G4 structures if they occur within G4 regions. However, there is currently a lack of software tools specifically designed to assess such effects. Here, we present an R/Bioconductor package named G4SNVHunter, which enables rapid detection of variants that may disrupt G4 structures. This tool, based on the core principles of the G4Hunter algorithm, can provide precise quantitative assessment of the propensity for G4 formation within genomic sequences. Specialized experimental methods can then be designed based on the results provided by G4SNVHunter to further verify the specific functions of the affected G4 structures, facilitating deeper insights into the biological impacts of genetic variants from the perspective of G4 structures. To showcase the functionality of the G4SNVHunter package, we analyzed the Neandertal and Denisovan archaic introgressed variants detected by the Sprime software, and identified approximately 5,800 variants located within G4 regions, among which around 230 may impair G4 structure formation propensity. The source code for the G4SNVHunter package has been publicly released under the MIT license at https://github.com/rongxinzh/G4SNVHunter and https://bioconductor.org/packages/devel/bioc/html/G4SNVHunter.html. Rongxin Zhang, Wenyong Zhu, Jean-Louis Mergny, Xiao Sun 0006 |
PLoS Comput. Biol. | 5 |
| 2023 | Improving the performance of single-cell RNA-seq data mining based on relative expression orderingsabstractThe advent of single-cell RNA-sequencing (scRNA-seq) provides an unprecedented opportunity to explore gene expression profiles at the single-cell level. However, gene expression values vary over time and under different conditions even within the same cell. There is an urgent need for more stable and reliable feature variables at the single-cell level to depict cell heterogeneity. Thus, we construct a new feature matrix called the delta rank matrix (DRM) from scRNA-seq data by integrating an a priori gene interaction network, which transforms the unreliable gene expression value into a stable gene interaction/edge value on a single-cell basis. This is the first time that a gene-level feature has been transformed into an interaction/edge-level for scRNA-seq data analysis based on relative expression orderings. Experiments on various scRNA-seq datasets have demonstrated that DRM performs better than the original gene expression matrix in cell clustering, cell identification and pseudo-trajectory reconstruction. More importantly, the DRM really achieves the fusion of gene expressions and gene interactions and provides a method of measuring gene interactions at the single-cell level. Thus, the DRM can be used to find changes in gene interactions among different cell types, which may open up a new way to analyze scRNA-seq data from an interaction perspective. In addition, DRM provides a new method to construct a cell-specific network for each single cell instead of a group of cells as in traditional network construction methods. DRM's exceptional performance is due to its extraction of rich gene-association information on biological systems and stable characterization of cells. Yuanyuan Chen 0014, Xiao Sun 0006 |
Briefings Bioinform. | 3 |
| 2023 | scAnno: a deconvolution strategy-based automatic cell type annotation tool for single-cell RNA-sequencing data setsabstractUndoubtedly, single-cell RNA sequencing (scRNA-seq) has changed the research landscape by providing insights into heterogeneous, complex and rare cell populations. Given that more such data sets will become available in the near future, their accurate assessment with compatible and robust models for cell type annotation is a prerequisite. Considering this, herein, we developed scAnno (scRNA-seq data annotation), an automated annotation tool for scRNA-seq data sets primarily based on the single-cell cluster levels, using a joint deconvolution strategy and logistic regression. We explicitly constructed a reference profile for human (30 cell types and 50 human tissues) and a reference profile for mouse (26 cell types and 50 mouse tissues) to support this novel methodology (scAnno). scAnno offers a possibility to obtain genes with high expression and specificity in a given cell type as cell type-specific genes (marker genes) by combining co-expression genes with seed genes as a core. Of importance, scAnno can accurately identify cell type-specific genes based on cell type reference expression profiles without any prior information. Particularly, in the peripheral blood mononuclear cell data set, the marker genes identified by scAnno showed cell type-specific expression, and the majority of marker genes matched exactly with those included in the CellMarker database. Besides validating the flexibility and interpretability of scAnno in identifying marker genes, we also proved its superiority in cell type annotation over other cell type annotation tools (SingleR, scPred, CHETAH and scmap-cluster) through internal validation of data sets (average annotation accuracy: 99.05%) and cross-platform data sets (average annotation accuracy: 95.56%). Taken together, we established the first novel methodology that utilizes a deconvolution strategy for automated cell typing and is capable of being a significant application in broader scRNA-seq analysis. scAnno is available at https://github.com/liuhong-jia/scAnno. Hongjia Liu, Huamei Li, Wenjuan Huang, Duo Pan, Xiao Sun 0006, Hongde Liu 0001 |
Briefings Bioinform. | 8 |
| 2022 | Relating Translation Efficiency to Protein Networks Provides Evolutionary Insights in Shewanella and Its Implications for Extracellular Electron TransferabstractShewanellaspecies are well-known for their extracellular electron transfer (EET) capacity, by which these microorganisms can transfer the electrons from intracellular environment to extracellular space for the reduction of the extracellular insoluble electron acceptors. Using a time-stamped data for the paired protein-mRNA, we investigate the impact of differential translation on the EET process ofShewanella oneidensisMR-1. Firstly, differentially translated proteins when O2levels are switched from high-O2to low-O2are identified by using a soft clustering method, 629 up-regulated translated proteins and 767 down-regulated translated proteins are considered to reflect the changes from inactivated to activated EET process. Then, we showed that the degrees of connectivity of differentially translated proteins were significantly larger than those of non-differentially translated proteins, and thereby these differentially translated proteins will be more important in the protein networks. After that, we networked these differentially translated proteins to construct the differentially translated sub-networks, and discussed the most important proteins that are involved in the EET process with the help of centralization analysis of these differentially translated networks. Furthermore, we also studied the differentially translated operonic genes. Taking together, this work searches the key proteins that potentially activated the EET process from a translational efficiency viewpoint. Dewu Ding, Xiao Sun 0006 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | Sample-specific perturbation of gene interactions identifies breast cancer subtypesabstractBreast cancer is a highly heterogeneous disease, and there are many forms of categorization for breast cancer based on gene expression profiles. Gene expression profiles are variables and may show differences if measured at different time points or under different conditions. In contrast, biological networks are relatively stable over time and under different conditions. In this study, we used a gene interaction network from a new point of view to explore the subtypes of breast cancer based on individual-specific edge perturbations measured by relative gene expression value. Our study reveals that there are four breast cancer subtypes based on gene interaction perturbations at the individual level. The new network-based subtypes of breast cancer show strong heterogeneity in prognosis, somatic mutations, phenotypic changes and enriched pathways. The network-based subtypes are closely related to the PAM50 subtypes and immunohistochemistry index. This work helps us to better understand the heterogeneity and mechanisms of breast cancer from a network perspective. Yuanyuan Chen 0014, Zixi Hu, Xiao Sun 0006 |
Briefings Bioinform. | 4 |
| 2021 | Classification of Mild Cognitive Impairment With Multimodal Data Using Both Labeled and Unlabeled SamplesabstractMild Cognitive Impairment (MCI) is a preclinical stage of Alzheimer's Disease (AD) and is clinical heterogeneity. The classification of MCI is crucial for the early diagnosis and treatment of AD. In this study, we investigated the potential of using both labeled and unlabeled samples from the Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort to classify MCI through the multimodal co-training method. We utilized both structural magnetic resonance imaging (sMRI) data and genotype data of 364 MCI samples including 228 labeled and 136 unlabeled MCI samples from the ADNI-1 cohort. First, the selected quantitative trait (QT) features from sMRI data and SNP features from genotype data were used to build two initial classifiers on 228 labeled MCI samples. Then, the co-training method was implemented to obtain new labeled samples from 136 unlabeled MCI samples. Finally, the random forest algorithm was used to obtain a combined classifier to classify MCI patients in the independent ADNI-2 dataset. The experimental results showed that our proposed framework obtains an accuracy of 85.50 percent and an AUC of 0.825 for MCI classification, respectively, which showed that the combined utilization of sMRI and SNP data through the co-training method could significantly improve the performances of MCI classification. Shaoxun Yuan, Haitao Li 0004, Xiao Sun 0006 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2019 | A Comparative Study of Network Motifs in the Integrated Transcriptional Regulation and Protein Interaction Networks of ShewanellaabstractThe Shewanella species shows a remarkable respiratory versatility with a great variety of extracellular electron acceptors (termed Extracellular Electron Transfer, EET). To explore relevant mechanisms from the network motif view, we constructed the integrated networks that combined transcriptional regulation interactions (TRIs) and protein-protein interactions (PPIs) for 13 Shewanella species, identified and compared the network motifs in these integrated networks. We found that the network motifs were evolutionary conserved in these integrated networks. The functional significance of the highly conserved motifs was discussed, especially the important ones that were potentially involved in the Shewanella EET processes. More importantly, we found that: 1) the motif co-regulated PPI took a role in the "standby mode" of protein utilization, which will be helpful for cells to rapidly response to environmental changes; and 2) the type II cofactors, which involved in the motif TRI interacting with a third protein, mainly carried out a signalling role in Shewanella oneidensis MR-1. Dewu Ding, Xiao Sun 0006 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2018 | Nucleosome Positioning of Intronless Genes in the Human GenomeabstractNucleosomes, the basic units of chromatin, are involved in transcription regulation and DNA replication. Intronless genes, which constitute 3 percent of the human genome, differ from intron-containing genes in evolution and function. Our analysis reveals that nucleosome positioning shows a distinct pattern in intronless and intron-containing genes. The nucleosome occupancy upstream of transcription start sites of intronless genes is lower than that of intron-containing genes. In contrast, high occupancy and well positioned nucleosomes are observed along the gene body of intronless genes, which is perfectly consistent with the barrier nucleosome model. Intronless genes have a significantly lower expression level than intron-containing genes and most of them are not expressed in CD4+ T cell lines and GM12878 cell lines, which results from their tissue specificity. However, the highly expressed genes are at the same expression level between the two types of genes. The highly expressed intronless genes require a higher density of RNA Pol II in an elongating state to compensate for the lack of introns. Additionally, 5' and 3' nucleosome depleted regions of highly expressed intronless genes are deeper than those of highly expressed intron-containing genes. Xiangfei Cheng, Yumin Nie, Yiru Zhang, Hongde Liu 0001, Xiao Sun 0006 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2015 | Accurate estimation of haplotype frequency from pooled sequencing data and cost-effective identification of rare haplotype carriers by overlapping pool sequencingabstractMOTIVATION: A variety of hypotheses have been proposed for finding the missing heritability of complex diseases in genome-wide association studies. Studies have focused on the value of haplotype to improve the power of detecting associations with disease. To facilitate haplotype-based association analysis, it is necessary to accurately estimate haplotype frequencies of pooled samples. RESULTS: Taking advantage of databases that contain prior haplotypes, we present Ehapp based on the algorithm for solving the system of linear equations to estimate the frequencies of haplotypes from pooled sequencing data. Effects of various factors in sequencing on the performance are evaluated using simulated data. Our method could estimate the frequencies of haplotypes with only about 3% average relative difference for pooled sequencing of the mixture of 10 haplotypes with total coverage of 50×. When unknown haplotypes exist, our method maintains excellent performance for haplotypes with actual frequencies >0.05. Comparisons with present method on simulated data in conjunction with publicly available Illumina sequencing data indicate that our method is state of the art for many sequencing study designs. We also demonstrate the feasibility of applying overlapping pool sequencing to identify rare haplotype carriers cost-effectively. AVAILABILITY AND IMPLEMENTATION: Ehapp (in Perl) for the Linux platforms is available online (http://bioinfo.seu.edu.cn/Ehapp/). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chang-Chang Cao, Xiao Sun 0006 |
Bioinform. | 2 |
| 2015 | DectICO: an alignment-free supervised metagenomic classification method based on feature extraction and dynamic selectionabstractBACKGROUND: Continual progress in next-generation sequencing allows for generating increasingly large metagenomes which are over time or space. Comparing and classifying the metagenomes with different microbial communities is critical. Alignment-free supervised classification is important for discriminating between the multifarious components of metagenomic samples, because it can be accomplished independently of known microbial genomes. RESULTS: We propose an alignment-free supervised metagenomic classification method called DectICO. The intrinsic correlation of oligonucleotides provides the feature set, which is selected dynamically using a kernel partial least squares algorithm, and the feature matrices extracted with this set are sequentially employed to train classifiers by support vector machine (SVM). We evaluated the classification performance of DectICO on three actual metagenomic sequencing datasets, two containing deep sequencing metagenomes and one of low coverage. Validation results show that DectICO is powerful, performs well based on long oligonucleotides (i.e., 6-mer to 8-mer), and is more stable and generalized than a sequence-composition-based method. The classifiers trained by our method are more accurate than non-dynamic feature selection methods and a recently published recursive-SVM-based classification approach. CONCLUSIONS: The alignment-free supervised classification method DectICO can accurately classify metagenomic samples without dependence on known microbial genomes. Selecting the ICO dynamically offers better stability and generality compared with sequence-composition-based classification algorithms. Our proposed method provides new insights in metagenomic sample classification. Fudong Cheng, Chang-Chang Cao, Xiao Sun 0006 |
BMC Bioinform. | 4 |
| 2015 | PRBP: Prediction of RNA-Binding Proteins Using a Random Forest Algorithm Combined with an RNA-Binding Residue PredictorabstractThe prediction of RNA-binding proteins is an incredibly challenging problem in computational biology. Although great progress has been made using various machine learning approaches with numerous features, the problem is still far from being solved. In this study, we attempt to predict RNA-binding proteins directly from amino acid sequences. A novel approach, PRBP predicts RNA-binding proteins using the information of predicted RNA-binding residues in conjunction with a random forest based method. For a given protein, we first predict its RNA-binding residues and then judge whether the protein binds RNA or not based on information from that prediction. If the protein cannot be identified by the information associated with its predicted RNA-binding residues, then a novel random forest predictor is used to determine if the query protein is a RNA-binding protein. We incorporated features of evolutionary information combined with physicochemical features (EIPP) and amino acid composition feature to establish the random forest predictor. Feature analysis showed that EIPP contributed the most to the prediction of RNA-binding proteins. The results also showed that the information from the RNA-binding residue prediction improved the overall performance of our RNA-binding protein prediction. It is anticipated that the PRBP method will become a useful tool for identifying RNA-binding proteins. A PRBP Web server implementation is freely available at http://www.cbi.seu.edu.cn/PRBP/. Xin Ma 0003, Xiao Sun 0006 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2014 | Quantitative group testing-based overlapping pool sequencing to identify rare variant carriersabstractBACKGROUND: Genome-wide association studies have revealed that rare variants are responsible for a large portion of the heritability of some complex human diseases. This highlights the increasing importance of detecting and screening for rare variants. Although the massively parallel sequencing technologies have greatly reduced the cost of DNA sequencing, the identification of rare variant carriers by large-scale re-sequencing remains prohibitively expensive because of the huge challenge of constructing libraries for thousands of samples. Recently, several studies have reported that techniques from group testing theory and compressed sensing could help identify rare variant carriers in large-scale samples with few pooled sequencing experiments and a dramatically reduced cost. RESULTS: Based on quantitative group testing, we propose an efficient overlapping pool sequencing strategy that allows the efficient recovery of variant carriers in numerous individuals with much lower costs than conventional methods. We used random k-set pool designs to mix samples, and optimized the design parameters according to an indicative probability. Based on a mathematical model of sequencing depth distribution, an optimal threshold was selected to declare a pool positive or negative. Then, using the quantitative information contained in the sequencing results, we designed a heuristic Bayesian probability decoding algorithm to identify variant carriers. Finally, we conducted in silico experiments to find variant carriers among 200 simulated Escherichia coli strains. With the simulated pools and publicly available Illumina sequencing data, our method correctly identified the variant carriers for 91.5-97.9% variants with the variant frequency ranging from 0.5 to 1.5%. CONCLUSIONS: Using the number of reads, variant carriers could be identified precisely even though samples were randomly selected and pooled. Our method performed better than the published DNA Sudoku design and compressed sequencing, especially in reducing the required data throughput and cost. Chang-Chang Cao, Xiao Sun 0006 |
BMC Bioinform. | 3 |
| 2012 | Sequence-Based Prediction of DNA-Binding Residues in Proteins with Conservation and Correlation InformationabstractThe recognition of DNA-binding residues in proteins is critical to our understanding of the mechanisms of DNA-protein interactions, gene expression, and for guiding drug design. Therefore, a prediction method DNABR (DNA Binding Residues) is proposed for predicting DNA-binding residues in protein sequences using the random forest (RF) classifier with sequence-based features. Two types of novel sequence features are proposed in this study, which reflect the information about the conservation of physicochemical properties of the amino acids, and the correlation of amino acids between different sequence positions in terms of physicochemical properties. The first type of feature uses the evolutionary information combined with the conservation of physicochemical properties of the amino acids while the second reflects the dependency effect of amino acids with regards to polarity charge and hydrophobic properties in the protein sequences. Those two features and an orthogonal binary vector which reflect the characteristics of 20 types of amino acids are used to build the DNABR, a model to predict DNA-binding residues in proteins. The DNABR model achieves a value of 0.6586 for Matthew’s correlation coefficient (MCC) and 93.04 percent overall accuracy (ACC) with a68.47 percent sensitivity (SE) and 98.16 percent specificity (SP), respectively. The comparisons with each feature demonstrate that these two novel features contribute most to the improvement in predictive ability. Furthermore, performance comparisons with other approaches clearly show that DNABR has an excellent prediction performance for detecting binding residues in putative DNA-binding protein. The DNABR web-server system is freely available at http://www.cbi.seu.edu.cn/DNABR/. Xin Ma 0003, Hongde Liu 0001, Jianming Xie, Xiao Sun 0006 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2009 | Prediction of DNA-binding residues in proteins from amino acid sequences using a random forest model with a hybrid featureabstractMOTIVATION: In this work, we aim to develop a computational approach for predicting DNA-binding sites in proteins from amino acid sequences. To avoid overfitting with this method, all available DNA-binding proteins from the Protein Data Bank (PDB) are used to construct the models. The random forest (RF) algorithm is used because it is fast and has robust performance for different parameter values. A novel hybrid feature is presented which incorporates evolutionary information of the amino acid sequence, secondary structure (SS) information and orthogonal binary vector (OBV) information which reflects the characteristics of 20 kinds of amino acids for two physical-chemical properties (dipoles and volumes of the side chains). The numbers of binding and non-binding residues in proteins are highly unbalanced, so a novel scheme is proposed to deal with the problem of imbalanced datasets by downsizing the majority class. RESULTS: The results show that the RF model achieves 91.41% overall accuracy with Matthew's correlation coefficient of 0.70 and an area under the receiver operating characteristic curve (AUC) of 0.913. To our knowledge, the RF method using the hybrid feature is currently the computationally optimal approach for predicting DNA-binding sites in proteins from amino acid sequences without using three-dimensional (3D) structural information. We have demonstrated that the prediction results are useful for understanding protein-DNA interactions. AVAILABILITY: DBindR web server implementation is freely available at http://www.cbi.seu.edu.cn/DBindR/DBindR.htm. Hongde Liu 0001, Xueye Duan, Xiao Sun 0006 |
Bioinform. | 7 |
| 2009 | Mechanism-anchored profiling derived from epigenetic networks predicts outcome in acute lymphoblastic leukemiaabstractBACKGROUND: Current outcome predictors based on "molecular profiling" rely on gene lists selected without consideration for their molecular mechanisms. This study was designed to demonstrate that we could learn about genes related to a specific mechanism and further use this knowledge to predict outcome in patients - a paradigm shift towards accurate "mechanism-anchored profiling". We propose a novel algorithm, PGnet, which predicts a tripartite mechanism-anchored network associated to epigenetic regulation consisting of phenotypes, genes and mechanisms. Genes termed as GEMs in this network meet all of the following criteria: (i) they are co-expressed with genes known to be involved in the biological mechanism of interest, (ii) they are also differentially expressed between distinct phenotypes relevant to the study, and (iii) as a biomodule, genes correlate with both the mechanism and the phenotype. RESULTS: This proof-of-concept study, which focuses on epigenetic mechanisms, was conducted in a well-studied set of 132 acute lymphoblastic leukemia (ALL) microarrays annotated with nine distinct phenotypes and three measures of response to therapy. We used established parametric and non parametric statistics to derive the PGnet tripartite network that consisted of 10 phenotypes and 33 significant clusters of GEMs comprising 535 distinct genes. The significance of PGnet was estimated from empirical p-values, and a robust subnetwork derived from ALL outcome data was produced by repeated random sampling. The evaluation of derived robust network to predict outcome (relapse of ALL) was significant (p = 3%), using one hundred three-fold cross-validations and the shrunken centroids classifier. CONCLUSION: To our knowledge, this is the first method predicting co-expression networks of genes associated with epigenetic mechanisms and to demonstrate its inherent capability to predict therapeutic outcome. This PGnet approach can be applied to any regulatory mechanisms including transcriptional or microRNA regulation in order to derive predictive molecular profiles that are mechanistically anchored. The implementation of PGnet in R is freely available at http://Lussierlab.org/publication/PGnet. Xinan Yang, James L. Chen, Jianming Xie, Xiao Sun 0006, Yves A. Lussier |
BMC Bioinform. | 5 |
| 2007 | Meta-analysis of several gene lists for distinct types of cancer: A simple way to reveal common prognostic markersabstractBACKGROUND: Although prognostic biomarkers specific for particular cancers have been discovered, microarray analysis of gene expression profiles, supported by integrative analysis algorithms, helps to identify common factors in molecular oncology. Similarities of Ordered Gene Lists (SOGL) is a recently proposed approach to meta-analysis suitable for identifying features shared by two data sets. Here we extend the idea of SOGL to the detection of significant prognostic marker genes from microarrays of multiple data sets. Three data sets for leukemia and the other six for different solid tumors are used to demonstrate our method, using established statistical techniques. RESULTS: We describe a set of significantly similar ordered gene lists, representing outcome comparisons for distinct types of cancer. This kind of similarity could improve the diagnostic accuracies of individual studies when SOGL is incorporated into the support vector machine algorithm. In particular, we investigate the similarities among three ordered gene lists pertaining to mesothelioma survival, prostate recurrence and glioma survival. The similarity-driving genes are related to the outcomes of patients with lung cancer with a hazard ratio of 4.47 (p = 0.035). Many of these genes are involved in breakdown of EMC proteins regulating angiogenesis, and may be used for further research on prognostic markers and molecular targets of gene therapy for cancers. CONCLUSION: The proposed method and its application show the potential of such meta-analyses in clinical studies of gene expression profiles. Xinan Yang, Xiao Sun 0006 |
BMC Bioinform. | 2 |
| 2006 | Support vector machine for classification of meiotic recombination hotspots and coldspots in Saccharomyces cerevisiaebased on codon compositionabstractBACKGROUND: Meiotic double-strand breaks occur at relatively high frequencies in some genomic regions (hotspots) and relatively low frequencies in others (coldspots). Hotspots and coldspots are receiving increasing attention in research into the mechanism of meiotic recombination. However, predicting hotspots and coldspots from DNA sequence information is still a challenging task. RESULTS: We present a novel method for classification of hot and cold ORFs located in hotspots and coldspots respectively in Saccharomyces cerevisiae, using support vector machine (SVM), which relies on codon composition differences. This method has achieved a high classification accuracy of 85.0%. Since codon composition is a fusion of codon usage bias and amino acid composition signals, the ability of these two kinds of sequence attributes to discriminate hot ORFs from cold ORFs was also investigated separately. Our results indicate that neither codon usage bias nor amino acid composition taken separately performed as well as codon composition. Moreover, our SVM based method was applied to the full genome: We predicted the hot/cold ORFs from the yeast genome by using cutoffs of recombination rate. We found that the performance of our method for predicting cold ORFs is not as good as that for predicting hot ORFs. Besides, we also observed a considerable correlation between meiotic recombination rate and amino acid composition of certain residues, which probably reflects the structural and functional dissimilarity between the hot and cold groups. CONCLUSION: We have introduced a SVM-based novel method to discriminate hot ORFs from cold ones. Applying codon composition as sequence attributes, we have achieved a high classification accuracy, which suggests that codon composition has strong potential to be used as sequence attributes in the prediction of hot and cold ORFs. Jianhong Weng, Xiao Sun 0006, Zuhong Lu |
BMC Bioinform. | 3 |