Guoli Ji

dblp:14/479 · DBLP profile ↗
← Back
36ranked-venue papers
10as first author
6since 2021 · last 2023
0000-0002-0474-5523ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 12 · 3 first-authorHuman-computer interaction and ubiquitous computing · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2023 Single-nucleus gene and gene set expression-based similarity network fusion identifies autism molecular subtypes
abstract
BACKGROUND: Autism spectrum disorder (ASD) is a complex neurodevelopmental disorder that is highly phenotypically and genetically heterogeneous. With the accumulation of biological sequencing data, more and more studies shift to molecular subtype-first approach, from identifying molecular subtypes based on genetic and molecular data to linking molecular subtypes with clinical manifestation, which can reduce heterogeneity before phenotypic profiling. RESULTS: In this study, we perform similarity network fusion to integrate gene and gene set expression data of multiple human brain cell types for ASD molecular subtype identification. Then we apply subtype-specific differential gene and gene set expression analyses to study expression patterns specific to molecular subtypes in each cell type. To demonstrate the biological and practical significance, we analyze the molecular subtypes, investigate their correlation with ASD clinical phenotype, and construct ASD molecular subtype prediction models. CONCLUSIONS: The identified molecular subtype-specific gene and gene set expression may be used to differentiate ASD molecular subtypes, facilitating the diagnosis and treatment of ASD. Our method provides an analytical pipeline for the identification of molecular subtypes and even disease subtypes of complex disorders.
Guoli Ji, Xilin Gao, Jinting Guan
BMC Bioinform.2
2022 scIAE: an integrative autoencoder-based ensemble classification framework for single-cell RNA-seq data
abstract
Single-cell RNA sequencing (scRNA-seq) allows quantitative analysis of gene expression at the level of single cells, beneficial to study cell heterogeneity. The recognition of cell types facilitates the construction of cell atlas in complex tissues or organisms, which is the basis of almost all downstream scRNA-seq data analyses. Using disease-related scRNA-seq data to perform the prediction of disease status can facilitate the specific diagnosis and personalized treatment of disease. Since single-cell gene expression data are high-dimensional and sparse with dropouts, we propose scIAE, an integrative autoencoder-based ensemble classification framework, to firstly perform multiple random projections and apply integrative and devisable autoencoders (integrating stacked, denoising and sparse autoencoders) to obtain compressed representations. Then base classifiers are built on the lower-dimensional representations and the predictions from all base models are integrated. The comparison of scIAE and common feature extraction methods shows that scIAE is effective and robust, independent of the choice of dimension, which is beneficial to subsequent cell classification. By testing scIAE on different types of data and comparing it with existing general and single-cell-specific classification methods, it is proven that scIAE has a great classification power in cell type annotation intradataset, across batches, across platforms and across species, and also disease status prediction. The architecture of scIAE is flexible and devisable, and it is available at https://github.com/JGuan-lab/scIAE.
Qingyang Yin, Jinting Guan, Guoli Ji
Briefings Bioinform.4
2021 scAPAtrap: identification and quantification of alternative polyadenylation sites from single-cell RNA-seq data
abstract
Alternative polyadenylation (APA) generates diverse mRNA isoforms, which contributes to transcriptome diversity and gene expression regulation by affecting mRNA stability, translation and localization in cells. The rapid development of 3' tag-based single-cell RNA-sequencing (scRNA-seq) technologies, such as CEL-seq and 10x Genomics, has led to the emergence of computational methods for identifying APA sites and profiling APA dynamics at single-cell resolution. However, existing methods fail to detect the precise location of poly(A) sites or sites with low read coverage. Moreover, they rely on priori genome annotation and can only detect poly(A) sites located within or near annotated genes. Here we proposed a tool called scAPAtrap for detecting poly(A) sites at the whole genome level in individual cells from 3' tag-based scRNA-seq data. scAPAtrap incorporates peak identification and poly(A) read anchoring, enabling the identification of the precise location of poly(A) sites, even for sites with low read coverage. Moreover, scAPAtrap can identify poly(A) sites without using priori genome annotation, which helps locate novel poly(A) sites in previously overlooked regions and improve genome annotation. We compared scAPAtrap with two latest methods, scAPA and Sierra, using scRNA-seq data from different experimental technologies and species. Results show that scAPAtrap identified poly(A) sites with higher accuracy and sensitivity than competing methods and could be used to explore APA dynamics among cell types or the heterogeneous APA isoform expression in individual cells. scAPAtrap is available at https://github.com/BMILAB/scAPAtrap.
Congting Ye, Wenbin Ye 0002, Guoli Ji
Briefings Bioinform.5
2021 QuantifyPoly(A): reshaping alternative polyadenylation landscapes of eukaryotes with weighted density peak clustering
abstract
The dynamic choice of different polyadenylation sites in a gene is referred to as alternative polyadenylation, which functions in many important biological processes. Large-scale messenger RNA 3' end sequencing has revealed that cleavage sites for polyadenylation are presented with microheterogeneity. To date, the conventional determination of polyadenylation site clusters is subjective and arbitrary, leading to inaccurate annotations. Here, we present a weighted density peak clustering method, QuantifyPoly(A), to accurately quantify genome-wide polyadenylation choices. Applying QuantifyPoly(A) on published 3' end sequencing datasets from both animals and plants, their polyadenylation profiles are reshaped into myriads of novel polyadenylation site clusters. Most of these novel polyadenylation site clusters show significantly dynamic usage across different biological samples or associate with binding sites of trans-acting factors. Upstream sequences of these clusters are enriched with polyadenylation signals UGUA, UAAA and/or AAUAAA in a species-dependent manner. Polyadenylation site clusters also exhibit species specificity, while plants ones generally show higher microheterogeneity than that of animals. QuantifyPoly(A) is broadly applicable to any types of 3' end sequencing data and species for accurate quantification and construction of the complex and dynamic polyadenylation landscape and enables us to decode alternative polyadenylation events invisible to conventional methods at a much higher resolution.
Congting Ye, Danhui Zhao, Wenbin Ye 0002, Guoli Ji, Qingshun Quinn Li, Juncheng Lin
Briefings Bioinform.5
2021 movAPA: modeling and visualization of dynamics of alternative polyadenylation across biological samples
abstract
MOTIVATION: Alternative polyadenylation (APA) has been widely recognized as a widespread mechanism modulated dynamically. Studies based on 3' end sequencing and/or RNA-seq have profiled poly(A) sites in various species with diverse pipelines, yet no unified and easy-to-use toolkit is available for comprehensive APA analyses. RESULTS: We developed an R package called movAPA for modeling and visualization of dynamics of alternative polyadenylation across biological samples. movAPA incorporates rich functions for preprocessing, annotation and statistical analyses of poly(A) sites, identification of poly(A) signals, profiling of APA dynamics and visualization. Particularly, seven metrics are provided for measuring the tissue-specificity or usages of APA sites across samples. Three methods are used for identifying 3' UTR shortening/lengthening events between conditions. APA site switching involving non-3' UTR polyadenylation can also be explored. Using poly(A) site data from rice and mouse sperm cells, we demonstrated the high scalability and flexibility of movAPA in profiling APA dynamics across tissues and single cells. AVAILABILITY AND IMPLEMENTATION: https://github.com/BMILAB/movAPA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wenbin Ye 0002, Hongjuan Fu, Congting Ye, Guoli Ji
Bioinform.5
2021 scLINE: A multi-network integration framework based on network embedding for representation of single-cell RNA-seq data
Huoyou Li, Xuesong Xiao, Lishan Ye, Guoli Ji
J. Biomed. Informatics5
2020 Gene Screening for Autism Based on Cell-type-specific Predictive Models
abstract
Autism spectrum disorder (ASD), with substantial genetic and phenotypic heterogeneity, is characterized by difficulties in social interaction and communication, and restricted and repetitive behaviors. Recent studies based on bulk RNA-seq data of brains from ASD patients have revealed the affected pathways in ASD, while the cell type heterogeneity of ASD is still needed to be explored. Gene prioritization studies can be conducted for screening gene candidates with high confidence, providing new insights for experimental studies. Based on the single-nucleus RNA-seq data of brains from ASD and healthy individuals, we identify cell-type-specific differential expressed genes by applying three kinds of methods for differential expression analysis and then construct cell-type-specific classification models for ASD adopting the algorithm of stochastic gradient boosting. We find layer 2/3 and 4 excitatory neurons, layer 5/6 cortico-cortical projection neurons, and protoplasmic astrocytes are vulnerable in ASD. Then we calculate gene importance to prioritize cell-type-specific differential expressed genes, and compare the top important genes across different cell types. Our results suggest that causal genes are distinct and dysregulated gene functions are different across brain cells in ASD. The constructed classification models can predict the diagnosis for a nucleus with given cell type, promoting the detection of ASD. The prioritized cell-type-specific genes may be used as potential ASD biomarkers, promoting the development of effective interventions.
Yiping Lin, Guoli Ji, Jinting Guan
BIBM3
2020 A survey on identification and quantification of alternative polyadenylation sites from RNA-seq data
abstract
Alternative polyadenylation (APA) has been implicated to play an important role in post-transcriptional regulation by regulating mRNA abundance, stability, localization and translation, which contributes considerably to transcriptome diversity and gene expression regulation. RNA-seq has become a routine approach for transcriptome profiling, generating unprecedented data that could be used to identify and quantify APA site usage. A number of computational approaches for identifying APA sites and/or dynamic APA events from RNA-seq data have emerged in the literature, which provide valuable yet preliminary results that should be refined to yield credible guidelines for the scientific community. In this review, we provided a comprehensive overview of the status of currently available computational approaches. We also conducted objective benchmarking analysis using RNA-seq data sets from different species (human, mouse and Arabidopsis) and simulated data sets to present a systematic evaluation of 11 representative methods. Our benchmarking study showed that the overall performance of all tools investigated is moderate, reflecting that there is still lot of scope to improve the prediction of APA site or dynamic APA events from RNA-seq data. Particularly, prediction results from individual tools differ considerably, and only a limited number of predicted APA sites or genes are common among different tools. Accordingly, we attempted to give some advice on how to assess the reliability of the obtained results. We also proposed practical recommendations on the appropriate method applicable to diverse scenarios and discussed implications and future directions relevant to profiling APA from RNA-seq data.
Moliang Chen, Guoli Ji, Hongjuan Fu, Qianmin Lin, Congting Ye, Wenbin Ye 0002, Yaru Su
Briefings Bioinform.2
2020 scHinter: imputing dropout events for single-cell RNA-seq data with limited sample size
abstract
MOTIVATION: Single-cell RNA-sequencing (scRNA-seq) is fast and becoming a powerful technique for studying dynamic gene regulation at unprecedented resolution. However, scRNA-seq data suffer from problems of extremely high dropout rate and cell-to-cell variability, demanding new methods to recover gene expression loss. Despite the availability of various dropout imputation approaches for scRNA-seq, most studies focus on data with a medium or large number of cells, while few studies have explicitly investigated the differential performance across different sample sizes or the applicability of the approach on small or imbalanced data. It is imperative to develop new imputation approaches with higher generalizability for data with various sample sizes. RESULTS: We proposed a method called scHinter for imputing dropout events for scRNA-seq with special emphasis on data with limited sample size. scHinter incorporates a voting-based ensemble distance and leverages the synthetic minority oversampling technique for random interpolation. A hierarchical framework is also embedded in scHinter to increase the reliability of the imputation for small samples. We demonstrated the ability of scHinter to recover gene expression measurements across a wide spectrum of scRNA-seq datasets with varied sample sizes. We comprehensively examined the impact of sample size and cluster number on imputation. Comprehensive evaluation of scHinter across diverse scRNA-seq datasets with imbalanced or limited sample size showed that scHinter achieved higher and more robust performance than competing approaches, including MAGIC, scImpute, SAVER and netSmooth. AVAILABILITY AND IMPLEMENTATION: Freely available for download at https://github.com/BMILAB/scHinter. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pengchao Ye, Wenbin Ye 0002, Congting Ye, Shuchao Li, Lishan Ye, Guoli Ji
Bioinform.6
2020 scDAPA: detection and visualization of dynamic alternative polyadenylation from single cell RNA-seq data
abstract
MOTIVATION: Alternative polyadenylation (APA) plays a key post-transcriptional regulatory role in mRNA stability and functions in eukaryotes. Single cell RNA-seq (scRNA-seq) is a powerful tool to discover cellular heterogeneity at gene expression level. Given 3' enriched strategy in library construction, the most commonly used scRNA-seq protocol-10× Genomics enables us to improve the study resolution of APA to the single cell level. However, currently there is no computational tool available for investigating APA profiles from scRNA-seq data. RESULTS: Here, we present a package scDAPA for detecting and visualizing dynamic APA from scRNA-seq data. Taking bam/sam files and cell cluster labels as inputs, scDAPA detects APA dynamics using a histogram-based method and the Wilcoxon rank-sum test, and visualizes candidate genes with dynamic APA. Benchmarking results demonstrated that scDAPA can effectively identify genes with dynamic APA among different cell groups from scRNA-seq data. AVAILABILITY AND IMPLEMENTATION: The scDAPA package is implemented in Shell and R, and is freely available at https://scdapa.sourceforge.io. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Congting Ye, Chen Yu 0007, Guoli Ji, Daniel R. Saban, Qingshun Quinn Li
Bioinform.5
2019 Human brain cell type-specific gene co-expression associated with autism spectrum disorder
abstract
Autism spectrum disorder (ASD) is a complex neuropsychiatric disorder with substantial phenotypic and genetic heterogeneity. Until now, about a thousand diverse genes have been associated with ASD, while it remains elusive that how disruptions in these different genes can lead to a common clinical phenotype. Therefore, it is essential to understand how ASD candidate genes relate to each other and identify potential shared molecular pathways. Human brain is a highly heterogeneous organ involving multiple cell types. Different functions in different types of cells may be dysregulated in ASD; investigating functional interactions between ASD candidate genes in normal human brain cells may shed new light on the genetic heterogeneity of ASD. To this end, we construct cell type-associated gene co-expression networks based on human brain cell gene expression data. Then we identify seven cell type-specific gene modules and analyze the specific gene functions in each cell type. We also identify six ASD-associated gene modules and study the dysregulated functions in ASD. Lastly, we obtain two ASD-associated cell type-specific gene modules for studying the cell type-specific aberrant functions in ASD. It is found that ASD-associated astrocytes-specific gene modules are relevant to endocytosis, neuron differentiation and cell projection organization, while ASD-associated neurons-specific gene modules are relevant to presynapse, glutamatergic synapse and neuron projection morphogenesis. Our method has been proven to be effective in discovering ASD-associated cell type-specific gene expression pattern. Our findings can promote the study of the heterogeneity of ASD in gene expression between different cell types, providing new insights into the molecular mechanisms underlying the pathogenesis of ASD.
Yiping Lin, Shuchao Li, Guoli Ji, Jinting Guan
BIBM3
2019 AStrap: identification of alternative splicing from transcript sequences without a reference genome
abstract
SUMMARY: Alternative splicing (AS) is a well-established mechanism for increasing transcriptome and proteome diversity, however, detecting AS events and distinguishing among AS types in organisms without available reference genomes remains challenging. We developed a de novo approach called AStrap for AS analysis without using a reference genome. AStrap identifies AS events by extensive pair-wise alignments of transcript sequences and predicts AS types by a machine-learning model integrating more than 500 assembled features. We evaluated AStrap using collected AS events from reference genomes of rice and human as well as single-molecule real-time sequencing data from Amborella trichopoda. Results show that AStrap can identify much more AS events with comparable or higher accuracy than the competing method. AStrap also possesses a unique feature of predicting AS types, which achieves an overall accuracy of ∼0.87 for different species. Extensive evaluation of AStrap using different parameters, sample sizes and machine-learning models on different species also demonstrates the robustness and flexibility of AStrap. AStrap could be a valuable addition to the community for the study of AS in non-model organisms with limited genetic resources. AVAILABILITY AND IMPLEMENTATION: AStrap is available for download at https://github.com/BMILAB/AStrap. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Guoli Ji, Wenbin Ye 0002, Yaru Su, Moliang Chen, Guangzao Huang
Bioinform.1
2019 Partial maximum correlation information: A new feature selection method for microarray data classification
Mingshun Yuan, Zijiang Yang 0001, Guoli Ji
Neurocomputing3
2019 Prediction of DNA-binding proteins by interaction fusion feature representation and selective ensemble
Wenjie You, Zijiang Yang 0001, Guangbao Guo, Xiu-Feng Wan, Guoli Ji
Knowl. Based Syst.5
2018 AEGS: identifying aberrantly expressed gene sets for differential variability analysis
abstract
Motivation: In gene expression studies, differential expression (DE) analysis has been widely used to identify genes with shifted expression mean between groups. Recently, differential variability (DV) analysis has been increasingly applied as analyzing changed expression variability (e.g. the changes in expression variance) between groups may reveal underlying genetic heterogeneity and undetected interactions, which has great implications in many fields of biology. An easy-to-use tool for DV analysis is needed. Results: We develop AEGS for DV analysis, to identify aberrantly expressed gene sets in diseased cases but not in controls. AEGS can rank individual genes in an aberrantly expressed gene set by each gene's relative contribution to the total degree of aberrant expression, prioritizing top genes. AEGS can be used for discovering gene sets with disease-specific expression variability changes. Availability and implementation: AEGS web server is accessible at http://bmi.xmu.edu.cn:8003/AEGS, where a stand-alone AEGS application can also be downloaded. Contact: [email protected].
Jinting Guan, Moliang Chen, Congting Ye, James J. Cai, Guoli Ji
Bioinform.5
2018 TSAPA: identification of tissue-specific alternative polyadenylation sites in plants
abstract
Summary: Alternative polyadenylation (APA) is now emerging as a widespread mechanism modulated tissue-specifically, which highlights the need to define tissue-specific poly(A) sites for profiling APA dynamics across tissues. We have developed an R package called TSAPA based on the machine learning model for identifying tissue-specific poly(A) sites in plants. A feature space including more than 200 features was assembled to specifically characterize poly(A) sites in plants. The classification model in TSAPA can be customized by selecting desirable features or classifiers. TSAPA is also capable of predicting tissue-specific poly(A) sites in unannotated intergenic regions. TSAPA will be a valuable addition to the community for studying dynamics of APA in plants. Availability and implementation: https://github.com/BMILAB/TSAPA. Supplementary information: Supplementary data are available at Bioinformatics online.
Guoli Ji, Moliang Chen, Wenbin Ye 0002, Congting Ye, Yaru Su
Bioinform.1
2018 APAtrap: identification and quantification of alternative polyadenylation sites from RNA-seq data
abstract
Motivation: Alternative polyadenylation (APA) has been increasingly recognized as a crucial mechanism that contributes to transcriptome diversity and gene expression regulation. As RNA-seq has become a routine protocol for transcriptome analysis, it is of great interest to leverage such unprecedented collection of RNA-seq data by new computational methods to extract and quantify APA dynamics in these transcriptomes. However, research progress in this area has been relatively limited. Conventional methods rely on either transcript assembly to determine transcript 3' ends or annotated poly(A) sites. Moreover, they can neither identify more than two poly(A) sites in a gene nor detect dynamic APA site usage considering more than two poly(A) sites. Results: We developed an approach called APAtrap based on the mean squared error model to identify and quantify APA sites from RNA-seq data. APAtrap is capable of identifying novel 3' UTRs and 3' UTR extensions, which contributes to locating potential poly(A) sites in previously overlooked regions and improving genome annotations. APAtrap also aims to tally all potential poly(A) sites and detect genes with differential APA site usages between conditions. Extensive comparisons of APAtrap with two other latest methods, ChangePoint and DaPars, using various RNA-seq datasets from simulation studies, human and Arabidopsis demonstrate the efficacy and flexibility of APAtrap for any organisms with an annotated genome. Availability and implementation: Freely available for download at https://apatrap.sourceforge.io. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Congting Ye, Yuqi Long, Guoli Ji, Qingshun Quinn Li
Bioinform.3
2017 An innovative one-class least squares support vector machine model based on continuous cognition
Guangzao Huang, Zijiang Yang 0001, Guoli Ji
Knowl. Based Syst.4
2017 Feature selection by maximizing correlation information for integrated high-dimensional protein data
Mingshun Yuan, Zijiang Yang 0001, Guangzao Huang, Guoli Ji
Pattern Recognit. Lett.4
2015 Genome-wide identification and predictive modeling of polyadenylation sites in eukaryotes
abstract
Polyadenylation [poly(A)] is a vital step in post-transcriptional processing of pre-mRNA. Alternative polyadenylation is a widespread mechanism of regulating gene expression in eukaryotes. Defining poly(A) sites contributes to the annotation of transcripts' ends and the study of gene regulatory mechanisms. Here, we survey methods for collecting poly(A) sites using high-throughput sequencing technologies and summarize the general processes for genome-wide poly(A) site identifications. We also compare the performances of various poly(A) site prediction models and discuss the relationship between poly(A) site identification from sequencing projects and predictive modeling. Moreover, we attempt to address some potential problems in current researches and propose future directions related to polyadenylation research.
Guoli Ji, Jinting Guan, Qingshun Quinn Li
Briefings Bioinform.1
2015 PASPA: a web server for mRNA poly(A) site predictions in plants and algae
abstract
Abstract Motivation: Polyadenylation is an essential process during eukaryotic gene expression. Prediction of poly(A) sites helps to define the 3′ end of genes, which is important for gene annotation and elucidating gene regulation mechanisms. However, due to limited knowledge of poly(A) signals, it is still challenging to predict poly(A) sites in plants and algae. PASPA is a web server for poly(A) site prediction in plants and algae, which integrates many in-house tools as add-ons to facilitate poly(A) site prediction, visualization and mining. This server can predict poly(A) sites for ten species, including seven previously poly(A) signal non-characterized species, with sensitivity and specificity in a range between 0.80 and 0.95. Availability and implementation: http://bmi.xmu.edu.cn/paspa Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Guoli Ji, Lei Li 0060, Qingshun Quinn Li
Bioinform.1
2014 Feature selection for high-dimensional multi-category data using PLS-based local recursive feature elimination
Wenjie You, Zijiang Yang 0001, Guoli Ji
Expert Syst. Appl.3
2014 Solving partially observable problems with inaccurate PSR models
Yunlong Liu 0003, Zijiang Yang 0001, Guoli Ji
Inf. Sci.3
2014 PLS-based recursive feature elimination for high-dimensional small sample
Wenjie You, Zijiang Yang 0001, Guoli Ji
Knowl. Based Syst.3
2013 Identification and predictive control for a circulation fluidized bed boiler
Guoli Ji, Jiangyin Huang, Kangkang Zhang, Yucai Zhu, Tianxiao Ji, Sun Zhou 0001
Knowl. Based Syst.1
2012 CBrowse: a SAM/BAM-based contig browser for transcriptome assembly visualization and analysis
abstract
SUMMARY: To address the impending need for exploring rapidly increased transcriptomics data generated for non-model organisms, we developed CBrowse, an AJAX-based web browser for visualizing and analyzing transcriptome assemblies and contigs. Designed in a standard three-tier architecture with a data pre-processing pipeline, CBrowse is essentially a Rich Internet Application that offers many seamlessly integrated web interfaces and allows users to navigate, sort, filter, search and visualize data smoothly. The pre-processing pipeline takes the contig sequence file in FASTA format and its relevant SAM/BAM file as the input; detects putative polymorphisms, simple sequence repeats and sequencing errors in contigs and generates image, JSON and database-compatible CSV text files that are directly utilized by different web interfaces. CBowse is a generic visualization and analysis tool that facilitates close examination of assembly quality, genetic polymorphisms, sequence repeats and/or sequencing errors in transcriptome sequencing projects. AVAILABILITY: CBrowse is distributed under the GNU General Public License, available at http://bioinfolab.muohio.edu/CBrowse/ CONTACT: [email protected] or [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Guoli Ji, Emily Schmidt, Douglas Lenox, Qi Liu 0019, Lin Liu 0001, Chun Liang
Bioinform.2
2012 A novel two-level nearest neighbor classification algorithm using an adaptive distance metric
Yunlong Gao 0001, Guoli Ji, Zijiang Yang 0001
Knowl. Based Syst.3
2012 A Dynamic AdaBoost Algorithm With Adaptive Changes of Loss Function
abstract
AdaBoost is a method to improve a given learning algorithm's classification accuracy by combining its hypotheses. Adaptivity, one of the significant advantages of AdaBoost, makes AdaBoost maximize the smallest margin so that AdaBoost has good generalization ability. However, when the samples with large negative margins are noisy or atypical, the maximized margin is actually a “hard margin.” The adaptive feature makes AdaBoost sensitive to the sampling fluctuations, and prone to overfitting. Therefore, the traditional schemes prevent AdaBoost from overfitting by heavily damping the influences of samples with large negative margins. However, the samples with large negative margins are not always noisy or atypical; thus, the traditional schemes of preventing overfitting may not be reasonable. In order to learn a classifier with high generalization performance and prevent overfitting, it is necessary to perform statistical analysis for the margins of training samples. Herein, Hoeffding inequality is adopted as a statistical tool to divide training samples into reliable samples and temporary unreliable samples. A new boosting algorithm, which is named DAdaBoost, is introduced to deal with reliable samples and temporary unreliable samples separately. Since DAdaBoost adjusts weighting scheme dynamically, the loss function of DAdaBoost is not fixed. In fact, it is a series of nonconvex functions that gradually approach the 0-1 function as the algorithm evolves. By defining a virtual classifier, the dynamic adjusted weighting scheme is well unified into the progress of DAdaBoost, and the upper bound of training error is deduced. The experiments on both synthetic and real world data show that DAdaBoost has many merits. Based on the experiments, we conclude that DAdaBoost can effectively prevent AdaBoost from overfitting.
Yunlong Gao 0001, Guoli Ji, Zijiang Yang 0001
IEEE Trans. Syst. Man Cybern. Part C2
2011 Identification of plant messenger RNA polyadenylation sites using length-variable second order Markov model
abstract
In this paper we adopted a length-variable second order Markov model to identify plant messenger RNA poly(A) sites, and provided a common method that only relies on the experimental sequences. The efficacy of our model is showed up to 92% sensitivity and 79% specificity. This method is particularly suitable for the prediction of the poly(A) site which is lack of biological priori knowledge and has poor conservative signal characteristic, as well as for the identification of the alternative poly(A) sites in different genetic regions. Compared with other algorithms, generalized hidden Markov model needed the signal distributions and AdaBoost required the construction of signal features around the sites, our model is more versatile.
Guoli Ji, Huanghui Zhang, Meishuang Tang
SMC1
2011 Using partial least squares and support vector machines for bankruptcy prediction
Zijiang Yang 0001, Wenjie You, Guoli Ji
Expert Syst. Appl.3
2011 Hybrid intelligent control scheme of a polymerization kettle for ACR production
Sun Zhou 0001, Guoli Ji, Zijiang Yang 0001
Knowl. Based Syst.2
2011 PLS-Based Gene Selection and Identification of Tumor-Specific Genes
abstract
In view of the characteristics of high-dimensional small sample, strong relevance, and high noise of the identification of tumor-specific genes on microarray, a novel partial least squares (PLS) based gene-selection method, which synthesizes genetic relatedness and is suitable for multicategory classification, is presented. Using the explanation difference of independent variables on dependent variable (class), we define three indicators for global gene selection, which takes into accounts the combined effects of all the genes and the correlation among the genes. Integrated with the linear kernel support vector classifier (SVC), the proposed method is tested by MIT acute myeloid leukemia/acute lymphoblastic leukemia (AML/ALL) and small round blue cell tumors (SRBCT) data sets. A subset of specific genes with small numbers and high identification are obtained. The results indicate that our proposed PLS-based method for tumor-specific genes selection is highly efficient. Compared to the literature, the selected specific genes from both two-category dataset AML/ALL and multicategory dataset SRBCT are credible. Further investigation shows that the proposed gene-selection method is robust. Overall, the proposed method can effectively solve feature-selection problem on high-dimensional small sample. At the same time, it has good performance for multicategory classification as well.
Guoli Ji, Zijiang Yang 0001, Wenjie You
IEEE Trans. Syst. Man Cybern. Part C1
2010 Messenger RNA Polyadenylation Site Recognition in Green Alga Chlamydomonas Reinhardtii
Guoli Ji, Qingshun Quinn Li, Jianti Zheng
ISNN (1)1
2010 A new method of information decision-making based on D-S evidence theory
abstract
D-S evidence theory is a method broadly applied in fusion for decision-making. However, this theory has some shortcomings in the formula of evidence combination with the exception that evidence of fully conflict can not be combined, then the probability validity is difficult to determine and sometimes the composed evidence is different from people's subjective judgments or some other issues. These confine the application of evidence to some extent. Some of them have the dubious credibility which affects the fusion result when Multi evidence are combined together. In order to expand the application of the formula of this theory and enhance the reliability of the fusion results, a new combination formula is introduced in this paper, which is also compared with other formulas in other literatures and finally the improved reliability of this combination formula is verified. At last, through data-mining of the decision-making information on a number of isolated points, a new method using combined evidence to make decisions is described. It's proven from the experimental results that the new combination method not only works well and effectively in the evidence of a high level of conflict but also is applicable to fusion for decision-making.
Junfeng Yao, Chengpeng Wu, Xiaobiao Xie, Guoli Ji, Prabir Bhattacharya
SMC5
2009 A Novel Method for Progressive Multiple Sequence Alignment Based on Lempel-Ziv
Guoli Ji, Congting Ye, Zijiang Yang 0001, Zhenya Guo
ICONIP (1)1
2007 Predictive modeling of plant messenger RNA polyadenylation sites
abstract
BACKGROUND: One of the essential processing events during pre-mRNA maturation is the post-transcriptional addition of a polyadenine [poly(A)] tail. The 3'-end poly(A) track protects mRNA from unregulated degradation, and indicates the integrity of mRNA through recognition by mRNA export and translation machinery. The position of a poly(A) site is predetermined by signals in the pre-mRNA sequence that are recognized by a complex of polyadenylation factors. These signals are generally tri-part sequence patterns around the cleavage site that serves as the future poly(A) site. In plants, there is little sequence conservation among these signal elements, which makes it difficult to develop an accurate algorithm to predict the poly(A) site of a given gene. We attempted to solve this problem. RESULTS: Based on our current working model and the profile of nucleotide sequence distribution of the poly(A) signals and around poly(A) sites in Arabidopsis, we have devised a Generalized Hidden Markov Model based algorithm to predict potential poly(A) sites. The high specificity and sensitivity of the algorithm were demonstrated by testing several datasets, and at the best combinations, both reach 97%. The accuracy of the program, called poly(A) site sleuth or PASS, has been demonstrated by the prediction of many validated poly(A) sites. PASS also predicted the changes of poly(A) site efficiency in poly(A) signal mutants that were constructed and characterized by traditional genetic experiments. The efficacy of PASS was demonstrated by predicting poly(A) sites within long genomic sequences. CONCLUSION: Based on the features of plant poly(A) signals, a computational model was built to effectively predict the poly(A) sites in Arabidopsis genes. The algorithm will be useful in gene annotation because a poly(A) site signifies the end of the transcript. This algorithm can also be used to predict alternative poly(A) sites in known genes, and will be useful in the design of transgenes for crop genetic engineering by predicting and eliminating undesirable poly(A) sites.
Guoli Ji, Jianti Zheng, Yingjia Shen, Ronghan Jiang, Johnny C. Loke, Kimberly M. Davis, Greg J. Reese, Qingshun Quinn Li
BMC Bioinform.1