EDBT 2026 Demo / reviewers in the wild / expert
Wei Chen 0064
dblp:181/2832-64
· DBLP profile ↗
22ranked-venue papers
3as first author
6since 2021 · last 2024
0000-0002-6857-7696ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ConvNeXt-MHC: improving MHC-peptide affinity prediction by structure-derived degenerate coding and the ConvNeXt modelabstractPeptide binding to major histocompatibility complex (MHC) proteins plays a critical role in T-cell recognition and the specificity of the immune response. Experimental validation such peptides is extremely resource-intensive. As a result, accurate computational prediction of binding peptides is highly important, particularly in the context of cancer immunotherapy applications, such as the identification of neoantigens. In recent years, there is a significant need to continually improve the existing prediction methods to meet the demands of this field. We developed ConvNeXt-MHC, a method for predicting MHC-I-peptide binding affinity. It introduces a degenerate encoding approach to enhance well-established panspecific methods and integrates transfer learning and semi-supervised learning methods into the cutting-edge deep learning framework ConvNeXt. Comprehensive benchmark results demonstrate that ConvNeXt-MHC outperforms state-of-the-art methods in terms of accuracy. We expect that ConvNeXt-MHC will help us foster new discoveries in the field of immunoinformatics in the distant future. We constructed a user-friendly website at http://www.combio-lezhang.online/predict/, where users can access our data and application. Le Zhang 0004, Wenkai Song, Tinghao Zhu, Yang Liu 0236, Wei Chen 0064 |
Briefings Bioinform. | 5 |
| 2022 | A merged molecular representation deep learning method for blood-brain barrier permeability predictionabstractThe ability of a compound to permeate across the blood-brain barrier (BBB) is a significant factor for central nervous system drug development. Thus, for speeding up the drug discovery process, it is crucial to perform high-throughput screenings to predict the BBB permeability of the candidate compounds. Although experimental methods are capable of determining BBB permeability, they are still cost-ineffective and time-consuming. To complement the shortcomings of existing methods, we present a deep learning-based multi-model framework model, called Deep-B3, to predict the BBB permeability of candidate compounds. In Deep-B3, the samples are encoded in three kinds of features, namely molecular descriptors and fingerprints, molecular graph and simplified molecular input line entry system (SMILES) text notation. The pre-trained models were built to extract latent features from the molecular graph and SMILES. These features depicted the compounds in terms of tabular data, image and text, respectively. The validation results yielded from the independent dataset demonstrated that the performance of Deep-B3 is superior to that of the state-of-the-art models. Hence, Deep-B3 holds the potential to become a useful tool for drug development. A freely available online web-server for Deep-B3 was established at http://cbcb.cdutcm.edu.cn/deepb3/, and the source code and dataset of Deep-B3 are available at https://github.com/GreatChenLab/Deep-B3. Qiang Tang 0013, Fulei Nie, Qi Zhao 0010, Wei Chen 0064 |
Briefings Bioinform. | 4 |
| 2022 | DeepLncPro: an interpretable convolutional neural network model for identifying long non-coding RNA promotersabstractLong non-coding RNA (lncRNA) plays important roles in a series of biological processes. The transcription of lncRNA is regulated by its promoter. Hence, accurate identification of lncRNA promoter will be helpful to understand its regulatory mechanisms. Since experimental techniques remain time consuming for gnome-wide promoter identification, developing computational tools to identify promoters are necessary. However, only few computational methods have been proposed for lncRNA promoter prediction and their performances still have room to be improved. In the present work, a convolutional neural network based model, called DeepLncPro, was proposed to identify lncRNA promoters in human and mouse. Comparative results demonstrated that DeepLncPro was superior to both state-of-the-art machine learning methods and existing models for identifying lncRNA promoters. Furthermore, DeepLncPro has the ability to extract and analyze transcription factor binding motifs from lncRNAs, which made it become an interpretable model. These results indicate that the DeepLncPro can server as a powerful tool for identifying lncRNA promoters. An open-source tool for DeepLncPro was provided at https://github.com/zhangtian-yang/DeepLncPro. Qiang Tang 0013, Fulei Nie, Qi Zhao 0010, Wei Chen 0064 |
Briefings Bioinform. | 5 |
| 2021 | Iterative feature representation algorithm to improve the predictive performance of N7-methylguanosine sitesabstractMOTIVATION: N7-methylguanosine (m7G) is an important epigenetic modification, playing an essential role in gene expression regulation. Therefore, accurate identification of m7G modifications will facilitate revealing and in-depth understanding their potential functional mechanisms. Although high-throughput experimental methods are capable of precisely locating m7G sites, they are still cost ineffective. Therefore, it's necessary to develop new methods to identify m7G sites. RESULTS: In this work, by using the iterative feature representation algorithm, we developed a machine learning based method, namely m7G-IFL, to identify m7G sites. To demonstrate its superiority, m7G-IFL was evaluated and compared with existing predictors. The results demonstrate that our predictor outperforms existing predictors in terms of accuracy for identifying m7G sites. By analyzing and comparing the features used in the predictors, we found that the positive and negative samples in our feature space were more separated than in existing feature space. This result demonstrates that our features extracted more discriminative information via the iterative feature learning process, and thus contributed to the predictive performance improvement. Chichi Dai, Pengmian Feng, Li-Zhen Cui 0001, Ran Su, Wei Chen 0064, Leyi Wei |
Briefings Bioinform. | 5 |
| 2021 | KNIndex: a comprehensive database of physicochemical properties for k-tuple nucleotidesabstractWith the development of high-throughput sequencing technology, the genomic sequences increased exponentially over the last decade. In order to decode these new genomic data, machine learning methods were introduced for genome annotation and analysis. Due to the requirement of most machines learning methods, the biological sequences must be represented as fixed-length digital vectors. In this representation procedure, the physicochemical properties of k-tuple nucleotides are important information. However, the values of the physicochemical properties of k-tuple nucleotides are scattered in different resources. To facilitate the studies on genomic sequences, we developed the first comprehensive database, namely KNIndex (https://knindex.pufengdu.org), for depositing and visualizing physicochemical properties of k-tuple nucleotides. Currently, the KNIndex database contains 182 properties including one for mononucleotide (DNA), 169 for dinucleotide (147 for DNA and 22 for RNA) and 12 for trinucleotide (DNA). KNIndex database also provides a user-friendly web-based interface for the users to browse, query, visualize and download the physicochemical properties of k-tuple nucleotides. With the built-in conversion and visualization functions, users are allowed to display DNA/RNA sequences as curves of multiple physicochemical properties. We wish that the KNIndex will facilitate the related studies in computational biology. Wen-Ya Zhang, Junhai Xu, Jun Wang 0081, Yuan-Ke Zhou, Wei Chen 0064, Pu-Feng Du |
Briefings Bioinform. | 5 |
| 2021 | Design powerful predictor for mRNA subcellular location prediction in Homo sapiensabstractMessenger RNAs (mRNAs) shoulder special responsibilities that transmit genetic code from DNA to discrete locations in the cytoplasm. The locating process of mRNA might provide spatial and temporal regulation of mRNA and protein functions. The situ hybridization and quantitative transcriptomics analysis could provide detail information about mRNA subcellular localization; however, they are time consuming and expensive. It is highly desired to develop computational tools for timely and effectively predicting mRNA subcellular location. In this work, by using binomial distribution and one-way analysis of variance, the optimal nonamer composition was obtained to represent mRNA sequences. Subsequently, a predictor based on support vector machine was developed to identify the mRNA subcellular localization. In 5-fold cross-validation, results showed that the accuracy is 90.12% for Homo sapiens (H. sapiens). The predictor may provide a reference for the study of mRNA localization mechanisms and mRNA translocation strategies. An online web server was established based on our models, which is available at http://lin-group.cn/server/iLoc-mRNA/. Zhao-Yue Zhang 0002, Hui Ding 0005, Dong Wang 0011, Wei Chen 0064, Hao Lin 0001 |
Briefings Bioinform. | 5 |
| 2020 | Evaluation of different computational methods on 5-methylcytosine sites identificationabstract5-Methylcytosine (m5C) plays an extremely important role in the basic biochemical process. With the great increase of identified m5C sites in a wide variety of organisms, their epigenetic roles become largely unknown. Hence, accurate identification of m5C site is a key step in understanding its biological functions. Over the past several years, more attentions have been paid on the identification of m5C sites in multiple species. In this work, we firstly summarized the current progresses in computational prediction of m5C sites and then constructed a more powerful and reliable model for identifying m5C sites. To train the model, we collected experimentally confirmed m5C data from Homo sapiens, Mus musculus, Saccharomyces cerevisiae and Arabidopsis thaliana, and compared the performances of different feature extraction methods and classification algorithms for optimizing prediction model. Based on the optimal model, a novel predictor called iRNA-m5C was developed for the recognition of m5C sites. Finally, we critically evaluated the performance of iRNA-m5C and compared it with existing methods. The result showed that iRNA-m5C could produce the best prediction performance. We hope that this paper could provide a guide on the computational identification of m5C site and also anticipate that the proposed iRNA-m5C will become a powerful tool for large scale identification of m5C sites. Hao Lv 0007, Zi-Mei Zhang, Jiu-Xin Tan, Wei Chen 0064, Hao Lin 0001 |
Briefings Bioinform. | 5 |
| 2020 | A comparison and assessment of computational method for identifying recombination hotspots in Saccharomyces cerevisiaeabstractMeiotic recombination is one of the most important driving forces of biological evolution, which is initiated by double-strand DNA breaks. Recombination has important roles in genome diversity and evolution. This review firstly provides a comprehensive survey of the 15 computational methods developed for identifying recombination hotspots in Saccharomyces cerevisiae. These computational methods were discussed and compared in terms of underlying algorithms, extracted features, predictive capability and practical utility. Subsequently, a more objective benchmark data set was constructed to develop a new predictor iRSpot-Pse6NC2.0 (http://lin-group.cn/server/iRSpot-Pse6NC2.0). To further demonstrate the generalization ability of these methods, we compared iRSpot-Pse6NC2.0 with existing methods on the chromosome XVI of S. cerevisiae. The results of the independent data set test demonstrated that the new predictor is superior to existing tools in the identification of recombination hotspots. The iRSpot-Pse6NC2.0 will become an important tool for identifying recombination hotspot. Wuritu Yang, Fu-Ying Dao, Hao Lv 0007, Hui Ding 0005, Wei Chen 0064, Hao Lin 0001 |
Briefings Bioinform. | 6 |
| 2020 | iMRM: a platform for simultaneously identifying multiple kinds of RNA modificationsabstractMOTIVATION: RNA modifications play critical roles in a series of cellular and developmental processes. Knowledge about the distributions of RNA modifications in the transcriptomes will provide clues to revealing their functions. Since experimental methods are time consuming and laborious for detecting RNA modifications, computational methods have been proposed for this aim in the past five years. However, there are some drawbacks for both experimental and computational methods in simultaneously identifying modifications occurred on different nucleotides. RESULTS: To address such a challenge, in this article, we developed a new predictor called iMRM, which is able to simultaneously identify m6A, m5C, m1A, ψ and A-to-I modifications in Homo sapiens, Mus musculus and Saccharomyces cerevisiae. In iMRM, the feature selection technique was used to pick out the optimal features. The results from both 10-fold cross-validation and jackknife test demonstrated that the performance of iMRM is superior to existing methods for identifying RNA modifications. AVAILABILITY AND IMPLEMENTATION: A user-friendly web server for iMRM was established at http://www.bioml.cn/XG_iRNA/home. The off-line command-line version is available at https://github.com/liukeweiaway/iMRM. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kewei Liu, Wei Chen 0064 |
Bioinform. | 2 |
| 2020 | DNA4mC-LIP: a linear integration method to identify N4-methylcytosine site in multiple speciesabstractMOTIVATION: DNA N4-methylcytosine (4mC) is a crucial epigenetic modification. However, the knowledge about its biological functions is limited. Effective and accurate identification of 4mC sites will be helpful to reveal its biological functions and mechanisms. Since experimental methods are cost and ineffective, a number of machine learning-based approaches have been proposed to detect 4mC sites. Although these methods yielded acceptable accuracy, there is still room for the improvement of the prediction performance and the stability of existing methods in practical applications. RESULTS: In this work, we first systematically assessed the existing methods based on an independent dataset. And then, we proposed DNA4mC-LIP, a linear integration method by combining existing predictors to identify 4mC sites in multiple species. The results obtained from independent dataset demonstrated that DNA4mC-LIP outperformed existing methods for identifying 4mC sites. To facilitate the scientific community, a web server for DNA4mC-LIP was developed. We anticipated that DNA4mC-LIP could serve as a powerful computational technique for identifying 4mC sites and facilitate the interpretation of 4mC mechanism. AVAILABILITY AND IMPLEMENTATION: http://i.uestc.edu.cn/DNA4mC-LIP/. CONTACT: [email protected] or [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qiang Tang 0013, Juanjuan Kang, Jiaqing Yuan, Hua Tang, Xianhai Li, Hao Lin 0001, Jian Huang 0004, Wei Chen 0064 |
Bioinform. | 8 |
| 2020 | VisFeature: a stand-alone program for visualizing and analyzing statistical features of biological sequencesabstractSUMMARY: Many efforts have been made in developing bioinformatics algorithms to predict functional attributes of genes and proteins from their primary sequences. One challenge in this process is to intuitively analyze and to understand the statistical features that have been selected by heuristic or iterative methods. In this paper, we developed VisFeature, which aims to be a helpful software tool that allows the users to intuitively visualize and analyze statistical features of all types of biological sequence, including DNA, RNA and proteins. VisFeature also integrates sequence data retrieval, multiple sequence alignments and statistical feature generation functions. AVAILABILITY AND IMPLEMENTATION: VisFeature is a desktop application that is implemented using JavaScript/Electron and R. The source codes of VisFeature are freely accessible from the GitHub repository (https://github.com/wangjun1996/VisFeature). The binary release, which includes an example dataset, can be freely downloaded from the same GitHub repository (https://github.com/wangjun1996/VisFeature/releases). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jun Wang 0081, Pu-Feng Du, Xin-Yu Xue, Guang-Ping Li 0005, Yuan-Ke Zhou, Hao Lin 0001, Wei Chen 0064 |
Bioinform. | 8 |
| 2019 | i6mA-Pred: identifying DNA N6-methyladenine sites in the rice genomeabstractMOTIVATION: DNA N6-methyladenine (6mA) is associated with a wide range of biological processes. Since the distribution of 6mA site in the genome is non-random, accurate identification of 6mA sites is crucial for understanding its biological functions. Although experimental methods have been proposed for this regard, they are still cost-ineffective for detecting 6mA site in genome-wide scope. Therefore, it is desirable to develop computational methods to facilitate the identification of 6mA site. RESULTS: In this study, a computational method called i6mA-Pred was developed to identify 6mA sites in the rice genome, in which the optimal nucleotide chemical properties obtained by the using feature selection technique were used to encode the DNA sequences. It was observed that the i6mA-Pred yielded an accuracy of 83.13% in the jackknife test. Meanwhile, the performance of i6mA-Pred was also superior to other methods. AVAILABILITY AND IMPLEMENTATION: A user-friendly web-server, i6mA-Pred is freely accessible at http://lin-group.cn/server/i6mA-Pred. Wei Chen 0064, Hao Lv 0007, Fulei Nie, Hao Lin 0001 |
Bioinform. | 1 |
| 2019 | Identify origin of replication in Saccharomyces cerevisiae using two-step feature selection techniqueabstractMOTIVATION: DNA replication is a key step to maintain the continuity of genetic information between parental generation and offspring. The initiation site of DNA replication, also called origin of replication (ORI), plays an extremely important role in the basic biochemical process. Thus, rapidly and effectively identifying the location of ORI in genome will provide key clues for genome analysis. Although biochemical experiments could provide detailed information for ORI, it requires high experimental cost and long experimental period. As good complements to experimental techniques, computational methods could overcome these disadvantages. RESULTS: Thus, in this study, we developed a predictor called iORI-PseKNC2.0 to identify ORIs in the Saccharomyces cerevisiae genome based on sequence information. The PseKNC including 90 physicochemical properties was proposed to formulate ORI and non-ORI samples. In order to improve the accuracy, a two-step feature selection was proposed to exclude redundant and noise information. As a result, the overall success rate of 88.53% was achieved in the 5-fold cross-validation test by using support vector machine. AVAILABILITY AND IMPLEMENTATION: Based on the proposed model, a user-friendly webserver was established and can be freely accessed at http://lin-group.cn/server/iORI-PseKNC2.0. The webserver will provide more convenience to most of wet-experimental scholars. Fu-Ying Dao, Hao Lv 0007, Chao-Qin Feng, Hui Ding 0005, Wei Chen 0064, Hao Lin 0001 |
Bioinform. | 6 |
| 2019 | iTerm-PseKNC: a sequence-based tool for predicting bacterial transcriptional terminatorsabstractMOTIVATION: Transcription termination is an important regulatory step of gene expression. If there is no terminator in gene, transcription could not stop, which will result in abnormal gene expression. Detecting such terminators can determine the operon structure in bacterial organisms and improve genome annotation. Thus, accurate identification of transcriptional terminators is essential and extremely important in the research of transcription regulations. RESULTS: In this study, we developed a new predictor called 'iTerm-PseKNC' based on support vector machine to identify transcription terminators. The binomial distribution approach was used to pick out the optimal feature subset derived from pseudo k-tuple nucleotide composition (PseKNC). The 5-fold cross-validation test results showed that our proposed method achieved an accuracy of 95%. To further evaluate the generalization ability of 'iTerm-PseKNC', the model was examined on independent datasets which are experimentally confirmed Rho-independent terminators in Escherichia coli and Bacillus subtilis genomes. As a result, all the terminators in E. coli and 87.5% of the terminators in B. subtilis were correctly identified, suggesting that the proposed model could become a powerful tool for bacterial terminator recognition. AVAILABILITY AND IMPLEMENTATION: For the convenience of most of wet-experimental researchers, the web-server for 'iTerm-PseKNC' was established at http://lin-group.cn/server/iTerm-PseKNC/, by which users can easily obtain their desired result without the need to go through the detailed mathematical equations involved. Chao-Qin Feng, Zhao-Yue Zhang 0002, Xiao-Juan Zhu, Wei Chen 0064, Hua Tang, Hao Lin 0001 |
Bioinform. | 5 |
| 2019 | iRNAD: a computational tool for identifying D modification sites in RNA sequenceabstractMOTIVATION: Dihydrouridine (D) is a common RNA post-transcriptional modification found in eukaryotes, bacteria and a few archaea. The modification can promote the conformational flexibility of individual nucleotide bases. And its levels are increased in cancerous tissues. Therefore, it is necessary to detect D in RNA for further understanding its functional roles. Since wet-experimental techniques for the aim are time-consuming and laborious, it is urgent to develop computational models to identify D modification sites in RNA. RESULTS: We constructed a predictor, called iRNAD, for identifying D modification sites in RNA sequence. In this predictor, the RNA samples derived from five species were encoded by nucleotide chemical property and nucleotide density. Support vector machine was utilized to perform the classification. The final model could produce the overall accuracy of 96.18% with the area under the receiver operating characteristic curve of 0.9839 in jackknife cross-validation test. Furthermore, we performed a series of validations from several aspects and demonstrated the robustness and reliability of the proposed model. AVAILABILITY AND IMPLEMENTATION: A user-friendly web-server called iRNAD can be freely accessible at http://lin-group.cn/server/iRNAD, which will provide convenience and guide to users for further studying D modification. Peng-Mian Feng, Wangren Qiu, Wei Chen 0064, Hao Lin 0001 |
Bioinform. | 5 |
| 2019 | Predicting protein structural classes for low-similarity sequences by evaluating different features
Xiao-Juan Zhu, Chao-Qin Feng, Hong-Yan Lai, Wei Chen 0064, Hao Lin 0001 |
Knowl. Based Syst. | 4 |
| 2019 | Identifying Sigma70 Promoters with Novel Pseudo Nucleotide CompositionabstractPromoters are DNA regulatory elements located directly upstream or at the 5' end of the transcription initiation site (TSS), which are in charge of gene transcription initiation. With the completion of a large number of microorganism genomics, it is urgent to predict promoters accurately in bacteria by using the computational method. In this work, a sequence-based predictor named "iPro70-PseZNC" was designed for identifying sigma70 promoters in prokaryote. In the predictor, the samples of DNA sequences are formulated by a novel pseudo nucleotide composition, called PseZNC, into which the multi-window Z-curve composition and six local DNA structural properties are incorporated. In the 5-fold cross-validation, the area under the curve of receiver operating characteristic of 0.909 was obtained on our benchmark dataset, indicating that the proposed predictor is promising and will provide an important guide in this area. Further studies showed that the performance of PseZNC is better than it of multi-window Z-curve composition. For the sake of convenience for researchers, a user-friendly online service was established and can be freely accessible at http://lin.uestc.edu.cn/server/iPro70-PseZNC. The PseZNC approach can be also extended to other DNA-related problems. Hao Lin 0001, Zhi-Yong Liang, Hua Tang, Wei Chen 0064 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2018 | iLoc-lncRNA: predict the subcellular location of lncRNAs by incorporating octamer composition into general PseKNCabstractMotivation: Long non-coding RNAs (lncRNAs) are a class of RNA molecules with more than 200 nucleotides. They have important functions in cell development and metabolism, such as genetic markers, genome rearrangements, chromatin modifications, cell cycle regulation, transcription and translation. Their functions are generally closely related to their localization in the cell. Therefore, knowledge about their subcellular locations can provide very useful clues or preliminary insight into their biological functions. Although biochemical experiments could determine the localization of lncRNAs in a cell, they are both time-consuming and expensive. Therefore, it is highly desirable to develop bioinformatics tools for fast and effective identification of their subcellular locations. Results: We developed a sequence-based bioinformatics tool called 'iLoc-lncRNA' to predict the subcellular locations of LncRNAs by incorporating the 8-tuple nucleotide features into the general PseKNC (Pseudo K-tuple Nucleotide Composition) via the binomial distribution approach. Rigorous jackknife tests have shown that the overall accuracy achieved by the new predictor on a stringent benchmark dataset is 86.72%, which is over 20% higher than that by the existing state-of-the-art predictor evaluated on the same tests. Availability and implementation: A user-friendly webserver has been established at http://lin-group.cn/server/iLoc-LncRNA, by which users can easily obtain their desired results. Supplementary information: Supplementary data are available at Bioinformatics online. Zhen-Dong Su 0001, Zhao-Yue Zhang 0002, Ya-Wei Zhao, Dong Wang 0011, Wei Chen 0064, Kuo-Chen Chou, Hao Lin 0001 |
Bioinform. | 6 |
| 2017 | iDNA4mC: identifying DNA N4-methylcytosine sites based on nucleotide chemical propertiesabstractMOTIVATION: DNA N4-methylcytosine (4mC) is an epigenetic modification. The knowledge about the distribution of 4mC is helpful for understanding its biological functions. Although experimental methods have been proposed to detect 4mC sites, they are expensive for performing genome-wide detections. Thus, it is necessary to develop computational methods for predicting 4mC sites. RESULTS: In this work, we developed iDNA4mC, the first webserver to identify 4mC sites, in which DNA sequences are encoded with both nucleotide chemical properties and nucleotide frequency. The predictive results of the rigorous jackknife test and cross species test demonstrated that the performance of iDNA4mC is quite promising and holds high potential to become a useful tool for identifying 4mC sites. AVAILABILITY AND IMPLEMENTATION: The user-friendly web-server, iDNA4mC, is freely accessible at http://lin.uestc.edu.cn/server/iDNA4mC. CONTACT: [email protected] or [email protected]. Wei Chen 0064, Peng-Mian Feng, Hui Ding 0005, Hao Lin 0001 |
Bioinform. | 1 |
| 2017 | Pro54DB: a database for experimentally verified sigma-54 promotersabstractSummary: In prokaryotes, the σ54 promoters are unique regulatory elements and have attracted much attention because they are in charge of the transcription of carbon and nitrogen-related genes and participate in numerous ancillary processes and environmental responses. All findings on σ54 promoters are favorable for a better understanding of their regulatory mechanisms in gene transcription and an accurate discovery of genes missed by the wet experimental evidences. In order to provide an up-to-date, interactive and extensible database for σ54 promoter, a free and easy accessed database called Pro54DB (σ54 promoter database) was built to collect information of σ54 promoter. In the current version, it has stored 210 experimental-confirmed σ54 promoters with 297 regulated genes in 43 species manually extracted from 133 publications, which is helpful for researchers in fields of bioinformatics and molecular biology. Availability and Implementation: Pro54DB is freely available on the web at http://lin.uestc.edu.cn/database/pro54db with all major browsers supported. Contacts: [email protected] or [email protected] Zhi-Yong Liang, Hong-Yan Lai, Huan-Huan Wei, Xin-Xin Chen, Ya-Wei Zhao, Zhen-Dong Su 0001, En-Ze Deng, Hua Tang, Wei Chen 0064, Hao Lin 0001 |
Bioinform. | 13 |
| 2015 | PseKNC-General: a cross-platform package for generating various modes of pseudo nucleotide compositionsabstractSUMMARY: The avalanche of genomic sequences generated in the post-genomic age requires efficient computational methods for rapidly and accurately identifying biological features from sequence information. Towards this goal, we developed a freely available and open-source package, called PseKNC-General (the general form of pseudo k-tuple nucleotide composition), that allows for fast and accurate computation of all the widely used nucleotide structural and physicochemical properties of both DNA and RNA sequences. PseKNC-General can generate several modes of pseudo nucleotide compositions, including conventional k-tuple nucleotide compositions, Moreau-Broto autocorrelation coefficient, Moran autocorrelation coefficient, Geary autocorrelation coefficient, Type I PseKNC and Type II PseKNC. In every mode, >100 physicochemical properties are available for choosing. Moreover, it is flexible enough to allow the users to calculate PseKNC with user-defined properties. The package can be run on Linux, Mac and Windows systems and also provides a graphical user interface. AVAILABILITY AND IMPLEMENTATION: The package is freely available at: http://lin.uestc.edu.cn/server/pseknc. Wei Chen 0064, Xitong Zhang, Jordan Brooker, Hao Lin 0001, Liqing Zhang 0002, Kuo-Chen Chou |
Bioinform. | 1 |
| 2014 | iNuc-PseKNC: a sequence-based predictor for predicting nucleosome positioning in genomes with pseudo k-tuple nucleotide compositionabstractMOTIVATION: Nucleosome positioning participates in many cellular activities and plays significant roles in regulating cellular processes. With the avalanche of genome sequences generated in the post-genomic age, it is highly desired to develop automated methods for rapidly and effectively identifying nucleosome positioning. Although some computational methods were proposed, most of them were species specific and neglected the intrinsic local structural properties that might play important roles in determining the nucleosome positioning on a DNA sequence. RESULTS: Here a predictor called 'iNuc-PseKNC' was developed for predicting nucleosome positioning in Homo sapiens, Caenorhabditis elegans and Drosophila melanogaster genomes, respectively. In the new predictor, the samples of DNA sequences were formulated by a novel feature-vector called 'pseudo k-tuple nucleotide composition', into which six DNA local structural properties were incorporated. It was observed by the rigorous cross-validation tests on the three stringent benchmark datasets that the overall success rates achieved by iNuc-PseKNC in predicting the nucleosome positioning of the aforementioned three genomes were 86.27%, 86.90% and 79.97%, respectively. Meanwhile, the results obtained by iNuc-PseKNC on various benchmark datasets used by the previous investigators for different genomes also indicated that the current predictor remarkably outperformed its counterparts. AVAILABILITY: A user-friendly web-server, iNuc-PseKNC is freely accessible at http://lin.uestc.edu.cn/server/iNuc-PseKNC. Shou-Hui Guo, En-Ze Deng, Liqin Xu, Hui Ding 0005, Hao Lin 0001, Wei Chen 0064, Kuo-Chen Chou |
Bioinform. | 6 |