Kenta Nakai

dblp:38/6619 · DBLP profile ↗
← Back
29ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-8721-8883ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 29 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 Establishing the Asia & Pacific Bioinformatics Joint Congress: a historic milestone in regional bioinformatics collaboration
abstract
In response to the need for greater cohesion among regional conferences, the Asia Pacific Bioinformatics Network (APBioNET) set out in 2015 to realize a long-held aspiration-a single, unifying bioinformatics "super conference" for the Asia & Pacific community. Nearly a decade of persistence, coordination, and coalition-building led to the inaugural Asia & Pacific Bioinformatics Joint Congress (APBJC2024) in Okinawa, Japan. Now established as a triennial event, APBJC stands as a testament to the power of collective vision and shared purpose, offering a unifying platform for regional collaboration and scientific exchange. Tagline: Bringing a Region Together: The Making of APBJC.
Asif M. Khan, Susumu Goto, Kenta Nakai, Limsoon Wong, Diane E. Kovats, Shinya Ikematsu, Yoshihiro Yamanishi, Nurul Salwanie Che Wahid, Pradeep Eranti, Yi-Ping Phoebe Chen, Tae-Min Kim, Shinn-Ying Ho, Jessica Cara Mar, Wataru Iwasaki 0001, Jayaraman Valadi, Prashanth Suravajhala, Christian Schönbach, Tin Wee Tan, Shoba Ranganathan, Kiyoko F. Aoki-Kinoshita
Briefings Bioinform.3
2025 Gene2role: a role-based gene embedding method for comparative analysis of signed gene regulatory networks
abstract
BACKGROUND: Understanding the dynamics of gene regulatory networks (GRNs) across various cellular states is crucial for deciphering the underlying mechanisms governing cell behavior and functionality. However, current comparative analytical methods, which often focus on simple topological information such as the degree of genes, are limited in their ability to fully capture the similarities and differences among the complex GRNs. RESULTS: We present Gene2role, a gene embedding approach that leverages multi-hop topological information from genes within signed GRNs. Initially, we demonstrated the effectiveness of Gene2role in capturing the intricate topological nuances of genes using GRNs inferred from four distinct data sources. Then, applying Gene2role to integrated GRNs allowed us to identify genes with significant topological changes across cell types or states, offering a fresh perspective beyond traditional differential gene expression analyses. Additionally, we quantified the stability of gene modules between two cellular states by measuring the changes in the gene embeddings within these modules. CONCLUSIONS: Our method augments the existing toolkit for probing the dynamic regulatory landscape, thereby opening new avenues for understanding gene behavior and interaction patterns across cellular transitions.
Wanzhe Xu, Fujio Toriumi, Kenta Nakai
BMC Bioinform.7
2024 HyGAnno: hybrid graph neural network-based cell type annotation for single-cell ATAC sequencing data
abstract
Reliable cell type annotations are crucial for investigating cellular heterogeneity in single-cell omics data. Although various computational approaches have been proposed for single-cell RNA sequencing (scRNA-seq) annotation, high-quality cell labels are still lacking in single-cell sequencing assay for transposase-accessible chromatin (scATAC-seq) data, because of extreme sparsity and inconsistent chromatin accessibility between datasets. Here, we present a novel automated cell annotation method that transfers cell type information from a well-labeled scRNA-seq reference to an unlabeled scATAC-seq target, via a parallel graph neural network, in a semi-supervised manner. Unlike existing methods that utilize only gene expression or gene activity features, HyGAnno leverages genome-wide accessibility peak features to facilitate the training process. In addition, HyGAnno reconstructs a reference-target cell graph to detect cells with low prediction reliability, according to their specific graph connectivity patterns. HyGAnno was assessed across various datasets, showcasing its strengths in precise cell annotation, generating interpretable cell embeddings, robustness to noisy reference data and adaptability to tumor tissues.
Martin Loza, Sung-Joon Park, Kenta Nakai
Briefings Bioinform.6
2023 CoraL: interpretable contrastive meta-learning for the prediction of cancer-associated ncRNA-encoded small peptides
abstract
NcRNA-encoded small peptides (ncPEPs) have recently emerged as promising targets and biomarkers for cancer immunotherapy. Therefore, identifying cancer-associated ncPEPs is crucial for cancer research. In this work, we propose CoraL, a novel supervised contrastive meta-learning framework for predicting cancer-associated ncPEPs. Specifically, the proposed meta-learning strategy enables our model to learn meta-knowledge from different types of peptides and train a promising predictive model even with few labeled samples. The results show that our model is capable of making high-confidence predictions on unseen cancer biomarkers with only five samples, potentially accelerating the discovery of novel cancer biomarkers for immunotherapy. Moreover, our approach remarkably outperforms existing deep learning models on 15 cancer-associated ncPEPs datasets, demonstrating its effectiveness and robustness. Interestingly, our model exhibits outstanding performance when extended for the identification of short open reading frames derived from ncPEPs, demonstrating the strong prediction ability of CoraL at the transcriptome level. Importantly, our feature interpretation analysis discovers unique sequential patterns as the fingerprint for each cancer-associated ncPEPs, revealing the relationship among certain cancer biomarkers that are validated by relevant literature and motif comparison. Overall, we expect CoraL to be a useful tool to decipher the pathogenesis of cancer and provide valuable information for cancer research. The dataset and source code of our proposed method can be found at https://github.com/Johnsunnn/CoraL.
Zhongshen Li, Junru Jin, Wentao Long, Haoqing Yu, Xin Gao 0001, Kenta Nakai, Quan Zou 0001, Leyi Wei
Briefings Bioinform.7
2022 Protein design via deep learning
abstract
Proteins with desired functions and properties are important in fields like nanotechnology and biomedicine. De novo protein design enables the production of previously unseen proteins from the ground up and is believed as a key point for handling real social challenges. Recent introduction of deep learning into design methods exhibits a transformative influence and is expected to represent a promising and exciting future direction. In this review, we retrospect the major aspects of current advances in deep-learning-based design procedures and illustrate their novelty in comparison with conventional knowledge-based approaches through noticeable cases. We not only describe deep learning developments in structure-based protein design and direct sequence design, but also highlight recent applications of deep reinforcement learning in protein design. The future perspectives on design goals, challenges and opportunities are also comprehensively discussed.
Wenze Ding, Kenta Nakai, Haipeng Gong
Briefings Bioinform.2
2022 Predicting protein-peptide binding residues via interpretable deep learning
abstract
SUMMARY: Identifying the protein-peptide binding residues is fundamentally important to understand the mechanisms of protein functions and explore drug discovery. Although several computational methods have been developed, most of them highly rely on third-party tools or complex data preprocessing for feature design, easily resulting in low computational efficacy and suffering from low predictive performance. To address the limitations, we propose PepBCL, a novel BERT (Bidirectional Encoder Representation from Transformers) -based contrastive learning framework to predict the protein-peptide binding residues based on protein sequences only. PepBCL is an end-to-end predictive model that is independent of feature engineering. Specifically, we introduce a well pre-trained protein language model that can automatically extract and learn high-latent representations of protein sequences relevant for protein structures and functions. Further, we design a novel contrastive learning module to optimize the feature representations of binding residues underlying the imbalanced dataset. We demonstrate that our proposed method significantly outperforms the state-of-the-art methods under benchmarking comparison, and achieves more robust performance. Moreover, we found that we further improve the performance via the integration of traditional features and our learnt features. Interestingly, the interpretable analysis of our model highlights the flexibility and adaptability of deep learning-based protein language model to capture both conserved and non-conserved sequential characteristics of peptide-binding residues. Finally, to facilitate the use of our method, we establish an online predictive platform as the implementation of the proposed PepBCL, which is now available at http://server.wei-group.net/PepBCL/. AVAILABILITY AND IMPLEMENTATION: https://github.com/Ruheng-W/PepBCL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ruheng Wang, Junru Jin, Quan Zou 0001, Kenta Nakai, Leyi Wei
Bioinform.4
2021 OpenContami: a web-based application for detecting microbial contaminants in next-generation sequencing data
abstract
SUMMARY: Microorganisms infect and contaminate eukaryotic cells during the course of biological experiments. Because microbes influence host cell biology and may therefore lead to erroneous conclusions, a computational platform that facilitates decontamination is indispensable. Recent studies show that next-generation sequencing (NGS) data can be used to identify the presence of exogenous microbial species. Previously, we proposed an algorithm to improve detection of microbes in NGS data. Here, we developed an online application, OpenContami, which allows researchers easy access to the algorithm via interactive web-based interfaces. We have designed the application by incorporating a database comprising analytical results from a large-scale public dataset and data uploaded by users. The database serves as a reference for assessing user data and provides a list of genera detected from negative blank controls as a 'blacklist', which is useful for studying human infectious diseases. OpenContami offers a comprehensive overview of exogenous species in NGS datasets; as such, it will increase our understanding of the impact of microbial contamination on biological and pathological traits. AVAILABILITY AND IMPLEMENTATION: OpenContami is freely available at: https://openlooper.hgc.jp/opencontami/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sung-Joon Park, Kenta Nakai
Bioinform.2
2021 A semi-supervised deep learning approach for predicting the functional effects of genomic non-coding variations
abstract
BACKGROUND: Understanding the functional effects of non-coding variants is important as they are often associated with gene-expression alteration and disease development. Over the past few years, many computational tools have been developed to predict their functional impact. However, the intrinsic difficulty in dealing with the scarcity of data leads to the necessity to further improve the algorithms. In this work, we propose a novel method, employing a semi-supervised deep-learning model with pseudo labels, which takes advantage of learning from both experimentally annotated and unannotated data. RESULTS: We prepared known functional non-coding variants with histone marks, DNA accessibility, and sequence context in GM12878, HepG2, and K562 cell lines. Applying our method to the dataset demonstrated its outstanding performance, compared with that of existing tools. Our results also indicated that the semi-supervised model with pseudo labels achieves higher predictive performance than the supervised model without pseudo labels. Interestingly, a model trained with the data in a certain cell line is unlikely to succeed in other cell lines, which implies the cell-type-specific nature of the non-coding variants. Remarkably, we found that DNA accessibility significantly contributes to the functional consequence of variants, which suggests the importance of open chromatin conformation prior to establishing the interaction of non-coding variants with gene regulation. CONCLUSIONS: The semi-supervised deep learning model coupled with pseudo labeling has advantages in studying with limited datasets, which is not unusual in biology. Our study provides an effective approach in finding non-coding mutations potentially associated with various biological phenomena, including human diseases.
Sung-Joon Park, Kenta Nakai
BMC Bioinform.3
2018 TimeXNet Web: identifying cellular response networks from diverse omics time-course data
abstract
Summary: Condition-specific time-course omics profiles are frequently used to study cellular response to stimuli and identify associated signaling pathways. However, few online tools allow users to analyze multiple types of high-throughput time-course data. TimeXNet Web is a web server that extracts a time-dependent gene/protein response network from time-course transcriptomic, proteomic or phospho-proteomic data, and an input interaction network. It classifies the given genes/proteins into time-dependent groups based on the time of their highest activity and identifies the most probable paths connecting genes/proteins in consecutive groups. The response sub-network is enriched in activated genes/proteins and contains novel regulators that do not show any observable change in the input data. Users can view the resultant response network and analyze it for functional enrichment. TimeXNet Web supports the analysis of high-throughput data from multiple species by providing high quality, weighted protein-protein interaction networks for 12 model organisms. Availability and implementation: http://txnet.hgc.jp/. Supplementary information: Supplementary data are available at Bioinformatics online.
Phit Ling Tan, Yosvany López, Kenta Nakai, Ashwini Patil
Bioinform.3
2016 A study on the application of topic models to motif finding algorithms
abstract
BACKGROUND: Topic models are statistical algorithms which try to discover the structure of a set of documents according to the abstract topics contained in them. Here we try to apply this approach to the discovery of the structure of the transcription factor binding sites (TFBS) contained in a set of biological sequences, which is a fundamental problem in molecular biology research for the understanding of transcriptional regulation. Here we present two methods that make use of topic models for motif finding. First, we developed an algorithm in which first a set of biological sequences are treated as text documents, and the k-mers contained in them as words, to then build a correlated topic model (CTM) and iteratively reduce its perplexity. We also used the perplexity measurement of CTMs to improve our previous algorithm based on a genetic algorithm and several statistical coefficients. RESULTS: The algorithms were tested with 56 data sets from four different species and compared to 14 other methods by the use of several coefficients both at nucleotide and site level. The results of our first approach showed a performance comparable to the other methods studied, especially at site level and in sensitivity scores, in which it scored better than any of the 14 existing tools. In the case of our previous algorithm, the new approach with the addition of the perplexity measurement clearly outperformed all of the other methods in sensitivity, both at nucleotide and site level, and in overall performance at site level. CONCLUSIONS: The statistics obtained show that the performance of a motif finding method based on the use of a CTM is satisfying enough to conclude that the application of topic models is a valid method for developing motif finding algorithms. Moreover, the addition of topic models to a previously developed method dramatically increased its performance, suggesting that this combined algorithm can be a useful tool to successfully predict motifs in different kinds of sets of DNA sequences.
Josep Basha Gutierrez, Kenta Nakai
BMC Bioinform.2
2013 Linking Transcriptional Changes over Time in Stimulated Dendritic Cells to Identify Gene Networks Activated during the Innate Immune Response
abstract
The innate immune response is primarily mediated by the Toll-like receptors functioning through the MyD88-dependent and TRIF-dependent pathways. Despite being widely studied, it is not yet completely understood and systems-level analyses have been lacking. In this study, we identified a high-probability network of genes activated during the innate immune response using a novel approach to analyze time-course gene expression profiles of activated immune cells in combination with a large gene regulatory and protein-protein interaction network. We classified the immune response into three consecutive time-dependent stages and identified the most probable paths between genes showing a significant change in expression at each stage. The resultant network contained several novel and known regulators of the innate immune response, many of which did not show any observable change in expression at the sampled time points. The response network shows the dominance of genes from specific functional classes during different stages of the immune response. It also suggests a role for the protein phosphatase 2a catalytic subunit α in the regulation of the immunoproteasome during the late phase of the response. In order to clarify the differences between the MyD88-dependent and TRIF-dependent pathways in the innate immune response, time-course gene expression profiles from MyD88-knockout and TRIF-knockout dendritic cells were analyzed. Their response networks suggest the dominance of the MyD88-dependent pathway in the innate immune response, and an association of the circadian regulators and immunoproteasomal degradation with the TRIF-dependent pathway. The response network presented here provides the most probable associations between genes expressed in the early and the late phases of the innate immune response, while taking into account the intermediate regulators. We propose that the method described here can also be used in the identification of time-dependent gene sub-networks in other biological systems.
Ashwini Patil, Yutaro Kumagai, Kuo-ching Liang, Yutaka Suzuki, Kenta Nakai
PLoS Comput. Biol.5
2011 Seed-Set Construction by Equi-entropy Partitioning for Efficient and Sensitive Short-Read Mapping
Kouichi Kimura, Asako Koike, Kenta Nakai
WABI3
2011 A regression analysis of gene expression in ES cells reveals two gene classes that are significantly different in epigenetic patterns
abstract
BACKGROUND: To understand the gene regulatory system that governs the self-renewal and pluripotency of embryonic stem cells (ESCs) is an important step for promoting regenerative medicine. In it, the role of several core transcription factors (TFs), such as Oct4, Sox2 and Nanog, has been intensively investigated, details of their involvement in the genome-wide gene regulation are still not well clarified. METHODS: We constructed a predictive model of genome-wide gene expression in mouse ESCs from publicly available ChIP-seq data of 12 core TFs. The tag sequences were remapped on the genome by various alignment tools. Then, the binding density of each TF is calculated from the genome-wide bona fide TF binding sites. The TF-binding data was combined with the data of several epigenetic states (DNA methylation, several histone modifications, and CpG island) of promoter regions. These data as well as the ordinary peak intensity data were used as predictors of a simple linear regression model that predicts absolute gene expression. We also developed a pipeline for analyzing the effects of predictors and their interactions. RESULTS: Through our analysis, we identified two classes of genes that are either well explained or inefficiently explained by our model. The latter class seems to be genes that are not directly regulated by the core TFs. The regulatory regions of these gene classes show apparently distinct patterns of DNA methylation, histone modifications, existence of CpG islands, and gene ontology terms, suggesting the relative importance of epigenetic effects. Furthermore, we identified statistically significant TF interactions correlated with the epigenetic modification patterns. CONCLUSIONS: Here, we proposed an improved prediction method in explaining the ESC-specific gene expression. Our study implies that the majority of genes are more or less directly regulated by the core TFs. In addition, our result is consistent with the general idea of relative importance of epigenetic effects in ESCs.
Sung-Joon Park, Kenta Nakai
BMC Bioinform.2
2010 Gradual transition from mosaic to global DNA methylation patterns during deuterostome evolution
abstract
BACKGROUND: DNA methylation by the Dnmt family occurs in vertebrates and invertebrates, including ascidians, and is thought to play important roles in gene regulation and genome stability, especially in vertebrates. However, the global methylation patterns of vertebrates and invertebrates are distinctive. Whereas almost all CpG sites are methylated in vertebrates, with the exception of those in CpG islands, the ascidian genome contains approximately equal amounts of methylated and unmethylated regions. Curiously, methylation status can be reliably estimated from the local frequency of CpG dinucleotides in the ascidian genome. Methylated and unmethylated regions tend to have few and many CpG sites, respectively, consistent with our knowledge of the methylation status of CpG islands and other regions in mammals. However, DNA methylation patterns and levels in vertebrates and invertebrates have not been analyzed in the same way. RESULTS: Using a new computational methodology based on the decomposition of the bimodal distributions of methylated and unmethylated regions, we estimated the extent of the global methylation patterns in a wide range of animals. We then examined the epigenetic changes in silico along the phylogenetic tree. We observed a gradual transition from fractional to global patterns of methylation in deuterostomes, rather than a clear demarcation between vertebrates and invertebrates. When we applied this methodology to six piscine genomes, some of which showed features similar to those of invertebrates. CONCLUSIONS: The mammalian global DNA methylation pattern was probably not acquired at an early stage of vertebrate evolution, but gradually expanded from that of a more ancient organism.
Kohji Okamura, Kazuaki A. Matsumoto, Kenta Nakai
BMC Bioinform.3
2010 InCoB2010 - 9th International Conference on Bioinformatics at Tokyo, Japan, September 26-28, 2010
abstract
The International Conference on Bioinformatics (InCoB), the annual conference of the Asia-Pacific Bioinformatics Network (APBioNet), is hosted in one of countries of the Asia-Pacific region. The 2010 conference was awarded to Japan and has attracted more than one hundred high-quality research paper submissions. Thorough peer reviewing resulted in 47 (43.5%) accepted papers out of 108 submissions. Submissions from Japan, R.O. Korea, P.R. China, Australia, Singapore and U.S.A totaled 43.8% and contributed to 57.4% of accepted papers. Manuscripts originating from Taiwan and India added up to 42.8% of submissions and 28.3% of acceptances. The fifteen articles published in this BMC Bioinformatics supplement cover disease informatics, structural bioinformatics and drug design, biological databases and software tools, signaling pathways, gene regulatory and biochemical networks, evolution and sequence analysis.
Christian Schönbach, Kenta Nakai, Tin Wee Tan, Shoba Ranganathan
BMC Bioinform.2
2006 Protein Subcellular Localisation Prediction with WoLF PSORT
Paul Horton, Keun-Joon Park, Takeshi Obayashi, Kenta Nakai
APBC4
2005 Prediction of Transcriptional Terminators in Bacillus subtilis and Related Species
abstract
In prokaryotes, genes belonging to the same operon are transcribed in a single mRNA molecule. Transcription starts as the RNA polymerase binds to the promoter and continues until it reaches a transcriptional terminator. Some terminators rely on the presence of the Rho protein, whereas others function independently of Rho. Such Rho-independent terminators consist of an inverted repeat followed by a stretch of thymine residues, allowing us to predict their presence directly from the DNA sequence. Unlike in Escherichia coli, the Rho protein is dispensable in Bacillus subtilis, suggesting a limited role for Rho-dependent termination in this organism and possibly in other Firmicutes. We analyzed 463 experimentally known terminating sequences in B. subtilis and found a decision rule to distinguish Rho-independent transcriptional terminators from non-terminating sequences. The decision rule allowed us to find the boundaries of operons in B. subtilis with a sensitivity and specificity of about 94%. Using the same decision rule, we found an average sensitivity of 94% for 57 bacteria belonging to the Firmicutes phylum, and a considerably lower sensitivity for other bacteria. Our analysis shows that Rho-independent termination is dominant for Firmicutes in general, and that the properties of the transcriptional terminators are conserved. Terminator prediction can be used to reliably predict the operon structure in these organisms, even in the absence of experimentally known operons. Genome-wide predictions of Rho-independent terminators for the 57 Firmicutes are available in the Supporting Information section.
Michiel J. L. de Hoon, Yuko Makita, Kenta Nakai, Satoru Miyano
PLoS Comput. Biol.3
2004 Finding Optimal Pairs of Cooperative and Competing Patterns with Bounded Distance
Shunsuke Inenaga, Hideo Bannai, Heikki Hyyrö, Ayumi Shinohara, Masayuki Takeda, Kenta Nakai, Satoru Miyano
Discovery Science6
2004 Finding Optimal Pairs of Patterns
Hideo Bannai, Heikki Hyyrö, Ayumi Shinohara, Masayuki Takeda, Kenta Nakai, Satoru Miyano
WABI5
2004 An O(N2) Algorithm for Discovering Optimal Boolean Pattern Pairs
abstract
We consider the problem of finding the optimal combination of string patterns, which characterizes a given set of strings that have a numeric attribute value assigned to each string. Pattern combinations are scored based on the correlation between their occurrences in the strings and the numeric attribute values. The aim is to find the combination of patterns which is best with respect to an appropriate scoring function. We present an O(N2) time algorithm for finding the optimal pair of substring patterns combined with Boolean functions, where N is the total length of the sequences. The algorithm looks for all possible Boolean combinations of the patterns, e.g., patterns of the form p and not q, which indicates that the pattern pair is considered to occur in a given string s, if p occurs in s, AND q does NOT occur in s. An efficient implementation using suffix arrays is presented, and we further show that the algorithm can be adapted to find the best k-pattern Boolean combination in O(Nk) time. The algorithm is applied to mRNA sequence data sets of moderate size combined with their turnover rates for the purpose of finding regulatory elements that cooperate, complement, or compete with each other in enhancing and/or silencing mRNA decay.
Hideo Bannai, Heikki Hyyrö, Ayumi Shinohara, Masayuki Takeda, Kenta Nakai, Satoru Miyano
IEEE ACM Trans. Comput. Biol. Bioinform.5
2003 MELINA: motif extraction from promoter regions of potentially co-regulated genes
abstract
Abstract Summary: ‘Melina’ assists users to compare the results of four public softwares for DNA motif extraction in order to both confirm the reliability of each finding and avoid missing important motifs. It is also useful to optimize the sensitivity of software with a series of different parameter settings. Availability: Melina is available at http://www.hgc.ims.u-tokyo.ac.jp/Melina/ Contact: [email protected] * To whom correspondence should be addressed.
Natalia V. Poluliakh, Toshihisa Takagi, Kenta Nakai
Bioinform.3
2002 Extensive feature detection of N-terminal protein sorting signals
abstract
MOTIVATION: The prediction of localization sites of various proteins is an important and challenging problem in the field of molecular biology. TargetP, by Emanuelsson et al. (J. Mol. Biol., 300, 1005-1016, 2000) is a neural network based system which is currently the best predictor in the literature for N-terminal sorting signals. One drawback of neural networks, however, is that it is generally difficult to understand and interpret how and why they make such predictions. In this paper, we aim to generate simple and interpretable rules as predictors, and still achieve a practical prediction accuracy. We adopt an approach which consists of an extensive search for simple rules and various attributes which is partially guided by human intuition. RESULTS: We have succeeded in finding rules whose prediction accuracies come close to that of TargetP, while still retaining a very simple and interpretable form. We also discuss and interpret the discovered rules.
Hideo Bannai, Yoshinori Tamada, Osamu Maruyama, Kenta Nakai, Satoru Miyano
Bioinform.4
2001 ORI-GENE: gene classification based on the evolutionary tree
abstract
MOTIVATION: Genome projects have produced large amounts of data on the sequences of new genes whose functions are as yet unknown. The functions of new genes are usually inferred by comparing their sequences with those of known genes, but evaluation of the sequence homology of individual genes does not make the most of the available sequence information. Therefore, new methods and tools for extracting more biological information from homology searches would be advantageous. RESULTS: We have developed a computational tool, ORI-GENE, to analyze the results of sequence homology searches from the perspective of the evolution of selected sets of new genes. ORI-GENE has a graphical interface and accomplishes two important tasks: first, based on the output of homology searches, it identifies species with similar genes and displays their pattern of distribution on the phylogenetic tree. This function enables one to infer the way in which a given gene may have propagated among species over time. Second, from the distribution patterns, it predicts the point at which a given gene may have been first acquired (i.e. its 'origin'), then classifies the gene on that basis. Because it makes use of available evolutionary information to show the way in which genes cluster among species, ORI-GENE should be an effective tool for the screening and classification of new genes revealed by genome analysis. AVAILABILITY: ORI-GENE is retrievable via the Internet at: http://www.rtc.riken.go.jp/jouhou/ORI-GENE.
Hideaki Mizuno, Yoshimasa Tanaka, Kenta Nakai, Akinori Sarai
Bioinform.3
1999 Modeling and predicting transcriptional units of <$O_SSF>Escherichia coli<$C_SSF>genes using hidden Markov models
abstract
MOTIVATION: The hidden Markov model (HMM) is a valuable technique for gene-finding, especially because its flexibility enables the inclusion of various sequence features. Recent programs for bacterial gene-finding include the information of ribosomal binding site (RBS) to improve the recognition accuracy of the start codon, using this feature. We report here our attempt to extend the model into the total transcriptional unit, enabling the prediction of operon structures. RESULTS: First, we improved the prediction accuracy of coding sequences (CDSs) by employing the models of 'typical', 'atypical' and 'negative (false-positive)' classes as well as the models of RBS and its downstream spacer. The sensitivity of exactly predicting the 204 experimentally confirmed CDSs reached 90.2% in an objective test. Based on the prediction result of CDSs, the positions of the promoters and terminators were predicted. Our model could exactly recognize 60% of 390 known transcriptional units. Thus, the accuracy and significance of this prediction problem is far from trivial. We would like to propose this problem as an open theme in bioinformatics because the ongoing or planned post-sequencing projects will produce much data for future improvements.
Tetsushi Yada, Mitsuteru Nakao, Yasushi Totoki, Kenta Nakai
Bioinform.4
1998 Automatic extraction of motifs represented in the hidden Markov model from a number of DNA sequences
abstract
MOTIVATION: Automatic extraction of motifs that occur frequently on a set of unaligned DNA sequences is useful for predicting the binding sites of unknown transcription factors. Several programs for this purpose have been released. However, in our opinion, they are not practical enough to be applied to a large number of upstream sequences. RESULTS: We propose a new program called YEBIS (Yet another Environment for the analysis of BIopolymer Sequences) which is capable of extracting a set of motifs, without any a priori knowledge, from a number of functionally related DNA sequences. Using the hidden Markov model, these motifs are represented in a more general form than other conventional methods, such as the weight matrix method. When applied to several sets of benchmark data, it was found that YEBIS had comparable capability to the existing methods, but was much faster. Moreover, it could extract all known motifs from the LTR sequences (long terminal repeat sequences) in a single run. Finally, it could be successfully applied to approximately 400 human promoter sequences and some of the extracted motifs turned out to be known cis-elements. Therefore, YEBIS could be a practical tool for exploring the upstream sequences of genomic ORFs, some of which are regulated in a similar fashion. AVAILABILITY: YEBIS will be distributed to academic users free of charge. All requests should be sent to the address below. CONTACT: E-MAIL: [email protected]
Tetsushi Yada, Yasushi Totoki, Masato Ishikawa, Kiyoshi Asai, Kenta Nakai
Bioinform.5
1997 Better Prediction of Protein Cellular Localization Sites with the it k Nearest Neighbors Classifier
Paul Horton, Kenta Nakai
ISMB2
1997 Functional Prediction of B. subtilis Genes from Their Regulatory Sequences
Tetsushi Yada, Yasushi Totoki, Takahiro Ishii, Kenta Nakai
ISMB4
1996 A Probabilistic Classification System for Predicting the Cellular Localization Sites of Proteins
Paul Horton, Kenta Nakai
ISMB2
1994 Gnome - an Internet-based sequence analysis tool
abstract
Gnome (GenomeNet Open Mail-service Environment) is a sequence analysis tool that enables an end-user to make use of several Internet- (mainly e-mail) based services with an easy-to-use graphical user interface. Users can conduct homology and motif searches, and database-entry retrieval against the latest databases by emitting search requests to and receiving their results form a search-server by e-mail. The search results are viewed and managed efficiently with this system. The Macintosh and X (Motif) versions of the Gnome client and the UNIX version of the Gnome server are available to academic users free of charge.
Kenta Nakai, T. Tokimori, A. Ogiwara, Ikuo Uchiyama, T. Niiyama
Comput. Appl. Biosci.1