VLDB 2026 Research / reviewers in the wild / expert
Susumu Goto
dblp:99/3539
· DBLP profile ↗
24ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0003-2989-8486ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
16 papers |
Bioinformatics and computational biology · 100% |
Topics — the 28 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
protein function prediction |
0.7 | 2 | 2020 | KofamKOALA: KEGG Ortholog assignment based on profile HMM and adaptive score threshold · Bioinform. 2020 Discriminating the reaction types of plant type III polyketide synthases · Bioinform. 2017 |
Bioinformatics and computational biology › data integration
biological database integration |
0.6 | 1 | 2022 | TogoID: an exploratory ID converter to bridge biological datasets · Bioinform. 2022 |
Bioinformatics and computational biology
identifier mapping |
0.6 | 1 | 2022 | TogoID: an exploratory ID converter to bridge biological datasets · Bioinform. 2022 |
Bioinformatics and computational biology › sequence analysis
profile hidden markov model |
0.4 | 1 | 2020 | KofamKOALA: KEGG Ortholog assignment based on profile HMM and adaptive score threshold · Bioinform. 2020 |
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
metabolic pathway analysis |
0.4 | 3 | 2014 | Metabolome-scale prediction of intermediate compounds in multistep metabolic pathways with a recursive supervised approach · Bioinform. 2014 Supervised de novo reconstruction of metabolic pathways from metabolome-scale compound sets · Bioinform. 2013 E-zyme: predicting potential EC numbers from the chemical transformation pattern of substrate-product pairs · Bioinform. 2009 |
Bioinformatics and computational biology
comparative genomics |
0.3 | 1 | 2017 | ViPTree: the viral proteomic tree server · Bioinform. 2017 |
Bioinformatics and computational biology › genomics › viral genomics
viral genome classification |
0.3 | 1 | 2017 | ViPTree: the viral proteomic tree server · Bioinform. 2017 |
Bioinformatics and computational biology › genomics
viral genomics |
0.3 | 1 | 2017 | ViPTree: the viral proteomic tree server · Bioinform. 2017 |
Bioinformatics and computational biology › drug discovery
drug-target interaction prediction |
0.3 | 2 | 2012 | Drug target prediction using adverse event report systems: a pharmacogenomic approach · Bioinform. 2012 Drug-target interaction prediction from chemical, genomic and pharmacological data in an integrated framework · Bioinform. 2010 |
Bioinformatics and computational biology › protein function prediction › enzyme function prediction
enzymatic reaction prediction |
0.2 | 1 | 2013 | Supervised de novo reconstruction of metabolic pathways from metabolome-scale compound sets · Bioinform. 2013 |
Bioinformatics and computational biology
drug discovery |
0.1 | 1 | 2012 | Drug target prediction using adverse event report systems: a pharmacogenomic approach · Bioinform. 2012 |
Bioinformatics and computational biology › drug discovery
drug repositioning |
0.1 | 1 | 2012 | Drug target prediction using adverse event report systems: a pharmacogenomic approach · Bioinform. 2012 |
Bioinformatics and computational biology › drug discovery
drug side effect prediction |
0.1 | 1 | 2012 | Relating drug-protein interaction network with drug side effects · Bioinform. 2012 |
Bioinformatics and computational biology › drug discovery
drug-target interaction |
0.1 | 1 | 2012 | Relating drug-protein interaction network with drug side effects · Bioinform. 2012 |
Bioinformatics and computational biology › genomics
pharmacogenomics |
0.1 | 1 | 2012 | Drug target prediction using adverse event report systems: a pharmacogenomic approach · Bioinform. 2012 |
Bioinformatics and computational biology › systems biology › systems medicine
systems pharmacology |
0.1 | 1 | 2012 | Relating drug-protein interaction network with drug side effects · Bioinform. 2012 |
Bioinformatics and computational biology
transcriptomics |
0.1 | 2 | 2009 | GeneRegionScan: a Bioconductor package for probe-level analysis of specific, small regions of the genome · Bioinform. 2009 Prediction of glycan structures from gene expression data based on glycosyltransferase reactions · Bioinform. 2005 |
Bioinformatics and computational biology › protein function prediction › enzyme function prediction
enzyme commission number prediction |
0.1 | 1 | 2009 | E-zyme: predicting potential EC numbers from the chemical transformation pattern of substrate-product pairs · Bioinform. 2009 |
Bioinformatics and computational biology › statistical genetics › quantitative trait locus mapping
expression quantitative trait loci analysis |
0.1 | 1 | 2009 | GeneRegionScan: a Bioconductor package for probe-level analysis of specific, small regions of the genome · Bioinform. 2009 |
Bioinformatics and computational biology
biological database |
0.1 | 1 | 2008 | varDB: a pathogen-specific sequence database of protein families involved in antigenic variation · Bioinform. 2008 |
Bioinformatics and computational biology › protein structure analysis
protein domain analysis |
0.1 | 1 | 2007 | The commonality of protein interaction networks determined in neurodegenerative disorders (NDDs) · Bioinform. 2007 |
Bioinformatics and computational biology › protein analysis › protein-protein interaction
protein-protein interaction network analysis |
0.1 | 1 | 2007 | The commonality of protein interaction networks determined in neurodegenerative disorders (NDDs) · Bioinform. 2007 |
Bioinformatics and computational biology › molecular informatics
glycoinformatics |
0.1 | 1 | 2005 | Prediction of glycan structures from gene expression data based on glycosyltransferase reactions · Bioinform. 2005 |
Bioinformatics and computational biology › sequence analysis
sequence similarity search |
0.1 | 1 | 2005 | Fast and accurate database homology search using upper bounds of local alignment scores · Bioinform. 2005 |
Bioinformatics and computational biology › sequence alignment
smith-waterman acceleration |
0.1 | 1 | 2005 | Fast and accurate database homology search using upper bounds of local alignment scores · Bioinform. 2005 |
Bioinformatics and computational biology › molecular informatics › cheminformatics
chemical database |
0.0 | 1 | 1998 | LIGAND: chemical database for enzyme reactions · Bioinform. 1998 |
Bioinformatics and computational biology
gene expression analysis |
0.0 | 1 | 2005 | Prediction of glycan structures from gene expression data based on glycosyltransferase reactions · Bioinform. 2005 |
Bioinformatics and computational biology › biological database
metabolic pathway database |
0.0 | 1 | 1998 | LIGAND: chemical database for enzyme reactions · Bioinform. 1998 |
Methods — techniques the papers use, named apart from their topics
profile hidden markov model · 0.7ontology construction · 0.6API development · 0.6adaptive score threshold · 0.4proteomic tree construction · 0.3mutual information · 0.3linear discriminant analysis · 0.3genomic alignment · 0.3recursive supervised learning · 0.2chemical substructure fingerprints · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Establishing the Asia & Pacific Bioinformatics Joint Congress: a historic milestone in regional bioinformatics collaborationabstractIn response to the need for greater cohesion among regional conferences, the Asia Pacific Bioinformatics Network (APBioNET) set out in 2015 to realize a long-held aspiration-a single, unifying bioinformatics "super conference" for the Asia & Pacific community. Nearly a decade of persistence, coordination, and coalition-building led to the inaugural Asia & Pacific Bioinformatics Joint Congress (APBJC2024) in Okinawa, Japan. Now established as a triennial event, APBJC stands as a testament to the power of collective vision and shared purpose, offering a unifying platform for regional collaboration and scientific exchange. Tagline: Bringing a Region Together: The Making of APBJC. Asif M. Khan, Susumu Goto, Kenta Nakai, Limsoon Wong, Diane E. Kovats, Shinya Ikematsu, Yoshihiro Yamanishi, Nurul Salwanie Che Wahid, Pradeep Eranti, Yi-Ping Phoebe Chen, Tae-Min Kim, Shinn-Ying Ho, Jessica Cara Mar, Wataru Iwasaki 0001, Jayaraman Valadi, Prashanth Suravajhala, Christian Schönbach, Tin Wee Tan, Shoba Ranganathan, Kiyoko F. Aoki-Kinoshita |
Briefings Bioinform. | 2 |
| 2024 | Enteropathway: the metabolic pathway database for the human gut microbiotaabstractThe human gut microbiota produces diverse, extensive metabolites that have the potential to affect host physiology. Despite significant efforts to identify metabolic pathways for producing these microbial metabolites, a comprehensive metabolic pathway database for the human gut microbiota is still lacking. Here, we present Enteropathway, a metabolic pathway database that integrates 3269 compounds, 3677 reactions, and 876 modules that were obtained from 1012 manually curated scientific literature. Notably, 698 modules of these modules are new entries and cannot be found in any other databases. The database is accessible from a web application (https://enteropathway.org) that offers a metabolic diagram for graphical visualization of metabolic pathways, a customization interface, and an enrichment analysis feature for highlighting enriched modules on the metabolic diagram. Overall, Enteropathway is a comprehensive reference database that can complement widely used databases, and a tool for visual and statistical analysis in human gut microbiota studies and was designed to help researchers pinpoint new insights into the complex interplay between microbiota and host metabolism. Hirotsugu Shiroma, Youssef Darzi, Etsuko Terajima, Zenichi Nakagawa, Hirotaka Tsuchikura, Naoki Tsukuda, Yuki Moriya, Shujiro Okuda, Susumu Goto, Takuji Yamada |
Briefings Bioinform. | 9 |
| 2022 | TogoID: an exploratory ID converter to bridge biological datasetsabstractMOTIVATION: Understanding life cannot be accomplished without making full use of biological data, which are scattered across databases of diverse categories in life sciences. To connect such data seamlessly, identifier (ID) conversion plays a key role. However, existing ID conversion services have disadvantages, such as covering only a limited range of biological categories of databases, not keeping up with the updates of the original databases and outputs being hard to interpret in the context of biological relations, especially when converting IDs in multiple steps. RESULTS: TogoID is an ID conversion service implementing unique features with an intuitive web interface and an application programming interface (API) for programmatic access. TogoID currently supports 65 datasets covering various biological categories. TogoID users can perform exploratory multistep conversions to find a path among IDs. To guide the interpretation of biological meanings in the conversions, we crafted an ontology that defines the semantics of the dataset relations. AVAILABILITY AND IMPLEMENTATION: The TogoID service is freely available on the TogoID website (https://togoid.dbcls.jp/) and the API is also provided to allow programmatic access. To encourage developers to add new dataset pairs, the system stores the configurations of pairs at the GitHub repository (https://github.com/togoid/togoid-config) and accepts the request of additional pairs. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shuya Ikeda, Hiromasa Ono, Tazro Ohta, Hirokazu Chiba, Yuki Naito, Yuki Moriya, Shuichi Kawashima, Yasunori Yamamoto, Shinobu Okamoto, Susumu Goto, Toshiaki Katayama |
Bioinform. | 10 |
| 2020 | KofamKOALA: KEGG Ortholog assignment based on profile HMM and adaptive score thresholdabstractSUMMARY: KofamKOALA is a web server to assign KEGG Orthologs (KOs) to protein sequences by homology search against a database of profile hidden Markov models (KOfam) with pre-computed adaptive score thresholds. KofamKOALA is faster than existing KO assignment tools with its accuracy being comparable to the best performing tools. Function annotation by KofamKOALA helps linking genes to KEGG resources such as the KEGG pathway maps and facilitates molecular network reconstruction. AVAILABILITY AND IMPLEMENTATION: KofamKOALA, KofamScan and KOfam are freely available from GenomeNet (https://www.genome.jp/tools/kofamkoala/). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Takuya Aramaki, Romain Blanc-Mathieu, Hisashi Endo, Koichi Ohkubo, Minoru Kanehisa, Susumu Goto, Hiroyuki Ogata |
Bioinform. | 6 |
| 2017 | ViPTree: the viral proteomic tree serverabstractSUMMARY: ViPTree is a web server provided through GenomeNet to generate viral proteomic trees for classification of viruses based on genome-wide similarities. Users can upload viral genomes sequenced either by genomics or metagenomics. ViPTree generates proteomic trees for the uploaded genomes together with flexibly selected reference viral genomes. ViPTree also serves as a platform to visually investigate genomic alignments and automatically annotated gene functions for the uploaded viral genomes, thus providing virus researchers the first choice for classifying and understanding newly sequenced viral genomes. AVAILABILITY AND IMPLEMENTATION: ViPTree is freely available at: http://www.genome.jp/viptree . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yosuke Nishimura, Takashi Yoshida, Megumi Kuronishi, Hideya Uehara, Hiroyuki Ogata, Susumu Goto |
Bioinform. | 6 |
| 2017 | Discriminating the reaction types of plant type III polyketide synthasesabstractMOTIVATION: Functional prediction of paralogs is challenging in bioinformatics because of rapid functional diversification after gene duplication events combined with parallel acquisitions of similar functions by different paralogs. Plant type III polyketide synthases (PKSs), producing various secondary metabolites, represent a paralogous family that has undergone gene duplication and functional alteration. Currently, there is no computational method available for the functional prediction of type III PKSs. RESULTS: We developed a plant type III PKS reaction predictor, pPAP, based on the recently proposed classification of type III PKSs. pPAP combines two kinds of similarity measures: one calculated by profile hidden Markov models (pHMMs) built from functionally and structurally important partial sequence regions, and the other based on mutual information between residue positions. pPAP targets PKSs acting on ring-type starter substrates, and classifies their functions into four reaction types. The pHMM approach discriminated two reaction types with high accuracy (97.5%, 39/40), but its accuracy decreased when discriminating three reaction types (87.8%, 43/49). When combined with a correlation-based approach, all 49 PKSs were correctly discriminated, and pPAP was still highly accurate (91.4%, 64/70) even after adding other reaction types. These results suggest pPAP, which is based on linear discriminant analyses of similarity measures, is effective for plant type III PKS function prediction. AVAILABILITY AND IMPLEMENTATION: pPAP is freely available at ftp://ftp.genome.jp/pub/tools/ppap/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yugo Shimizu, Hiroyuki Ogata, Susumu Goto |
Bioinform. | 3 |
| 2014 | Metabolome-scale prediction of intermediate compounds in multistep metabolic pathways with a recursive supervised approachabstractMOTIVATION: Metabolic pathway analysis is crucial not only in metabolic engineering but also in rational drug design. However, the biosynthetic/biodegradation pathways are known only for a small portion of metabolites, and a vast amount of pathways remain uncharacterized. Therefore, an important challenge in metabolomics is the de novo reconstruction of potential reaction networks on a metabolome-scale. RESULTS: In this article, we develop a novel method to predict the multistep reaction sequences for de novo reconstruction of metabolic pathways in the reaction-filling framework. We propose a supervised approach to learn what we refer to as 'multistep reaction sequence likeness', i.e. whether a compound-compound pair is possibly converted to each other by a sequence of enzymatic reactions. In the algorithm, we propose a recursive procedure of using step-specific classifiers to predict the intermediate compounds in the multistep reaction sequences, based on chemical substructure fingerprints/descriptors of compounds. We further demonstrate the usefulness of our proposed method on the prediction of enzymatic reaction networks from a metabolome-scale compound set and discuss characteristic features of the extracted chemical substructure transformation patterns in multistep reaction sequences. Our comprehensively predicted reaction networks help to fill the metabolic gap and to infer new reaction sequences in metabolic pathways. AVAILABILITY AND IMPLEMENTATION: Materials are available for free at http://web.kuicr.kyoto-u.ac.jp/supp/kot/ismb2014/ Masaaki Kotera, Yasuo Tabei, Yoshihiro Yamanishi, Ai Muto, Yuki Moriya, Toshiaki Tokimatsu, Susumu Goto |
Bioinform. | 7 |
| 2013 | Supervised de novo reconstruction of metabolic pathways from metabolome-scale compound setsabstractMOTIVATION: The metabolic pathway is an important biochemical reaction network involving enzymatic reactions among chemical compounds. However, it is assumed that a large number of metabolic pathways remain unknown, and many reactions are still missing even in known pathways. Therefore, the most important challenge in metabolomics is the automated de novo reconstruction of metabolic pathways, which includes the elucidation of previously unknown reactions to bridge the metabolic gaps. RESULTS: In this article, we develop a novel method to reconstruct metabolic pathways from a large compound set in the reaction-filling framework. We define feature vectors representing the chemical transformation patterns of compound-compound pairs in enzymatic reactions using chemical fingerprints. We apply a sparsity-induced classifier to learn what we refer to as 'enzymatic-reaction likeness', i.e. whether compound pairs are possibly converted to each other by enzymatic reactions. The originality of our method lies in the search for potential reactions among many compounds at a time, in the extraction of reaction-related chemical transformation patterns and in the large-scale applicability owing to the computational efficiency. In the results, we demonstrate the usefulness of our proposed method on the de novo reconstruction of 134 metabolic pathways in Kyoto Encyclopedia of Genes and Genomes (KEGG). Our comprehensively predicted reaction networks of 15 698 compounds enable us to suggest many potential pathways and to increase research productivity in metabolomics. AVAILABILITY: Softwares are available on request. Supplementary material are available at http://web.kuicr.kyoto-u.ac.jp/supp/kot/ismb2013/. Masaaki Kotera, Yasuo Tabei, Yoshihiro Yamanishi, Toshiaki Tokimatsu, Susumu Goto |
Bioinform. | 5 |
| 2012 | Relating drug-protein interaction network with drug side effectsabstractMOTIVATION: Identifying the emergence and underlying mechanisms of drug side effects is a challenging task in the drug development process. This underscores the importance of system-wide approaches for linking different scales of drug actions; namely drug-protein interactions (molecular scale) and side effects (phenotypic scale) toward side effect prediction for uncharacterized drugs. RESULTS: We performed a large-scale analysis to extract correlated sets of targeted proteins and side effects, based on the co-occurrence of drugs in protein-binding profiles and side effect profiles, using sparse canonical correlation analysis. The analysis of 658 drugs with the two profiles for 1368 proteins and 1339 side effects led to the extraction of 80 correlated sets. Enrichment analyses using KEGG and Gene Ontology showed that most of the correlated sets were significantly enriched with proteins that are involved in the same biological pathways, even if their molecular functions are different. This allowed for a biologically relevant interpretation regarding the relationship between drug-targeted proteins and side effects. The extracted side effects can be regarded as possible phenotypic outcomes by drugs targeting the proteins that appear in the same correlated set. The proposed method is expected to be useful for predicting potential side effects of new drug candidate compounds based on their protein-binding profiles. SUPPLEMENTARY INFORMATION: Datasets and all results are available at http://web.kuicr.kyoto-u.ac.jp/supp/smizutan/target-effect/. AVAILABILITY: Software is available at the above supplementary website. CONTACT: [email protected], or [email protected]. Sayaka Mizutani, Edouard Pauwels, Véronique Stoven, Susumu Goto, Yoshihiro Yamanishi |
Bioinform. | 4 |
| 2012 | Drug target prediction using adverse event report systems: a pharmacogenomic approachabstractMOTIVATION: Unexpected drug activities derived from off-targets are usually undesired and harmful; however, they can occasionally be beneficial for different therapeutic indications. There are many uncharacterized drugs whose target proteins (including the primary target and off-targets) remain unknown. The identification of all potential drug targets has become an important issue in drug repositioning to reuse known drugs for new therapeutic indications. RESULTS: We defined pharmacological similarity for all possible drugs using the US Food and Drug Administration's (FDA's) adverse event reporting system (AERS) and developed a new method to predict unknown drug-target interactions on a large scale from the integration of pharmacological similarity of drugs and genomic sequence similarity of target proteins in the framework of a pharmacogenomic approach. The proposed method was applicable to a large number of drugs and it was useful especially for predicting unknown drug-target interactions that could not be expected from drug chemical structures. We made a comprehensive prediction for potential off-targets of 1874 drugs with known targets and potential target profiles of 2519 drugs without known targets, which suggests many potential drug-target interactions that were not predicted by previous chemogenomic or pharmacogenomic approaches. AVAILABILITY: Softwares are available upon request. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Datasets and all results are available at http://cbio.ensmp.fr/~yyamanishi/aers/. Masataka Takarabe, Masaaki Kotera, Yosuke Nishimura, Susumu Goto, Yoshihiro Yamanishi |
Bioinform. | 4 |
| 2011 | MUCHA: multiple chemical alignment algorithm to identify building block substructures of orphan secondary metabolitesabstractBACKGROUND: In contrast to the increasing number of the successful genome projects, there still remain many orphan metabolites for which their synthesis processes are unknown. Metabolites, including these orphan metabolites, can be classified into groups that share the same core substructures, originated from the same biosynthetic pathways. It is known that many metabolites are synthesized by adding up building blocks to existing metabolites. Therefore, it is proposed that, for any given group of metabolites, finding the core substructure and the branched substructures can help predict their biosynthetic pathway. There already have been many reports on the multiple graph alignment techniques to find the conserved chemical substructures in relatively small molecules. However, they are optimized for ligand binding and are not suitable for metabolomic studies. RESULTS: We developed an efficient multiple graph alignment method named as MUCHA (Multiple Chemical Alignment), specialized for finding metabolic building blocks. This method showed the strength in finding metabolic building blocks with preserving the relative positions among the substructures, which is not achieved by simply applying the frequent graph mining techniques. Compared with the combined pairwise alignments, this proposed MUCHA method generally reduced computational costs with improving the quality of the alignment. CONCLUSIONS: MUCHA successfully find building blocks of secondary metabolites, and has a potential to complement to other existing methods to reconstruct metabolic networks using reaction patterns. Masaaki Kotera, Toshiaki Tokimatsu, Minoru Kanehisa, Susumu Goto |
BMC Bioinform. | 4 |
| 2010 | Drug-target interaction prediction from chemical, genomic and pharmacological data in an integrated frameworkabstractMOTIVATION: In silico prediction of drug-target interactions from heterogeneous biological data is critical in the search for drugs and therapeutic targets for known diseases such as cancers. There is therefore a strong incentive to develop new methods capable of detecting these potential drug-target interactions efficiently. RESULTS: In this article, we investigate the relationship between the chemical space, the pharmacological space and the topology of drug-target interaction networks, and show that drug-target interactions are more correlated with pharmacological effect similarity than with chemical structure similarity. We then develop a new method to predict unknown drug-target interactions from chemical, genomic and pharmacological data on a large scale. The proposed method consists of two steps: (i) prediction of pharmacological effects from chemical structures of given compounds and (ii) inference of unknown drug-target interactions based on the pharmacological effect similarity in the framework of supervised bipartite graph inference. The originality of the proposed method lies in the prediction of potential pharmacological similarity for any drug candidate compounds and in the integration of chemical, genomic and pharmacological data in a unified framework. In the results, we make predictions for four classes of important drug-target interactions involving enzymes, ion channels, GPCRs and nuclear receptors. Our comprehensively predicted drug-target interaction networks enable us to suggest many potential drug-target interactions and to increase research productivity toward genomic drug discovery. SUPPLEMENTARY INFORMATION: Datasets and all prediction results are available at http://cbio.ensmp.fr/~yyamanishi/pharmaco/. AVAILABILITY: Softwares are available upon request. Yoshihiro Yamanishi, Masaaki Kotera, Minoru Kanehisa, Susumu Goto |
Bioinform. | 4 |
| 2009 | GeneRegionScan: a Bioconductor package for probe-level analysis of specific, small regions of the genomeabstractSUMMARY: Whole-genome microarrays allow us to interrogate the entire transcriptome of a cell. Affymetrix microarrays are constructed using several probes that match to different regions of a gene and a summarization step reduces this complexity into a single value, representing the expression level of the gene or the expression level of an exon in the case of exon arrays. However, this simplification eliminates information that might be useful when focusing on specific genes of interest. To address these limitations, we present a software package for the R platform that allows detailed analysis of expression at the probe level. The package matches the probe sequences against a target gene sequence (either mRNA or DNA) and shows the expression levels of each probe along the gene. It also features functions to fit a linear regression based on several genetic models that enables study of the relationship between gene expression and genotype. AVAILABILITY AND IMPLEMENTATION: The software is implemented as a platform-independent R package available through the Bioconductor repository at http://www.bioconductor.org/. It is licensed as GPL 2.0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lasse Folkersen, Diego Diez, Craig E. Wheelock, Jesper Z. Haeggström, Susumu Goto, Per Eriksson, Anders Gabrielsen |
Bioinform. | 5 |
| 2009 | E-zyme: predicting potential EC numbers from the chemical transformation pattern of substrate-product pairsabstractMOTIVATION: The IUBMB's Enzyme Nomenclature system, commonly known as the Enzyme Commission (EC) numbers, plays key roles in classifying enzymatic reactions and in linking the enzyme genes or proteins to reactions in metabolic pathways. There are numerous reactions known to be present in various pathways but without any official EC numbers, most of which have no hope to be given ones because of the lack of the published articles on enzyme assays. RESULTS: In this article we propose a new method to predict the potential EC numbers to given reactant pairs (substrates and products) or uncharacterized reactions, and a web-server named E-zyme as an application. This technology is based on our original biochemical transformation pattern which we call an 'RDM pattern', and consists of three steps: (i) graph alignment of a query reactant pair (substrates and products) for computing the query RDM pattern, (ii) multi-layered partial template matching by comparing the query RDM pattern with template patterns related with known EC numbers and (iii) weighted major voting scheme for selecting appropriate EC numbers. As the result, cross-validation experiments show that the proposed method achieves both high coverage and high prediction accuracy at a practical level, and consistently outperforms the previous method. AVAILABILITY: The E-zyme system is available at http://www.genome.jp/tools/e-zyme/. Yoshihiro Yamanishi, Masahiro Hattori, Masaaki Kotera, Susumu Goto, Minoru Kanehisa |
Bioinform. | 4 |
| 2008 | varDB: a pathogen-specific sequence database of protein families involved in antigenic variationabstractUNLABELLED: Infectious diseases are a major threat to global public health and prosperity. The causative agents consist of a suite of pathogens, ranging from bacteria to viruses, including fungi, helminthes and protozoa. Although these organisms are extremely varied in their biological structure and interactions with the host, they share similar methods of evading the host immune system. Antigenic variation and drift are mechanisms by which pathogens change their exposed epitopes while maintaining protein function. Accordingly, these traits enable pathogens to establish chronic infections in the host. The varDB database was developed to serve as a central repository of protein and nucleotide sequences as well as associated features (e.g. field isolate data, clinical parameters, etc.) involved in antigenic variation. The data currently contained in varDB were mined from GenBank as well as multiple specialized data repositories (e.g. PlasmoDB, GiardiaDB). Family members and ortholog groups were identified using a hierarchical search strategy, including literature/author-based searches and HMM profiles. Included in the current release are>29,00 sequences from 39 gene families from 25 different pathogens. This resource will enable researchers to compare antigenic variation within and across taxa with the goal of identifying common mechanisms of pathogenicity to assist in the fight against a range of devastating diseases. AVAILABILITY: varDB is freely accessible at http://www.vardb.org/ C. Nelson Hayes, Diego Diez, Nicolas Joannin, Wataru Honda, Minoru Kanehisa, Mats Wahlgren, Craig E. Wheelock, Susumu Goto |
Bioinform. | 8 |
| 2007 | The commonality of protein interaction networks determined in neurodegenerative disorders (NDDs)abstractMOTIVATION: Neurodegenerative disorders (NDDs) are progressive and fatal disorders, which are commonly characterized by the intracellular or extracellular presence of abnormal protein aggregates. The identification and verification of proteins interacting with causative gene products are effective ways to understand their physiological and pathological functions. The objective of this research is to better understand common molecular pathogenic mechanisms in NDDs by employing protein-protein interaction networks, the domain characteristics commonly identified in NDDs and correlation among NDDs based on domain information. RESULTS: By reviewing published literatures in PubMed, we created pathway maps in Kyoto Encyclopedia of Genes and Genomes (KEGG) for the protein-protein interactions in six NDDs: Alzheimer's disease (AD), Parkinson's disease (PD), amyotrophic lateral sclerosis (ALS), Huntington's disease (HD), dentatorubral-pallidoluysian atrophy (DRPLA) and prion disease (PRION). We also collected data on 201 interacting proteins and 13 compounds with 282 interactions from the literature. We found 19 proteins common to these six NDDs. These common proteins were mainly involved in the apoptosis and MAPK signaling pathways. We expanded the interaction network by adding protein interaction data from the Human Protein Reference Database and gene expression data from the Human Gene Expression Index Database. We then carried out domain analysis on the extended network and found the characteristic domains, such as 14-3-3 protein, phosphotyrosine interaction domain and caspase domain, for the common proteins. Moreover, we found a relatively high correlation between AD, PD, HD and PRION, but not ALS or DRPLA, in terms of the protein domain distributions. AVAILABILITY: http://www.genome.jp/kegg/pathway/hsa/hsa01510.html (KEGG pathway maps for NDDs). Vachiranee Limviphuvadh, Seigo Tanaka, Susumu Goto, Kunihiro Ueda, Minoru Kanehisa |
Bioinform. | 3 |
| 2007 | Mining prokaryotic genomes for unknown amino acids: a stop-codon-based approachabstractBACKGROUND: Selenocysteine and pyrrolysine are the 21st and 22nd amino acids, which are genetically encoded by stop codons. Since a number of microbial genomes have been completely sequenced to date, it is tempting to ask whether the 23rd amino acid is left undiscovered in these genomes. Recently, a computational study addressed this question and reported that no tRNA gene for unknown amino acid was found in genome sequences available. However, performance of the tRNA prediction program on an unknown tRNA family, which may have atypical sequence and structure, is unclear, thereby rendering their result inconclusive. A protein-level study will provide independent insight into the novel amino acid. RESULTS: Assuming that the 23rd amino acid is also encoded by a stop codon, we systematically predicted proteins that contain stop-codon-encoded amino acids from 191 prokaryotic genomes. Since our prediction method relies only on the conservation patterns of primary sequences, it also provides an opportunity to search novel selenoproteins and other readthrough proteins. It successfully recovered many of currently known selenoproteins and pyrrolysine proteins. However, no promising candidate for the 23rd amino acid was detected, and only one novel selenoprotein was predicted. CONCLUSION: Our result suggests that the unknown amino acid encoded by stop codons does not exist, or its phylogenetic distribution is rather limited, which is in agreement with the previous study on tRNA. The method described here can be used in future studies to explore novel readthrough events from complete genomes, which are rapidly growing. Masashi Fujita, Hisaaki Mihara, Susumu Goto, Nobuyoshi Esaki, Minoru Kanehisa |
BMC Bioinform. | 3 |
| 2007 | Regulation of metabolic networks by small molecule metabolitesabstractBACKGROUND: The ability to regulate metabolism is a fundamental process in living systems. We present an analysis of one of the mechanisms by which metabolic regulation occurs: enzyme inhibition and activation by small molecules. We look at the network properties of this regulatory system and the relationship between the chemical properties of regulatory molecules. RESULTS: We find that many features of the regulatory network, such as the degree and clustering coefficient, closely match those of the underlying metabolic network. While these global features are conserved across several organisms, we do find local differences between regulation in E. coli and H. sapiens which reflect their different lifestyles. Chemical structure appears to play an important role in determining a compounds suitability for use in regulation. Chemical structure also often determines how groups of similar compounds can regulate sets of enzymes. These groups of compounds and the enzymes they regulate form modules that mirror the modules and pathways of the underlying metabolic network. We also show how knowledge of chemical structure and regulation could be used to predict regulatory interactions for drugs. CONCLUSION: The metabolic regulatory network shares many of the global properties of the metabolic network, but often varies at the level of individual compounds. Chemical structure is a key determinant in deciding how a compound is used in regulation and for defining modules within the regulatory system. Alex Gutteridge, Minoru Kanehisa, Susumu Goto |
BMC Bioinform. | 3 |
| 2006 | Extraction of phylogenetic network modules from the metabolic networkabstractBACKGROUND: In bio-systems, genes, proteins and compounds are related to each other, thus forming complex networks. Although each organism has its individual network, some organisms contain common sub-networks based on function. Given a certain sub-network, the distribution of organisms common to it represents the diversity of its function. RESULTS: We extracted such "common" sub-networks, defined as "phylogenetic network modules," using phylogenetic profiles and cluster analysis. The enzymes in the same "phylogenetic network module" have similar phylogenetic profiles and related functions. These modules are shown to be phylogenetic building blocks. Furthermore, the network of the modules illustrated hierarchical feature as well as the network of enzymes involved in the metabolism. CONCLUSION: We conclude that phylogenetic network modules are evolutionary conserved functional units in the metabolic network. We claim that our concept of phylogenetic modules provides a more accurate understanding of the evolution of biological networks. Takuji Yamada, Minoru Kanehisa, Susumu Goto |
BMC Bioinform. | 3 |
| 2005 | Fast and accurate database homology search using upper bounds of local alignment scoresabstractMOTIVATION: It is widely recognized that homology search and ortholog clustering are very useful for analyzing biological sequences. However, recent growth of sequence database size makes homolog detection difficult, and rapid and accurate methods are required. RESULTS: We present a novel method for fast and accurate homology detection, assuming that the Smith-Waterman (SW) scores between all similar sequence pairs in a target database are computed and stored. In this method, SW alignment is computed only if the upper bound, which is derived from our novel inequality, is higher than the given threshold. In contrast to other methods such as FASTA and BLAST, this method is guaranteed to find all sequences whose scores against the query are higher than the specified threshold. Results of computational experiments suggest that the method is dozens of times faster than SSEARCH if genome sequence data of closely related species are available. Masumi Itoh, Susumu Goto, Tatsuya Akutsu, Minoru Kanehisa |
Bioinform. | 2 |
| 2005 | Prediction of glycan structures from gene expression data based on glycosyltransferase reactionsabstractMOTIVATION: Glycan chains are synthesized by a combination of several kinds of glycosyltransferases (GTs). Thus, once we know the repertoire of GTs in the genome, in the transcriptome or in the proteome, it should in principle be possible to predict the repertoire of possible glycan structures in an organism or at a specific stage of the cell. Here, we show that a repertoire of glycan structures can be predicted from the set of GTs in the transcriptome. That is, using knowledge about glycan structure characteristics, we can predict glycan structures from incomplete or noisy data such as DNA microarray data. RESULTS: First, we constructed a reaction pattern library consisting of bond-formation patterns of GT reactions and investigated the co-occurrence frequencies of all reaction patterns in the glycan database. This was followed by the prediction of glycan structures using this library and a co-occurrence score. A penalty score was also implemented in the prediction method. Then we examined the performance of prediction by the leave-one-out cross validation method using individual reaction pattern profiles in the KEGG GLYCAN database as virtual expression profiles. The accuracy of prediction was 81%. Finally, we applied the prediction method to real expression data. Using expression profiles from the human carcinoma cell, glycan structures with sialic acid and sialyl Lewis X epitope were predicted, which corresponded well with experimental results. Shin Kawano, Kosuke Hashimoto, Takashi Miyama, Susumu Goto, Minoru Kanehisa |
Bioinform. | 4 |
| 1998 | Metabolic Pathway Interface to Molecular Biology DatabasesabstractWe present results of providing database support to biomedicine via federation of SDB Cooperation/Integration based upon the KEGG GUI for molecular biology. The federation provides a common link to three molecular biology databases. The added value of the federation is freedom from consulting multiple references to ascertain the full set of enzymatic reactions in a metabolic pathway, and the option of selecting multiple queries to submit to the federated SDBs. Each of the SDBs is extensive, but incomplete. The union of the SDBs, implemented transparently by the federation, is more complete. Each SDB provides a different approach to the options available for data presentation and a different set of Web server tools for data analysis. Thus, an important part of the added value of the federation is the cross-fertilization available in the union of the molecular biological content, the presentation of data, and the tools available for analysis. Barry Zeeberg, Kevin Watanabe, Susumu Goto, Ross A. Overbeek, Larry Kerschberg, George Michaels |
SSDBM | 3 |
| 1998 | LIGAND: chemical database for enzyme reactionsabstractMOTIVATION: The existing molecular biology databases focus on the sequence and structural aspects of biological macromolecules, i.e. DNAs, RNAs and proteins. However, in order to understand the functional aspects, it is essential to computerize the interaction of these molecules. Furthermore, living cells contain additional molecules, such as metabolic compounds and metal ions, that may also be considered as parts of the basic building blocks of life, but are not well organized in public databases. LIGAND chemical database is our attempt to solve these problems, at least for enzymatic reactions. RESULTS: LIGAND consists of two sections: ENZYME and COMPOUND. The ENZYME section is an extension of previous studies (Suyama et al. , Comput. Applic. Biosci., 9, 9-15, 1993), and it is a flat-file representation of 3303 enzymes and 2976 enzymatic reactions in the chemical equation format that can be parsed by machine. The COMPOUND section has been newly constructed for information on the nomenclature and chemical structures of compounds. It contains 5383 chemical compounds. Both ENZYME and COMPOUND entries contain rich cross-reference information, most of which is automatically generated by the DBGET/LinkDB system, thus providing the linkage between chemical and biological databases. LIGAND is updated daily, tightly coupled with the KEGG metabolic pathway database, and forms the basis for reconstruction and computation of pathways. AVAILABILITY: LIGAND can be accessed through the DBGET/LinkDB and KEGG systems in the Japanese GenomeNet database service via http://www.genome.ad.jp/. The flat-file format of the LIGAND database can be downloaded by anonymous FTP via ftp://kegg. genome.adjp/molecules/ligand/. CONTACT: [email protected]; [email protected]; [email protected] Susumu Goto, Takaaki Nishioka, Minoru Kanehisa |
Bioinform. | 1 |
| 1993 | Object-Oriented Database with Rule-Based Query Interface for Genomic Computation
Susumu Goto, Norihiro Sakamoto, Toshihisa Takagi |
DASFAA | 1 |