Yoshihiro Yamanishi

dblp:87/1818 · DBLP profile ↗
← Back
46ranked-venue papers
7as first author
18since 2021 · last 2025
0000-0003-2279-8773ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 29 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 8 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 InstGAN: Instant Actor-Critic-Driven GAN for De Novo Molecule Generation and Property Optimization
abstract
Deep generative models, such as generative adversarial networks (GANs), have been employed for de~novo molecular generation in drug discovery. Most prior studies have utilized reinforcement learning (RL) algorithms, particularly Monte Carlo tree search (MCTS), to handle the discrete nature of molecular representations in GANs. However, due to the inherent instability in training GANs and RL models, along with the high computational cost associated with MCTS sampling, MCTS RL-based GANs struggle to scale to large chemical databases. To tackle these challenges, this study introduces a novel GAN based on actor-critic RL with instant and global rewards, called InstGAN, to generate molecules at the token-level with multi-property optimization. Furthermore, maximized information entropy is leveraged to alleviate the mode collapse. The experimental results demonstrate that InstGAN outperforms other baselines, achieves comparable performance to state-of-the-art models, and efficiently generates molecules with multi-property optimization. The code is available at: https://github.com/tang777777/InstGAN.
Huidong Tang, Chen Li 0027, Sayaka Kamei, Yoshihiro Yamanishi, Yasuhiko Morimoto
IJCAI4
2025 Gx2Mol: De Novo Generation of Hit-Like Molecules from Gene Expression Profiles
Chen Li 0027, Yoshihiro Yamanishi
ECML/PKDD (3)2
2025 AI-driven transcriptome profile-guided hit molecule generation
abstract
D e n o v o generation of bioactive and drug-like hit molecules is a pivotal goal in computer-aided drug discovery. While artificial intelligence (AI) has proven adept at generating molecules with desired chemical properties, previous studies often overlook the influence of disease-specific cellular environments. This study introduces GxVAEs, a novel AI-driven deep generative model designed to produce hit molecules from transcriptome profiles using dual variational autoencoders (VAEs). The first VAE, ProfileVAE, extracts latent features from transcriptome profiles to guide the second VAE, MolVAE, in generating hit molecules. GxVAEs aim to bridge the gap between molecule generation and the biological context of disease, producing molecules that are biologically relevant within specific cellular environments or pathological conditions. Experimental results and case studies focused on hit molecule generation demonstrate that GxVAEs surpass current state-of-the-art methods, in terms of reproducibility of known ligands. This approach is expected to effectively find potential molecular structures with bioactivities across diverse disease contexts.
Chen Li 0027, Yoshihiro Yamanishi
Artif. Intell.2
2025 Establishing the Asia & Pacific Bioinformatics Joint Congress: a historic milestone in regional bioinformatics collaboration
abstract
In response to the need for greater cohesion among regional conferences, the Asia Pacific Bioinformatics Network (APBioNET) set out in 2015 to realize a long-held aspiration-a single, unifying bioinformatics "super conference" for the Asia & Pacific community. Nearly a decade of persistence, coordination, and coalition-building led to the inaugural Asia & Pacific Bioinformatics Joint Congress (APBJC2024) in Okinawa, Japan. Now established as a triennial event, APBJC stands as a testament to the power of collective vision and shared purpose, offering a unifying platform for regional collaboration and scientific exchange. Tagline: Bringing a Region Together: The Making of APBJC.
Asif M. Khan, Susumu Goto, Kenta Nakai, Limsoon Wong, Diane E. Kovats, Shinya Ikematsu, Yoshihiro Yamanishi, Nurul Salwanie Che Wahid, Pradeep Eranti, Yi-Ping Phoebe Chen, Tae-Min Kim, Shinn-Ying Ho, Jessica Cara Mar, Wataru Iwasaki 0001, Jayaraman Valadi, Prashanth Suravajhala, Christian Schönbach, Tin Wee Tan, Shoba Ranganathan, Kiyoko F. Aoki-Kinoshita
Briefings Bioinform.8
2025 SSL-VQ: vector-quantized variational autoencoders for semi-supervised prediction of therapeutic targets across diverse diseases
abstract
MOTIVATION: Identifying effective therapeutic targets poses a challenge in drug discovery, especially for uncharacterized diseases without known therapeutic targets (e.g. rare diseases, intractable diseases). RESULTS: This study presents a novel machine learning approach using multimodal vector-quantized variational autoencoders (VQ-VAEs) for predicting therapeutic target molecules across diseases. To address the lack of known therapeutic target-disease associations, we incorporate the information on uncharacterized diseases without known targets or uncharacterized proteins without known indications (applicable diseases) in the semi-supervised learning (SSL) framework. The method integrates disease-specific and protein perturbation profiles with genetic perturbations (e.g. gene knockdowns and gene overexpressions) at the transcriptome level. Cross-cell representation learning, facilitated by VQ-VAEs, was performed to extract informative features from protein perturbation profiles across diverse human cell types. Concurrently, cross-disease representation learning was performed, leveraging VQ-VAE, to extract informative features reflecting disease states from disease-specific profiles. The model's applicability to uncharacterized diseases or proteins is enhanced by considering the consistency between disease-specific and patient-specific signatures. The efficacy of the method is demonstrated across three practical scenarios for 79 diseases: target repositioning for target-disease pairs, new target prediction for uncharacterized diseases, and new indication prediction for uncharacterized proteins. This method is expected to be valuable for identifying therapeutic targets across various diseases. AVAILABILITY AND IMPLEMENTATION: Code: github.com/YamanishiLab/SSL-VQ and Data: 10.5281/zenodo.14644837.
Satoko Namba, Noriko Yuyama Otani, Yoshihiro Yamanishi
Bioinform.4
2024 GxVAEs: Two Joint VAEs Generate Hit Molecules from Gene Expression Profiles
abstract
The de novo generation of hit-like molecules that show bioactivity and drug-likeness is an important task in computer-aided drug discovery. Although artificial intelligence can generate molecules with desired chemical properties, most previous studies have ignored the influence of disease-related cellular environments. This study proposes a novel deep generative model called GxVAEs to generate hit-like molecules from gene expression profiles by leveraging two joint variational autoencoders (VAEs). The first VAE, ProfileVAE, extracts latent features from gene expression profiles. The extracted features serve as the conditions that guide the second VAE, which is called MolVAE, in generating hit-like molecules. GxVAEs bridge the gap between molecular generation and the cellular environment in a biological system, and produce molecules that are biologically meaningful in the context of specific diseases. Experiments and case studies on the generation of therapeutic molecules show that GxVAEs outperforms current state-of-the-art baselines and yield hit-like molecules with potential bioactivity and drug-like properties. We were able to successfully generate the potential molecular structures with therapeutic effects for various diseases from patients’ disease profiles.
Chen Li 0027, Yoshihiro Yamanishi
AAAI2
2024 TenGAN: Pure Transformer Encoders Make an Efficient Discrete GAN for De Novo Molecular Generation
abstract
Deep generative models for de novo molecular generation using discrete data, such as the simplified molecular-input line-entry system (SMILES) strings, have attracted widespread attention in drug design. However, training instability often plagues generative adversarial networks (GANs), leading to problems such as mode collapse and low diversity. This study proposes a pure transformer encoder-based GAN (TenGAN) to solve these issues. The generator and discriminator of TenGAN are variants of the transformer encoders and are combined with reinforcement learning (RL) to generate molecules with the desired chemical properties. Besides, data augmentation of the variant SMILES is leveraged for the TenGAN training to learn the semantics and syntax of SMILES strings. Additionally, we introduce an enhanced variant of TenGAN, named Ten(W)GAN, which incorporates mini-batch discrimination and Wasserstein GAN to improve the ability to generate molecules. The experimental results and ablation studies on the QM9 and ZINC datasets showed that the proposed models generated highly valid and novel molecules with the desired chemical properties in a computationally efficient manner.
Chen Li 0027, Yoshihiro Yamanishi
AISTATS2
2024 DIRECTEUR: transcriptome-based prediction of small molecules that replace transcription factors for direct cell conversion
abstract
MOTIVATION: Direct reprogramming (DR) is a process that directly converts somatic cells to target cells. Although DR via small molecules is safer than using transcription factors (TFs) in terms of avoidance of tumorigenic risk, the determination of DR-inducing small molecules is challenging. RESULTS: Here we present a novel in silico method, DIRECTEUR, to predict small molecules that replace TFs for DR. We extracted DR-characteristic genes using transcriptome profiles of cells in which DR was induced by TFs, and performed a variant of simulated annealing to explore small molecule combinations with similar gene expression patterns with DR-inducing TFs. We applied DIRECTEUR to predicting combinations of small molecules that convert fibroblasts into neurons or cardiomyocytes, and were able to reproduce experimentally verified and functionally related molecules inducing the corresponding conversions. The proposed method is expected to be useful for practical applications in regenerative medicine. AVAILABILITY AND IMPLEMENTATION: The code and data are available at the following link: https://github.com/HamanoLaboratory/DIRECTEUR.git.
Momoko Hamano, Toru Nakamura, Ryoku Ito, Yuki Shimada, Michio Iwata, Jun-ichi Takeshita, Ryohei Eguchi, Yoshihiro Yamanishi
Bioinform.8
2024 TRAITER: transformer-guided diagnosis and prognosis of heart failure using cell nuclear morphology and DNA damage marker
abstract
MOTIVATION: Heart failure (HF), a major cause of morbidity and mortality, necessitates precise diagnostic and prognostic methods. RESULTS: This study presents a novel deep learning approach, Transformer-based Analysis of Images of Tissue for Effective Remedy (TRAITER), for HF diagnosis and prognosis. Using image segmentation techniques and a Vision Transformer, TRAITER predicts HF likelihood from cardiac tissue cell nuclear morphology images and the potential for left ventricular reverse remodeling (LVRR) from dual-stained images with cell nuclei and DNA damage markers. In HF prediction using 31 158 images from 9 patients, TRAITER achieved 83.1% accuracy. For LVRR prediction with 231 840 images from 46 patients, TRAITER attained 84.2% accuracy for individual images and 92.9% for individual patients. TRAITER outperformed other neural network models in terms of receiver operating characteristics, and precision-recall curves. Our method promises to advance personalized HF medicine decision-making. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at the following link: https://github.com/HamanoLaboratory/predict-of-HF-and-LVRR.
Hiromu Hayashi, Toshiyuki Ko, Zhehao Dai, Kanna Fujita, Seitaro Nomura, Hiroki Kiyoshima, Shinya Ishihara, Momoko Hamano, Issei Komuro, Yoshihiro Yamanishi
Bioinform.10
2024 Quantitative evaluation of molecular generation performance of graph-based GANs
Jinli Zhang, Zongli Jiang, Man Wu, Chen Li 0027, Yoshihiro Yamanishi
Softw. Qual. J.6
2023 Mode Collapse Alleviation of Reinforcement Learning-based GANs in Drug Design
abstract
De novo drug design is a challenging task that involves understanding the principles of chemistry, chemical properties, and the rules that govern molecular interactions. Deep learning-based generative models, such as MolGAN, offer a promising approach for generating new molecules with the desired chemical properties from molecular graphs. Such models often combine a discrete generative adversarial network (GAN) and reinforcement learning (RL) to produce highly valid and novel molecules. However, the severe mode collapse problem leads to low performance. This study aims to alleviate and investigate the effect of multiple factors on mode collapse. We conducted experiments on different sampling methods, training epochs, and datasets of various volumes and evaluated the experimental results using performance metrics such as validity, uniqueness, novelty, and diversity. The experimental results demonstrate that noise sampling distributions, training epochs, and training data volumes affect performance. The experimental results provide a direction for mitigating the mode collapse problem for RL-based discrete GANs.
Zongli Jiang, Jinli Zhang, Man Wu, Chen Li 0027, Yoshihiro Yamanishi
BIBM6
2023 SpotGAN: A Reverse-Transformer GAN Generates Scaffold-Constrained Molecules with Property Optimization
Chen Li 0027, Yoshihiro Yamanishi
ECML/PKDD (1)2
2023 EarlGAN: An enhanced actor-critic reinforcement learning agent-driven GAN for de novo drug design
Huidong Tang, Chen Li 0027, Huachong Yu, Sayaka Kamei, Yoshihiro Yamanishi, Yasuhiko Morimoto
Pattern Recognit. Lett.6
2022 Transformer-based Objective-reinforced Generative Adversarial Network to Generate Desired Molecules
abstract
Deep generative models of sequence-structure data have attracted widespread attention in drug discovery. However, such models cannot fully extract the semantic features of molecules from sequential representations. Moreover, mode collapse reduces the diversity of the generated molecules. This paper proposes a transformer-based objective-reinforced generative adversarial network (TransORGAN) to generate molecules. TransORGAN leverages a transformer architecture as a generator and uses a stochastic policy gradient for reinforcement learning to generate plausible molecules with rich semantic features. The discriminator grants rewards that guide the policy update of the generator, while an objective-reinforced penalty encourages the generation of diverse molecules. Experiments were performed using the ZINC chemical dataset, and the results demonstrated the usefulness of TransORGAN in terms of uniqueness, novelty, and diversity of the generated molecules.
Chen Li 0027, Chikashige Yamanaka, Kazuma Kaitoh, Yoshihiro Yamanishi
IJCAI4
2022 TRANSDIRE: data-driven direct reprogramming by a pioneer factor-guided trans-omics approach
abstract
MOTIVATION: Direct reprogramming involves the direct conversion of fully differentiated mature cell types into various other cell types while bypassing an intermediate pluripotent state (e.g. induced pluripotent stem cells). Cell differentiation by direct reprogramming is determined by two types of transcription factors (TFs): pioneer factors (PFs) and cooperative TFs. PFs have the distinct ability to open chromatin aggregations, assemble a collective of cooperative TFs and activate gene expression. The experimental determination of two types of TFs is extremely difficult and costly. RESULTS: In this study, we developed a novel computational method, TRANSDIRE (TRANS-omics-based approach for DIrect REprogramming), to predict the TFs that induce direct reprogramming in various human cell types using multiple omics data. In the algorithm, potential PFs were predicted based on low signal chromatin regions, and the cooperative TFs were predicted through a trans-omics analysis of genomic data (e.g. enhancers), transcriptome data (e.g. gene expression profiles in human cells), epigenome data (e.g. chromatin immunoprecipitation sequencing data) and interactome data. We applied the proposed methods to the reconstruction of TFs that induce direct reprogramming from fibroblasts to six other cell types: hepatocytes, cartilaginous cells, neurons, cardiomyocytes, pancreatic cells and Paneth cells. We demonstrated that the methods successfully predicted TFs for most cell conversions with high accuracy. Thus, the proposed methods are expected to be useful for various practical applications in regenerative medicine. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at the following website: http://figshare.com/s/b653781a5b9e6639972b. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ryohei Eguchi, Momoko Hamano, Michio Iwata, Toru Nakamura, Shinya Oki, Yoshihiro Yamanishi
Bioinform.6
2022 Small compound-based direct cell conversion with combinatorial optimization of pathway regulations
abstract
MOTIVATION: Direct cell conversion, direct reprogramming (DR), is an innovative technology that directly converts source cells to target cells without bypassing induced pluripotent stem cells. The use of small compounds (e.g. drugs) for DR can help avoid carcinogenic risk induced by gene transfection; however, experimentally identifying small compounds remains challenging because of combinatorial explosion. RESULTS: In this article, we present a new computational method, COMPRENDRE (combinatorial optimization of pathway regulations for direct reprograming), to elucidate the mechanism of small compound-based DR and predict new combinations of small compounds for DR. We estimated the potential target proteins of DR-inducing small compounds and identified a set of target pathways involving DR. We identified multiple DR-related pathways that have not previously been reported to induce neurons or cardiomyocytes from fibroblasts. To overcome the problem of combinatorial explosion, we developed a variant of a simulated annealing algorithm to identify the best set of compounds that can regulate DR-related pathways. Consequently, the proposed method enabled to predict new DR-inducing candidate combinations with fewer compounds and to successfully reproduce experimentally verified compounds inducing the direct conversion from fibroblasts to neurons or cardiomyocytes. The proposed method is expected to be useful for practical applications in regenerative medicine. AVAILABILITY AND IMPLEMENTATION: The code supporting the current study is available at the http://labo.bio.kyutech.ac.jp/~yamani/comprendre. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Toru Nakamura, Michio Iwata, Momoko Hamano, Ryohei Eguchi, Jun-ichi Takeshita, Yoshihiro Yamanishi
Bioinform.6
2022 From drug repositioning to target repositioning: prediction of therapeutic targets using genetically perturbed transcriptomic signatures
abstract
MOTIVATION: A critical element of drug development is the identification of therapeutic targets for diseases. However, the depletion of therapeutic targets is a serious problem. RESULTS: In this study, we propose the novel concept of target repositioning, an extension of the concept of drug repositioning, to predict new therapeutic targets for various diseases. Predictions were performed by a trans-disease analysis which integrated genetically perturbed transcriptomic signatures (knockdown of 4345 genes and overexpression of 3114 genes) and disease-specific gene transcriptomic signatures of 79 diseases. The trans-disease method, which takes into account similarities among diseases, enabled us to distinguish the inhibitory from activatory targets and to predict the therapeutic targetability of not only proteins with known target-disease associations but also orphan proteins without known associations. Our proposed method is expected to be useful for understanding the commonality of mechanisms among diseases and for therapeutic target identification in drug discovery. AVAILABILITY AND IMPLEMENTATION: Supplemental information and software are available at the following website [http://labo.bio.kyutech.ac.jp/~yamani/target_repositioning/]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Satoko Namba, Michio Iwata, Yoshihiro Yamanishi
Bioinform.3
2022 Epigenetic landscape of drug responses revealed through large-scale ChIP-seq data analyses
abstract
BACKGROUND: Elucidating the modes of action (MoAs) of drugs and drug candidate compounds is critical for guiding translation from drug discovery to clinical application. Despite the development of several data-driven approaches for predicting chemical-disease associations, the molecular cues that organize the epigenetic landscape of drug responses remain poorly understood. RESULTS: With the use of a computational method, we attempted to elucidate the epigenetic landscape of drug responses, in terms of transcription factors (TFs), through large-scale ChIP-seq data analyses. In the algorithm, we systematically identified TFs that regulate the expression of chemically induced genes by integrating transcriptome data from chemical induction experiments and almost all publicly available ChIP-seq data (consisting of 13,558 experiments). By relating the resultant chemical-TF associations to a repository of associated proteins for a wide range of diseases, we made a comprehensive prediction of chemical-TF-disease associations, which could then be used to account for drug MoAs. Using this approach, we predicted that: (1) cisplatin promotes the anti-tumor activity of TP53 family members but suppresses the cancer-inducing function of MYCs; (2) inhibition of RELA and E2F1 is pivotal for leflunomide to exhibit antiproliferative activity; and (3) CHD8 mediates valproic acid-induced autism. CONCLUSIONS: Our proposed approach has the potential to elucidate the MoAs for both approved drugs and candidate compounds from an epigenetic perspective, thereby revealing new therapeutic targets, and to guide the discovery of unexpected therapeutic effects, side effects, and novel targets and actions.
Zhaonan Zou, Michio Iwata, Yoshihiro Yamanishi, Shinya Oki
BMC Bioinform.3
2020 Network-based characterization of disease-disease relationships in terms of drugs and therapeutic targets
abstract
MOTIVATION: Disease states are distinguished from each other in terms of differing clinical phenotypes, but characteristic molecular features are often common to various diseases. Similarities between diseases can be explained by characteristic gene expression patterns. However, most disease-disease relationships remain uncharacterized. RESULTS: In this study, we proposed a novel approach for network-based characterization of disease-disease relationships in terms of drugs and therapeutic targets. We performed large-scale analyses of omics data and molecular interaction networks for 79 diseases, including adrenoleukodystrophy, leukaemia, Alzheimer's disease, asthma, atopic dermatitis, breast cancer, cystic fibrosis and inflammatory bowel disease. We quantified disease-disease similarities based on proximities of abnormally expressed genes in various molecular networks, and showed that similarities between diseases could be explained by characteristic molecular network topologies. Furthermore, we developed a kernel matrix regression algorithm to predict the commonalities of drugs and therapeutic targets among diseases. Our comprehensive prediction strategy indicated many new associations among phenotypically diverse diseases. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Midori Iida, Michio Iwata, Yoshihiro Yamanishi
Bioinform.3
2020 Dual graph convolutional neural network for predicting chemical networks
abstract
BACKGROUND: Predicting of chemical compounds is one of the fundamental tasks in bioinformatics and chemoinformatics, because it contributes to various applications in metabolic engineering and drug discovery. The recent rapid growth of the amount of available data has enabled applications of computational approaches such as statistical modeling and machine learning method. Both a set of chemical interactions and chemical compound structures are represented as graphs, and various graph-based approaches including graph convolutional neural networks have been successfully applied to chemical network prediction. However, there was no efficient method that can consider the two different types of graphs in an end-to-end manner. RESULTS: We give a new formulation of the chemical network prediction problem as a link prediction problem in a graph of graphs (GoG) which can represent the hierarchical structure consisting of compound graphs and an inter-compound graph. We propose a new graph convolutional neural network architecture called dual graph convolutional network that learns compound representations from both the compound graphs and the inter-compound network in an end-to-end manner. CONCLUSIONS: Experiments using four chemical networks with different sparsity levels and degree distributions shows that our dual graph convolution approach achieves high prediction performance in relatively dense networks, while the performance becomes inferior on extremely-sparse networks.
Shonosuke Harada, Hirotaka Akita, Masashi Tsubaki, Yukino Baba, Ichigaku Takigawa, Yoshihiro Yamanishi, Hisashi Kashima
BMC Bioinform.6
2020 Space-Efficient Feature Maps for String Alignment Kernels
abstract
Abstract String kernels are attractive data analysis tools for analyzing string data. Among them, alignment kernels are known for their high prediction accuracies in string classifications when tested in combination with SVM in various applications. However, alignment kernels have a crucial drawback in that they scale poorly due to their quadratic computation complexity in the number of input strings, which limits large-scale applications in practice. We address this need by presenting the first approximation for string alignment kernels, which we call space-efficient feature maps for edit distance with moves (SFMEDM), by leveraging a metric embedding named edit-sensitive parsing and feature maps (FMs) of random Fourier features (RFFs) for large-scale string analyses. The original FMs for RFFs consume a huge amount of memory proportional to the dimension d of input vectors and the dimension D of output vectors, which prohibits its large-scale applications. We present novel space-efficient feature maps (SFMs) of RFFs for a space reduction from O(dD) of the original FMs to O(d) of SFMs with a theoretical guarantee with respect to concentration bounds. We experimentally test SFMEDM on its ability to learn SVM for large-scale string classifications with various massive string data, and we demonstrate the superior performance of SFMEDM with respect to prediction accuracy, scalability and computation efficiency.
Yasuo Tabei, Yoshihiro Yamanishi, Rasmus Pagh
Data Sci. Eng.2
2019 Space-Efficient Feature Maps for String Alignment Kernels
abstract
String kernels are attractive data analysis tools for analyzing string data. Among them, alignment kernels are known for their high prediction accuracies in string classifications when tested in combination with SVM in various applications. However, alignment kernels have a crucial drawback in that they scale poorly due to their quadratic computation complexity in the number of input strings, which limits large-scale applications in practice. We address this need by presenting the first approximation for string alignment kernels, which we call space-efficient feature maps for edit distance with moves (SFMEDM), by leveraging a metric embedding named edit sensitive parsing (ESP) and feature maps (FMs) of random Fourier features (RFFs). The original FMs for RFFs consume a huge amount of memory proportional to the dimension d of input vectors and the dimension D of output vectors. Thus, we present novel space-efficient feature maps (SFMs) of RFFs for a space reduction from O(dD) of the original FMs to O(d) of SFMs with a theoretical guarantee with respect to concentration bounds. We experimentally test SFMEDM on its ability to learn SVM for large-scale string classifications with various massive string data, and we demonstrate the superior performance of SFMEDM with respect to prediction accuracy, scalability and computation efficiency.
Yasuo Tabei, Yoshihiro Yamanishi, Rasmus Pagh
ICDM2
2019 Predicting drug-induced transcriptome responses of a wide range of human cell lines by a novel tensor-train decomposition algorithm
abstract
MOTIVATION: Genome-wide identification of the transcriptomic responses of human cell lines to drug treatments is a challenging issue in medical and pharmaceutical research. However, drug-induced gene expression profiles are largely unknown and unobserved for all combinations of drugs and human cell lines, which is a serious obstacle in practical applications. RESULTS: Here, we developed a novel computational method to predict unknown parts of drug-induced gene expression profiles for various human cell lines and predict new drug therapeutic indications for a wide range of diseases. We proposed a tensor-train weighted optimization (TT-WOPT) algorithm to predict the potential values for unknown parts in tensor-structured gene expression data. Our results revealed that the proposed TT-WOPT algorithm can accurately reconstruct drug-induced gene expression data for a range of human cell lines in the Library of Integrated Network-based Cellular Signatures. The results also revealed that in comparison with the use of original gene expression profiles, the use of imputed gene expression profiles improved the accuracy of drug repositioning. We also performed a comprehensive prediction of drug indications for diseases with gene expression profiles, which suggested many potential drug indications that were not predicted by previous approaches. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Michio Iwata, Longhao Yuan, Qibin Zhao, Yasuo Tabei, Francois Berenger, Ryusuke Sawada, Sayaka Akiyoshi, Momoko Hamano, Yoshihiro Yamanishi
Bioinform.9
2019 Guest Editorial for the 16th Asia Pacific Bioinformatics Conference
abstract
The eight papers in this special section were presented at the 16th Asia Pacific Bioinformatics Conference (APBC2018), which was held in Yokohama, Japan, 15-17 January 2018. The aim of this conference is to provide an international forum for researchers, professionals, and industrial practitioners to share their knowledge and ideas of how to surf the tidal wave of information in the area of bioinformatics and computational biology.
Yoshihiro Yamanishi, Yasubumi Sakakibara, Yi-Ping Phoebe Chen
IEEE ACM Trans. Comput. Biol. Bioinform.1
2016 Scalable Partial Least Squares Regression on Grammar-Compressed Data Matrices
abstract
With massive high-dimensional data now commonplace in research and industry, there is a strong and growing demand for more scalable computational techniques for data analysis and knowledge discovery. Key to turning these data into knowledge is the ability to learn statistical models with high interpretability. Current methods for learning statistical models either produce models that are not interpretable or have prohibitive computational costs when applied to massive data. In this paper we address this need by presenting a scalable algorithm for partial least squares regression (PLS), which we call compression-based PLS (cPLS), to learn predictive linear models with a high interpretability from massive high-dimensional data. We propose a novel grammar-compressed representation of data matrices that supports fast row and column access while the data matrix is in a compressed form. The original data matrix is grammar-compressed and then the linear model in PLS is learned on the compressed data matrix, which results in a significant reduction in working space, greatly improving scalability. We experimentally test cPLS on its ability to learn linear models for classification, regression and feature extraction with various massive high-dimensional data, and show that cPLS performs superiorly in terms of prediction accuracy, computational efficiency, and interpretability.
Yasuo Tabei, Hiroto Saigo, Yoshihiro Yamanishi, Simon J. Puglisi
KDD3
2016 Simultaneous prediction of enzyme orthologs from chemical transformation patterns for de novo metabolic pathway reconstruction
abstract
MOTIVATION: Metabolic pathways are an important class of molecular networks consisting of compounds, enzymes and their interactions. The understanding of global metabolic pathways is extremely important for various applications in ecology and pharmacology. However, large parts of metabolic pathways remain unknown, and most organism-specific pathways contain many missing enzymes. RESULTS: In this study we propose a novel method to predict the enzyme orthologs that catalyze the putative reactions to facilitate the de novo reconstruction of metabolic pathways from metabolome-scale compound sets. The algorithm detects the chemical transformation patterns of substrate-product pairs using chemical graph alignments, and constructs a set of enzyme-specific classifiers to simultaneously predict all the enzyme orthologs that could catalyze the putative reactions of the substrate-product pairs in the joint learning framework. The originality of the method lies in its ability to make predictions for thousands of enzyme orthologs simultaneously, as well as its extraction of enzyme-specific chemical transformation patterns of substrate-product pairs. We demonstrate the usefulness of the proposed method by applying it to some ten thousands of metabolic compounds, and analyze the extracted chemical transformation patterns that provide insights into the characteristics and specificities of enzymes. The proposed method will open the door to both primary (central) and secondary metabolism in genomics research, increasing research productivity to tackle a wide variety of environmental and public health matters. CONTACT: : [email protected].
Yasuo Tabei, Yoshihiro Yamanishi, Masaaki Kotera
Bioinform.2
2015 Metabolome-scale de novo pathway reconstruction using regioisomer-sensitive graph alignments
abstract
MOTIVATION: Recent advances in mass spectrometry and related metabolomics technologies have enabled the rapid and comprehensive analysis of numerous metabolites. However, biosynthetic and biodegradation pathways are only known for a small portion of metabolites, with most metabolic pathways remaining uncharacterized. RESULTS: In this study, we developed a novel method for supervised de novo metabolic pathway reconstruction with an improved graph alignment-based approach in the reaction-filling framework. We proposed a novel chemical graph alignment algorithm, which we called PACHA (Pairwise Chemical Aligner), to detect the regioisomer-sensitive connectivities between the aligned substructures of two compounds. Unlike other existing graph alignment methods, PACHA can efficiently detect only one common subgraph between two compounds. Our results show that the proposed method outperforms previous descriptor-based methods or existing graph alignment-based methods in the enzymatic reaction-likeness prediction for isomer-enriched reactions. It is also useful for reaction annotation that assigns potential reaction characteristics such as EC (Enzyme Commission) numbers and PIERO (Enzymatic Reaction Ontology for Partial Information) terms to substrate-product pairs. Finally, we conducted a comprehensive enzymatic reaction-likeness prediction for all possible uncharacterized compound pairs, suggesting potential metabolic pathways for newly predicted substrate-product pairs.
Yoshihiro Yamanishi, Yasuo Tabei, Masaaki Kotera
Bioinform.1
2014 Metabolome-scale prediction of intermediate compounds in multistep metabolic pathways with a recursive supervised approach
abstract
MOTIVATION: Metabolic pathway analysis is crucial not only in metabolic engineering but also in rational drug design. However, the biosynthetic/biodegradation pathways are known only for a small portion of metabolites, and a vast amount of pathways remain uncharacterized. Therefore, an important challenge in metabolomics is the de novo reconstruction of potential reaction networks on a metabolome-scale. RESULTS: In this article, we develop a novel method to predict the multistep reaction sequences for de novo reconstruction of metabolic pathways in the reaction-filling framework. We propose a supervised approach to learn what we refer to as 'multistep reaction sequence likeness', i.e. whether a compound-compound pair is possibly converted to each other by a sequence of enzymatic reactions. In the algorithm, we propose a recursive procedure of using step-specific classifiers to predict the intermediate compounds in the multistep reaction sequences, based on chemical substructure fingerprints/descriptors of compounds. We further demonstrate the usefulness of our proposed method on the prediction of enzymatic reaction networks from a metabolome-scale compound set and discuss characteristic features of the extracted chemical substructure transformation patterns in multistep reaction sequences. Our comprehensively predicted reaction networks help to fill the metabolic gap and to infer new reaction sequences in metabolic pathways. AVAILABILITY AND IMPLEMENTATION: Materials are available for free at http://web.kuicr.kyoto-u.ac.jp/supp/kot/ismb2014/
Masaaki Kotera, Yasuo Tabei, Yoshihiro Yamanishi, Ai Muto, Yuki Moriya, Toshiaki Tokimatsu, Susumu Goto
Bioinform.3
2013 Succinct interval-splitting tree for scalable similarity search of compound-protein pairs with property constraints
abstract
Analyzing functional interactions between small compounds and proteins is indispensable in genomic drug discovery. Since rich information on various compound-protein inter- actions is available in recent molecular databases, strong demands for making best use of such databases require to in- vent powerful methods to help us find new functional compound-protein pairs on a large scale. We present the succinct interval-splitting tree algorithm (SITA) that efficiently per- forms similarity search in databases for compound-protein pairs with respect to both binary fingerprints and real-valued properties. SITA achieves both time and space efficiency by developing the data structure called interval-splitting trees, which enables to efficiently prune the useless portions of search space, and by incorporating the ideas behind wavelet tree, a succinct data structure to compactly represent trees. We experimentally test SITA on the ability to retrieve similar compound-protein pairs/substrate-product pairs for a query from large databases with over 200 million compound- protein pairs/substrate-product pairs and show that SITA performs better than other possible approaches.
Yasuo Tabei, Akihiro Kishimoto, Masaaki Kotera, Yoshihiro Yamanishi
KDD4
2013 Supervised de novo reconstruction of metabolic pathways from metabolome-scale compound sets
abstract
MOTIVATION: The metabolic pathway is an important biochemical reaction network involving enzymatic reactions among chemical compounds. However, it is assumed that a large number of metabolic pathways remain unknown, and many reactions are still missing even in known pathways. Therefore, the most important challenge in metabolomics is the automated de novo reconstruction of metabolic pathways, which includes the elucidation of previously unknown reactions to bridge the metabolic gaps. RESULTS: In this article, we develop a novel method to reconstruct metabolic pathways from a large compound set in the reaction-filling framework. We define feature vectors representing the chemical transformation patterns of compound-compound pairs in enzymatic reactions using chemical fingerprints. We apply a sparsity-induced classifier to learn what we refer to as 'enzymatic-reaction likeness', i.e. whether compound pairs are possibly converted to each other by enzymatic reactions. The originality of our method lies in the search for potential reactions among many compounds at a time, in the extraction of reaction-related chemical transformation patterns and in the large-scale applicability owing to the computational efficiency. In the results, we demonstrate the usefulness of our proposed method on the de novo reconstruction of 134 metabolic pathways in Kyoto Encyclopedia of Genes and Genomes (KEGG). Our comprehensively predicted reaction networks of 15 698 compounds enable us to suggest many potential pathways and to increase research productivity in metabolomics. AVAILABILITY: Softwares are available on request. Supplementary material are available at http://web.kuicr.kyoto-u.ac.jp/supp/kot/ismb2013/.
Masaaki Kotera, Yasuo Tabei, Yoshihiro Yamanishi, Toshiaki Tokimatsu, Susumu Goto
Bioinform.3
2012 Relating drug-protein interaction network with drug side effects
abstract
MOTIVATION: Identifying the emergence and underlying mechanisms of drug side effects is a challenging task in the drug development process. This underscores the importance of system-wide approaches for linking different scales of drug actions; namely drug-protein interactions (molecular scale) and side effects (phenotypic scale) toward side effect prediction for uncharacterized drugs. RESULTS: We performed a large-scale analysis to extract correlated sets of targeted proteins and side effects, based on the co-occurrence of drugs in protein-binding profiles and side effect profiles, using sparse canonical correlation analysis. The analysis of 658 drugs with the two profiles for 1368 proteins and 1339 side effects led to the extraction of 80 correlated sets. Enrichment analyses using KEGG and Gene Ontology showed that most of the correlated sets were significantly enriched with proteins that are involved in the same biological pathways, even if their molecular functions are different. This allowed for a biologically relevant interpretation regarding the relationship between drug-targeted proteins and side effects. The extracted side effects can be regarded as possible phenotypic outcomes by drugs targeting the proteins that appear in the same correlated set. The proposed method is expected to be useful for predicting potential side effects of new drug candidate compounds based on their protein-binding profiles. SUPPLEMENTARY INFORMATION: Datasets and all results are available at http://web.kuicr.kyoto-u.ac.jp/supp/smizutan/target-effect/. AVAILABILITY: Software is available at the above supplementary website. CONTACT: [email protected], or [email protected].
Sayaka Mizutani, Edouard Pauwels, Véronique Stoven, Susumu Goto, Yoshihiro Yamanishi
Bioinform.5
2012 Identification of chemogenomic features from drug-target interaction networks using interpretable classifiers
abstract
MOTIVATION: Drug effects are mainly caused by the interactions between drug molecules and their target proteins including primary targets and off-targets. Identification of the molecular mechanisms behind overall drug-target interactions is crucial in the drug design process. RESULTS: We develop a classifier-based approach to identify chemogenomic features (the underlying associations between drug chemical substructures and protein domains) that are involved in drug-target interaction networks. We propose a novel algorithm for extracting informative chemogenomic features by using L(1) regularized classifiers over the tensor product space of possible drug-target pairs. It is shown that the proposed method can extract a very limited number of chemogenomic features without loosing the performance of predicting drug-target interactions and the extracted features are biologically meaningful. The extracted substructure-domain association network enables us to suggest ligand chemical fragments specific for each protein domain and ligand core substructures important for a wide range of protein families. AVAILABILITY: Softwares are available at the supplemental website. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Datasets and all results are available at http://cbio.ensmp.fr/~yyamanishi/l1binary/ .
Yasuo Tabei, Edouard Pauwels, Véronique Stoven, Kazuhiro Takemoto, Yoshihiro Yamanishi
Bioinform.5
2012 Drug target prediction using adverse event report systems: a pharmacogenomic approach
abstract
MOTIVATION: Unexpected drug activities derived from off-targets are usually undesired and harmful; however, they can occasionally be beneficial for different therapeutic indications. There are many uncharacterized drugs whose target proteins (including the primary target and off-targets) remain unknown. The identification of all potential drug targets has become an important issue in drug repositioning to reuse known drugs for new therapeutic indications. RESULTS: We defined pharmacological similarity for all possible drugs using the US Food and Drug Administration's (FDA's) adverse event reporting system (AERS) and developed a new method to predict unknown drug-target interactions on a large scale from the integration of pharmacological similarity of drugs and genomic sequence similarity of target proteins in the framework of a pharmacogenomic approach. The proposed method was applicable to a large number of drugs and it was useful especially for predicting unknown drug-target interactions that could not be expected from drug chemical structures. We made a comprehensive prediction for potential off-targets of 1874 drugs with known targets and potential target profiles of 2519 drugs without known targets, which suggests many potential drug-target interactions that were not predicted by previous chemogenomic or pharmacogenomic approaches. AVAILABILITY: Softwares are available upon request. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Datasets and all results are available at http://cbio.ensmp.fr/~yyamanishi/aers/.
Masataka Takarabe, Masaaki Kotera, Yosuke Nishimura, Susumu Goto, Yoshihiro Yamanishi
Bioinform.5
2011 Predicting drug side-effect profiles: a chemical fragment-based approach
abstract
BACKGROUND: Drug side-effects, or adverse drug reactions, have become a major public health concern. It is one of the main causes of failure in the process of drug development, and of drug withdrawal once they have reached the market. Therefore, in silico prediction of potential side-effects early in the drug discovery process, before reaching the clinical stages, is of great interest to improve this long and expensive process and to provide new efficient and safe therapies for patients. RESULTS: In the present work, we propose a new method to predict potential side-effects of drug candidate molecules based on their chemical structures, applicable on large molecular databanks. A unique feature of the proposed method is its ability to extract correlated sets of chemical substructures (or chemical fragments) and side-effects. This is made possible using sparse canonical correlation analysis (SCCA). In the results, we show the usefulness of the proposed method by predicting 1385 side-effects in the SIDER database from the chemical structures of 888 approved drugs. These predictions are performed with simultaneous extraction of correlated ensembles formed by a set of chemical substructures shared by drugs that are likely to have a set of side-effects. We also conduct a comprehensive side-effect prediction for many uncharacterized drug molecules stored in DrugBank, and were able to confirm interesting predictions using independent source of information. CONCLUSIONS: The proposed method is expected to be useful in various stages of the drug development process.
Edouard Pauwels, Véronique Stoven, Yoshihiro Yamanishi
BMC Bioinform.3
2010 Drug-target interaction prediction from chemical, genomic and pharmacological data in an integrated framework
abstract
MOTIVATION: In silico prediction of drug-target interactions from heterogeneous biological data is critical in the search for drugs and therapeutic targets for known diseases such as cancers. There is therefore a strong incentive to develop new methods capable of detecting these potential drug-target interactions efficiently. RESULTS: In this article, we investigate the relationship between the chemical space, the pharmacological space and the topology of drug-target interaction networks, and show that drug-target interactions are more correlated with pharmacological effect similarity than with chemical structure similarity. We then develop a new method to predict unknown drug-target interactions from chemical, genomic and pharmacological data on a large scale. The proposed method consists of two steps: (i) prediction of pharmacological effects from chemical structures of given compounds and (ii) inference of unknown drug-target interactions based on the pharmacological effect similarity in the framework of supervised bipartite graph inference. The originality of the proposed method lies in the prediction of potential pharmacological similarity for any drug candidate compounds and in the integration of chemical, genomic and pharmacological data in a unified framework. In the results, we make predictions for four classes of important drug-target interactions involving enzymes, ion channels, GPCRs and nuclear receptors. Our comprehensively predicted drug-target interaction networks enable us to suggest many potential drug-target interactions and to increase research productivity toward genomic drug discovery. SUPPLEMENTARY INFORMATION: Datasets and all prediction results are available at http://cbio.ensmp.fr/~yyamanishi/pharmaco/. AVAILABILITY: Softwares are available upon request.
Yoshihiro Yamanishi, Masaaki Kotera, Minoru Kanehisa, Susumu Goto
Bioinform.1
2009 On Pairwise Kernels: An Efficient Alternative and Generalization Analysis
Hisashi Kashima, Satoshi Oyama, Yoshihiro Yamanishi, Koji Tsuda
PAKDD3
2009 Link Propagation: A Fast Semi-supervised Learning Algorithm for Link Prediction
abstract
We propose Link Propagation as a new semi-supervised learning method for link prediction problems, where the task is to predict unknown parts of the network structure by using auxiliary information such as node similarities. Since the proposed method can fill in missing parts of tensors, it is applicable to multi-relational domains, allowing us to handle multiple types of links simultaneously. We also give a novel efficient algorithm for Link Propagation based on an accelerated conjugate gradient method.
Hisashi Kashima, Tsuyoshi Kato, Yoshihiro Yamanishi, Masashi Sugiyama, Koji Tsuda
SDM3
2009 Supervised prediction of drug-target interactions using bipartite local models
abstract
MOTIVATION: In silico prediction of drug-target interactions from heterogeneous biological data is critical in the search for drugs for known diseases. This problem is currently being attacked from many different points of view, a strong indication of its current importance. Precisely, being able to predict new drug-target interactions with both high precision and accuracy is the holy grail, a fundamental requirement for in silico methods to be useful in a biological setting. This, however, remains extremely challenging due to, amongst other things, the rarity of known drug-target interactions. RESULTS: We propose a novel supervised inference method to predict unknown drug-target interactions, represented as a bipartite graph. We use this method, known as bipartite local models to first predict target proteins of a given drug, then to predict drugs targeting a given protein. This gives two independent predictions for each putative drug-target interaction, which we show can be combined to give a definitive prediction for each interaction. We demonstrate the excellent performance of the proposed method in the prediction of four classes of drug-target interaction networks involving enzymes, ion channels, G protein-coupled receptors (GPCRs) and nuclear receptors in human. This enables us to suggest a number of new potential drug-target interactions. AVAILABILITY: An implementation of the proposed algorithm is available upon request from the authors. Datasets and all prediction results are available at http://cbio.ensmp.fr/~yyamanishi/bipartitelocal/.
Kevin Bleakley, Yoshihiro Yamanishi
Bioinform.2
2009 Simultaneous inference of biological networks of multiple species from genome-wide data and evolutionary information: a semi-supervised approach
abstract
MOTIVATION: The existing supervised methods for biological network inference work on each of the networks individually based only on intra-species information such as gene expression data. We believe that it will be more effective to use genomic data and cross-species evolutionary information from different species simultaneously, rather than to use the genomic data alone. RESULTS: We created a new semi-supervised learning method called Link Propagation for inferring biological networks of multiple species based on genome-wide data and evolutionary information. The new method was applied to simultaneous reconstruction of three metabolic networks of Caenorhabditis elegans, Helicobacter pylori and Saccharomyces cerevisiae, based on gene expression similarities and amino acid sequence similarities. The experimental results proved that the new simultaneous network inference method consistently improves the predictive performance over the individual network inferences, and it also outperforms in accuracy and speed other established methods such as the pairwise support vector machine. AVAILABILITY: The software and data are available at http://cbio.ensmp.fr/~yyamanishi/LinkPropagation/.
Hisashi Kashima, Yoshihiro Yamanishi, Tsuyoshi Kato, Masashi Sugiyama, Koji Tsuda
Bioinform.2
2009 E-zyme: predicting potential EC numbers from the chemical transformation pattern of substrate-product pairs
abstract
MOTIVATION: The IUBMB's Enzyme Nomenclature system, commonly known as the Enzyme Commission (EC) numbers, plays key roles in classifying enzymatic reactions and in linking the enzyme genes or proteins to reactions in metabolic pathways. There are numerous reactions known to be present in various pathways but without any official EC numbers, most of which have no hope to be given ones because of the lack of the published articles on enzyme assays. RESULTS: In this article we propose a new method to predict the potential EC numbers to given reactant pairs (substrates and products) or uncharacterized reactions, and a web-server named E-zyme as an application. This technology is based on our original biochemical transformation pattern which we call an 'RDM pattern', and consists of three steps: (i) graph alignment of a query reactant pair (substrates and products) for computing the query RDM pattern, (ii) multi-layered partial template matching by comparing the query RDM pattern with template patterns related with known EC numbers and (iii) weighted major voting scheme for selecting appropriate EC numbers. As the result, cross-validation experiments show that the proposed method achieves both high coverage and high prediction accuracy at a practical level, and consistently outperforms the previous method. AVAILABILITY: The E-zyme system is available at http://www.genome.jp/tools/e-zyme/.
Yoshihiro Yamanishi, Masahiro Hattori, Masaaki Kotera, Susumu Goto, Minoru Kanehisa
Bioinform.1
2008 Prediction of drug-target interaction networks from the integration of chemical and genomic spaces
abstract
MOTIVATION: The identification of interactions between drugs and target proteins is a key area in genomic drug discovery. Therefore, there is a strong incentive to develop new methods capable of detecting these potential drug-target interactions efficiently. RESULTS: In this article, we characterize four classes of drug-target interaction networks in humans involving enzymes, ion channels, G-protein-coupled receptors (GPCRs) and nuclear receptors, and reveal significant correlations between drug structure similarity, target sequence similarity and the drug-target interaction network topology. We then develop new statistical methods to predict unknown drug-target interaction networks from chemical structure and genomic sequence information simultaneously on a large scale. The originality of the proposed method lies in the formalization of the drug-target interaction inference as a supervised learning problem for a bipartite graph, the lack of need for 3D structure information of the target proteins, and in the integration of chemical and genomic spaces into a unified space that we call 'pharmacological space'. In the results, we demonstrate the usefulness of our proposed method for the prediction of the four classes of drug-target interaction networks. Our comprehensively predicted drug-target interaction networks enable us to suggest many potential drug-target interactions and to increase research productivity toward genomic drug discovery. AVAILABILITY: Softwares are available upon request. SUPPLEMENTARY INFORMATION: Datasets and all prediction results are available at http://web.kuicr.kyoto-u.ac.jp/supp/yoshi/drugtarget/.
Yoshihiro Yamanishi, Michihiro Araki, Alex Gutteridge, Wataru Honda, Minoru Kanehisa
ISMB1
2008 Supervised Bipartite Graph Inference
abstract
We formulate the problem of bipartite graph inference as a supervised learning problem, and propose a new method to solve it from the viewpoint of distance metric learning. The method involves the learning of two mappings of the heterogeneous objects to a unified Euclidean space representing the network topology of the bipartite graph, where the graph is easy to infer. The algorithm can be formulated as an optimization problem in a reproducing kernel Hilbert space. We report encouraging results on the problem of compound-protein interaction network reconstruction from chemical structure data and genomic sequence data.
Yoshihiro Yamanishi
NIPS1
2007 Glycan classification with tree kernels
abstract
MOTIVATION: Glycans are covalent assemblies of sugar that play crucial roles in many cellular processes. Recently, comprehensive data about the structure and function of glycans have been accumulated, therefore the need for methods and algorithms to analyze these data is growing fast. RESULTS: This article presents novel methods for classifying glycans and detecting discriminative glycan motifs with support vector machines (SVM). We propose a new class of tree kernels to measure the similarity between glycans. These kernels are based on the comparison of tree substructures, and take into account several glycan features such as the sugar type, the sugar bound type or layer depth. The proposed methods are tested on their ability to classify human glycans into four blood components: leukemia cells, erythrocytes, plasma and serum. They are shown to outperform a previously published method. We also applied a feature selection approach to extract glycan motifs which are characteristic of each blood component. We confirmed that some leukemia-specific glycan motifs detected by our method corresponded to several results in the literature. AVAILABILITY: Softwares are available upon request. SUPPLEMENTARY INFORMATION: Datasets are available at the following website: http://web.kuicr.kyoto-u.ac.jp/supp/yoshi/glycankernel/
Yoshihiro Yamanishi, Francis R. Bach, Jean-Philippe Vert
Bioinform.1
2006 Partial correlation coefficient between distance matrices as a new indicator of protein-protein interactions
abstract
MOTIVATION: The computational prediction of protein-protein interactions is currently a major issue in bioinformatics. Recently, a variety of co-evolution-based methods have been investigated toward this goal. In this study, we introduced a partial correlation coefficient as a new measure for the degree of co-evolution between proteins, and proposed its use to predict protein-protein interactions. RESULTS: The accuracy of the prediction by the proposed method was compared with those of the original mirror tree method and the projection method previously developed by our group. We found that the partial correlation coefficient effectively reduces the number of false positives, as compared with other methods, although the number of false negatives increased in the prediction by the partial correlation coefficient. AVAILABILITY: The R script for the prediction of protein-protein interactions reported in this manuscript is available at http://timpani.genome.ad.jp/~parco/
Tetsuya Sato 0003, Yoshihiro Yamanishi, Katsuhisa Horimoto, Minoru Kanehisa, Hiroyuki Toh
Bioinform.2
2005 The inference of protein-protein interactions by co-evolutionary analysis is improved by excluding the information about the phylogenetic relationships
abstract
MOTIVATION: The prediction of protein-protein interactions is currently an important issue in bioinformatics. The mirror tree method uses evolutionary information to predict protein-protein interactions. However, it has been recognized that predictions by the mirror tree method lead to many false positives. The incentive of our study was to solve this problem by improving the method of extracting the co-evolutionary information regarding the protein pairs. RESULTS: We developed a novel method to predict protein-protein interactions from co-evolutionary information in the framework of the mirror tree method. The originality is the use of the projection operator to exclude the information about the phylogenetic relationships among the source organisms from the distance matrix. Each distance matrix was transformed into a vector for the operation. The vector is referred to as a 'phylogenetic vector'. We have proposed three ways to extract the phylogenetic information: (1) using the 16S rRNA from the same source organisms as the proteins under consideration, (2) averaging the phylogenetic vectors and (3) analyzing the principal components of the phylogenetic vectors. We examined the performance of the proposed methods to predict interacting protein pairs from Escherichia coli, using experimentally verified data. Our method was successful, and it drastically reduced the number of false positives in the prediction. AVAILABILITY: The R script for the prediction of protein-protein interactions reported in this manuscript is available at http://timpani.genome.ad.jp/~proj/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: The information is also available at the same site as the R script.
Tetsuya Sato 0003, Yoshihiro Yamanishi, Minoru Kanehisa, Hiroyuki Toh
Bioinform.2
2004 Supervised Graph Inference
abstract
We formulate the problem of graph inference where part of the graph is known as a supervised learning problem, and propose an algorithm to solve it. The method involves the learning of a mapping of the vertices to a Euclidean space where the graph is easy to infer, and can be formu- lated as an optimization problem in a reproducing kernel Hilbert space. We report encouraging results on the problem of metabolic network re- construction from genomic data.
Jean-Philippe Vert, Yoshihiro Yamanishi
NIPS2