EDBT 2026 Demo / reviewers in the wild / expert
Satoru Miyano
dblp:73/2912
· DBLP profile ↗
122ranked-venue papers
13as first author
8since 2021 · last 2026
0000-0002-1753-6616ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 88 · 8 since 2021Theory of computation · 23 · 13 first-authorArtificial intelligence and machine learning · 14Databases, data management, data science and information retrieval · 7 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MetaCCI: meta cell-cell interaction inference and its application to CCIs characteristics of MDSabstractMOTIVATION: Cell-cell interactions (CCIs) are fundamental to multicellular organisms and play crucial roles in diverse biological processes and disease mechanisms. Understanding CCIs is vital for deciphering disease pathogenesis and developing therapeutic strategies. Although numerous computational methods have been developed to infer CCIs from complex biological data, most existing approaches rely primarily on single-gene expression levels and ligand-receptor databases, often failing to capture the nuanced network-wide changes characteristic of disease states. RESULT: We propose MetaCCI, a novel computational strategy that integrates meta-information into CCI inference by extending the traditional gene expression-based analysis to a gene regulatory network framework. MetaCCI meticulously combines established ligand-receptor pairs with quantitative insights into gene behavior within complex gene networks, enabling the precise extraction of relevant targets for CCI inference. Subsequently, CCI inference was performed using an eigen cell co-expression network, providing a more holistic view of cell-cell communication. Monte Carlo simulations demonstrated that MetaCCI consistently outperforms existing methods in CCI inference. We applied MetaCCI to characterize cell-cell communication in Myelodysplastic Syndromes (MDS). Our results identified distinct interaction patterns in MDS compared with normal cell populations, specifically highlighting the loss of CCIs between "Dendritic cells and Hematopoietic precursor cells" and between "Dendritic cells and Hematopoietic multipotent progenitor cells" as characteristic features of MDS. Furthermore, FABP5, CD63, and HMGB1 were identified as MDS-specific markers. These findings suggest that diminished CCIs involving dendritic cells, hematopoietic precursor cells, and multipotent progenitor cells are pivotal to MDS pathogenesis. AVAILABILITY AND IMPLEMENTATION: The MetaCCI software is freely available at https://github.com/HeewonGitHub/MetaCCI. An archived version of the software and example datasets used in this study is available at Zenodo: https://doi.org/10.5281/zenodo.20101527. Heewon Park, Seiya Imoto, Satoru Miyano |
Bioinform. | 3 |
| 2025 | Powerful gene network enrichment analysis and its application to severe COVID-19 gene networkabstractUnderstanding complex disease mechanisms requires research methods beyond individual gene analysis to capture the coordinated behavior of genes within regulatory networks. Traditional gene set enrichment approaches such as over-representation analysis and gene set enrichment analysis focus primarily on gene lists and often overlook the intricate network structures that control cellular processes. Although a gene network enrichment analysis strategy (GbNEA) has been proposed, this method assesses enrichment significance via phenotype permutation and the Kolmogorov-Smirnov test, which lowers statistical power and increases the computational burden due to repeated gene network re-estimation. To overcome these limitations, we developed a novel approach, powerful gene network enrichment analysis (PGNEA), which characterizes gene networks by integrating gene expression, regulatory effects, and hubness. PGNEA evaluates the enrichment of phenotype-specific gene networks by quantifying differences in gene activity patterns and assesses statistical significance by evaluating permutation of gene activity rather than phenotype permutation. This approach exhibits significantly enhanced computational efficiency and statistical sensitivity. We demonstrated the advantages of PGNEA through Monte Carlo simulations and applied it to whole-blood RNA-seq data obtained from the Japan COVID-19 Task Force. PGNEA successfully identified viral infection-related pathways enriched in severe COVID-19 gene networks, including those linked to "COVID-19," "HIV-1 infection," "Hepatitis B," "Influenza A," "Measles," and "Kaposi sarcoma-associated herpesvirus infection." Notably, key molecular markers such as PIK3, NF-B family members, FOXA, JUN, and CXCL8 were identified, with strong and consistent molecular interplays between CXCL8 and NFKBIA. These findings underscore the potential of PGNEA as an efficient tool for identifying biologically meaningful pathways and network-level mechanisms associated with various phenotypes, including severe viral infections. Heewon Park, Seiya Imoto, Satoru Miyano |
Briefings Bioinform. | 3 |
| 2023 | Acceleration of BAM I/O on distributed file systemsabstractRapid advances in high-throughput sequencers have made it possible to obtain large amounts of whole genome data quickly and inexpensively. As the amount of data increases, the increase in computation time has become a serious problem. One of the main causes of this problem is file I/O performance. Most pipelines do not implement file I/O suitable for distributed file systems, which are common storage systems in parallel computers used for large-scale analysis. In this study, we developed an I/O system adapted to distributed file systems. In the proposed system, storage access frequency was suppressed by I/O buffers, and thread parallelization by OpenMP was employed. We tested the developed system and a high parallel speedup was achieved. Availability: The developed system is freely available from https://github.com/SatoshiITO/GAPS Satoru Miyano, Kenji Ono |
BIBM | 2 |
| 2022 | Identification of bacteriophage genome sequences with representation learningabstractMOTIVATION: Bacteriophages/phages are the viruses that infect and replicate within bacteria and archaea, and rich in human body. To investigate the relationship between phages and microbial communities, the identification of phages from metagenome sequences is the first step. Currently, there are two main methods for identifying phages: database-based (alignment-based) methods and alignment-free methods. Database-based methods typically use a large number of sequences as references; alignment-free methods usually learn the features of the sequences with machine learning and deep learning models. RESULTS: We propose INHERIT which uses a deep representation learning model to integrate both database-based and alignment-free methods, combining the strengths of both. Pre-training is used as an alternative way of acquiring knowledge representations from existing databases, while the BERT-style deep learning framework retains the advantage of alignment-free methods. We compare INHERIT with four existing methods on a third-party benchmark dataset. Our experiments show that INHERIT achieves a better performance with the F1-score of 0.9932. In addition, we find that pre-training two species separately helps the non-alignment deep learning model make more accurate predictions. AVAILABILITY AND IMPLEMENTATION: The codes of INHERIT are now available in: https://github.com/Celestial-Bai/INHERIT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zeheng Bai, Satoru Miyano, Rui Yamaguchi, Kosuke Fujimoto, Satoshi Uematsu, Seiya Imoto |
Bioinform. | 3 |
| 2022 | RoDiCE: robust differential protein co-expression analysis for cancer complexomeabstractMOTIVATION: The full spectrum of abnormalities in cancer-associated protein complexes remains largely unknown. Comparing the co-expression structure of each protein complex between tumor and healthy cells may provide insights regarding cancer-specific protein dysfunction. However, the technical limitations of mass spectrometry-based proteomics, including contamination with biological protein variants, causes noise that leads to non-negligible over- (or under-) estimating co-expression. RESULTS: We propose a robust algorithm for identifying protein complex aberrations in cancer based on differential protein co-expression testing. Our method based on a copula is sufficient for improving identification accuracy with noisy data compared to conventional linear correlation-based approaches. As an application, we use large-scale proteomic data from renal cancer to show that important protein complexes, regulatory signaling pathways and drug targets can be identified. The proposed approach surpasses traditional linear correlations to provide insights into higher-order differential co-expression structures. AVAILABILITY AND IMPLEMENTATION: https://github.com/ymatts/RoDiCE. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yusuke Matsui 0002, Yuichi Abe, Kohei Uno, Satoru Miyano |
Bioinform. | 4 |
| 2022 | PredictiveNetwork: predictive gene network estimation with application to gastric cancer drug response-predictive network analysisabstractBACKGROUND: Gene regulatory networks have garnered a large amount of attention to understand disease mechanisms caused by complex molecular network interactions. These networks have been applied to predict specific clinical characteristics, e.g., cancer, pathogenicity, and anti-cancer drug sensitivity. However, in most previous studies using network-based prediction, the gene networks were estimated first, and predicted clinical characteristics based on pre-estimated networks. Thus, the estimated networks cannot describe clinical characteristic-specific gene regulatory systems. Furthermore, existing computational methods were developed from algorithmic and mathematics viewpoints, without considering network biology. RESULTS: To effectively predict clinical characteristics and estimate gene networks that provide critical insights into understanding the biological mechanisms involved in a clinical characteristic, we propose a novel strategy for predictive gene network estimation. The proposed strategy simultaneously performs gene network estimation and prediction of the clinical characteristic. In this strategy, the gene network is estimated with minimal network estimation and prediction errors. We incorporate network biology by assuming that neighboring genes in a network have similar biological functions, while hub genes play key roles in biological processes. Thus, the proposed method provides interpretable prediction results and enables us to uncover biologically reliable marker identification. Monte Carlo simulations shows the effectiveness of our method for feature selection in gene estimation and prediction with excellent prediction accuracy. We applied the proposed strategy to construct gastric cancer drug-responsive networks. CONCLUSION: We identified gastric drug response predictive markers and drug sensitivity/resistance-specific markers, AKR1B10, AKR1C3, ANXA10, and ZNF165, based on GDSC data analysis. Our results for identifying drug sensitive and resistant specific molecular interplay are strongly supported by previous studies. We expect that the proposed strategy will be a useful tool for uncovering crucial molecular interactions involved a specific biological mechanism, such as cancer progression or acquired drug resistance. Heewon Park, Seiya Imoto, Satoru Miyano |
BMC Bioinform. | 3 |
| 2021 | On the application of BERT models for nanopore methylation detectionabstractDNA methylation is a common nucleotide modification, which is associated with various biological processes, such as gene expression and aging. Nanopore sequencing provides a direct detecting approach through searching specific current signal shifts. Recently, model-based approaches, especially those using deep learning models, have achieved significant performance improvements on nanopore methylation detection. In this work, we explore using the non-recurrent neural network structure of Bidirectional Encoder Representations from Transformers (BERT) for the task, which provides an alternative fast inference model to the state-of-the-art bi-directional Recurrent Neural Network (biRNN). In addition, we propose a refined BERT model with relative position representation and center hidden units concatenation, which takes account of the task-specific characters into modeling. We evaluate the proposed models on the R9 benchmark datasets of different motifs and methyltransferases. The experiment results show that the refined BERT model can achieve competitive or even better results than the state-of-the-art biRNN model, while the model inference speed is faster. Kiyoshi Yamaguchi, Sera Hatakeyama, Yoichi Furukawa, Satoru Miyano, Rui Yamaguchi, Seiya Imoto |
BIBM | 5 |
| 2021 | Enhancing breakpoint resolution with deep segmentation model: A general refinement method for read-depth based structural variant callersabstractRead-depths (RDs) are frequently used in identifying structural variants (SVs) from sequencing data. For existing RD-based SV callers, it is difficult for them to determine breakpoints in single-nucleotide resolution due to the noisiness of RD data and the bin-based calculation. In this paper, we propose to use the deep segmentation model UNet to learn base-wise RD patterns surrounding breakpoints of known SVs. We integrate model predictions with an RD-based SV caller to enhance breakpoints in single-nucleotide resolution. We show that UNet can be trained with a small amount of data and can be applied both in-sample and cross-sample. An enhancement pipeline named RDBKE significantly increases the number of SVs with more precise breakpoints on simulated and real data. The source code of RDBKE is freely available at https://github.com/yaozhong/deepIntraSV. Seiya Imoto, Satoru Miyano, Rui Yamaguchi |
PLoS Comput. Biol. | 3 |
| 2020 | Neoantimon: a multifunctional R package for identification of tumor-specific neoantigensabstractSUMMARY: It is known that some mutant peptides, such as those resulting from missense mutations and frameshift insertions, can bind to the major histocompatibility complex and be presented to antitumor T cells on the surface of a tumor cell. These peptides are termed neoantigen, and it is important to understand this process for cancer immunotherapy. Here, we introduce an R package termed Neoantimon that can predict a list of potential neoantigens from a variety of mutations, which include not only somatic point mutations but insertions, deletions and structural variants. Beyond the existing applications, Neoantimon is capable of attaching and reflecting several additional information, e.g. wild-type binding capability, allele specific RNA expression levels, single nucleotide polymorphism information and combinations of mutations to filter out infeasible peptides as neoantigen. AVAILABILITY AND IMPLEMENTATION: The R package is available at http://github/hase62/Neoantimon. Takanori Hasegawa, Shuto Hayashi, Eigo Shimizu, Shinichi Mizuno, Atsushi Niida, Rui Yamaguchi, Satoru Miyano, Hidewaki Nakagawa, Seiya Imoto |
Bioinform. | 7 |
| 2020 | Nanopore basecalling from a perspective of instance segmentationabstractBACKGROUND: Nanopore sequencing is a rapidly developing third-generation sequencing technology, which can generate long nucleotide reads of molecules within a portable device in real-time. Through detecting the change of ion currency signals during a DNA/RNA fragment's pass through a nanopore, genotypes are determined. Currently, the accuracy of nanopore basecalling has a higher error rate than the basecalling of short-read sequencing. Through utilizing deep neural networks, the-state-of-the art nanopore basecallers achieve basecalling accuracy in a range from 85% to 95%. RESULT: In this work, we proposed a novel basecalling approach from a perspective of instance segmentation. Different from previous approaches of doing typical sequence labeling, we formulated the basecalling problem as a multi-label segmentation task. Meanwhile, we proposed a refined U-net model which we call UR-net that can model sequential dependencies for a one-dimensional segmentation task. The experiment results show that the proposed basecaller URnano achieves competitive results on the in-species data, compared to the recently proposed CTC-featured basecallers. CONCLUSION: Our results show that formulating the basecalling problem as a one-dimensional segmentation task is a promising approach, which does basecalling and segmentation jointly. Arda Akdemir, Georg Tremmel, Seiya Imoto, Satoru Miyano, Tetsuo Shibuya, Rui Yamaguchi |
BMC Bioinform. | 5 |
| 2019 | A Bayesian model integration for mutation calling through data partitioningabstractMOTIVATION: Detection of somatic mutations from tumor and matched normal sequencing data has become among the most important analysis methods in cancer research. Some existing mutation callers have focused on additional information, e.g. heterozygous single-nucleotide polymorphisms (SNPs) nearby mutation candidates or overlapping paired-end read information. However, existing methods cannot take multiple information sources into account simultaneously. Existing Bayesian hierarchical model-based methods construct two generative models, the tumor model and error model, and limited information sources have been modeled. RESULTS: We proposed a Bayesian model integration framework named as partitioning-based model integration. In this framework, through introducing partitions for paired-end reads based on given information sources, we integrate existing generative models and utilize multiple information sources. Based on that, we constructed a novel Bayesian hierarchical model-based method named as OHVarfinDer. In both the tumor model and error model, we introduced partitions for a set of paired-end reads that cover a mutation candidate position, and applied a different generative model for each category of paired-end reads. We demonstrated that our method can utilize both heterozygous SNP information and overlapping paired-end read information effectively in simulation datasets and real datasets. AVAILABILITY AND IMPLEMENTATION: https://github.com/takumorizo/OHVarfinDer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Takuya Moriyama, Seiya Imoto, Shuto Hayashi, Yuichi Shiraishi, Satoru Miyano, Rui Yamaguchi |
Bioinform. | 5 |
| 2019 | Virtual Grid Engine: a simulated grid engine environment for large-scale supercomputersabstractBACKGROUND: Supercomputers have become indispensable infrastructures in science and industries. In particular, most state-of-the-art scientific results utilize massively parallel supercomputers ranked in TOP500. However, their use is still limited in the bioinformatics field due to the fundamental fact that the asynchronous parallel processing service of Grid Engine is not provided on them. To encourage the use of massively parallel supercomputers in bioinformatics, we developed middleware called Virtual Grid Engine, which enables software pipelines to automatically perform their tasks as MPI programs. RESULT: We conducted basic tests to check the time required to assign jobs to workers by VGE. The results showed that the overhead of the employed algorithm was 246 microseconds and our software can manage thousands of jobs smoothly on the K computer. We also tried a practical test in the bioinformatics field. This test included two tasks, the split and BWA alignment of input FASTQ data. 25,055 nodes (2,000,440 cores) were used for this calculation and accomplished it in three hours. CONCLUSION: We considered that there were four important requirements for this kind of software, non-privilege server program, multiple job handling, dependency control, and usability. We carefully designed and checked all requirements. And this software fulfilled all the requirements and achieved good performance in a large scale analysis. Masaaki Yadome, Tatsuo Nishiki, Shigeru Ishiduki, Hikaru Inoue, Rui Yamaguchi, Satoru Miyano |
BMC Bioinform. | 7 |
| 2018 | Virtual Grid Engine: Accelerating thousands of omics sample analyses using large-scale supercomputers
Masaaki Yadome, Tatsuo Nishiki, Shigeru Ishiduki, Hikaru Inoue, Rui Yamaguchi, Satoru Miyano |
BIBM | 7 |
| 2017 | Reconstruction of high read-depth signals from low-depth whole genome sequencing data using deep learningabstractMotivation: Next-generation sequencing (NGS) technologies using DNA, RNA, or methylation sequencing are prevailing tools used in modern genome research. For DNA sequencing, whole genome sequencing (WGS) and whole exome sequencing (WES) are two typical applications with a different preference on the trade-off between sequencing depth and base coverage. Although sequencing costs have been greatly reduced, the sequence depth used in WGS is relatively lower than WES (e.g., ~35× vs. 100×~). In addition, biases and batch effects may exist in different stages of a NGS experiment. Using low-depth and biased WGS data for downstream analyses is more sensitive to the bias problem and makes it even more difficult to uncover real biological signals in the data. In this work, we focused on reconstructing high read-depth signals from low-depth WGS data. We make use of a pair of WGS data with different read-depth for the same sample and learn a mapping from low-depth signals to high-depth in the given platform. Results: We explored three different reconstruction models from shallow to deep. Our experimental results show that by only using the read depth information, deeper models do not perform far better than a linear regression model. Through incorporating additional information, such as GC-content, mappability and nucleotide sequence information, the performance of convolutional neural network (CNN) models can be further improved. We made use of the reconstructed read-depth signals in downstream analysis to identify copy number variation segments for single sample. The experiment results show that segments that are not detected using low-depth data, can be detected with the reconstructed signals by the CNN model using extra biological information. Seiya Imoto, Satoru Miyano, Rui Yamaguchi |
BIBM | 3 |
| 2017 | phyC: Clustering cancer evolutionary treesabstractMulti-regional sequencing provides new opportunities to investigate genetic heterogeneity within or between common tumors from an evolutionary perspective. Several state-of-the-art methods have been proposed for reconstructing cancer evolutionary trees based on multi-regional sequencing data to develop models of cancer evolution. However, there have been few studies on comparisons of a set of cancer evolutionary trees. We propose a clustering method (phyC) for cancer evolutionary trees, in which sub-groups of the trees are identified based on topology and edge length attributes. For interpretation, we also propose a method for evaluating the sub-clonal diversity of trees in the clusters, which provides insight into the acceleration of sub-clonal expansion. Simulation showed that the proposed method can detect true clusters with sufficient accuracy. Application of the method to actual multi-regional sequencing data of clear cell renal carcinoma and non-small cell lung cancer allowed for the detection of clusters related to cancer type or phenotype. phyC is implemented with R(≥3.2.2) and is available from https://github.com/ymatts/phyC. Yusuke Matsui 0002, Atsushi Niida, Ryutaro Uchi, Koshi Mimori, Satoru Miyano, Teppei Shimamura |
PLoS Comput. Biol. | 5 |
| 2017 | A Novel Adaptive Penalized Logistic Regression for Uncovering Biomarker Associated with Anti-Cancer Drug SensitivityabstractWe propose a novel adaptive penalized logistic regression modeling strategy based on Wilcoxon rank sum test (WRST) to effectively uncover driver genes in classification. In order to incorporate significance of gene in classification, we first measure significance of each gene by gene ranking method based on WRST, and then the adaptive L1-type penalty is discriminately imposed on each gene depending on the measured importance degree of gene. The incorporating significance of genes into adaptive logistic regression enables us to impose a large amount of penalty on low ranking genes, and thus noise genes are easily deleted from the model and we can effectively identify driver genes. Monte Carlo experiments and real world example are conducted to investigate effectiveness of the proposed approach. In Sanger data analysis, we introduce a strategy to identify expression modules indicating gene regulatory mechanisms via the principal component analysis (PCA), and perform logistic regression modeling based on not a single gene but gene expression modules. We can see through Monte Carlo experiments and real world example that the proposed adaptive penalized logistic regression outperforms feature selection and classification compared with existing L1-type regularization. The discriminately imposed penalty based on WRST effectively performs crucial gene selection, and thus our method can improve classification accuracy without interruption of noise genes. Furthermore, it can be seen through Sanger data analysis that the method for gene expression modules based on principal components and their loading scores provides interpretable results in biological viewpoints. Heewon Park, Yuichi Shiraishi, Seiya Imoto, Satoru Miyano |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2016 | OVarCall: Bayesian Mutation Calling Method Utilizing Overlapping Paired-End Reads
Takuya Moriyama, Yuichi Shiraishi, Kenichi Chiba, Rui Yamaguchi, Seiya Imoto, Satoru Miyano |
ISBRA | 6 |
| 2016 | D3M: detection of differential distributions of methylation levelsabstractMOTIVATION: DNA methylation is an important epigenetic modification related to a variety of diseases including cancers. We focus on the methylation data from Illumina's Infinium HumanMethylation450 BeadChip. One of the key issues of methylation analysis is to detect the differential methylation sites between case and control groups. Previous approaches describe data with simple summary statistics or kernel function, and then use statistical tests to determine the difference. However, a summary statistics-based approach cannot capture complicated underlying structure, and a kernel function-based approach lacks interpretability of results. RESULTS: We propose a novel method D(3)M, for detection of differential distribution of methylation, based on distribution-valued data. Our method can detect the differences in high-order moments, such as shapes of underlying distributions in methylation profiles, based on the Wasserstein metric. We test the significance of the difference between case and control groups and provide an interpretable summary of the results. The simulation results show that the proposed method achieves promising accuracy and shows favorable results compared with previous methods. Glioblastoma multiforme and lower grade glioma data from The Cancer Genome Atlas show that our method supports recent biological advances and suggests new insights. AVAILABILITY AND IMPLEMENTATION: R implemented code is freely available from https://github.com/ymatts/D3M/ CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yusuke Matsui 0002, Masahiro Mizuta, Satoru Miyano, Teppei Shimamura |
Bioinform. | 4 |
| 2015 | High performance computing of a fusion gene detection pipeline on the K computerabstractRecently developed high-throughput sequencers can generate a huge amount of omics data, and TOP500-class supercomputers are required to analyze such large datasets. However, these supercomputers do not usually support grid engines, which are commonly used in bioinformatics, making it necessary to parallelize the software. Parallelization and optimization require domain and specific knowledge, which poses a challenge for most bioinformaticians. We here propose a simple methodology for the parallelization of pipeline software. To demonstrate the efficacy of the methodology, we employed the Genomon-fusion as a sample software pipeline and ported it onto the K computer by applying our method. Simultaneous analysis of a massive amount of samples was performed using the K computer in a very short period of time. Yuichi Shiraishi, Teppei Shimamura, Kenichi Chiba, Satoru Miyano |
BIBM | 5 |
| 2015 | Binary Contingency Table Method for Analyzing Gene Mutation in Cancer Genome
Emi Ayada, Atsushi Niida, Takanori Hasegawa, Satoru Miyano, Seiya Imoto |
ISBRA | 4 |
| 2015 | Genomon ITDetector: a tool for somatic internal tandem duplication detection from cancer genome sequencing dataabstractSUMMARY: Somatic internal tandem duplications (ITDs) are known to play important roles in cancer pathogenesis. Although recent advances in high-throughput sequencing technologies have enabled genome-wide detection of various types of genomic mutations, including single nucleotide variants, indels and structural variations, only a few studies have focused on ITDs. We have developed an analytical tool called 'Genomon ITDetector' for genome-wide detection of somatic ITDs. After evaluating the sensitivity and precision of the proposed approach using synthetic data, we have demonstrated that it can successfully detect not only common ITDs involving FLT3, but also a number of ITDs affecting other putative driver genes in acute myeloid leukemia exome sequencing data. Availability and implementaion: Genomon ITDetector is freely available at https://github.com/ken0-1n/Genomon-ITDetector. Kenichi Chiba, Yuichi Shiraishi, Yasunobu Nagata, Seiya Imoto, Seishi Ogawa, Satoru Miyano |
Bioinform. | 7 |
| 2014 | Parameter estimation in multi-compartment SIR model
Masaya M. Saito, Seiya Imoto, Rui Yamaguchi, Satoru Miyano, Tomoyuki Higuchi |
FUSION | 4 |
| 2014 | HapMuC: somatic mutation calling using heterozygous germ line variants near candidate mutationsabstractMOTIVATION: Identifying somatic changes from tumor and matched normal sequences has become a standard approach in cancer research. More specifically, this requires accurate detection of somatic point mutations with low allele frequencies in impure and heterogeneous cancer samples. Although haplotype phasing information derived by using heterozygous germ line variants near candidate mutations would improve accuracy, no somatic mutation caller that uses such information is currently available. RESULTS: We propose a Bayesian hierarchical method, termed HapMuC, in which power is increased by using available information on heterozygous germ line variants located near candidate mutations. We first constructed two generative models (the mutation model and the error model). In the generative models, we prepared candidate haplotypes, considering a heterozygous germ line variant if available, and the observed reads were realigned to the haplotypes. We then inferred the haplotype frequencies and computed the marginal likelihoods using a variational Bayesian algorithm. Finally, we derived a Bayes factor for evaluating the possibility of the existence of somatic mutations. We also demonstrated that our algorithm has superior specificity and sensitivity compared with existing methods, as determined based on a simulation, the TCGA Mutation Calling Benchmark 4 datasets and data from the COLO-829 cell line. AVAILABILITY AND IMPLEMENTATION: The HapMuC source code is available from http://github.com/usuyama/hapmuc. Naoto Usuyama, Yuichi Shiraishi, Yusuke Sato, Haruki Kume, Yukio Homma, Seishi Ogawa, Satoru Miyano, Seiya Imoto |
Bioinform. | 7 |
| 2014 | A feature selection method using improved regularized linear discriminant analysis
Alok Sharma, Kuldip K. Paliwal, Seiya Imoto, Satoru Miyano |
Mach. Vis. Appl. | 4 |
| 2013 | Estimation of abrupt changes in sentinel observation data of influenza epidemics in Japan
Masaya M. Saito, Seiya Imoto, Rui Yamaguchi, Satoru Miyano, Tomoyuki Higuchi |
FUSION | 4 |
| 2013 | XiP: a computational environment to create, extend and share workflowsabstractUNLABELLED: XiP (eXtensible integrative Pipeline) is a flexible, editable and modular environment with a user-friendly interface that does not require previous advanced programming skills to run, construct and edit workflows. XiP allows the construction of workflows by linking components written in both R and Java, the analysis of high-throughput data in grid engine systems and also the development of customized pipelines that can be encapsulated in a package and distributed. XiP already comes with several ready-to-use pipeline flows for the most common genomic and transcriptomic analysis and ∼300 computational components. AVAILABILITY: XiP is open source, freely available under the Lesser General Public License (LGPL) and can be downloaded from http://xip.hgc.jp. Masao Nagasaki, André Fujita, Yayoi Sekiya, Ayumu Saito, Emi Ikeda, Chen Li 0007, Satoru Miyano |
Bioinform. | 7 |
| 2013 | A strategy to select suitable physicochemical attributes of amino acids for protein fold recognitionabstractBACKGROUND: Assigning a protein into one of its folds is a transitional step for discovering three dimensional protein structure, which is a challenging task in bimolecular (biological) science. The present research focuses on: 1) the development of classifiers, and 2) the development of feature extraction techniques based on syntactic and/or physicochemical properties. RESULTS: Apart from the above two main categories of research, we have shown that the selection of physicochemical attributes of the amino acids is an important step in protein fold recognition and has not been explored adequately. We have presented a multi-dimensional successive feature selection (MD-SFS) approach to systematically select attributes. The proposed method is applied on protein sequence data and an improvement of around 24% in fold recognition has been noted when selecting attributes appropriately. CONCLUSION: The MD-SFS has been applied successfully in selecting physicochemical attributes of the amino acids. The selected attributes show improved protein fold recognition performance. Alok Sharma, Kuldip K. Paliwal, Abdollah Dehzangi, James G. Lyons, Seiya Imoto, Satoru Miyano |
BMC Bioinform. | 6 |
| 2012 | Identifiability of local transmissibility parameters in agent-based pandemic simulation
Masaya M. Saito, Seiya Imoto, Rui Yamaguchi, Satoru Miyano, Tomoyuki Higuchi |
FUSION | 4 |
| 2012 | IRView: a database and viewer for protein interacting regionsabstractUNLABELLED: Protein-protein interactions (PPIs) are mediated through specific regions on proteins. Some proteins have two or more protein interacting regions (IRs) and some IRs are competitively used for interactions with different proteins. IRView currently contains data for 3417 IRs in human and mouse proteins. The data were obtained from different sources and combined with annotated region data from InterPro. Information on non-synonymous single nucleotide polymorphism sites and variable regions owing to alternative mRNA splicing is also included. The IRView web interface displays all IR data, including user-uploaded data, on reference sequences so that the positional relationship between IRs can be easily understood. IRView should be useful for analyzing underlying relationships between the proteins behind the PPI networks. AVAILABILITY: IRView is publicly available on the web at http://ir.hgc.jp/ Shigeo Fujimori, Naoya Hirai, Kazuyo Masuoka, Tomohiro Oshikubo, Tatsuhiro Yamashita, Takanori Washio, Ayumu Saito, Masao Nagasaki, Satoru Miyano, Etsuko Miyamoto-Sato |
Bioinform. | 9 |
| 2012 | Statistical model-based testing to evaluate the recurrence of genomic aberrationsabstractMOTIVATION: In cancer genomes, chromosomal regions harboring cancer genes are often subjected to genomic aberrations like copy number alteration and loss of heterozygosity. Given this, finding recurrent genomic aberrations is considered an apt approach for screening cancer genes. Although several permutation-based tests have been proposed for this purpose, none of them are designed to find recurrent aberrations from the genomic dataset without paired normal sample controls. Their application to unpaired genomic data may lead to false discoveries, because they retrieve pseudo-aberrations that exist in normal genomes as polymorphisms. RESULTS: We develop a new parametric method named parametric aberration recurrence test (PART) to test for the recurrence of genomic aberrations. The introduction of Poisson-binomial statistics allow us to compute small P-values more efficiently and precisely than the previously proposed permutation-based approach. Moreover, we extended PART to cover unpaired data (PART-up) so that there is a statistical basis for analyzing unpaired genomic data. PART-up uses information from unpaired normal sample controls to remove pseudo-aberrations in unpaired genomic data. Using PART-up, we successfully predict recurrent genomic aberrations in cancer cell line samples whose paired normal sample controls are unavailable. This article thus proposes a powerful statistical framework for the identification of driver aberrations, which would be applicable to ever-increasing amounts of cancer genomic data seen in the era of next generation sequencing. AVAILABILITY: Our implementations of PART and PART-up are available from http://www.hgc.jp/~niiyan/PART/manual.html. Atsushi Niida, Seiya Imoto, Teppei Shimamura, Satoru Miyano |
Bioinform. | 4 |
| 2012 | ChopSticks: High-resolution analysis of homozygous deletions by exploiting concordant read pairsabstractBACKGROUND: Structural variations (SVs) in genomes are commonly observed even in healthy individuals and play key roles in biological functions. To understand their functional impact or to infer molecular mechanisms of SVs, they have to be characterized with the maximum resolution. However, high-resolution analysis is a difficult task because it requires investigation of the complex structures involved in an enormous number of alignments of next-generation sequencing (NGS) reads and genome sequences that contain errors. RESULTS: We propose a new method called ChopSticks that improves the resolution of SV detection for homozygous deletions even when the depth of coverage is low. Conventional methods based on read pairs use only discordant pairs to localize the positions of deletions, where a discordant pair is a read pair whose alignment has an aberrant strand or distance. In contrast, our method exploits concordant reads as well. We theoretically proved that when the depth of coverage approaches zero or infinity, the expected resolution of our method is asymptotically equal to that of methods based only on discordant pairs under double coverage. To confirm the effectiveness of ChopSticks, we conducted computational experiments against both simulated NGS reads and real NGS sequences. The resolution of deletion calls by other methods was significantly improved, thus demonstrating the usefulness of ChopSticks. CONCLUSIONS: ChopSticks can generate high-resolution deletion calls of homozygous deletions using information independent of other methods, and it is therefore useful to examine the functional impact of SVs or to infer SV generation mechanisms. Tomohiro Yasuda, Shin Suzuki, Masao Nagasaki, Satoru Miyano |
BMC Bioinform. | 4 |
| 2012 | Identifying Gene Pathways Associated with Cancer Characteristics via Sparse Statistical MethodsabstractWe propose a statistical method for uncovering gene pathways that characterize cancer heterogeneity. To incorporate knowledge of the pathways into the model, we define a set of activities of pathways from microarray gene expression data based on the Sparse Probabilistic Principal Component Analysis (SPPCA). A pathway activity logistic regression model is then formulated for cancer phenotype. To select pathway activities related to binary cancer phenotypes, we use the elastic net for the parameter estimation and derive a model selection criterion for selecting tuning parameters included in the model estimation. Our proposed method can also reverse-engineer gene networks based on the identified multiple pathways that enables us to discover novel gene-gene associations relating with the cancer phenotypes. We illustrate the whole process of the proposed method through the analysis of breast cancer gene expression data. Shuichi Kawano, Teppei Shimamura, Atsushi Niida, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, Ryo Yoshida, Cristin G. Print, Satoru Miyano |
IEEE ACM Trans. Comput. Biol. Bioinform. | 9 |
| 2012 | A Top-r Feature Selection Algorithm for Microarray Gene Expression DataabstractMost of the conventional feature selection algorithms have a drawback whereby a weakly ranked gene that could perform well in terms of classification accuracy with an appropriate subset of genes will be left out of the selection. Considering this shortcoming, we propose a feature selection algorithm in gene expression data analysis of sample classifications. The proposed algorithm first divides genes into subsets, the sizes of which are relatively small (roughly of size h), then selects informative smaller subsets of genes (of size r < h) from a subset and merges the chosen genes with another gene subset (of size r) to update the gene subset. We repeat this process until all subsets are merged into one informative subset. We illustrate the effectiveness of the proposed algorithm by analyzing three distinct gene expression data sets. Our method shows promising classification accuracy for all the test data sets. We also show the relevance of the selected genes in terms of their biological functions. Alok Sharma, Seiya Imoto, Satoru Miyano |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2011 | Estimation of macroscopic parameter in agent-based pandemic simulation
Masaya M. Saito, Seiya Imoto, Rui Yamaguchi, Satoru Miyano, Tomoyuki Higuchi |
FUSION | 4 |
| 2011 | Comprehensive Pharmacogenomic Pathway Screening by Data Assimilation
Takanori Hasegawa, Rui Yamaguchi, Masao Nagasaki, Seiya Imoto, Satoru Miyano |
ISBRA | 5 |
| 2011 | CSO validator: improving manual curation workflow for biological pathwaysabstractSUMMARY: Manual curation and validation of large-scale biological pathways are required to obtain high-quality pathway databases. In a typical curation process, model validation and model update based on appropriate feedback are repeated and requires considerable cooperation of scientists. We have developed a CSO (Cell System Ontology) validator to reduce the repetition and time during the curation process. This tool assists in quickly obtaining agreement among curators and domain experts and in providing a consistent and accurate pathway database. AVAILABILITY: The tool is available on http://csovalidator.csml.org. CONTACT: [email protected]. Euna Jeong, Masao Nagasaki, Emi Ikeda, Yayoi Sekiya, Ayumu Saito, Satoru Miyano |
Bioinform. | 6 |
| 2011 | MIRACH: efficient model checker for quantitative biological pathway modelsabstractUNLABELLED: Model checking is playing an increasingly important role in systems biology as larger and more complex biological pathways are being modeled. In this article we report the release of an efficient model checker MIRACH 1.0, which supports any model written in popular formats such as CSML and SBML. MIRACH is integrated with a Petri-net-based simulation engine, enabling efficient online (on-the-fly) checking. In our experiment, by using Levchenko et al. model, we reveal that timesaving gains by using MIRACH easily surpass 400% compared with its offline-based counterpart. AVAILABILITY AND IMPLEMENTATION: MIRACH 1.0 was developed using Java and thus executable on any platform installed with JDK 6.0 (not JRE 6.0) or later. MIRACH 1.0, along with its source codes, documentation and examples are available at http://sourceforge.net/projects/mirach/ under the LGPLv3 license. Chuan Hock Koh, Masao Nagasaki, Ayumu Saito, Chen Li 0007, Limsoon Wong, Satoru Miyano |
Bioinform. | 6 |
| 2011 | Systems biology model repository for macrophage pathway simulationabstractSUMMARY: The Macrophage Pathway Knowledgebase (MACPAK) is a computational system that allows biomedical researchers to query and study the dynamic behaviors of macrophage molecular pathways. It integrates the knowledge of 230 reviews that were carefully checked by specialists for their accuracy and then converted to 230 dynamic mathematical pathway models. MACPAK comprises a total of 24 009 entities and 12 774 processes and is described in the Cell System Markup Language (CSML), an XML format that runs on the Cell Illustrator platform and can be visualized with a customized Cytoscape for further analysis. AVAILABILITY: MACPAK can be accessed via an interactive web site at http://macpak.csml.org. The CSML pathway models are available under the Creative Commons license. Masao Nagasaki, Ayumu Saito, André Fujita, Georg Tremmel, Kazuko Ueno, Emi Ikeda, Euna Jeong, Satoru Miyano |
Bioinform. | 8 |
| 2011 | A rank-based statistical test for measuring synergistic effects between two gene setsabstractMOTIVATION: Due to recent advances in high-throughput technologies, data on various types of genomic annotation have accumulated. These data will be crucially helpful for elucidating the combinatorial logic of transcription. Although several approaches have been proposed for inferring cooperativity among multiple factors, most approaches are haunted by the issues of normalization and threshold values. RESULTS: In this article, we propose a rank-based non-parametric statistical test for measuring the effects between two gene sets. This method is free from the issues of normalization and threshold value determination for gene expression values. Furthermore, we have proposed an efficient Markov chain Monte Carlo method for calculating an approximate significance value of synergy. We have applied this approach for detecting synergistic combinations of transcription factor binding motifs and histone modifications. AVAILABILITY: C implementation of the method is available from http://www.hgc.jp/~yshira/software/rankSynergy.zip. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuichi Shiraishi, Mariko Okada, Satoru Miyano |
Bioinform. | 3 |
| 2011 | SiGN-SSM: open source parallel software for estimating gene networks with state space modelsabstractUNLABELLED: SiGN-SSM is an open-source gene network estimation software able to run in parallel on PCs and massively parallel supercomputers. The software estimates a state space model (SSM), that is a statistical dynamic model suitable for analyzing short time and/or replicated time series gene expression profiles. SiGN-SSM implements a novel parameter constraint effective to stabilize the estimated models. Also, by using a supercomputer, it is able to determine the gene network structure by a statistical permutation test in a practical time. SiGN-SSM is applicable not only to analyzing temporal regulatory dependencies between genes, but also to extracting the differentially regulated genes from time series expression profiles. AVAILABILITY: SiGN-SSM is distributed under GNU Affero General Public Licence (GNU AGPL) version 3 and can be downloaded at http://sign.hgc.jp/signssm/. The pre-compiled binaries for some architectures are available in addition to the source code. The pre-installed binaries are also available on the Human Genome Center supercomputer system. The online manual and the supplementary information of SiGN-SSM is available on our web site. CONTACT: [email protected]. Yoshinori Tamada, Rui Yamaguchi, Seiya Imoto, Osamu Hirose, Ryo Yoshida, Masao Nagasaki, Satoru Miyano |
Bioinform. | 7 |
| 2011 | Ontology-based instance data validation for high-quality curated biological pathwaysabstractBACKGROUND: Modeling in systems biology is vital for understanding the complexity of biological systems across scales and predicting system-level behaviors. To obtain high-quality pathway databases, it is essential to improve the efficiency of model validation and model update based on appropriate feedback. RESULTS: We have developed a new method to guide creating novel high-quality biological pathways, using a rule-based validation. Rules are defined to correct models against biological semantics and improve models for dynamic simulation. In this work, we have defined 40 rules which constrain event-specific participants and the related features and adding missing processes based on biological events. This approach is applied to data in Cell System Ontology which is a comprehensive ontology that represents complex biological pathways with dynamics and visualization. The experimental results show that the relatively simple rules can efficiently detect errors made during curation, such as misassignment and misuse of ontology concepts and terms in curated models. CONCLUSIONS: A new rule-based approach has been developed to facilitate model validation and model complementation. Our rule-based validation embedding biological semantics enables us to provide high-quality curated biological pathways. This approach can serve as a preprocessing step for model integration, exchange and extraction data, and simulation. Euna Jeong, Masao Nagasaki, Kazuko Ueno, Satoru Miyano |
BMC Bioinform. | 4 |
| 2011 | ClipCrop: a tool for detecting structural variations with single-base resolution using soft-clipping informationabstractBACKGROUND: Structural variations (SVs) change the structure of the genome and are therefore the causes of various diseases. Next-generation sequencing allows us to obtain a multitude of sequence data, some of which can be used to infer the position of SVs. METHODS: We developed a new method and implementation named ClipCrop for detecting SVs with single-base resolution using soft-clipping information. A soft-clipped sequence is an unmatched fragment in a partially mapped read. To assess the performance of ClipCrop with other SV-detecting tools, we generated various patterns of simulation data - SV lengths, read lengths, and the depth of coverage of short reads - with insertions, deletions, tandem duplications, inversions and single nucleotide alterations in a human chromosome. For comparison, we selected BreakDancer, CNVnator and Pindel, each of which adopts a different approach to detect SVs, e.g. discordant pair approach, depth of coverage approach and split read approach, respectively. RESULTS: Our method outperformed BreakDancer and CNVnator in both discovering rate and call accuracy in any type of SV. Pindel offered a similar performance as our method, but our method crucially outperformed for detecting small duplications. From our experiments, ClipCrop infer reliable SVs for the data set with more than 50 bases read lengths and 20x depth of coverage, both of which are reasonable values in current NGS data set. CONCLUSIONS: ClipCrop can detect SVs with higher discovering rate and call accuracy than any other tool in our simulation data set. Shin Suzuki, Tomohiro Yasuda, Yuichi Shiraishi, Satoru Miyano, Masao Nagasaki |
BMC Bioinform. | 4 |
| 2011 | Parallel Algorithm for Learning Optimal Bayesian Network Structure
Yoshinori Tamada, Seiya Imoto, Satoru Miyano |
J. Mach. Learn. Res. | 3 |
| 2011 | Hybrid Petri net based modeling for biological pathway simulation
Hiroshi Matsuno, Masao Nagasaki, Satoru Miyano |
Nat. Comput. | 3 |
| 2011 | High Performance Hybrid Functional Petri Net Simulations of Biological Pathway Models on CUDAabstractHybrid functional Petri nets are a wide-spread tool for representing and simulating biological models. Due to their potential of providing virtual drug testing environments, biological simulations have a growing impact on pharmaceutical research. Continuous research advancements in biology and medicine lead to exponentially increasing simulation times, thus raising the demand for performance accelerations by efficient and inexpensive parallel computation solutions. Recent developments in the field of general-purpose computation on graphics processing units (GPGPU) enabled the scientific community to port a variety of compute intensive algorithms onto the graphics processing unit (GPU). This work presents the first scheme for mapping biological hybrid functional Petri net models, which can handle both discrete and continuous entities, onto compute unified device architecture (CUDA) enabled GPUs. GPU accelerated simulations are observed to run up to 18 times faster than sequential implementations. Simulating the cell boundary formation by Delta-Notch signaling on a CUDA enabled GPU results in a speedup of approximately 7x for a model containing 1,600 cells. Georgios Chalkidis, Masao Nagasaki, Satoru Miyano |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2011 | Inferring Contagion in Regulatory NetworksabstractSeveral gene regulatory network models containing concepts of directionality at the edges have been proposed. However, only a few reports have an interpretable definition of directionality. Here, differently from the standard causality concept defined by Pearl, we introduce the concept of contagion in order to infer directionality at the edges, i.e., asymmetries in gene expression dependences of regulatory networks. Moreover, we present a bootstrap algorithm in order to test the contagion concept. This technique was applied in simulated data and, also, in an actual large sample of biological data. Literature review has confirmed some genes identified by contagion as actually belonging to the TP53 pathway. André Fujita, João R. Sato, Marcos Angelo Almeida Demasi, Rui Yamaguchi, Teppei Shimamura, Carlos Eduardo Ferreira, Mari Cleide Sogayar, Satoru Miyano |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2011 | Estimating Genome-Wide Gene Networks Using Nonparametric Bayesian Network Models on Massively Parallel ComputersabstractWe present a novel algorithm to estimate genome-wide gene networks consisting of more than 20,000 genes from gene expression data using nonparametric Bayesian networks. Due to the difficulty of learning Bayesian network structures, existing algorithms cannot be applied to more than a few thousand genes. Our algorithm overcomes this limitation by repeatedly estimating subnetworks in parallel for genes selected by neighbor node sampling. Through numerical simulation, we confirmed that our algorithm outperformed a heuristic algorithm in a shorter time. We applied our algorithm to microarray data from human umbilical vein endothelial cells (HUVECs) treated with siRNAs, to construct a human genome-wide gene network, which we compared to a small gene network estimated for the genes extracted using a traditional bioinformatics method. The results showed that our genome-wide gene network contains many features of the small network, as well as others that could not be captured during the small network estimation. The results also revealed master-regulator genes that are not in the small network but that control many of the genes in the small network. These analyses were impossible to realize without our proposed algorithm. Yoshinori Tamada, Seiya Imoto, Hiromitsu Araki, Masao Nagasaki, Cristin G. Print, Stephen D. Charnock-Jones, Satoru Miyano |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2010 | Identifying Hidden Confounders in Gene Networks by Bayesian NetworksabstractIn the estimation of gene networks from microarray gene expression data, we propose a statistical method for quantification of the hidden confounders in gene networks, which were possibly removed from the set of genes on the gene networks or are novel biological elements that are not measured by microarrays. Due to high computational cost of the structural learning of Bayesian networks and the limited source of the microarray data, it is usual to perform gene selection prior to the estimation of gene networks. Therefore, there exist missing genes that decrease accuracy and interpretability of the estimated gene networks. The proposed method can identify hidden confounders based on the conflicts of the estimated local Bayesian network structures and estimate their ideal profiles based on the proposed Bayesian networks with hidden variables with an EM algorithm. From the estimated ideal profiles, we can identify genes which are missing in the network or suggest the existence of the novel biological elements if the ideal profiles are not significantly correlated with any expression profiles of genes. To the best of our knowledge, this research is the first study to theoretically characterize missing genes in gene networks and practically utilize this information to refine network estimation. Tomoya Higashigaki, Kaname Kojima, Rui Yamaguchi, Masato Inoue, Seiya Imoto, Satoru Miyano |
BIBE | 6 |
| 2010 | Discovering functional gene pathways associated with cancer heterogeneity via sparse supervised learningabstractWe propose a statistical method for uncovering gene pathways that characterize cancer heterogeneity. To incorporate knowledge of the pathways into the model, we define a set of activities of pathways from microarray gene expression data based on the sparse probabilistic principal component analysis. A pathway activity logistic regression model is then formulated for cancer phenotype. To select pathway activities related to binary cancer phenotypes, we use the elastic net for the parameter estimation and derive a model selection criterion for selecting tuning parameters included in the model estimation. Our proposed method can also reverse-engineer gene networks based on the identified multiple pathways that enables us to discover novel gene-gene associations relating with the cancer phenotypes. We illustrate the whole process of the proposed method through the analysis of breast cancer gene expression data. Shuichi Kawano, Teppei Shimamura, Atsushi Niida, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, Ryo Yoshida, Cristin G. Print, Satoru Miyano |
BIBM | 9 |
| 2010 | A fast and robust statistical test based on likelihood ratio with Bartlett correction to identify Granger causality between gene setsabstractAbstract Summary: We propose a likelihood ratio test (LRT) with Bartlett correction in order to identify Granger causality between sets of time series gene expression data. The performance of the proposed test is compared to a previously published bootstrap-based approach. LRT is shown to be significantly faster and statistically powerful even within non-Normal distributions. An R package named gGranger containing an implementation for both Granger causality identification tests is also provided. Availability: http://dnagarden.ims.u-tokyo.ac.jp/afujita/en/doku.php?id=ggranger. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. André Fujita, Kaname Kojima, Alexandre Galvão Patriota, João R. Sato, Patricia Severino, Satoru Miyano |
Bioinform. | 6 |
| 2010 | DA 1.0: parameter estimation of biological pathways using data assimilation approachabstractSUMMARY: Data assimilation (DA) is a computational approach that estimates unknown parameters in a pathway model using time-course information. Particle filtering, the underlying method used, is a well-established statistical method that approximates the joint posterior distributions of parameters by using sequentially generated Monte Carlo samples. In this article, we report the release of Java-based software (DA 1.0) with an intuitive and user-friendly interface to allow users to carry out parameters estimation using DA. AVAILABILITY AND IMPLEMENTATION: DA 1.0 was developed using Java and thus would be executable on any platform installed with JDK 6.0 (not JRE 6.0) or later. DA 1.0 is freely available for academic users and can be launched or downloaded from http://da.csml.org. Chuan Hock Koh, Masao Nagasaki, Ayumu Saito, Limsoon Wong, Satoru Miyano |
Bioinform. | 5 |
| 2010 | Model-free unsupervised gene set screening based on information enrichment in expression profilesabstractMOTIVATION: A number of unsupervised gene set screening methods have recently been developed for search of putative functional gene sets based on their expression profiles. Most of the methods statistically evaluate whether the expression profiles of each gene set are fit to assumed models: e.g. co-expression across all samples or a subgroup of samples. However, it is possible that they fail to capture informative gene sets whose expression profiles are not fit to the assumed models. RESULTS: To overcome this limitation, we propose a model-free unsupervised gene set screening method, Matrix Information Enrichment Analysis (MIEA). Without assuming any specific models, MIEA screens gene sets based on information richness of their expression profiles. We extensively compared the performance of MIEA to those of other unsupervised gene set screening methods, using various types of simulated and real data. The benchmark tests demonstrated that MIEA can detect singular expression profiles that the other methods fail to find, and performs broadly well for various types of input data. Taken together, this study introduces MIEA as a broadly applicable gene set screening tool for mining regulatory programs from transcriptome data. Atsushi Niida, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, André Fujita, Teppei Shimamura, Satoru Miyano |
Bioinform. | 7 |
| 2010 | Inferring dynamic gene networks under varying conditions for transcriptomic network comparisonabstractMOTIVATION: Elucidating the differences between cellular responses to various biological conditions or external stimuli is an important challenge in systems biology. Many approaches have been developed to reverse engineer a cellular system, called gene network, from time series microarray data in order to understand a transcriptomic response under a condition of interest. Comparative topological analysis has also been applied based on the gene networks inferred independently from each of the multiple time series datasets under varying conditions to find critical differences between these networks. However, these comparisons often lead to misleading results, because each network contains considerable noise due to the limited length of the time series. RESULTS: We propose an integrated approach for inferring multiple gene networks from time series expression data under varying conditions. To the best of our knowledge, our approach is the first reverse-engineering method that is intended for transcriptomic network comparison between varying conditions. Furthermore, we propose a state-of-the-art parameter estimation method, relevance-weighted recursive elastic net, for providing higher precision and recall than existing reverse-engineering methods. We analyze experimental data of MCF-7 human breast cancer cells stimulated by epidermal growth factor or heregulin with several doses and provide novel biological hypotheses through network comparison. AVAILABILITY: The software NETCOMP is available at http://bonsai.ims.u-tokyo.ac.jp/ approximately shima/NETCOMP/. Teppei Shimamura, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, Satoru Miyano |
Bioinform. | 5 |
| 2010 | An efficient biological pathway layout algorithm combining grid-layout and spring embedder for complicated cellular location informationabstractBACKGROUND: Graph drawing is one of the important techniques for understanding biological regulations in a cell or among cells at the pathway level. Among many available layout algorithms, the spring embedder algorithm is widely used not only for pathway drawing but also for circuit placement and www visualization and so on because of the harmonized appearance of its results. For pathway drawing, location information is essential for its comprehension. However, complex shapes need to be taken into account when torus-shaped location information such as nuclear inner membrane, nuclear outer membrane, and plasma membrane is considered. Unfortunately, the spring embedder algorithm cannot easily handle such information. In addition, crossings between edges and nodes are usually not considered explicitly. RESULTS: We proposed a new grid-layout algorithm based on the spring embedder algorithm that can handle location information and provide layouts with harmonized appearance. In grid-layout algorithms, the mapping of nodes to grid points that minimizes a cost function is searched. By imposing positional constraints on grid points, location information including complex shapes can be easily considered. Our layout algorithm includes the spring embedder cost as a component of the cost function. We further extend the layout algorithm to enable dynamic update of the positions and sizes of compartments at each step. CONCLUSIONS: The new spring embedder-based grid-layout algorithm and a spring embedder algorithm are applied to three biological pathways; endothelial cell model, Fas-induced apoptosis model, and C. elegans cell fate simulation model. From the positional constraints, all the results of our algorithm satisfy location information, and hence, more comprehensible layouts are obtained as compared to the spring embedder algorithm. From the comparison of the number of crossings, the results of the grid-layout-based algorithm tend to contain more crossings than those of the spring embedder algorithm due to the positional constraints. For a fair comparison, we also apply our proposed method without positional constraints. This comparison shows that these results contain less crossings than those of the spring embedder algorithm. We also compared layouts of the proposed algorithm with and without compartment update and verified that latter can reach better local optima. Kaname Kojima, Masao Nagasaki, Satoru Miyano |
BMC Bioinform. | 3 |
| 2010 | Optimal Search on Clustered Structural Constraint for Learning Bayesian Network Structure
Kaname Kojima, Eric Perrier, Seiya Imoto, Satoru Miyano |
J. Mach. Learn. Res. | 4 |
| 2009 | Computational Predictions for Functional Proteins Working after Cleaved in Apoptotic PathwayabstractProtein cleavage by a caspase enzyme is observed in many pathways including apoptosis pathway, which induces protein modifications such as N-myristoylation. By exhaustive search of NCBI GenBank database based on the information in PeptideCutter website, we have identified a protein very likely to control apoptosis after N-myristoylation which is caused by the cleavage by caspase-8. To find more rules for sequences cleaved by caspase enzymes, we conducted computational experiments using the machine learning system BONSAI, and extracted some characteristic sequences involving serine-threonine kinase motif. This suggests the new possibility that the machine learning technique finds the significant sequence such as the recognition site of the enzyme. Chigusa Miyakawa, Manabu Sugii, Hiroshi Matsuno, Satoru Miyano |
CISIS | 4 |
| 2009 | Better Decomposition Heuristics for the Maximum-Weight Connected Graph Problem Using Betweenness Centrality
Takanori Yamamoto, Hideo Bannai, Masao Nagasaki, Satoru Miyano |
Discovery Science | 4 |
| 2009 | The impact of measurement errors in the identification of regulatory networksabstractBACKGROUND: There are several studies in the literature depicting measurement error in gene expression data and also, several others about regulatory network models. However, only a little fraction describes a combination of measurement error in mathematical regulatory networks and shows how to identify these networks under different rates of noise. RESULTS: This article investigates the effects of measurement error on the estimation of the parameters in regulatory networks. Simulation studies indicate that, in both time series (dependent) and non-time series (independent) data, the measurement error strongly affects the estimated parameters of the regulatory network models, biasing them as predicted by the theory. Moreover, when testing the parameters of the regulatory network models, p-values computed by ignoring the measurement error are not reliable, since the rate of false positives are not controlled under the null hypothesis. In order to overcome these problems, we present an improved version of the Ordinary Least Square estimator in independent (regression models) and dependent (autoregressive models) data when the variables are subject to noises. Moreover, measurement error estimation procedures for microarrays are also described. Simulation results also show that both corrected methods perform better than the standard ones (i.e., ignoring measurement error). The proposed methodologies are illustrated using microarray data from lung cancer patients and mouse liver time series data. CONCLUSIONS: Measurement error dangerously affects the identification of regulatory network models, thus, they must be reduced or taken into account in order to avoid erroneous conclusions. This could be one of the reasons for high biological false positive rates identified in actual regulatory network models. André Fujita, Alexandre Galvão Patriota, João R. Sato, Satoru Miyano |
BMC Bioinform. | 4 |
| 2009 | BFL: a node and edge betweenness based fast layout algorithm for large scale networksabstractBACKGROUND: Network visualization would serve as a useful first step for analysis. However, current graph layout algorithms for biological pathways are insensitive to biologically important information, e.g. subcellular localization, biological node and graph attributes, or/and not available for large scale networks, e.g. more than 10000 elements. RESULTS: To overcome these problems, we propose the use of a biologically important graph metric, betweenness, a measure of network flow. This metric is highly correlated with many biological phenomena such as lethality and clusters. We devise a new fast parallel algorithm calculating betweenness to minimize the preprocessing cost. Using this metric, we also invent a node and edge betweenness based fast layout algorithm (BFL). BFL places the high-betweenness nodes to optimal positions and allows the low-betweenness nodes to reach suboptimal positions. Furthermore, BFL reduces the runtime by combining a sequential insertion algorim with betweenness. For a graph with n nodes, this approach reduces the expected runtime of the algorithm to O(n2) when considering edge crossings, and to O(n log n) when considering only density and edge lengths. CONCLUSION: Our BFL algorithm is compared against fast graph layout algorithms and approaches requiring intensive optimizations. For gene networks, we show that our algorithm is faster than all layout algorithms tested while providing readability on par with intensive optimization algorithms. We achieve a 1.4 second runtime for a graph with 4000 nodes and 12000 edges on a standard desktop computer. Tatsunori B. Hashimoto, Masao Nagasaki, Kaname Kojima, Satoru Miyano |
BMC Bioinform. | 4 |
| 2009 | Network-Based Predictions and Simulations by Biological State Space Models: Search for Drug Mode of Action
Rui Yamaguchi, Seiya Imoto, Satoru Miyano |
J. Comput. Sci. Technol. | 3 |
| 2008 | Preface
Alvis Brazma, Satoru Miyano, Tatsuya Akutsu |
APBC | 2 |
| 2008 | Partial Order-Based Bayesian Network Learning Algorithm for Estimating Gene NetworksabstractFor learning Bayesian network structure from data, order-based algorithms such as K2 algorithm are widely used.In this paper, we consider a problem of constructing the order of nodes in such algorithms based on prior knowledge of gene networks. However, in many cases the prior knowledge is given as partial order of genes and we need to extend the order-based algorithm to partial order-based one. By extending our prior work we propose an efficient partial order-based algorithm for estimating gene networks based on Bayesian networks. The computational complexity of the proposed algorithm is shown. Kazuyuki Numata, Seiya Imoto, Satoru Miyano |
BIBM | 3 |
| 2008 | Statistical inference of transcriptional module-based gene networks from time course gene expression profiles by using state space modelsabstractMOTIVATION: Statistical inference of gene networks by using time-course microarray gene expression profiles is an essential step towards understanding the temporal structure of gene regulatory mechanisms. Unfortunately, most of the current studies have been limited to analysing a small number of genes because the length of time-course gene expression profiles is fairly short. One promising approach to overcome such a limitation is to infer gene networks by exploring the potential transcriptional modules which are sets of genes sharing a common function or involved in the same pathway. RESULTS: In this article, we present a novel approach based on the state space model to identify the transcriptional modules and module-based gene networks simultaneously. The state space model has the potential to infer large-scale gene networks, e.g. of order 10(3), from time-course gene expression profiles. Particularly, we succeeded in the identification of a cell cycle system by using the gene expression profiles of Saccharomyces cerevisiae in which the length of the time-course and number of genes were 24 and 4382, respectively. However, when analysing shorter time-course data, e.g. of length 10 or less, the parameter estimations of the state space model often fail due to overfitting. To extend the applicability of the state space model, we provide an approach to use the technical replicates of gene expression profiles, which are often measured in duplicate or triplicate. The use of technical replicates is important for achieving highly-efficient inferences of gene networks with short time-course data. The potential of the proposed method has been demonstrated through the time-course analysis of the gene expression profiles of human umbilical vein endothelial cells (HUVECs) undergoing growth factor deprivation-induced apoptosis. AVAILABILITY: Supplementary Information and the software (TRANS-MNET) are available at http://daweb.ism.ac.jp/~yoshidar/software/ssm/. Osamu Hirose, Ryo Yoshida, Seiya Imoto, Rui Yamaguchi, Tomoyuki Higuchi, Stephen D. Charnock-Jones, Cristin G. Print, Satoru Miyano |
Bioinform. | 8 |
| 2008 | Fast grid layout algorithm for biological networks with sweep calculationabstractMOTIVATION: Properly drawn biological networks are of great help in the comprehension of their characteristics. The quality of the layouts for retrieved biological networks is critical for pathway databases. However, since it is unrealistic to manually draw biological networks for every retrieval, automatic drawing algorithms are essential. Grid layout algorithms handle various biological properties such as aligning vertices having the same attributes and complicated positional constraints according to their subcellular localizations; thus, they succeed in providing biologically comprehensible layouts. However, existing grid layout algorithms are not suitable for real-time drawing, which is one of requisites for applications to pathway databases, due to their high-computational cost. In addition, they do not consider edge directions and their resulting layouts lack traceability for biochemical reactions and gene regulations, which are the most important features in biological networks. RESULTS: We devise a new calculation method termed sweep calculation and reduce the time complexity of the current grid layout algorithms through its encoding and decoding processes. We conduct practical experiments by using 95 pathway models of various sizes from TRANSPATH and show that our new grid layout algorithm is much faster than existing grid layout algorithms. For the cost function, we introduce a new component that penalizes undesirable edge directions to avoid the lack of traceability in pathways due to the differences in direction between in-edges and out-edges of each vertex. AVAILABILITY: Java implementations of our layout algorithms are available in Cell Illustrator. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kaname Kojima, Masao Nagasaki, Satoru Miyano |
Bioinform. | 3 |
| 2008 | Bayesian learning of biological pathways on genomic data assimilationabstractMOTIVATION: Mathematical modeling and simulation, based on biochemical rate equations, provide us a rigorous tool for unraveling complex mechanisms of biological pathways. To proceed to simulation experiments, it is an essential first step to find effective values of model parameters, which are difficult to measure from in vivo and in vitro experiments. Furthermore, once a set of hypothetical models has been created, any statistical criterion is needed to test the ability of the constructed models and to proceed to model revision. RESULTS: The aim of our research is to present a new statistical technology towards data-driven construction of in silico biological pathways. The method starts with a knowledge-based modeling with hybrid functional Petri net. It then proceeds to the Bayesian learning of model parameters for which experimental data are available. This process exploits quantitative measurements of evolving biochemical reactions, e.g. gene expression data. Another important issue that we consider is statistical evaluation and comparison of the constructed hypothetical pathways. For this purpose, we have developed a new Bayesian information-theoretic measure that assesses the predictability and the biological robustness of in silico pathways. AVAILABILITY: The FORTRAN source codes are available at the URL http://daweb.ism.ac.jpyoshidar/GDA/ SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ryo Yoshida, Masao Nagasaki, Rui Yamaguchi, Seiya Imoto, Satoru Miyano, Tomoyuki Higuchi |
Bioinform. | 5 |
| 2008 | ExonMiner: Web service for analysis of GeneChip Exon array dataabstractBACKGROUND: Some splicing isoform-specific transcriptional regulations are related to disease. Therefore, detection of disease specific splice variations is the first step for finding disease specific transcriptional regulations. Affymetrix Human Exon 1.0 ST Array can measure exon-level expression profiles that are suitable to find differentially expressed exons in genome-wide scale. However, exon array produces massive datasets that are more than we can handle and analyze on personal computer. RESULTS: We have developed ExonMiner that is the first all-in-one web service for analysis of exon array data to detect transcripts that have significantly different splicing patterns in two cells, e.g. normal and cancer cells. ExonMiner can perform the following analyses: (1) data normalization, (2) statistical analysis based on two-way ANOVA, (3) finding transcripts with significantly different splice patterns, (4) efficient visualization based on heatmaps and barplots, and (5) meta-analysis to detect exon level biomarkers. We implemented ExonMiner on a supercomputer system in order to perform genome-wide analysis for more than 300,000 transcripts in exon array data, which has the potential to reveal the aberrant splice variations in cancer cells as exon level biomarkers. CONCLUSION: ExonMiner is well suited for analysis of exon array data and does not require any installation of software except for internet browsers. What all users need to do is to access the ExonMiner URL http://ae.hgc.jp/exonminer. Users can analyze full dataset of exon array data within hours by high-level statistical analysis with sound theoretical basis that finds aberrant splice variants as biomarkers. Kazuyuki Numata, Ryo Yoshida, Masao Nagasaki, Ayumu Saito, Seiya Imoto, Satoru Miyano |
BMC Bioinform. | 6 |
| 2007 | A Structure Learning Algorithm for Inference of Gene Networks from Microarray Gene Expression Data Using Bayesian NetworksabstractEstimation of gene networks based on microarray gene expression data is an important problem in systems biology. In this paper we use Bayesian networks as a mathematical model for reverse-engineering gene networks from microarray data. In such a case, structural learning of Bayesian networks is known as an NP-hard problem and we need to use heuristic algorithms to find better network structures. Recently, several algorithms have been proposed to estimate optimal Bayesian network structure, but the number of genes included in the network is limited less than 30 or so. In order to apply Bayesian network approach to drug target gene discovery, we need to consider gene networks with several hundreds of genes. Therefore we need to develop more efficient algorithms to learn Bayesian network structure based on observed data. In this paper we propose an efficient structural learning algorithm for Bayesian networks by extending K2 algorithm that is one of the standard learning algorithms in Bayesian networks. We conduct Monte Carlo simulations to examine the effectiveness of the proposed algorithm by comparing with greedy hill-climbing algorithm. We also show the application of yeast gene network estimation based on the proposed algorithm. Kazuyuki Numata, Seiya Imoto, Satoru Miyano |
BIBE | 3 |
| 2007 | Computational Genome-Wide Discovery of Aberrant Splice Variations with Exon Expression ProfilesabstractAlternative splicing plays a prominent role in eukaryotic gene regulations that allow a single gene to generate the multiple mRNA products. The recent advent of GeneChipregHuman Exon 1.0 ST Array enables us to measure the exon expression profiles of human cells on a genome-wide scale. With this advent, analysis of functional gene regulation could be extended to detect not only differentially expressed genes, but also specific splicing events that occur in target cells, but not in normal controls. We address some statistical issues for the identification of biomarker splice variations with exon expression data. The proposed method involves the following steps: (1) Whole transcript analysis with the nonparametric analysis of variance (ANOVA) to identify potential biomarkers that present specific splice variations. (2) Meta-analysis for discriminating non-specific splice variations that are caused by clinical heterogeneity in the collected samples. In the analysis of human cells, controlling non-specific splicing factors is essential for success in the detection of biomarker splice variations because splice patterns are possibly affected by inter-individual differences in the collected samples. We demonstrate its utility and perform a whole transcript analysis of exon expression profiles of colorectal carcinoma. Ryo Yoshida, Kazuyuki Numata, Seiya Imoto, Masao Nagasaki, Atsushi Doi, Kazuko Ueno, Satoru Miyano |
BIBE | 7 |
| 2007 | Statistical Absolute Evaluation of Gene Ontology Terms with Gene Expression Data
Pramod K. Gupta, Ryo Yoshida, Seiya Imoto, Rui Yamaguchi, Satoru Miyano |
ISBRA | 5 |
| 2007 | An efficient grid layout algorithm for biological networks utilizing various biological attributesabstractBACKGROUND: Clearly visualized biopathways provide a great help in understanding biological systems. However, manual drawing of large-scale biopathways is time consuming. We proposed a grid layout algorithm that can handle gene-regulatory networks and signal transduction pathways by considering edge-edge crossing, node-edge crossing, distance measure between nodes, and subcellular localization information from Gene Ontology. Consequently, the layout algorithm succeeded in drastically reducing these crossings in the apoptosis model. However, for larger-scale networks, we encountered three problems: (i) the initial layout is often very far from any local optimum because nodes are initially placed at random, (ii) from a biological viewpoint, human layouts still exceed automatic layouts in understanding because except subcellular localization, it does not fully utilize biological information of pathways, and (iii) it employs a local search strategy in which the neighborhood is obtained by moving one node at each step, and automatic layouts suggest that simultaneous movements of multiple nodes are necessary for better layouts, while such extension may face worsening the time complexity. RESULTS: We propose a new grid layout algorithm. To address problem (i), we devised a new force-directed algorithm whose output is suitable as the initial layout. For (ii), we considered that an appropriate alignment of nodes having the same biological attribute is one of the most important factors of the comprehension, and we defined a new score function that gives an advantage to such configurations. For solving problem (iii), we developed a search strategy that considers swapping nodes as well as moving a node, while keeping the order of the time complexity. Though a naïve implementation increases by one order, the time complexity, we solved this difficulty by devising a method that caches differences between scores of a layout and its possible updates. CONCLUSION: Layouts of the new grid layout algorithm are compared with that of the previous algorithm and human layout in an endothelial cell model, three times as large as the apoptosis model. The total cost of the result from the new grid layout algorithm is similar to that of the human layout. In addition, its convergence time is drastically reduced (40% reduction). Kaname Kojima, Masao Nagasaki, Euna Jeong, Mitsuru Kato, Satoru Miyano |
BMC Bioinform. | 5 |
| 2007 | AYUMS: an algorithm for completely automatic quantitation based on LC-MS/MS proteome data and its application to the analysis of signal transductionabstractBACKGROUND: Comprehensive description of the behavior of cellular components in a quantitative manner is essential for systematic understanding of biological events. Recent LC-MS/MS (tandem mass spectrometry coupled with liquid chromatography) technology, in combination with the SILAC (Stable Isotope Labeling by Amino acids in Cell culture) method, has enabled us to make relative quantitation at the proteome level. The recent report by Blagoev et al. (Nat. Biotechnol., 22, 1139-1145, 2004) indicated that this method was also applicable for the time-course analysis of cellular signaling events. Relative quatitation can easily be performed by calculating the ratio of peak intensities corresponding to differentially labeled peptides in the MS spectrum. As currently available software requires some GUI applications and is time-consuming, it is not suitable for processing large-scale proteome data. RESULTS: To resolve this difficulty, we developed an algorithm that automatically detects the peaks in each spectrum. Using this algorithm, we developed a software tool named AYUMS that automatically identifies the peaks corresponding to differentially labeled peptides, compares these peaks, calculates each of the peak ratios in mixed samples, and integrates them into one data sheet. This software has enabled us to dramatically save time for generation of the final report. CONCLUSION: AYUMS is a useful software tool for comprehensive quantitation of the proteome data generated by LC-MS/MS analysis. This software was developed using Java and runs on Linux, Windows, and Mac OS X. Please contact [email protected] if you are interested in the application. The project web page is http://www.csml.org/ayums/. Ayumu Saito, Masao Nagasaki, Masaaki Oyama, Hiroko Kozuka-Hata, Kentaro Semba, Sumio Sugano, Tadashi Yamamoto, Satoru Miyano |
BMC Bioinform. | 8 |
| 2007 | On the complexity of deriving position specific score matrices from positive and negative sequences
Tatsuya Akutsu, Hideo Bannai, Satoru Miyano, Sascha Ott |
Discret. Appl. Math. | 3 |
| 2006 | ArrayCluster: an analytic tool for clustering, data visualization and module finder on gene expression profilesabstractSUMMARY: One of the significant challenges in gene expression analysis is to find unknown subtypes of several diseases at the molecular levels. This task can be addressed by grouping gene expression patterns of the collected samples on the basis of a large number of genes. Application of commonly used clustering methods to such a dataset however are likely to fail owing to over-learning, because the number of samples to be grouped is much smaller than the data dimension which is equal to the number of genes involved in the dataset. To overcome such difficulty, we developed a novel model-based clustering method, referred to as the mixed factors analysis. The ArrayCluster is a freely available software to perform the mixed factors analysis. It provides us some analytic tools for clustering DNA microarray experiments, data visualization and an automatic detector for module transcriptional of genes that are relevant to the calibrated molecular subtypes and so on. Ryo Yoshida, Tomoyuki Higuchi, Seiya Imoto, Satoru Miyano |
Bioinform. | 4 |
| 2005 | A new regulatory interaction suggested by simulations for circadian genetic control mechanism in mammals
Hiroshi Matsuno, Shin-Ichi T. Inouye, Yasuki Okitsu, Yasushi Fujii, Satoru Miyano |
APBC | 5 |
| 2005 | Estimating Gene Networks from Expression Data and Binding Location Data via Boolean Networks
Osamu Hirose, Naoki Nariai, Yoshinori Tamada, Hideo Bannai, Seiya Imoto, Satoru Miyano |
ICCSA (3) | 6 |
| 2005 | Superiority of network motifs over optimal networks and an application to the revelation of gene network evolutionabstractMOTIVATION: Estimating the network of regulative interactions between genes from gene expression measurements is a major challenge. Recently, we have shown that for gene networks of up to around 35 genes, optimal network models can be computed. However, even optimal gene network models will in general contain false edges, since the expression data will not unambiguously point to a single network. RESULTS: In order to overcome this problem, we present a computational method to enumerate the most likely m networks and to extract a widely common subgraph (denoted as gene network motif) from these. We apply the method to bacterial gene expression data and extensively compare estimation results to knowledge. Our results reveal that gene network motifs are in significantly better agreement to biological knowledge than optimal network models. We also confirm this observation in a series of estimations using synthetic microarray data and compare estimations by our method with previous estimations for yeast. Furthermore, we use our method to estimate similarities and differences of the gene networks that regulate tryptophan metabolism in two related species and thereby demonstrate the analysis of gene network evolution. AVAILABILITY: Commercial license negotiable with Gene Networks Inc. ([email protected]) CONTACT: [email protected] Sascha Ott, Annika Hansen, SunYong Kim, Satoru Miyano |
Bioinform. | 4 |
| 2005 | Prediction of Transcriptional Terminators in Bacillus subtilis and Related SpeciesabstractIn prokaryotes, genes belonging to the same operon are transcribed in a single mRNA molecule. Transcription starts as the RNA polymerase binds to the promoter and continues until it reaches a transcriptional terminator. Some terminators rely on the presence of the Rho protein, whereas others function independently of Rho. Such Rho-independent terminators consist of an inverted repeat followed by a stretch of thymine residues, allowing us to predict their presence directly from the DNA sequence. Unlike in Escherichia coli, the Rho protein is dispensable in Bacillus subtilis, suggesting a limited role for Rho-dependent termination in this organism and possibly in other Firmicutes. We analyzed 463 experimentally known terminating sequences in B. subtilis and found a decision rule to distinguish Rho-independent transcriptional terminators from non-terminating sequences. The decision rule allowed us to find the boundaries of operons in B. subtilis with a sensitivity and specificity of about 94%. Using the same decision rule, we found an average sensitivity of 94% for 57 bacteria belonging to the Firmicutes phylum, and a considerably lower sensitivity for other bacteria. Our analysis shows that Rho-independent termination is dominant for Firmicutes in general, and that the properties of the transcriptional terminators are conserved. Terminator prediction can be used to reliably predict the operon structure in these organisms, even in the absence of experimentally known operons. Genome-wide predictions of Rho-independent terminators for the 57 Firmicutes are available in the Supporting Information section. Michiel J. L. de Hoon, Yuko Makita, Kenta Nakai, Satoru Miyano |
PLoS Comput. Biol. | 4 |
| 2004 | Integrating Biopathway Databases for Large-scale Modeling and Simulation
Masao Nagasaki, Atsushi Doi, Hiroshi Matsuno, Satoru Miyano |
APBC | 4 |
| 2004 | Case-Control Study of Binary Disease Trait Considering Interactions between SNPs and Environmental Effects using Logistic RegressionabstractIn this paper, we propose a combination of logistic regression and genetic algorithm for the association study of the binary disease trait. We use a logistic regression model to describe the relation of multiple SNPs, environments and the target binary trait. The logistic regression model can capture the continuous effects of environments without categorization, which causes the loss of the information. To construct an accurate prediction rule for binary trait, we adopted Akaike information criterion (AIC) to find the most effective set of SNPs and environments. That is, the set of SNPs and environments that gives the smallest AIC is chosen as the optimal set. Since the number of combinations of SNPs and environments is usually huge, we propose the use of the genetic algorithm for choosing the optimal SNPs and environments in the sense of AIC. We show the effectiveness of the proposed method through the analysis of the case/control populations of diabetes patients. We succeeded in finding an efficient set to predict types of diabetes and some SNPs which have strong interactions to age while it is not significant as a single locus. Reiichiro Nakamichi, Seiya Imoto, Satoru Miyano |
BIBE | 3 |
| 2004 | Finding Optimal Pairs of Cooperative and Competing Patterns with Bounded Distance
Shunsuke Inenaga, Hideo Bannai, Heikki Hyyrö, Ayumi Shinohara, Masayuki Takeda, Kenta Nakai, Satoru Miyano |
Discovery Science | 7 |
| 2004 | Finding Optimal Pairs of Patterns
Hideo Bannai, Heikki Hyyrö, Ayumi Shinohara, Masayuki Takeda, Kenta Nakai, Satoru Miyano |
WABI | 6 |
| 2004 | An O(N2) Algorithm for Discovering Optimal Boolean Pattern PairsabstractWe consider the problem of finding the optimal combination of string patterns, which characterizes a given set of strings that have a numeric attribute value assigned to each string. Pattern combinations are scored based on the correlation between their occurrences in the strings and the numeric attribute values. The aim is to find the combination of patterns which is best with respect to an appropriate scoring function. We present an O(N2) time algorithm for finding the optimal pair of substring patterns combined with Boolean functions, where N is the total length of the sequences. The algorithm looks for all possible Boolean combinations of the patterns, e.g., patterns of the form p and not q, which indicates that the pattern pair is considered to occur in a given string s, if p occurs in s, AND q does NOT occur in s. An efficient implementation using suffix arrays is presented, and we further show that the algorithm can be adapted to find the best k-pattern Boolean combination in O(Nk) time. The algorithm is applied to mRNA sequence data sets of moderate size combined with their turnover rates for the purpose of finding regulatory elements that cooperate, complement, or compete with each other in enhancing and/or silencing mRNA decay. Hideo Bannai, Heikki Hyyrö, Ayumi Shinohara, Masayuki Takeda, Kenta Nakai, Satoru Miyano |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2003 | Inferring gene networks from time series microarray data using dynamic Bayesian networksabstractDynamic Bayesian networks (DBNs) are considered as a promising model for inferring gene networks from time series microarray data. DBNs have overtaken Bayesian networks (BNs) as DBNs can construct cyclic regulations using time delay information. In this paper, a general framework for DBN modelling is outlined. Both discrete and continuous DBN models are constructed systematically and criteria for learning network structures are introduced from a Bayesian statistical viewpoint. This paper reviews the applications of DBNs over the past years. Real data applications for Saccharomyces cerevisiae time series gene expression data are also shown. SunYong Kim, Seiya Imoto, Satoru Miyano |
Briefings Bioinform. | 3 |
| 2003 | Identification of genetic networks by strategic gene disruptions and gene overexpressions under a boolean model
Tatsuya Akutsu, Satoru Kuhara, Osamu Maruyama, Satoru Miyano |
Theor. Comput. Sci. | 4 |
| 2003 | A simple greedy algorithm for finding functional relations: efficient implementation and average case analysis
Tatsuya Akutsu, Satoru Miyano, Satoru Kuhara |
Theor. Comput. Sci. | 2 |
| 2002 | On the Complexity of Deriving Position Specific Score Matrices from Examples
Tatsuya Akutsu, Hideo Bannai, Satoru Miyano, Sascha Ott |
CPM | 3 |
| 2002 | Inferring Gene Regulatory Networks from Time-Ordered Gene Expression Data Using Differential Equations
Michiel J. L. de Hoon, Seiya Imoto, Satoru Miyano |
Discovery Science | 3 |
| 2002 | Toward Drawing an Atlas of Hypothesis Classes: Approximating a Hypothesis via Another Hypothesis Model
Osamu Maruyama, Takayoshi Shoudai, Satoru Miyano |
Discovery Science | 3 |
| 2002 | Extensive feature detection of N-terminal protein sorting signalsabstractMOTIVATION: The prediction of localization sites of various proteins is an important and challenging problem in the field of molecular biology. TargetP, by Emanuelsson et al. (J. Mol. Biol., 300, 1005-1016, 2000) is a neural network based system which is currently the best predictor in the literature for N-terminal sorting signals. One drawback of neural networks, however, is that it is generally difficult to understand and interpret how and why they make such predictions. In this paper, we aim to generate simple and interpretable rules as predictors, and still achieve a practical prediction accuracy. We adopt an approach which consists of an extensive search for simple rules and various attributes which is partially guided by human intuition. RESULTS: We have succeeded in finding rules whose prediction accuracies come close to that of TargetP, while still retaining a very simple and interpretable form. We also discuss and interpret the discovered rules. Hideo Bannai, Yoshinori Tamada, Osamu Maruyama, Kenta Nakai, Satoru Miyano |
Bioinform. | 5 |
| 2002 | Statistical analysis of a small set of time-ordered gene expression data using linear splinesabstractAbstract Motivation: Recently, the temporal response of genes to changes in their environment has been investigated using cDNA microarray technology by measuring the gene expression levels at a small number of time points. Conventional techniques for time series analysis are not suitable for such a short series of time-ordered data. The analysis of gene expression data has therefore usually been limited to a fold-change analysis, instead of a systematic statistical approach. Methods: We use the maximum likelihood method together with Akaike's Information Criterion to fit linear splines to a small set of time-ordered gene expression data in order to infer statistically meaningful information from the measurements. The significance of measured gene expression data is assessed using Student's t-test. Results: Previous gene expression measurements of the cyanobacterium Synechocystis sp. PCC6803 were reanalyzed using linear splines. The temporal response was identified of many genes that had been missed by a fold-change analysis. Based on our statistical analysis, we found that about four gene expression measurements or more are needed at each time point. Availability: An extension module for Python to calculate linear spline functions is available at http://bonsai.ims.u-tokyo.ac.jp/~mdehoon. This software package (with patent pending) is free of charge for academic use only. Contact: [email protected] * To whom correspondence should be addressed. Michiel J. L. de Hoon, Seiya Imoto, Satoru Miyano |
Bioinform. | 3 |
| 2002 | Fast algorithm for extracting multiple unordered short motifs using bit operations
Osamu Maruyama, Hideo Bannai, Yoshinori Tamada, Satoru Kuhara, Satoru Miyano |
Inf. Sci. | 5 |
| 2001 | VML: A View Modeling Language for Computational Knowledge Discovery
Hideo Bannai, Yoshinori Tamada, Osamu Maruyama, Satoru Miyano |
Discovery Science | 4 |
| 2001 | Learning Conformation Rules
Osamu Maruyama, Takayoshi Shoudai, Emiko Furuichi, Satoru Kuhara, Satoru Miyano |
Discovery Science | 5 |
| 2001 | ProDDO: a database of disordered proteins from the Protein Data Bank (PDB)abstractAbstract Summary: ProDDO represents a ‘pre-screened’ database that denotes disorder (or possible disorder) in proteins from the PDB. Availability: ProDDO is available at http://bonsai.ims.u-tokyo.ac.jp/~klsim/database.html Contact: [email protected] * To whom correspondence should be addressed. Present address: Kentucky Center for Structural Biology, Department of Biochemistry, University of Kentucky, 800 Rose Street, Lexington, KY 40536-0298, USA. Kim Lan Sim, Tomoyuki Uchida, Satoru Miyano |
Bioinform. | 3 |
| 2000 | A Simple Greedy Algorithm for Finding Functional Relations: Efficient Implementation and Average Case Anaylsis
Tatsuya Akutsu, Satoru Miyano, Satoru Kuhara |
Discovery Science | 2 |
| 2000 | Algorithms for identifying Boolean networks and related biological networks based on matrix multiplication and fingerprint functionabstractDue to the recent progress of the DNA microarray technology, a large number of gene expression profile data are being produced. How to analyze gene expression data is an important topic in computational molecular biology Several studies have been done using the Boolean network as a model of a genetic network This paper proposes efficient algorithms for identifying Boolean networks of bounded indegree and related biological networks, where identification of a Boolean network can be formalized as a problem of identifying many Boolean functions simultaneously. For the identification of a Boolean network, an O(mnD+1) time naive algorithm and a simple O(mnD) time algorithm are known, where n denotes the number of nodes, m denotes the number of examples, and D denotes the maximum indegree. This paper presents an improved O(mw-2nD + mnD+w-3) time Monte-Carlo type randomized algorithm, where w is the exponent of matrix multiplication (currently, w < 2376). The algorithm is obtained by combining fast matrix multiplication with the randomized fingerprint function for string matching. Although the algorithm and its analysis are simple, the result is non-trivial and the technique can be applied to several related problems. Tatsuya Akutsu, Satoru Miyano, Satoru Kuhara |
RECOMB | 2 |
| 2000 | Inferring qualitative relations in genetic networks and metabolic pathwaysabstractMOTIVATION: Inferring genetic network architecture from time series data of gene expression patterns is an important topic in bioinformatics. Although inference algorithms based on the Boolean network were proposed, the Boolean network was not sufficient as a model of a genetic network. RESULTS: First, a Boolean network model with noise is proposed, together with an inference algorithm for it. Next, a qualitative network model is proposed, in which regulation rules are represented as qualitative rules and embedded in the network structure. Algorithms are also presented for inferring qualitative relations from time series data. Then, an algorithm for inferring S-systems (synergistic and saturable systems) from time series data is presented, where S-systems are based on a particular kind of nonlinear differential equation and have been applied to the analysis of various biological systems. Theoretical results are shown for Boolean networks with noises and simple qualitative networks. Computational results are shown for Boolean networks with noises and S-systems, where real data are not used because the proposed models are still conceptual and the quantity and quality of currently available data are not enough for the application of the proposed methods. Tatsuya Akutsu, Satoru Miyano, Satoru Kuhara |
Bioinform. | 2 |
| 1999 | A Definition of Discovery in Terms of Generalized Descriptional Complexity
Hideo Bannai, Satoru Miyano |
Discovery Science | 2 |
| 1999 | Designing Views in HypothesisCreator: System for Assisting in Discovery
Osamu Maruyama, Tomoyuki Uchida, Kim Lan Sim, Satoru Miyano |
Discovery Science | 4 |
| 1999 | On the Approximation of Protein Threading
Tatsuya Akutsu, Satoru Miyano |
Theor. Comput. Sci. | 2 |
| 1999 | Foreword: Genome Informatics
Satoru Miyano |
Theor. Comput. Sci. | 1 |
| 1998 | Toward Genomic Hypothesis Creator: View Designer for Discovery
Osamu Maruyama, Tomoyuki Uchida, Takayoshi Shoudai, Satoru Miyano |
Discovery Science | 4 |
| 1998 | Identification of Gene Regulatory Networks by Strategic Gene Disruptions and Gene Overexpressions
Tatsuya Akutsu, Satoru Kuhara, Osamu Maruyama, Satoru Miyano |
SODA | 4 |
| 1997 | On the approximation of protein threadingabstractIn this paper, we study the protein threading problem, which was proposed for finding a folded 3D protein structure from an amino acid sequence.Since this problem was already proved to be NP-hard by Lathrop, we study polynomial time approximation algorithms.First we show that the protein threading problem is MAX SNP-hard.Next we show that the protein threading problem can be approximated within a factor 4 for a special case in which a graph representing interaction between residues (amino acids) is planar.This case corresponds to a P-sheet substructure, which appears in most protein structures. Tatsuya Akutsu, Satoru Miyano |
RECOMB | 2 |
| 1996 | Extracting Best Consensus Motifs from Positive and Negative Examples
Erika Tateishi, Osamu Maruyama, Satoru Miyano |
STACS | 3 |
| 1996 | Inferring a Tree from Walks
Osamu Maruyama, Satoru Miyano |
Theor. Comput. Sci. | 2 |
| 1995 | Algorithmic Problems Arising from Genome Informatics (Abstract)
Satoru Miyano |
ISAAC | 1 |
| 1995 | BONSAI Garden: Parallel Knowledge Discovery System for Amino Acid Sequences
Takayoshi Shoudai, Michael Lappe, Satoru Miyano, Ayumi Shinohara, Takeo Okazaki, Setsuo Arikawa, Tomoyuki Uchida, Shinichi Shimozono, Takeshi Shinohara, Satoru Kuhara |
ISMB | 3 |
| 1995 | Graph Inference from a Walk for TRees of Bounded Degree 3 is NP-Complete
Osamu Maruyama, Satoru Miyano |
MFCS | 2 |
| 1995 | Using Maximal Independent Sets to Solve Problems in Parallel
Takayoshi Shoudai, Satoru Miyano |
Theor. Comput. Sci. | 2 |
| 1992 | Inferring a Tree from Walks
Osamu Maruyama, Satoru Miyano |
MFCS | 2 |
| 1991 | Delta_2^P-Complete Lexicographically First Maximal Subgraph Problems
Satoru Miyano |
Theor. Comput. Sci. | 1 |
| 1989 | The Lexicographically First Maximum Subgraph Problems: P-Completeness and NC Algorithms
Satoru Miyano |
Math. Syst. Theory | 1 |
| 1988 | Δ2p-Complete Lexicographically First Maximal Subgraph Problems
Satoru Miyano |
MFCS | 1 |
| 1988 | A Parallelizable Lexicographically First Maximal Edge-Induced Subgraph Problem
Satoru Miyano |
Inf. Process. Lett. | 1 |
| 1987 | The Lexicographically First Maximal Subgraph Problems: P-Completeness and NC Algorithms
Satoru Miyano |
ICALP | 1 |
| 1984 | Remarks on Two-Way Automata with Weak-Counters
Satoru Miyano |
Inf. Process. Lett. | 1 |
| 1984 | Alternating Finite Automata on omega-Words
Satoru Miyano, Takeshi Hayashi 0004 |
Theor. Comput. Sci. | 1 |
| 1983 | Remarks on Multihead Pushdown Automata and Multihead Stack Automata
Satoru Miyano |
J. Comput. Syst. Sci. | 1 |
| 1982 | A Hierarchy Theorem for Multihead Stack-Counter Automata
Satoru Miyano |
Acta Informatica | 1 |
| 1982 | Two-Way Deterministic Multi-Weak-Counter Machines
Satoru Miyano |
Theor. Comput. Sci. | 1 |
| 1980 | One-Way Weak-Stack-Counter Automata
Satoru Miyano |
J. Comput. Syst. Sci. | 1 |