Rui Yamaguchi

dblp:92/7043 · DBLP profile ↗
← Back
30ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0003-1224-227XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 26 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4
YearPublicationVenuePosition
2025 RVINN: a flexible modeling for inferring dynamic transcriptional and post-transcriptional regulation using physics-informed neural networks
abstract
SUMMARY: Dynamic gene expression is controlled by transcriptional and post-transcriptional regulation. Recent studies on transcriptional bursting and buffering have increasingly highlighted the dynamic gene regulatory mechanisms. However, direct measurement techniques still face various constraints and require complementary methodologies, which are both comprehensive and versatile. To address this issue, inference approaches based on transcriptome data and differential equation models representing the messenger RNA lifecycle have been proposed. However, the inference of complex dynamics under diverse experimental conditions and biological scenarios remains challenging. In this study, we developed a flexible modeling using physics-informed neural networks and demonstrated its performance using simulation and experimental data. Our model has the ability to computationally revalidate and visualize dynamic biological phenomena, such as transcriptional ripple, co-bursting, and buffering in a breast cancer cell line. Furthermore, our results suggest putative molecular mechanisms underlying these phenomena. We propose a novel approach for inferring transcriptional and post-transcriptional regulation and expect to offer valuable insights for experimental and systems biology. AVAILABILITY AND IMPLEMENTATION: https://github.com/omuto/RVINN.
Osamu Muto, Zhongliang Guo 0003, Rui Yamaguchi
Bioinform.3
2022 Identification of bacteriophage genome sequences with representation learning
abstract
MOTIVATION: Bacteriophages/phages are the viruses that infect and replicate within bacteria and archaea, and rich in human body. To investigate the relationship between phages and microbial communities, the identification of phages from metagenome sequences is the first step. Currently, there are two main methods for identifying phages: database-based (alignment-based) methods and alignment-free methods. Database-based methods typically use a large number of sequences as references; alignment-free methods usually learn the features of the sequences with machine learning and deep learning models. RESULTS: We propose INHERIT which uses a deep representation learning model to integrate both database-based and alignment-free methods, combining the strengths of both. Pre-training is used as an alternative way of acquiring knowledge representations from existing databases, while the BERT-style deep learning framework retains the advantage of alignment-free methods. We compare INHERIT with four existing methods on a third-party benchmark dataset. Our experiments show that INHERIT achieves a better performance with the F1-score of 0.9932. In addition, we find that pre-training two species separately helps the non-alignment deep learning model make more accurate predictions. AVAILABILITY AND IMPLEMENTATION: The codes of INHERIT are now available in: https://github.com/Celestial-Bai/INHERIT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zeheng Bai, Satoru Miyano, Rui Yamaguchi, Kosuke Fujimoto, Satoshi Uematsu, Seiya Imoto
Bioinform.4
2021 On the application of BERT models for nanopore methylation detection
abstract
DNA methylation is a common nucleotide modification, which is associated with various biological processes, such as gene expression and aging. Nanopore sequencing provides a direct detecting approach through searching specific current signal shifts. Recently, model-based approaches, especially those using deep learning models, have achieved significant performance improvements on nanopore methylation detection. In this work, we explore using the non-recurrent neural network structure of Bidirectional Encoder Representations from Transformers (BERT) for the task, which provides an alternative fast inference model to the state-of-the-art bi-directional Recurrent Neural Network (biRNN). In addition, we propose a refined BERT model with relative position representation and center hidden units concatenation, which takes account of the task-specific characters into modeling. We evaluate the proposed models on the R9 benchmark datasets of different motifs and methyltransferases. The experiment results show that the refined BERT model can achieve competitive or even better results than the state-of-the-art biRNN model, while the model inference speed is faster.
Kiyoshi Yamaguchi, Sera Hatakeyama, Yoichi Furukawa, Satoru Miyano, Rui Yamaguchi, Seiya Imoto
BIBM6
2021 Halcyon: an accurate basecaller exploiting an encoder-decoder model with monotonic attention
abstract
MOTIVATION: In recent years, nanopore sequencing technology has enabled inexpensive long-read sequencing, which promises reads longer than a few thousand bases. Such long-read sequences contribute to the precise detection of structural variations and accurate haplotype phasing. However, deciphering precise DNA sequences from noisy and complicated nanopore raw signals remains a crucial demand for downstream analyses based on higher-quality nanopore sequencing, although various basecallers have been introduced to date. RESULTS: To address this need, we developed a novel basecaller, Halcyon, that incorporates neural-network techniques frequently used in the field of machine translation. Our model employs monotonic-attention mechanisms to learn semantic correspondences between nucleotides and signal levels without any pre-segmentation against input signals. We evaluated performance with a human whole-genome sequencing dataset and demonstrated that Halcyon outperformed existing third-party basecallers and achieved competitive performance against the latest Oxford Nanopore Technologies' basecallers. AVAILABILITYAND IMPLEMENTATION: The source code (halcyon) can be found at https://github.com/relastle/halcyon.
Hiroki Konishi, Rui Yamaguchi, Kiyoshi Yamaguchi, Yoichi Furukawa, Seiya Imoto
Bioinform.2
2021 Enhancing breakpoint resolution with deep segmentation model: A general refinement method for read-depth based structural variant callers
abstract
Read-depths (RDs) are frequently used in identifying structural variants (SVs) from sequencing data. For existing RD-based SV callers, it is difficult for them to determine breakpoints in single-nucleotide resolution due to the noisiness of RD data and the bin-based calculation. In this paper, we propose to use the deep segmentation model UNet to learn base-wise RD patterns surrounding breakpoints of known SVs. We integrate model predictions with an RD-based SV caller to enhance breakpoints in single-nucleotide resolution. We show that UNet can be trained with a small amount of data and can be applied both in-sample and cross-sample. An enhancement pipeline named RDBKE significantly increases the number of SVs with more precise breakpoints on simulated and real data. The source code of RDBKE is freely available at https://github.com/yaozhong/deepIntraSV.
Seiya Imoto, Satoru Miyano, Rui Yamaguchi
PLoS Comput. Biol.4
2020 Neoantimon: a multifunctional R package for identification of tumor-specific neoantigens
abstract
SUMMARY: It is known that some mutant peptides, such as those resulting from missense mutations and frameshift insertions, can bind to the major histocompatibility complex and be presented to antitumor T cells on the surface of a tumor cell. These peptides are termed neoantigen, and it is important to understand this process for cancer immunotherapy. Here, we introduce an R package termed Neoantimon that can predict a list of potential neoantigens from a variety of mutations, which include not only somatic point mutations but insertions, deletions and structural variants. Beyond the existing applications, Neoantimon is capable of attaching and reflecting several additional information, e.g. wild-type binding capability, allele specific RNA expression levels, single nucleotide polymorphism information and combinations of mutations to filter out infeasible peptides as neoantigen. AVAILABILITY AND IMPLEMENTATION: The R package is available at http://github/hase62/Neoantimon.
Takanori Hasegawa, Shuto Hayashi, Eigo Shimizu, Shinichi Mizuno, Atsushi Niida, Rui Yamaguchi, Satoru Miyano, Hidewaki Nakagawa, Seiya Imoto
Bioinform.6
2020 Nanopore basecalling from a perspective of instance segmentation
abstract
BACKGROUND: Nanopore sequencing is a rapidly developing third-generation sequencing technology, which can generate long nucleotide reads of molecules within a portable device in real-time. Through detecting the change of ion currency signals during a DNA/RNA fragment's pass through a nanopore, genotypes are determined. Currently, the accuracy of nanopore basecalling has a higher error rate than the basecalling of short-read sequencing. Through utilizing deep neural networks, the-state-of-the art nanopore basecallers achieve basecalling accuracy in a range from 85% to 95%. RESULT: In this work, we proposed a novel basecalling approach from a perspective of instance segmentation. Different from previous approaches of doing typical sequence labeling, we formulated the basecalling problem as a multi-label segmentation task. Meanwhile, we proposed a refined U-net model which we call UR-net that can model sequential dependencies for a one-dimensional segmentation task. The experiment results show that the proposed basecaller URnano achieves competitive results on the in-species data, compared to the recently proposed CTC-featured basecallers. CONCLUSION: Our results show that formulating the basecalling problem as a one-dimensional segmentation task is a promising approach, which does basecalling and segmentation jointly.
Arda Akdemir, Georg Tremmel, Seiya Imoto, Satoru Miyano, Tetsuo Shibuya, Rui Yamaguchi
BMC Bioinform.7
2019 A Bayesian model integration for mutation calling through data partitioning
abstract
MOTIVATION: Detection of somatic mutations from tumor and matched normal sequencing data has become among the most important analysis methods in cancer research. Some existing mutation callers have focused on additional information, e.g. heterozygous single-nucleotide polymorphisms (SNPs) nearby mutation candidates or overlapping paired-end read information. However, existing methods cannot take multiple information sources into account simultaneously. Existing Bayesian hierarchical model-based methods construct two generative models, the tumor model and error model, and limited information sources have been modeled. RESULTS: We proposed a Bayesian model integration framework named as partitioning-based model integration. In this framework, through introducing partitions for paired-end reads based on given information sources, we integrate existing generative models and utilize multiple information sources. Based on that, we constructed a novel Bayesian hierarchical model-based method named as OHVarfinDer. In both the tumor model and error model, we introduced partitions for a set of paired-end reads that cover a mutation candidate position, and applied a different generative model for each category of paired-end reads. We demonstrated that our method can utilize both heterozygous SNP information and overlapping paired-end read information effectively in simulation datasets and real datasets. AVAILABILITY AND IMPLEMENTATION: https://github.com/takumorizo/OHVarfinDer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Takuya Moriyama, Seiya Imoto, Shuto Hayashi, Yuichi Shiraishi, Satoru Miyano, Rui Yamaguchi
Bioinform.6
2019 Virtual Grid Engine: a simulated grid engine environment for large-scale supercomputers
abstract
BACKGROUND: Supercomputers have become indispensable infrastructures in science and industries. In particular, most state-of-the-art scientific results utilize massively parallel supercomputers ranked in TOP500. However, their use is still limited in the bioinformatics field due to the fundamental fact that the asynchronous parallel processing service of Grid Engine is not provided on them. To encourage the use of massively parallel supercomputers in bioinformatics, we developed middleware called Virtual Grid Engine, which enables software pipelines to automatically perform their tasks as MPI programs. RESULT: We conducted basic tests to check the time required to assign jobs to workers by VGE. The results showed that the overhead of the employed algorithm was 246 microseconds and our software can manage thousands of jobs smoothly on the K computer. We also tried a practical test in the bioinformatics field. This test included two tasks, the split and BWA alignment of input FASTQ data. 25,055 nodes (2,000,440 cores) were used for this calculation and accomplished it in three hours. CONCLUSION: We considered that there were four important requirements for this kind of software, non-privilege server program, multiple job handling, dependency control, and usability. We carefully designed and checked all requirements. And this software fulfilled all the requirements and achieved good performance in a large scale analysis.
Masaaki Yadome, Tatsuo Nishiki, Shigeru Ishiduki, Hikaru Inoue, Rui Yamaguchi, Satoru Miyano
BMC Bioinform.6
2019 Capturing the differences between humoral immunity in the normal and tumor environments from repertoire-seq of B-cell receptors using supervised machine learning
abstract
BACKGROUND: The recent success of immunotherapy in treating tumors has attracted increasing interest in research related to the adaptive immune system in the tumor microenvironment. Recent advances in next-generation sequencing technology enabled the sequencing of whole T-cell receptors (TCRs) and B-cell receptors (BCRs)/immunoglobulins (Igs) in the tumor microenvironment. Since BCRs/Igs in tumor tissues have high affinities for tumor-specific antigens, the patterns of their amino acid sequences and other sequence-independent features such as the number of somatic hypermutations (SHMs) may differ between the normal and tumor microenvironments. However, given the high diversity of BCRs/Igs and the rarity of recurrent sequences among individuals, it is far more difficult to capture such differences in BCR/Ig sequences than in TCR sequences. The aim of this study was to explore the possibility of discriminating BCRs/Igs in tumor and in normal tissues, by capturing these differences using supervised machine learning methods applied to RNA sequences of BCRs/Igs. RESULTS: RNA sequences of BCRs/Igs were obtained from matched normal and tumor specimens from 90 gastric cancer patients. BCR/Ig-features obtained in Rep-Seq were used to classify individual BCR/Ig sequences into normal or tumor classes. Different machine learning models using various features were constructed as well as gradient boosting machine (GBM) classifier combining these models. The results demonstrated that BCR/Ig sequences between normal and tumor microenvironments exhibit their differences. Next, by using a GBM trained to classify individual BCR/Ig sequences, we tried to classify sets of BCR/Ig sequences into normal or tumor classes. As a result, an area under the curve (AUC) value of 0.826 was achieved, suggesting that BCR/Ig repertoires have distinct sequence-level features in normal and tumor tissues. CONCLUSIONS: To the best of our knowledge, this is the first study to show that BCR/Ig sequences derived from tumor and normal tissues have globally distinct patterns, and that these tissues can be effectively differentiated using BCR/Ig repertoires.
Hiroki Konishi, Daisuke Komura, Hiroto Katoh, Shinichiro Atsumi, Hirotomo Koda, Asami Yamamoto, Yasuyuki Seto, Masashi Fukayama, Rui Yamaguchi, Seiya Imoto, Shumpei Ishikawa
BMC Bioinform.9
2018 Virtual Grid Engine: Accelerating thousands of omics sample analyses using large-scale supercomputers
Masaaki Yadome, Tatsuo Nishiki, Shigeru Ishiduki, Hikaru Inoue, Rui Yamaguchi, Satoru Miyano
BIBM6
2017 Reconstruction of high read-depth signals from low-depth whole genome sequencing data using deep learning
abstract
Motivation: Next-generation sequencing (NGS) technologies using DNA, RNA, or methylation sequencing are prevailing tools used in modern genome research. For DNA sequencing, whole genome sequencing (WGS) and whole exome sequencing (WES) are two typical applications with a different preference on the trade-off between sequencing depth and base coverage. Although sequencing costs have been greatly reduced, the sequence depth used in WGS is relatively lower than WES (e.g., ~35× vs. 100×~). In addition, biases and batch effects may exist in different stages of a NGS experiment. Using low-depth and biased WGS data for downstream analyses is more sensitive to the bias problem and makes it even more difficult to uncover real biological signals in the data. In this work, we focused on reconstructing high read-depth signals from low-depth WGS data. We make use of a pair of WGS data with different read-depth for the same sample and learn a mapping from low-depth signals to high-depth in the given platform. Results: We explored three different reconstruction models from shallow to deep. Our experimental results show that by only using the read depth information, deeper models do not perform far better than a linear regression model. Through incorporating additional information, such as GC-content, mappability and nucleotide sequence information, the performance of convolutional neural network (CNN) models can be further improved. We made use of the reconstructed read-depth signals in downstream analysis to identify copy number variation segments for single sample. The experiment results show that segments that are not detected using low-depth data, can be detected with the reconstructed signals by the CNN model using extra biological information.
Seiya Imoto, Satoru Miyano, Rui Yamaguchi
BIBM4
2016 OVarCall: Bayesian Mutation Calling Method Utilizing Overlapping Paired-End Reads
Takuya Moriyama, Yuichi Shiraishi, Kenichi Chiba, Rui Yamaguchi, Seiya Imoto, Satoru Miyano
ISBRA4
2014 Parameter estimation in multi-compartment SIR model
Masaya M. Saito, Seiya Imoto, Rui Yamaguchi, Satoru Miyano, Tomoyuki Higuchi
FUSION3
2013 Estimation of abrupt changes in sentinel observation data of influenza epidemics in Japan
Masaya M. Saito, Seiya Imoto, Rui Yamaguchi, Satoru Miyano, Tomoyuki Higuchi
FUSION3
2012 Identifiability of local transmissibility parameters in agent-based pandemic simulation
Masaya M. Saito, Seiya Imoto, Rui Yamaguchi, Satoru Miyano, Tomoyuki Higuchi
FUSION3
2012 Identifying Gene Pathways Associated with Cancer Characteristics via Sparse Statistical Methods
abstract
We propose a statistical method for uncovering gene pathways that characterize cancer heterogeneity. To incorporate knowledge of the pathways into the model, we define a set of activities of pathways from microarray gene expression data based on the Sparse Probabilistic Principal Component Analysis (SPPCA). A pathway activity logistic regression model is then formulated for cancer phenotype. To select pathway activities related to binary cancer phenotypes, we use the elastic net for the parameter estimation and derive a model selection criterion for selecting tuning parameters included in the model estimation. Our proposed method can also reverse-engineer gene networks based on the identified multiple pathways that enables us to discover novel gene-gene associations relating with the cancer phenotypes. We illustrate the whole process of the proposed method through the analysis of breast cancer gene expression data.
Shuichi Kawano, Teppei Shimamura, Atsushi Niida, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, Ryo Yoshida, Cristin G. Print, Satoru Miyano
IEEE ACM Trans. Comput. Biol. Bioinform.5
2011 Estimation of macroscopic parameter in agent-based pandemic simulation
Masaya M. Saito, Seiya Imoto, Rui Yamaguchi, Satoru Miyano, Tomoyuki Higuchi
FUSION3
2011 Comprehensive Pharmacogenomic Pathway Screening by Data Assimilation
Takanori Hasegawa, Rui Yamaguchi, Masao Nagasaki, Seiya Imoto, Satoru Miyano
ISBRA2
2011 SiGN-SSM: open source parallel software for estimating gene networks with state space models
abstract
UNLABELLED: SiGN-SSM is an open-source gene network estimation software able to run in parallel on PCs and massively parallel supercomputers. The software estimates a state space model (SSM), that is a statistical dynamic model suitable for analyzing short time and/or replicated time series gene expression profiles. SiGN-SSM implements a novel parameter constraint effective to stabilize the estimated models. Also, by using a supercomputer, it is able to determine the gene network structure by a statistical permutation test in a practical time. SiGN-SSM is applicable not only to analyzing temporal regulatory dependencies between genes, but also to extracting the differentially regulated genes from time series expression profiles. AVAILABILITY: SiGN-SSM is distributed under GNU Affero General Public Licence (GNU AGPL) version 3 and can be downloaded at http://sign.hgc.jp/signssm/. The pre-compiled binaries for some architectures are available in addition to the source code. The pre-installed binaries are also available on the Human Genome Center supercomputer system. The online manual and the supplementary information of SiGN-SSM is available on our web site. CONTACT: [email protected].
Yoshinori Tamada, Rui Yamaguchi, Seiya Imoto, Osamu Hirose, Ryo Yoshida, Masao Nagasaki, Satoru Miyano
Bioinform.2
2011 Inferring Contagion in Regulatory Networks
abstract
Several gene regulatory network models containing concepts of directionality at the edges have been proposed. However, only a few reports have an interpretable definition of directionality. Here, differently from the standard causality concept defined by Pearl, we introduce the concept of contagion in order to infer directionality at the edges, i.e., asymmetries in gene expression dependences of regulatory networks. Moreover, we present a bootstrap algorithm in order to test the contagion concept. This technique was applied in simulated data and, also, in an actual large sample of biological data. Literature review has confirmed some genes identified by contagion as actually belonging to the TP53 pathway.
André Fujita, João R. Sato, Marcos Angelo Almeida Demasi, Rui Yamaguchi, Teppei Shimamura, Carlos Eduardo Ferreira, Mari Cleide Sogayar, Satoru Miyano
IEEE ACM Trans. Comput. Biol. Bioinform.4
2010 Identifying Hidden Confounders in Gene Networks by Bayesian Networks
abstract
In the estimation of gene networks from microarray gene expression data, we propose a statistical method for quantification of the hidden confounders in gene networks, which were possibly removed from the set of genes on the gene networks or are novel biological elements that are not measured by microarrays. Due to high computational cost of the structural learning of Bayesian networks and the limited source of the microarray data, it is usual to perform gene selection prior to the estimation of gene networks. Therefore, there exist missing genes that decrease accuracy and interpretability of the estimated gene networks. The proposed method can identify hidden confounders based on the conflicts of the estimated local Bayesian network structures and estimate their ideal profiles based on the proposed Bayesian networks with hidden variables with an EM algorithm. From the estimated ideal profiles, we can identify genes which are missing in the network or suggest the existence of the novel biological elements if the ideal profiles are not significantly correlated with any expression profiles of genes. To the best of our knowledge, this research is the first study to theoretically characterize missing genes in gene networks and practically utilize this information to refine network estimation.
Tomoya Higashigaki, Kaname Kojima, Rui Yamaguchi, Masato Inoue, Seiya Imoto, Satoru Miyano
BIBE3
2010 Discovering functional gene pathways associated with cancer heterogeneity via sparse supervised learning
abstract
We propose a statistical method for uncovering gene pathways that characterize cancer heterogeneity. To incorporate knowledge of the pathways into the model, we define a set of activities of pathways from microarray gene expression data based on the sparse probabilistic principal component analysis. A pathway activity logistic regression model is then formulated for cancer phenotype. To select pathway activities related to binary cancer phenotypes, we use the elastic net for the parameter estimation and derive a model selection criterion for selecting tuning parameters included in the model estimation. Our proposed method can also reverse-engineer gene networks based on the identified multiple pathways that enables us to discover novel gene-gene associations relating with the cancer phenotypes. We illustrate the whole process of the proposed method through the analysis of breast cancer gene expression data.
Shuichi Kawano, Teppei Shimamura, Atsushi Niida, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, Ryo Yoshida, Cristin G. Print, Satoru Miyano
BIBM5
2010 Model-free unsupervised gene set screening based on information enrichment in expression profiles
abstract
MOTIVATION: A number of unsupervised gene set screening methods have recently been developed for search of putative functional gene sets based on their expression profiles. Most of the methods statistically evaluate whether the expression profiles of each gene set are fit to assumed models: e.g. co-expression across all samples or a subgroup of samples. However, it is possible that they fail to capture informative gene sets whose expression profiles are not fit to the assumed models. RESULTS: To overcome this limitation, we propose a model-free unsupervised gene set screening method, Matrix Information Enrichment Analysis (MIEA). Without assuming any specific models, MIEA screens gene sets based on information richness of their expression profiles. We extensively compared the performance of MIEA to those of other unsupervised gene set screening methods, using various types of simulated and real data. The benchmark tests demonstrated that MIEA can detect singular expression profiles that the other methods fail to find, and performs broadly well for various types of input data. Taken together, this study introduces MIEA as a broadly applicable gene set screening tool for mining regulatory programs from transcriptome data.
Atsushi Niida, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, André Fujita, Teppei Shimamura, Satoru Miyano
Bioinform.3
2010 Inferring dynamic gene networks under varying conditions for transcriptomic network comparison
abstract
MOTIVATION: Elucidating the differences between cellular responses to various biological conditions or external stimuli is an important challenge in systems biology. Many approaches have been developed to reverse engineer a cellular system, called gene network, from time series microarray data in order to understand a transcriptomic response under a condition of interest. Comparative topological analysis has also been applied based on the gene networks inferred independently from each of the multiple time series datasets under varying conditions to find critical differences between these networks. However, these comparisons often lead to misleading results, because each network contains considerable noise due to the limited length of the time series. RESULTS: We propose an integrated approach for inferring multiple gene networks from time series expression data under varying conditions. To the best of our knowledge, our approach is the first reverse-engineering method that is intended for transcriptomic network comparison between varying conditions. Furthermore, we propose a state-of-the-art parameter estimation method, relevance-weighted recursive elastic net, for providing higher precision and recall than existing reverse-engineering methods. We analyze experimental data of MCF-7 human breast cancer cells stimulated by epidermal growth factor or heregulin with several doses and provide novel biological hypotheses through network comparison. AVAILABILITY: The software NETCOMP is available at http://bonsai.ims.u-tokyo.ac.jp/ approximately shima/NETCOMP/.
Teppei Shimamura, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, Satoru Miyano
Bioinform.3
2009 Network-Based Predictions and Simulations by Biological State Space Models: Search for Drug Mode of Action
Rui Yamaguchi, Seiya Imoto, Satoru Miyano
J. Comput. Sci. Technol.1
2008 Statistical inference of transcriptional module-based gene networks from time course gene expression profiles by using state space models
abstract
MOTIVATION: Statistical inference of gene networks by using time-course microarray gene expression profiles is an essential step towards understanding the temporal structure of gene regulatory mechanisms. Unfortunately, most of the current studies have been limited to analysing a small number of genes because the length of time-course gene expression profiles is fairly short. One promising approach to overcome such a limitation is to infer gene networks by exploring the potential transcriptional modules which are sets of genes sharing a common function or involved in the same pathway. RESULTS: In this article, we present a novel approach based on the state space model to identify the transcriptional modules and module-based gene networks simultaneously. The state space model has the potential to infer large-scale gene networks, e.g. of order 10(3), from time-course gene expression profiles. Particularly, we succeeded in the identification of a cell cycle system by using the gene expression profiles of Saccharomyces cerevisiae in which the length of the time-course and number of genes were 24 and 4382, respectively. However, when analysing shorter time-course data, e.g. of length 10 or less, the parameter estimations of the state space model often fail due to overfitting. To extend the applicability of the state space model, we provide an approach to use the technical replicates of gene expression profiles, which are often measured in duplicate or triplicate. The use of technical replicates is important for achieving highly-efficient inferences of gene networks with short time-course data. The potential of the proposed method has been demonstrated through the time-course analysis of the gene expression profiles of human umbilical vein endothelial cells (HUVECs) undergoing growth factor deprivation-induced apoptosis. AVAILABILITY: Supplementary Information and the software (TRANS-MNET) are available at http://daweb.ism.ac.jp/~yoshidar/software/ssm/.
Osamu Hirose, Ryo Yoshida, Seiya Imoto, Rui Yamaguchi, Tomoyuki Higuchi, Stephen D. Charnock-Jones, Cristin G. Print, Satoru Miyano
Bioinform.4
2008 Bayesian learning of biological pathways on genomic data assimilation
abstract
MOTIVATION: Mathematical modeling and simulation, based on biochemical rate equations, provide us a rigorous tool for unraveling complex mechanisms of biological pathways. To proceed to simulation experiments, it is an essential first step to find effective values of model parameters, which are difficult to measure from in vivo and in vitro experiments. Furthermore, once a set of hypothetical models has been created, any statistical criterion is needed to test the ability of the constructed models and to proceed to model revision. RESULTS: The aim of our research is to present a new statistical technology towards data-driven construction of in silico biological pathways. The method starts with a knowledge-based modeling with hybrid functional Petri net. It then proceeds to the Bayesian learning of model parameters for which experimental data are available. This process exploits quantitative measurements of evolving biochemical reactions, e.g. gene expression data. Another important issue that we consider is statistical evaluation and comparison of the constructed hypothetical pathways. For this purpose, we have developed a new Bayesian information-theoretic measure that assesses the predictability and the biological robustness of in silico pathways. AVAILABILITY: The FORTRAN source codes are available at the URL http://daweb.ism.ac.jpyoshidar/GDA/ SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ryo Yoshida, Masao Nagasaki, Rui Yamaguchi, Seiya Imoto, Satoru Miyano, Tomoyuki Higuchi
Bioinform.3
2007 Statistical Absolute Evaluation of Gene Ontology Terms with Gene Expression Data
Pramod K. Gupta, Ryo Yoshida, Seiya Imoto, Rui Yamaguchi, Satoru Miyano
ISBRA4
2005 Estimating Gene Networks with cDNA Microarray Data Using State-Space Models
Rui Yamaguchi, Satoru Yamashita, Tomoyuki Higuchi
ICCSA (3)1