Seiya Imoto

dblp:91/5869 · DBLP profile ↗
← Back
71ranked-venue papers
0as first author
24since 2021 · last 2026
0000-0002-2989-308XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 62 · 24 since 2021Artificial intelligence and machine learning · 6Databases, data management, data science and information retrieval · 4
YearPublicationVenuePosition
2026 Comprehensive review and assessment of multi-species splicing variant prediction: task-specific deep learning models and genomic foundation models
abstract
Alternative splicing generates transcriptomic and proteomic diversity essential for eukaryotic complexity, yet genetic variants disrupting the splicing code underlie numerous human diseases. Deep learning (DL) models and genomic foundation models (GFMs) have achieved outstanding accuracy for predicting splicing variant effects in humans. However, their transferability to non-human species remains poorly understood, limiting applications in agricultural genomics, comparative biology, and non-model organism research, where experimentally validated variant datasets are limited or lacking. In this study, we comprehensively reviewed 35 computational approaches in terms of their architectural characteristics for splicing site and variant prediction and analysis. We systematically benchmarked the performance of 10 representative models for splicing variant prediction across human, rat, pig, and chicken, including four task-specific DL models and six GFMs, using our manually assembled benchmark datasets. Our benchmarking results revealed a substantial cross-species performance decrease (~21%-33% in the area under the receiver operating characteristic curve - AUROC) using task-specific models from human to non-human species datasets. We then applied a supervised adaptation to frozen GFM embeddings (DNABERT-2, Evo 2, Genos) by adding a lightweight classifier (i.e. a multi-layer perceptron) and reduced the cross-species performance decrease for rat and pig (8.56%-23.84% in AUROC), while performance on chicken was very close to human (decline within 1%, even exceeding by 0.52% when using the Evo 2 embedding). We proposed several directions to improve the prediction performance of splicing variants, including feature representation transfer and multi-modal fusion integrating global context, universal embeddings, and species-aware conditioning. We hope our comprehensive review and performance benchmarking can provide useful computational insights for further advancement of splicing variant prediction.
Yinuo Sun, Xiaoyu Wang 0016, Yuheng Jia, Seiya Imoto, Fuyi Li, Chen Li 0021, Jiangning Song
Briefings Bioinform.4
2026 MetaCCI: meta cell-cell interaction inference and its application to CCIs characteristics of MDS
abstract
MOTIVATION: Cell-cell interactions (CCIs) are fundamental to multicellular organisms and play crucial roles in diverse biological processes and disease mechanisms. Understanding CCIs is vital for deciphering disease pathogenesis and developing therapeutic strategies. Although numerous computational methods have been developed to infer CCIs from complex biological data, most existing approaches rely primarily on single-gene expression levels and ligand-receptor databases, often failing to capture the nuanced network-wide changes characteristic of disease states. RESULT: We propose MetaCCI, a novel computational strategy that integrates meta-information into CCI inference by extending the traditional gene expression-based analysis to a gene regulatory network framework. MetaCCI meticulously combines established ligand-receptor pairs with quantitative insights into gene behavior within complex gene networks, enabling the precise extraction of relevant targets for CCI inference. Subsequently, CCI inference was performed using an eigen cell co-expression network, providing a more holistic view of cell-cell communication. Monte Carlo simulations demonstrated that MetaCCI consistently outperforms existing methods in CCI inference. We applied MetaCCI to characterize cell-cell communication in Myelodysplastic Syndromes (MDS). Our results identified distinct interaction patterns in MDS compared with normal cell populations, specifically highlighting the loss of CCIs between "Dendritic cells and Hematopoietic precursor cells" and between "Dendritic cells and Hematopoietic multipotent progenitor cells" as characteristic features of MDS. Furthermore, FABP5, CD63, and HMGB1 were identified as MDS-specific markers. These findings suggest that diminished CCIs involving dendritic cells, hematopoietic precursor cells, and multipotent progenitor cells are pivotal to MDS pathogenesis. AVAILABILITY AND IMPLEMENTATION: The MetaCCI software is freely available at https://github.com/HeewonGitHub/MetaCCI. An archived version of the software and example datasets used in this study is available at Zenodo: https://doi.org/10.5281/zenodo.20101527.
Heewon Park, Seiya Imoto, Satoru Miyano
Bioinform.2
2025 STForte: tissue context-specific encoding and consistency-aware spatial imputation for spatially resolved transcriptomics
abstract
Encoding spatially resolved transcriptomics (SRT) data serves to identify the biological semantics of RNA expression within the tissue while preserving spatial characteristics. Depending on the analytical scenario, one may focus on different contextual structures of tissues. For instance, anatomical regions reveal consistent patterns by focusing on spatial homogeneity, while elucidating complex tumor micro-environments requires more expression heterogeneity. However, current spatial encoding methods lack consideration of the tissue context. Meanwhile, most developed SRT technologies are still limited in providing exact patterns of intact tissues due to limitations such as low resolution or missed measurements. Here, we propose STForte, a novel pairwise graph autoencoder-based approach with cross-reconstruction and adversarial distribution matching, to model the spatial homogeneity and expression heterogeneity of SRT data. STForte extracts interpretable latent encodings, enabling downstream analysis by accurately portraying various tissue contexts. Moreover, STForte allows spatial imputation using only spatial consistency to restore the biological patterns of unobserved locations or low-quality cells, thereby providing fine-grained views to enhance the SRT analysis. Extensive evaluations of datasets under different scenarios and SRT platforms demonstrate that STForte is a scalable and versatile tool for providing enhanced insights into spatial data analysis.
Yuxuan Pang, Chunxuan Wang, Seiya Imoto, Tzong-Yi Lee
Briefings Bioinform.5
2025 Powerful gene network enrichment analysis and its application to severe COVID-19 gene network
abstract
Understanding complex disease mechanisms requires research methods beyond individual gene analysis to capture the coordinated behavior of genes within regulatory networks. Traditional gene set enrichment approaches such as over-representation analysis and gene set enrichment analysis focus primarily on gene lists and often overlook the intricate network structures that control cellular processes. Although a gene network enrichment analysis strategy (GbNEA) has been proposed, this method assesses enrichment significance via phenotype permutation and the Kolmogorov-Smirnov test, which lowers statistical power and increases the computational burden due to repeated gene network re-estimation. To overcome these limitations, we developed a novel approach, powerful gene network enrichment analysis (PGNEA), which characterizes gene networks by integrating gene expression, regulatory effects, and hubness. PGNEA evaluates the enrichment of phenotype-specific gene networks by quantifying differences in gene activity patterns and assesses statistical significance by evaluating permutation of gene activity rather than phenotype permutation. This approach exhibits significantly enhanced computational efficiency and statistical sensitivity. We demonstrated the advantages of PGNEA through Monte Carlo simulations and applied it to whole-blood RNA-seq data obtained from the Japan COVID-19 Task Force. PGNEA successfully identified viral infection-related pathways enriched in severe COVID-19 gene networks, including those linked to "COVID-19," "HIV-1 infection," "Hepatitis B," "Influenza A," "Measles," and "Kaposi sarcoma-associated herpesvirus infection." Notably, key molecular markers such as PIK3, NF-B family members, FOXA, JUN, and CXCL8 were identified, with strong and consistent molecular interplays between CXCL8 and NFKBIA. These findings underscore the potential of PGNEA as an efficient tool for identifying biologically meaningful pathways and network-level mechanisms associated with various phenotypes, including severe viral infections.
Heewon Park, Seiya Imoto, Satoru Miyano
Briefings Bioinform.2
2025 PharaCon: a new framework for identifying bacteriophages via conditional representation learning
abstract
MOTIVATION: Identifying bacteriophages (phages) within metagenomic sequences is essential for understanding microbial community dynamics. Transformer-based foundation models have been successfully employed to address various biological challenges. However, these models are typically pre-trained with self-supervised tasks that do not consider label variance in the pre-training data. This presents a challenge for phage identification as pre-training on mixed bacterial and phage data may lead to information bias due to the imbalance between bacterial and phage samples. RESULTS: To overcome this limitation, we proposed a novel conditional BERT framework that incorporates label classes as special tokens during pre-training. Specifically, our conditional BERT model attaches labels directly during tokenization, introducing label constraints into the model's input. Additionally, we introduced a new fine-tuning scheme that enables the conditional BERT to be effectively utilized for classification tasks. This framework allows the BERT model to acquire label-specific contextual representations from mixed sequence data during pre-training and applies the conditional BERT as a classifier during fine-tuning, and we named the fine-tuned model as PharaCon. We evaluated PharaCon against several existing methods on both simulated sequence datasets and real metagenomic contig datasets. The results demonstrate PharaCon's effectiveness and efficiency in phage identification, highlighting the advantages of incorporating label information during both pre-training and fine-tuning. AVAILABILITY AND IMPLEMENTATION: The source code and associated data can be accessed at https://github.com/Celestial-Bai/PharaCon.
Zeheng Bai, Yuxuan Pang, Seiya Imoto
Bioinform.4
2025 Gene behaviors-based network enrichment analysis and its application to reveal immune disease pathways enriched with COVID-19 severity-specific gene networks
abstract
MOTIVATION: Gene network analysis is essential for understanding the complex mechanisms underlying diseases, which often involve disruptions in molecular networks rather than individual genes. Despite the availability of large-scale omics datasets and computational tools for gene network analysis, interpretation of the biological relevance of these extensive networks remains challenging. RESULTS: We propose a novel computational strategy, gene behaviors-based network enrichment analysis, which systematically identifies functional pathways enriched in phenotype-specific gene networks. Our novel method incorporates comprehensive network characteristics, i.e. gene expression levels, edge strengths, and structural patterns of edges, to rank genes based on activity and assess pathway enrichment, effectively identifying functional pathways enriched within these networks. Through simulation studies, our strategy demonstrated superior performance compared with that of existing methods in identifying enriched pathways. We applied this strategy to whole-blood RNA-seq data from 1102 COVID-19 samples provided by the Japan COVID-19 Task Force. The analysis revealed immune disease pathways enriched with COVID-19 severity-specific gene networks, including "Systemic lupus erythematosus" in asymptomatic and severe samples and "Inflammatory bowel disease," "Primary immunodeficiency," and "Rheumatoid arthritis" in mild samples. Key biomarkers of COVID-19, such as CXCL8, S100A9, and HLA class I genes, have been identified as critical hub genes and the main players within these networks. AVAILABILITY AND IMPLEMENTATION: Code is available in Figshare (https://doi.org/10.6084/m9.figshare.29093648.v3).
Heewon Park, Seiya Imoto, Satory Miyano
Bioinform.2
2025 Histo-Genomic Knowledge Association for Cancer Prognosis From Histopathology Whole Slide Images
abstract
Histo-genomic multi-modal methods have emerged as a powerful paradigm, demonstrating significant potential for cancer prognosis. However, genome sequencing, unlike histopathology imaging, is still not widely accessible in underdeveloped regions, limiting the application of these multi-modal approaches in clinical settings. To address this, we propose a novel Genome-informed Hyper-Attention Network, termed G-HANet, which is capable of effectively learning the histo-genomic associations during training to elevate uni-modal whole slide image (WSI)-based inference for the first time. Compared with the potential knowledge distillation strategy for this setting (i.e., distilling a multi-modal network to a uni-modal network), our end-to-end model is superior in training efficiency and learning cross-modal interactions. Specifically, the network comprises cross-modal associating branch (CAB) and hyper-attention survival branch (HSB). Through the genomic data reconstruction from WSIs, CAB effectively distills the associations between functional genotypes and morphological phenotypes and offers insights into the gene expression profiles in the feature space. Subsequently, HSB leverages the distilled histo-genomic associations as well as the generated morphology-based weights to achieve the hyper-attention modeling of the patients from both histopathology and genomic perspectives to improve cancer prognosis. Extensive experiments are conducted on five TCGA benchmarking datasets and the results demonstrate that G-HANet significantly outperforms the state-of-the-art WSI-based methods and achieves competitive performance with genome-based and multi-modal methods. G-HANet is expected to be explored as a useful tool by the research community to address the current bottleneck of insufficient histo-genomic data pairing in the context of cancer prognosis and precision oncology. The code is available at https://github.com/ZacharyWang-007/G-HANet.
Zhikang Wang, Yingxue Xu, Seiya Imoto, Hao Chen 0011, Jiangning Song
IEEE Trans. Medical Imaging4
2024 Deep learning approaches for non-coding genetic variant effect prediction: current progress and future prospects
abstract
Recent advancements in high-throughput sequencing technologies have significantly enhanced our ability to unravel the intricacies of gene regulatory processes. A critical challenge in this endeavor is the identification of variant effects, a key factor in comprehending the mechanisms underlying gene regulation. Non-coding variants, constituting over 90% of all variants, have garnered increasing attention in recent years. The exploration of gene variant impacts and regulatory mechanisms has spurred the development of various deep learning approaches, providing new insights into the global regulatory landscape through the analysis of extensive genetic data. Here, we provide a comprehensive overview of the development of the non-coding variants models based on bulk and single-cell sequencing data and their model-based interpretation and downstream tasks. This review delineates the popular sequencing technologies for epigenetic profiling and deep learning approaches for discerning the effects of non-coding variants. Additionally, we summarize the limitations of current approaches in variant effect prediction research and outline opportunities for improvement. We anticipate that our study will offer a practical and useful guide for the bioinformatic community to further advance the unraveling of genetic variant effects.
Xiaoyu Wang 0016, Fuyi Li, Seiya Imoto, Hsin-Hui Shen, Shanshan Li 0008, Yuming Guo 0001, Jian Yang 0005, Jiangning Song
Briefings Bioinform.4
2024 biotextgraph: graphical summarization of functional similarities from textual information
abstract
SUMMARY: Functional interpretation of biological entities such as differentially expressed genes is one of the fundamental analyses in bioinformatics. The task can be addressed by using biological pathway databases with enrichment analysis (EA). However, textual description of biological entities in public databases is less explored and integrated in existing tools and it has a potential to reveal new mechanisms. Here, we present a new R package biotextgraph for graphical summarization of omics' textual description data which enables assessment of functional similarities of the lists of biological entities. We illustrate application examples of annotating gene identifiers in addition to EA. The results suggest that the visualization based on words and inspection of biological entities with text can reveal a set of biologically meaningful terms that could not be obtained by using biological pathway databases alone. The results suggest the usefulness of the package in the routine analysis of omics-related data. The package also offers a web-based application for convenient querying. AVAILABILITY AND IMPLEMENTATION: The package, documentation, and web server are available at: https://github.com/noriakis/biotextgraph.
Noriaki Sato, Zuguang Gu, Seiya Imoto
Bioinform.4
2024 Dual-stream multi-dependency graph neural network enables precise cancer survival analysis
abstract
Histopathology image-based survival prediction aims to provide a precise assessment of cancer prognosis and can inform personalized treatment decision-making in order to improve patient outcomes. However, existing methods cannot automatically model the complex correlations between numerous morphologically diverse patches in each whole slide image (WSI), thereby preventing them from achieving a more profound understanding and inference of the patient status. To address this, here we propose a novel deep learning framework, termed dual-stream multi-dependency graph neural network (DM-GNN), to enable precise cancer patient survival analysis. Specifically, DM-GNN is structured with the feature updating and global analysis branches to better model each WSI as two graphs based on morphological affinity and global co-activating dependencies. As these two dependencies depict each WSI from distinct but complementary perspectives, the two designed branches of DM-GNN can jointly achieve the multi-view modeling of complex correlations between the patches. Moreover, DM-GNN is also capable of boosting the utilization of dependency information during graph construction by introducing the affinity-guided attention recalibration module as the readout function. This novel module offers increased robustness against feature perturbation, thereby ensuring more reliable and stable predictions. Extensive benchmarking experiments on five TCGA datasets demonstrate that DM-GNN outperforms other state-of-the-art methods and offers interpretable prediction insights based on the morphological depiction of high-attention patches. Overall, DM-GNN represents a powerful and auxiliary tool for personalized cancer prognosis from histopathology images and has great potential to assist clinicians in making personalized treatment decisions and improving patient outcomes.
Zhikang Wang, Jiani Ma, Chris Bain, Seiya Imoto, Pietro Liò, Hongmin Cai, Hao Chen 0011, Jiangning Song
Medical Image Anal.5
2023 Imputing time-series microbiome abundance profiles with diffusion model
abstract
With the advancements in genome sequencing technologies, the understanding of gut microbiome has been expanded substantially. Recently, the study of time-series microbiome has garnered increased attention, aiming to elucidate the temporal shifts in microbiome and their subsequent implications for human health. However, time-series microbiome data often contain missing time points, complicating its analysis. In this study, we propose a novel method of imputation using a diffusion model to solve this challenging problem. We train a diffusion model on both the 16S rRNA and WGS sequencing data and impute the missing values within relative abundance profiles at selected time points. Compared with the previous approaches, the proposed method can achieve a lower mean absolute error of imputed values across various missing ratios, especially in 16S rRNA data. In addition, in a downstream predictive task, the performance on specific outcomes can be improved using the imputed values by the diffusion model. Our method provides an effective and stable solution for processing time-series microbiome data with a substantial proportion of missing values.
Misato Seki, Seiya Imoto
BIBM3
2023 iAMPCN: a deep-learning approach for identifying antimicrobial peptides and their functional activities
abstract
Antimicrobial peptides (AMPs) are short peptides that play crucial roles in diverse biological processes and have various functional activities against target organisms. Due to the abuse of chemical antibiotics and microbial pathogens' increasing resistance to antibiotics, AMPs have the potential to be alternatives to antibiotics. As such, the identification of AMPs has become a widely discussed topic. A variety of computational approaches have been developed to identify AMPs based on machine learning algorithms. However, most of them are not capable of predicting the functional activities of AMPs, and those predictors that can specify activities only focus on a few of them. In this study, we first surveyed 10 predictors that can identify AMPs and their functional activities in terms of the features they employed and the algorithms they utilized. Then, we constructed comprehensive AMP datasets and proposed a new deep learning-based framework, iAMPCN (identification of AMPs based on CNNs), to identify AMPs and their related 22 functional activities. Our experiments demonstrate that iAMPCN significantly improved the prediction performance of AMPs and their corresponding functional activities based on four types of sequence features. Benchmarking experiments on the independent test datasets showed that iAMPCN outperformed a number of state-of-the-art approaches for predicting AMPs and their functional activities. Furthermore, we analyzed the amino acid preferences of different AMP activities and evaluated the model on datasets of varying sequence redundancy thresholds. To facilitate the community-wide identification of AMPs and their corresponding functional types, we have made the source codes of iAMPCN publicly available at https://github.com/joy50706/iAMPCN/tree/master. We anticipate that iAMPCN can be explored as a valuable tool for identifying potential AMPs with specific functional activities for further experimental validation.
Jing Xu 0008, Fuyi Li, Chen Li 0021, Cornelia B. Landersdorfer, Hsin-Hui Shen, Anton Y. Peleg, Jian Li 0052, Seiya Imoto, Jianhua Yao 0001, Tatsuya Akutsu, Jiangning Song
Briefings Bioinform.9
2023 Zero-shot-capable identification of phage-host relationships with whole-genome sequence representation by contrastive learning
abstract
Accurately identifying phage-host relationships from their genome sequences is still challenging, especially for those phages and hosts with less homologous sequences. In this work, focusing on identifying the phage-host relationships at the species and genus level, we propose a contrastive learning based approach to learn whole-genome sequence embeddings that can take account of phage-host interactions (PHIs). Contrastive learning is used to make phages infecting the same hosts close to each other in the new representation space. Specifically, we rephrase whole-genome sequences with frequency chaos game representation (FCGR) and learn latent embeddings that 'encapsulate' phages and host relationships through contrastive learning. The contrastive learning method works well on the imbalanced dataset. Based on the learned embeddings, a proposed pipeline named CL4PHI can predict known hosts and unseen hosts in training. We compare our method with two recently proposed state-of-the-art learning-based methods on their benchmark datasets. The experiment results demonstrate that the proposed method using contrastive learning improves the prediction accuracy on known hosts and demonstrates a zero-shot prediction capability on unseen hosts. In terms of potential applications, the rapid pace of genome sequencing across different species has resulted in a vast amount of whole-genome sequencing data that require efficient computational methods for identifying phage-host interactions. The proposed approach is expected to address this need by efficiently processing whole-genome sequences of phages and prokaryotic hosts and capturing features related to phage-host relationships for genome sequence representation. This approach can be used to accelerate the discovery of phage-host interactions and aid in the development of phage-based therapies for infectious diseases.
Zeheng Bai, Kosuke Fujimoto, Satoshi Uematsu, Seiya Imoto
Briefings Bioinform.6
2023 PFresGO: an attention mechanism-based deep-learning approach for protein annotation by integrating gene ontology inter-relationships
abstract
MOTIVATION: The rapid accumulation of high-throughput sequence data demands the development of effective and efficient data-driven computational methods to functionally annotate proteins. However, most current approaches used for functional annotation simply focus on the use of protein-level information but ignore inter-relationships among annotations. RESULTS: Here, we established PFresGO, an attention-based deep-learning approach that incorporates hierarchical structures in Gene Ontology (GO) graphs and advances in natural language processing algorithms for the functional annotation of proteins. PFresGO employs a self-attention operation to capture the inter-relationships of GO terms, updates its embedding accordingly and uses a cross-attention operation to project protein representations and GO embedding into a common latent space to identify global protein sequence patterns and local functional residues. We demonstrate that PFresGO consistently achieves superior performance across GO categories when compared with 'state-of-the-art' methods. Importantly, we show that PFresGO can identify functionally important residues in protein sequences by assessing the distribution of attention weightings. PFresGO should serve as an effective tool for the accurate functional annotation of proteins and functional domains within proteins. AVAILABILITY AND IMPLEMENTATION: PFresGO is available for academic purposes at https://github.com/BioColLab/PFresGO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tong Pan, Chen Li 0021, Yue Bi, Zhikang Wang, Robin B. Gasser, Anthony W. Purcell, Tatsuya Akutsu, Geoffrey I. Webb, Seiya Imoto, Jiangning Song
Bioinform.9
2023 ggkegg: analysis and visualization of KEGG data utilizing the grammar of graphics
abstract
SUMMARY: The Kyoto Encyclopedia of Genes and Genomes (KEGG) database serves as a valuable systems biology resource and is widely utilized in diverse research fields. However, existing software does not allow flexible visualization and network analyses of the vast and complex KEGG data. We developed ggkegg, an R package that integrates KEGG information with ggplot2 and ggraph. ggkegg enables enhanced visualization and network analyses of KEGG data. We demonstrate the utility of the package by providing examples of its application in single-cell, bulk transcriptome, and microbiome analyses. ggkegg may empower researchers to analyze complex biological networks and present their results effectively. AVAILABILITY AND IMPLEMENTATION: The package and user documentation are available at: https://github.com/noriakis/ggkegg.
Noriaki Sato, Miho Uematsu, Kosuke Fujimoto, Satoshi Uematsu, Seiya Imoto
Bioinform.5
2023 Targeting tumor heterogeneity: multiplex-detection-based multiple instance learning for whole slide image classification
abstract
MOTIVATION: Multiple instance learning (MIL) is a powerful technique to classify whole slide images (WSIs) for diagnostic pathology. The key challenge of MIL on WSI classification is to discover the critical instances that trigger the bag label. However, tumor heterogeneity significantly hinders the algorithm's performance. RESULTS: Here, we propose a novel multiplex-detection-based multiple instance learning (MDMIL) which targets tumor heterogeneity by multiplex detection strategy and feature constraints among samples. Specifically, the internal query generated after the probability distribution analysis and the variational query optimized throughout the training process are utilized to detect potential instances in the form of internal and external assistance, respectively. The multiplex detection strategy significantly improves the instance-mining capacity of the deep neural network. Meanwhile, a memory-based contrastive loss is proposed to reach consistency on various phenotypes in the feature space. The novel network and loss function jointly achieve high robustness towards tumor heterogeneity. We conduct experiments on three computational pathology datasets, e.g. CAMELYON16, TCGA-NSCLC, and TCGA-RCC. Benchmarking experiments on the three datasets illustrate that our proposed MDMIL approach achieves superior performance over several existing state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: MDMIL is available for academic purposes at https://github.com/ZacharyWang-007/MDMIL.
Zhikang Wang, Yue Bi, Tong Pan, Xiaoyu Wang 0016, Chris Bain, Richard Bassed, Seiya Imoto, Jianhua Yao 0001, Roger J. Daly, Jiangning Song
Bioinform.7
2023 Investigation of the BERT model on nucleotide sequences with non-standard pre-training and evaluation of different k-mer embeddings
abstract
MOTIVATION: In recent years, pre-training with the transformer architecture has gained significant attention. While this approach has led to notable performance improvements across a variety of downstream tasks, the underlying mechanisms by which pre-training models influence these tasks, particularly in the context of biological data, are not yet fully elucidated. RESULTS: In this study, focusing on the pre-training on nucleotide sequences, we decompose a pre-training model of Bidirectional Encoder Representations from Transformers (BERT) into its embedding and encoding modules to analyze what a pre-trained model learns from nucleotide sequences. Through a comparative study of non-standard pre-training at both the data and model levels, we find that a typical BERT model learns to capture overlapping-consistent k-mer embeddings for its token representation within its embedding module. Interestingly, using the k-mer embeddings pre-trained on random data can yield similar performance in downstream tasks, when compared with those using the k-mer embeddings pre-trained on real biological sequences. We further compare the learned k-mer embeddings with other established k-mer representations in downstream tasks of sequence-based functional prediction. Our experimental results demonstrate that the dense representation of k-mers learned from pre-training can be used as a viable alternative to one-hot encoding for representing nucleotide sequences. Furthermore, integrating the pre-trained k-mer embeddings with simpler models can achieve competitive performance in two typical downstream tasks. AVAILABILITY AND IMPLEMENTATION: The source code and associated data can be accessed at https://github.com/yaozhong/bert_investigation.
Zeheng Bai, Seiya Imoto
Bioinform.3
2022 Identification of bacteriophage genome sequences with representation learning
abstract
MOTIVATION: Bacteriophages/phages are the viruses that infect and replicate within bacteria and archaea, and rich in human body. To investigate the relationship between phages and microbial communities, the identification of phages from metagenome sequences is the first step. Currently, there are two main methods for identifying phages: database-based (alignment-based) methods and alignment-free methods. Database-based methods typically use a large number of sequences as references; alignment-free methods usually learn the features of the sequences with machine learning and deep learning models. RESULTS: We propose INHERIT which uses a deep representation learning model to integrate both database-based and alignment-free methods, combining the strengths of both. Pre-training is used as an alternative way of acquiring knowledge representations from existing databases, while the BERT-style deep learning framework retains the advantage of alignment-free methods. We compare INHERIT with four existing methods on a third-party benchmark dataset. Our experiments show that INHERIT achieves a better performance with the F1-score of 0.9932. In addition, we find that pre-training two species separately helps the non-alignment deep learning model make more accurate predictions. AVAILABILITY AND IMPLEMENTATION: The codes of INHERIT are now available in: https://github.com/Celestial-Bai/INHERIT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zeheng Bai, Satoru Miyano, Rui Yamaguchi, Kosuke Fujimoto, Satoshi Uematsu, Seiya Imoto
Bioinform.7
2022 PredictiveNetwork: predictive gene network estimation with application to gastric cancer drug response-predictive network analysis
abstract
BACKGROUND: Gene regulatory networks have garnered a large amount of attention to understand disease mechanisms caused by complex molecular network interactions. These networks have been applied to predict specific clinical characteristics, e.g., cancer, pathogenicity, and anti-cancer drug sensitivity. However, in most previous studies using network-based prediction, the gene networks were estimated first, and predicted clinical characteristics based on pre-estimated networks. Thus, the estimated networks cannot describe clinical characteristic-specific gene regulatory systems. Furthermore, existing computational methods were developed from algorithmic and mathematics viewpoints, without considering network biology. RESULTS: To effectively predict clinical characteristics and estimate gene networks that provide critical insights into understanding the biological mechanisms involved in a clinical characteristic, we propose a novel strategy for predictive gene network estimation. The proposed strategy simultaneously performs gene network estimation and prediction of the clinical characteristic. In this strategy, the gene network is estimated with minimal network estimation and prediction errors. We incorporate network biology by assuming that neighboring genes in a network have similar biological functions, while hub genes play key roles in biological processes. Thus, the proposed method provides interpretable prediction results and enables us to uncover biologically reliable marker identification. Monte Carlo simulations shows the effectiveness of our method for feature selection in gene estimation and prediction with excellent prediction accuracy. We applied the proposed strategy to construct gastric cancer drug-responsive networks. CONCLUSION: We identified gastric drug response predictive markers and drug sensitivity/resistance-specific markers, AKR1B10, AKR1C3, ANXA10, and ZNF165, based on GDSC data analysis. Our results for identifying drug sensitive and resistant specific molecular interplay are strongly supported by previous studies. We expect that the proposed strategy will be a useful tool for uncovering crucial molecular interactions involved a specific biological mechanism, such as cancer progression or acquired drug resistance.
Heewon Park, Seiya Imoto, Satoru Miyano
BMC Bioinform.2
2021 Discovering microbe functionality in human disease with a gene-ontology-aware model
abstract
With the development of sequencing technology, a large amount of metagenome-wide association studies have been designed. To find associations between the microbiome and human disease, many microbiome features including microbiome relative abundance and Gene Ontology (GO) are utilized. However, few studies describe the microbe functionality in human diseases due to the complexity of summarizing a large amount of microbe functionality. To investigate the functional roles of microbiome in human diseases, here we constructed a feature selecting model utilizing microbiome relative abundance and GO. We proposed two GO features extraction methods utilizing GO hierarchical structure information and evaluate the effectiveness of these GO features in type II diabetes and liver cirrhosis using typical machine learning models including logistic regression, support vector machine, random forest, and Naive Bayes. Furthermore, we design a method to summarize the crucial microbe functionality in human diseases by constructing the relation network between important GO terms and microbiome species extracted by our feature selecting model.
Seiya Imoto
BIBM3
2021 On the application of BERT models for nanopore methylation detection
abstract
DNA methylation is a common nucleotide modification, which is associated with various biological processes, such as gene expression and aging. Nanopore sequencing provides a direct detecting approach through searching specific current signal shifts. Recently, model-based approaches, especially those using deep learning models, have achieved significant performance improvements on nanopore methylation detection. In this work, we explore using the non-recurrent neural network structure of Bidirectional Encoder Representations from Transformers (BERT) for the task, which provides an alternative fast inference model to the state-of-the-art bi-directional Recurrent Neural Network (biRNN). In addition, we propose a refined BERT model with relative position representation and center hidden units concatenation, which takes account of the task-specific characters into modeling. We evaluate the proposed models on the R9 benchmark datasets of different motifs and methyltransferases. The experiment results show that the refined BERT model can achieve competitive or even better results than the state-of-the-art biRNN model, while the model inference speed is faster.
Kiyoshi Yamaguchi, Sera Hatakeyama, Yoichi Furukawa, Satoru Miyano, Rui Yamaguchi, Seiya Imoto
BIBM7
2021 Halcyon: an accurate basecaller exploiting an encoder-decoder model with monotonic attention
abstract
MOTIVATION: In recent years, nanopore sequencing technology has enabled inexpensive long-read sequencing, which promises reads longer than a few thousand bases. Such long-read sequences contribute to the precise detection of structural variations and accurate haplotype phasing. However, deciphering precise DNA sequences from noisy and complicated nanopore raw signals remains a crucial demand for downstream analyses based on higher-quality nanopore sequencing, although various basecallers have been introduced to date. RESULTS: To address this need, we developed a novel basecaller, Halcyon, that incorporates neural-network techniques frequently used in the field of machine translation. Our model employs monotonic-attention mechanisms to learn semantic correspondences between nucleotides and signal levels without any pre-segmentation against input signals. We evaluated performance with a human whole-genome sequencing dataset and demonstrated that Halcyon outperformed existing third-party basecallers and achieved competitive performance against the latest Oxford Nanopore Technologies' basecallers. AVAILABILITYAND IMPLEMENTATION: The source code (halcyon) can be found at https://github.com/relastle/halcyon.
Hiroki Konishi, Rui Yamaguchi, Kiyoshi Yamaguchi, Yoichi Furukawa, Seiya Imoto
Bioinform.5
2021 HEAL: an automated deep learning framework for cancer histopathology image analysis
abstract
MOTIVATION: Digital pathology supports analysis of histopathological images using deep learning methods at a large-scale. However, applications of deep learning in this area have been limited by the complexities of configuration of the computational environment and of hyperparameter optimization, which hinder deployment and reduce reproducibility. RESULTS: Here, we propose HEAL, a deep learning-based automated framework for easy, flexible and multi-faceted histopathological image analysis. We demonstrate its utility and functionality by performing two case studies on lung cancer and one on colon cancer. Leveraging the capability of Docker, HEAL represents an ideal end-to-end tool to conduct complex histopathological analysis and enables deep learning in a broad range of applications for cancer image analysis. AVAILABILITY AND IMPLEMENTATION: The docker image of HEAL is available at https://hub.docker.com/r/docurdt/heal and related documentation and datasets are available at http://heal.erc.monash.edu.au. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yanan Wang 0003, Nicolas Coudray, Yun Zhao 0004, Fuyi Li, Changyuan Hu, Seiya Imoto, Aristotelis Tsirigos, Geoffrey I. Webb, Roger J. Daly, Jiangning Song
Bioinform.7
2021 Enhancing breakpoint resolution with deep segmentation model: A general refinement method for read-depth based structural variant callers
abstract
Read-depths (RDs) are frequently used in identifying structural variants (SVs) from sequencing data. For existing RD-based SV callers, it is difficult for them to determine breakpoints in single-nucleotide resolution due to the noisiness of RD data and the bin-based calculation. In this paper, we propose to use the deep segmentation model UNet to learn base-wise RD patterns surrounding breakpoints of known SVs. We integrate model predictions with an RD-based SV caller to enhance breakpoints in single-nucleotide resolution. We show that UNet can be trained with a small amount of data and can be applied both in-sample and cross-sample. An enhancement pipeline named RDBKE significantly increases the number of SVs with more precise breakpoints on simulated and real data. The source code of RDBKE is freely available at https://github.com/yaozhong/deepIntraSV.
Seiya Imoto, Satoru Miyano, Rui Yamaguchi
PLoS Comput. Biol.2
2020 Neoantimon: a multifunctional R package for identification of tumor-specific neoantigens
abstract
SUMMARY: It is known that some mutant peptides, such as those resulting from missense mutations and frameshift insertions, can bind to the major histocompatibility complex and be presented to antitumor T cells on the surface of a tumor cell. These peptides are termed neoantigen, and it is important to understand this process for cancer immunotherapy. Here, we introduce an R package termed Neoantimon that can predict a list of potential neoantigens from a variety of mutations, which include not only somatic point mutations but insertions, deletions and structural variants. Beyond the existing applications, Neoantimon is capable of attaching and reflecting several additional information, e.g. wild-type binding capability, allele specific RNA expression levels, single nucleotide polymorphism information and combinations of mutations to filter out infeasible peptides as neoantigen. AVAILABILITY AND IMPLEMENTATION: The R package is available at http://github/hase62/Neoantimon.
Takanori Hasegawa, Shuto Hayashi, Eigo Shimizu, Shinichi Mizuno, Atsushi Niida, Rui Yamaguchi, Satoru Miyano, Hidewaki Nakagawa, Seiya Imoto
Bioinform.9
2020 Nanopore basecalling from a perspective of instance segmentation
abstract
BACKGROUND: Nanopore sequencing is a rapidly developing third-generation sequencing technology, which can generate long nucleotide reads of molecules within a portable device in real-time. Through detecting the change of ion currency signals during a DNA/RNA fragment's pass through a nanopore, genotypes are determined. Currently, the accuracy of nanopore basecalling has a higher error rate than the basecalling of short-read sequencing. Through utilizing deep neural networks, the-state-of-the art nanopore basecallers achieve basecalling accuracy in a range from 85% to 95%. RESULT: In this work, we proposed a novel basecalling approach from a perspective of instance segmentation. Different from previous approaches of doing typical sequence labeling, we formulated the basecalling problem as a multi-label segmentation task. Meanwhile, we proposed a refined U-net model which we call UR-net that can model sequential dependencies for a one-dimensional segmentation task. The experiment results show that the proposed basecaller URnano achieves competitive results on the in-species data, compared to the recently proposed CTC-featured basecallers. CONCLUSION: Our results show that formulating the basecalling problem as a one-dimensional segmentation task is a promising approach, which does basecalling and segmentation jointly.
Arda Akdemir, Georg Tremmel, Seiya Imoto, Satoru Miyano, Tetsuo Shibuya, Rui Yamaguchi
BMC Bioinform.4
2019 A Bayesian model integration for mutation calling through data partitioning
abstract
MOTIVATION: Detection of somatic mutations from tumor and matched normal sequencing data has become among the most important analysis methods in cancer research. Some existing mutation callers have focused on additional information, e.g. heterozygous single-nucleotide polymorphisms (SNPs) nearby mutation candidates or overlapping paired-end read information. However, existing methods cannot take multiple information sources into account simultaneously. Existing Bayesian hierarchical model-based methods construct two generative models, the tumor model and error model, and limited information sources have been modeled. RESULTS: We proposed a Bayesian model integration framework named as partitioning-based model integration. In this framework, through introducing partitions for paired-end reads based on given information sources, we integrate existing generative models and utilize multiple information sources. Based on that, we constructed a novel Bayesian hierarchical model-based method named as OHVarfinDer. In both the tumor model and error model, we introduced partitions for a set of paired-end reads that cover a mutation candidate position, and applied a different generative model for each category of paired-end reads. We demonstrated that our method can utilize both heterozygous SNP information and overlapping paired-end read information effectively in simulation datasets and real datasets. AVAILABILITY AND IMPLEMENTATION: https://github.com/takumorizo/OHVarfinDer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Takuya Moriyama, Seiya Imoto, Shuto Hayashi, Yuichi Shiraishi, Satoru Miyano, Rui Yamaguchi
Bioinform.2
2019 Capturing the differences between humoral immunity in the normal and tumor environments from repertoire-seq of B-cell receptors using supervised machine learning
abstract
BACKGROUND: The recent success of immunotherapy in treating tumors has attracted increasing interest in research related to the adaptive immune system in the tumor microenvironment. Recent advances in next-generation sequencing technology enabled the sequencing of whole T-cell receptors (TCRs) and B-cell receptors (BCRs)/immunoglobulins (Igs) in the tumor microenvironment. Since BCRs/Igs in tumor tissues have high affinities for tumor-specific antigens, the patterns of their amino acid sequences and other sequence-independent features such as the number of somatic hypermutations (SHMs) may differ between the normal and tumor microenvironments. However, given the high diversity of BCRs/Igs and the rarity of recurrent sequences among individuals, it is far more difficult to capture such differences in BCR/Ig sequences than in TCR sequences. The aim of this study was to explore the possibility of discriminating BCRs/Igs in tumor and in normal tissues, by capturing these differences using supervised machine learning methods applied to RNA sequences of BCRs/Igs. RESULTS: RNA sequences of BCRs/Igs were obtained from matched normal and tumor specimens from 90 gastric cancer patients. BCR/Ig-features obtained in Rep-Seq were used to classify individual BCR/Ig sequences into normal or tumor classes. Different machine learning models using various features were constructed as well as gradient boosting machine (GBM) classifier combining these models. The results demonstrated that BCR/Ig sequences between normal and tumor microenvironments exhibit their differences. Next, by using a GBM trained to classify individual BCR/Ig sequences, we tried to classify sets of BCR/Ig sequences into normal or tumor classes. As a result, an area under the curve (AUC) value of 0.826 was achieved, suggesting that BCR/Ig repertoires have distinct sequence-level features in normal and tumor tissues. CONCLUSIONS: To the best of our knowledge, this is the first study to show that BCR/Ig sequences derived from tumor and normal tissues have globally distinct patterns, and that these tissues can be effectively differentiated using BCR/Ig repertoires.
Hiroki Konishi, Daisuke Komura, Hiroto Katoh, Shinichiro Atsumi, Hirotomo Koda, Asami Yamamoto, Yasuyuki Seto, Masashi Fukayama, Rui Yamaguchi, Seiya Imoto, Shumpei Ishikawa
BMC Bioinform.10
2017 Reconstruction of high read-depth signals from low-depth whole genome sequencing data using deep learning
abstract
Motivation: Next-generation sequencing (NGS) technologies using DNA, RNA, or methylation sequencing are prevailing tools used in modern genome research. For DNA sequencing, whole genome sequencing (WGS) and whole exome sequencing (WES) are two typical applications with a different preference on the trade-off between sequencing depth and base coverage. Although sequencing costs have been greatly reduced, the sequence depth used in WGS is relatively lower than WES (e.g., ~35× vs. 100×~). In addition, biases and batch effects may exist in different stages of a NGS experiment. Using low-depth and biased WGS data for downstream analyses is more sensitive to the bias problem and makes it even more difficult to uncover real biological signals in the data. In this work, we focused on reconstructing high read-depth signals from low-depth WGS data. We make use of a pair of WGS data with different read-depth for the same sample and learn a mapping from low-depth signals to high-depth in the given platform. Results: We explored three different reconstruction models from shallow to deep. Our experimental results show that by only using the read depth information, deeper models do not perform far better than a linear regression model. Through incorporating additional information, such as GC-content, mappability and nucleotide sequence information, the performance of convolutional neural network (CNN) models can be further improved. We made use of the reconstructed read-depth signals in downstream analysis to identify copy number variation segments for single sample. The experiment results show that segments that are not detected using low-depth data, can be detected with the reconstructed signals by the CNN model using extra biological information.
Seiya Imoto, Satoru Miyano, Rui Yamaguchi
BIBM2
2017 A Novel Adaptive Penalized Logistic Regression for Uncovering Biomarker Associated with Anti-Cancer Drug Sensitivity
abstract
We propose a novel adaptive penalized logistic regression modeling strategy based on Wilcoxon rank sum test (WRST) to effectively uncover driver genes in classification. In order to incorporate significance of gene in classification, we first measure significance of each gene by gene ranking method based on WRST, and then the adaptive L1-type penalty is discriminately imposed on each gene depending on the measured importance degree of gene. The incorporating significance of genes into adaptive logistic regression enables us to impose a large amount of penalty on low ranking genes, and thus noise genes are easily deleted from the model and we can effectively identify driver genes. Monte Carlo experiments and real world example are conducted to investigate effectiveness of the proposed approach. In Sanger data analysis, we introduce a strategy to identify expression modules indicating gene regulatory mechanisms via the principal component analysis (PCA), and perform logistic regression modeling based on not a single gene but gene expression modules. We can see through Monte Carlo experiments and real world example that the proposed adaptive penalized logistic regression outperforms feature selection and classification compared with existing L1-type regularization. The discriminately imposed penalty based on WRST effectively performs crucial gene selection, and thus our method can improve classification accuracy without interruption of noise genes. Furthermore, it can be seen through Sanger data analysis that the method for gene expression modules based on principal components and their loading scores provides interpretable results in biological viewpoints.
Heewon Park, Yuichi Shiraishi, Seiya Imoto, Satoru Miyano
IEEE ACM Trans. Comput. Biol. Bioinform.3
2016 OVarCall: Bayesian Mutation Calling Method Utilizing Overlapping Paired-End Reads
Takuya Moriyama, Yuichi Shiraishi, Kenichi Chiba, Rui Yamaguchi, Seiya Imoto, Satoru Miyano
ISBRA5
2015 Binary Contingency Table Method for Analyzing Gene Mutation in Cancer Genome
Emi Ayada, Atsushi Niida, Takanori Hasegawa, Satoru Miyano, Seiya Imoto
ISBRA5
2015 Genomon ITDetector: a tool for somatic internal tandem duplication detection from cancer genome sequencing data
abstract
SUMMARY: Somatic internal tandem duplications (ITDs) are known to play important roles in cancer pathogenesis. Although recent advances in high-throughput sequencing technologies have enabled genome-wide detection of various types of genomic mutations, including single nucleotide variants, indels and structural variations, only a few studies have focused on ITDs. We have developed an analytical tool called 'Genomon ITDetector' for genome-wide detection of somatic ITDs. After evaluating the sensitivity and precision of the proposed approach using synthetic data, we have demonstrated that it can successfully detect not only common ITDs involving FLT3, but also a number of ITDs affecting other putative driver genes in acute myeloid leukemia exome sequencing data. Availability and implementaion: Genomon ITDetector is freely available at https://github.com/ken0-1n/Genomon-ITDetector.
Kenichi Chiba, Yuichi Shiraishi, Yasunobu Nagata, Seiya Imoto, Seishi Ogawa, Satoru Miyano
Bioinform.5
2014 Parameter estimation in multi-compartment SIR model
Masaya M. Saito, Seiya Imoto, Rui Yamaguchi, Satoru Miyano, Tomoyuki Higuchi
FUSION2
2014 HapMuC: somatic mutation calling using heterozygous germ line variants near candidate mutations
abstract
MOTIVATION: Identifying somatic changes from tumor and matched normal sequences has become a standard approach in cancer research. More specifically, this requires accurate detection of somatic point mutations with low allele frequencies in impure and heterogeneous cancer samples. Although haplotype phasing information derived by using heterozygous germ line variants near candidate mutations would improve accuracy, no somatic mutation caller that uses such information is currently available. RESULTS: We propose a Bayesian hierarchical method, termed HapMuC, in which power is increased by using available information on heterozygous germ line variants located near candidate mutations. We first constructed two generative models (the mutation model and the error model). In the generative models, we prepared candidate haplotypes, considering a heterozygous germ line variant if available, and the observed reads were realigned to the haplotypes. We then inferred the haplotype frequencies and computed the marginal likelihoods using a variational Bayesian algorithm. Finally, we derived a Bayes factor for evaluating the possibility of the existence of somatic mutations. We also demonstrated that our algorithm has superior specificity and sensitivity compared with existing methods, as determined based on a simulation, the TCGA Mutation Calling Benchmark 4 datasets and data from the COLO-829 cell line. AVAILABILITY AND IMPLEMENTATION: The HapMuC source code is available from http://github.com/usuyama/hapmuc.
Naoto Usuyama, Yuichi Shiraishi, Yusuke Sato, Haruki Kume, Yukio Homma, Seishi Ogawa, Satoru Miyano, Seiya Imoto
Bioinform.8
2014 A feature selection method using improved regularized linear discriminant analysis
Alok Sharma, Kuldip K. Paliwal, Seiya Imoto, Satoru Miyano
Mach. Vis. Appl.3
2013 Estimation of abrupt changes in sentinel observation data of influenza epidemics in Japan
Masaya M. Saito, Seiya Imoto, Rui Yamaguchi, Satoru Miyano, Tomoyuki Higuchi
FUSION2
2013 A strategy to select suitable physicochemical attributes of amino acids for protein fold recognition
abstract
BACKGROUND: Assigning a protein into one of its folds is a transitional step for discovering three dimensional protein structure, which is a challenging task in bimolecular (biological) science. The present research focuses on: 1) the development of classifiers, and 2) the development of feature extraction techniques based on syntactic and/or physicochemical properties. RESULTS: Apart from the above two main categories of research, we have shown that the selection of physicochemical attributes of the amino acids is an important step in protein fold recognition and has not been explored adequately. We have presented a multi-dimensional successive feature selection (MD-SFS) approach to systematically select attributes. The proposed method is applied on protein sequence data and an improvement of around 24% in fold recognition has been noted when selecting attributes appropriately. CONCLUSION: The MD-SFS has been applied successfully in selecting physicochemical attributes of the amino acids. The selected attributes show improved protein fold recognition performance.
Alok Sharma, Kuldip K. Paliwal, Abdollah Dehzangi, James G. Lyons, Seiya Imoto, Satoru Miyano
BMC Bioinform.5
2012 Identifiability of local transmissibility parameters in agent-based pandemic simulation
Masaya M. Saito, Seiya Imoto, Rui Yamaguchi, Satoru Miyano, Tomoyuki Higuchi
FUSION2
2012 Statistical model-based testing to evaluate the recurrence of genomic aberrations
abstract
MOTIVATION: In cancer genomes, chromosomal regions harboring cancer genes are often subjected to genomic aberrations like copy number alteration and loss of heterozygosity. Given this, finding recurrent genomic aberrations is considered an apt approach for screening cancer genes. Although several permutation-based tests have been proposed for this purpose, none of them are designed to find recurrent aberrations from the genomic dataset without paired normal sample controls. Their application to unpaired genomic data may lead to false discoveries, because they retrieve pseudo-aberrations that exist in normal genomes as polymorphisms. RESULTS: We develop a new parametric method named parametric aberration recurrence test (PART) to test for the recurrence of genomic aberrations. The introduction of Poisson-binomial statistics allow us to compute small P-values more efficiently and precisely than the previously proposed permutation-based approach. Moreover, we extended PART to cover unpaired data (PART-up) so that there is a statistical basis for analyzing unpaired genomic data. PART-up uses information from unpaired normal sample controls to remove pseudo-aberrations in unpaired genomic data. Using PART-up, we successfully predict recurrent genomic aberrations in cancer cell line samples whose paired normal sample controls are unavailable. This article thus proposes a powerful statistical framework for the identification of driver aberrations, which would be applicable to ever-increasing amounts of cancer genomic data seen in the era of next generation sequencing. AVAILABILITY: Our implementations of PART and PART-up are available from http://www.hgc.jp/~niiyan/PART/manual.html.
Atsushi Niida, Seiya Imoto, Teppei Shimamura, Satoru Miyano
Bioinform.2
2012 Identifying Gene Pathways Associated with Cancer Characteristics via Sparse Statistical Methods
abstract
We propose a statistical method for uncovering gene pathways that characterize cancer heterogeneity. To incorporate knowledge of the pathways into the model, we define a set of activities of pathways from microarray gene expression data based on the Sparse Probabilistic Principal Component Analysis (SPPCA). A pathway activity logistic regression model is then formulated for cancer phenotype. To select pathway activities related to binary cancer phenotypes, we use the elastic net for the parameter estimation and derive a model selection criterion for selecting tuning parameters included in the model estimation. Our proposed method can also reverse-engineer gene networks based on the identified multiple pathways that enables us to discover novel gene-gene associations relating with the cancer phenotypes. We illustrate the whole process of the proposed method through the analysis of breast cancer gene expression data.
Shuichi Kawano, Teppei Shimamura, Atsushi Niida, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, Ryo Yoshida, Cristin G. Print, Satoru Miyano
IEEE ACM Trans. Comput. Biol. Bioinform.4
2012 A Top-r Feature Selection Algorithm for Microarray Gene Expression Data
abstract
Most of the conventional feature selection algorithms have a drawback whereby a weakly ranked gene that could perform well in terms of classification accuracy with an appropriate subset of genes will be left out of the selection. Considering this shortcoming, we propose a feature selection algorithm in gene expression data analysis of sample classifications. The proposed algorithm first divides genes into subsets, the sizes of which are relatively small (roughly of size h), then selects informative smaller subsets of genes (of size r < h) from a subset and merges the chosen genes with another gene subset (of size r) to update the gene subset. We repeat this process until all subsets are merged into one informative subset. We illustrate the effectiveness of the proposed algorithm by analyzing three distinct gene expression data sets. Our method shows promising classification accuracy for all the test data sets. We also show the relevance of the selected genes in terms of their biological functions.
Alok Sharma, Seiya Imoto, Satoru Miyano
IEEE ACM Trans. Comput. Biol. Bioinform.2
2011 Estimation of macroscopic parameter in agent-based pandemic simulation
Masaya M. Saito, Seiya Imoto, Rui Yamaguchi, Satoru Miyano, Tomoyuki Higuchi
FUSION2
2011 Comprehensive Pharmacogenomic Pathway Screening by Data Assimilation
Takanori Hasegawa, Rui Yamaguchi, Masao Nagasaki, Seiya Imoto, Satoru Miyano
ISBRA4
2011 SiGN-SSM: open source parallel software for estimating gene networks with state space models
abstract
UNLABELLED: SiGN-SSM is an open-source gene network estimation software able to run in parallel on PCs and massively parallel supercomputers. The software estimates a state space model (SSM), that is a statistical dynamic model suitable for analyzing short time and/or replicated time series gene expression profiles. SiGN-SSM implements a novel parameter constraint effective to stabilize the estimated models. Also, by using a supercomputer, it is able to determine the gene network structure by a statistical permutation test in a practical time. SiGN-SSM is applicable not only to analyzing temporal regulatory dependencies between genes, but also to extracting the differentially regulated genes from time series expression profiles. AVAILABILITY: SiGN-SSM is distributed under GNU Affero General Public Licence (GNU AGPL) version 3 and can be downloaded at http://sign.hgc.jp/signssm/. The pre-compiled binaries for some architectures are available in addition to the source code. The pre-installed binaries are also available on the Human Genome Center supercomputer system. The online manual and the supplementary information of SiGN-SSM is available on our web site. CONTACT: [email protected].
Yoshinori Tamada, Rui Yamaguchi, Seiya Imoto, Osamu Hirose, Ryo Yoshida, Masao Nagasaki, Satoru Miyano
Bioinform.3
2011 Parallel Algorithm for Learning Optimal Bayesian Network Structure
Yoshinori Tamada, Seiya Imoto, Satoru Miyano
J. Mach. Learn. Res.2
2011 Estimating exogenous variables in data with more variables than observations
Yasuhiro Sogawa, Shohei Shimizu, Teppei Shimamura, Aapo Hyvärinen, Takashi Washio, Seiya Imoto
Neural Networks6
2011 Estimating Genome-Wide Gene Networks Using Nonparametric Bayesian Network Models on Massively Parallel Computers
abstract
We present a novel algorithm to estimate genome-wide gene networks consisting of more than 20,000 genes from gene expression data using nonparametric Bayesian networks. Due to the difficulty of learning Bayesian network structures, existing algorithms cannot be applied to more than a few thousand genes. Our algorithm overcomes this limitation by repeatedly estimating subnetworks in parallel for genes selected by neighbor node sampling. Through numerical simulation, we confirmed that our algorithm outperformed a heuristic algorithm in a shorter time. We applied our algorithm to microarray data from human umbilical vein endothelial cells (HUVECs) treated with siRNAs, to construct a human genome-wide gene network, which we compared to a small gene network estimated for the genes extracted using a traditional bioinformatics method. The results showed that our genome-wide gene network contains many features of the small network, as well as others that could not be captured during the small network estimation. The results also revealed master-regulator genes that are not in the small network but that control many of the genes in the small network. These analyses were impossible to realize without our proposed algorithm.
Yoshinori Tamada, Seiya Imoto, Hiromitsu Araki, Masao Nagasaki, Cristin G. Print, Stephen D. Charnock-Jones, Satoru Miyano
IEEE ACM Trans. Comput. Biol. Bioinform.2
2010 Identifying Hidden Confounders in Gene Networks by Bayesian Networks
abstract
In the estimation of gene networks from microarray gene expression data, we propose a statistical method for quantification of the hidden confounders in gene networks, which were possibly removed from the set of genes on the gene networks or are novel biological elements that are not measured by microarrays. Due to high computational cost of the structural learning of Bayesian networks and the limited source of the microarray data, it is usual to perform gene selection prior to the estimation of gene networks. Therefore, there exist missing genes that decrease accuracy and interpretability of the estimated gene networks. The proposed method can identify hidden confounders based on the conflicts of the estimated local Bayesian network structures and estimate their ideal profiles based on the proposed Bayesian networks with hidden variables with an EM algorithm. From the estimated ideal profiles, we can identify genes which are missing in the network or suggest the existence of the novel biological elements if the ideal profiles are not significantly correlated with any expression profiles of genes. To the best of our knowledge, this research is the first study to theoretically characterize missing genes in gene networks and practically utilize this information to refine network estimation.
Tomoya Higashigaki, Kaname Kojima, Rui Yamaguchi, Masato Inoue, Seiya Imoto, Satoru Miyano
BIBE5
2010 Discovering functional gene pathways associated with cancer heterogeneity via sparse supervised learning
abstract
We propose a statistical method for uncovering gene pathways that characterize cancer heterogeneity. To incorporate knowledge of the pathways into the model, we define a set of activities of pathways from microarray gene expression data based on the sparse probabilistic principal component analysis. A pathway activity logistic regression model is then formulated for cancer phenotype. To select pathway activities related to binary cancer phenotypes, we use the elastic net for the parameter estimation and derive a model selection criterion for selecting tuning parameters included in the model estimation. Our proposed method can also reverse-engineer gene networks based on the identified multiple pathways that enables us to discover novel gene-gene associations relating with the cancer phenotypes. We illustrate the whole process of the proposed method through the analysis of breast cancer gene expression data.
Shuichi Kawano, Teppei Shimamura, Atsushi Niida, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, Ryo Yoshida, Cristin G. Print, Satoru Miyano
BIBM4
2010 Discovery of Exogenous Variables in Data with More Variables Than Observations
Yasuhiro Sogawa, Shohei Shimizu, Aapo Hyvärinen, Takashi Washio, Teppei Shimamura, Seiya Imoto
ICANN (1)6
2010 Model-free unsupervised gene set screening based on information enrichment in expression profiles
abstract
MOTIVATION: A number of unsupervised gene set screening methods have recently been developed for search of putative functional gene sets based on their expression profiles. Most of the methods statistically evaluate whether the expression profiles of each gene set are fit to assumed models: e.g. co-expression across all samples or a subgroup of samples. However, it is possible that they fail to capture informative gene sets whose expression profiles are not fit to the assumed models. RESULTS: To overcome this limitation, we propose a model-free unsupervised gene set screening method, Matrix Information Enrichment Analysis (MIEA). Without assuming any specific models, MIEA screens gene sets based on information richness of their expression profiles. We extensively compared the performance of MIEA to those of other unsupervised gene set screening methods, using various types of simulated and real data. The benchmark tests demonstrated that MIEA can detect singular expression profiles that the other methods fail to find, and performs broadly well for various types of input data. Taken together, this study introduces MIEA as a broadly applicable gene set screening tool for mining regulatory programs from transcriptome data.
Atsushi Niida, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, André Fujita, Teppei Shimamura, Satoru Miyano
Bioinform.2
2010 Inferring dynamic gene networks under varying conditions for transcriptomic network comparison
abstract
MOTIVATION: Elucidating the differences between cellular responses to various biological conditions or external stimuli is an important challenge in systems biology. Many approaches have been developed to reverse engineer a cellular system, called gene network, from time series microarray data in order to understand a transcriptomic response under a condition of interest. Comparative topological analysis has also been applied based on the gene networks inferred independently from each of the multiple time series datasets under varying conditions to find critical differences between these networks. However, these comparisons often lead to misleading results, because each network contains considerable noise due to the limited length of the time series. RESULTS: We propose an integrated approach for inferring multiple gene networks from time series expression data under varying conditions. To the best of our knowledge, our approach is the first reverse-engineering method that is intended for transcriptomic network comparison between varying conditions. Furthermore, we propose a state-of-the-art parameter estimation method, relevance-weighted recursive elastic net, for providing higher precision and recall than existing reverse-engineering methods. We analyze experimental data of MCF-7 human breast cancer cells stimulated by epidermal growth factor or heregulin with several doses and provide novel biological hypotheses through network comparison. AVAILABILITY: The software NETCOMP is available at http://bonsai.ims.u-tokyo.ac.jp/ approximately shima/NETCOMP/.
Teppei Shimamura, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, Satoru Miyano
Bioinform.2
2010 Optimal Search on Clustered Structural Constraint for Learning Bayesian Network Structure
Kaname Kojima, Eric Perrier, Seiya Imoto, Satoru Miyano
J. Mach. Learn. Res.3
2009 Gene set-based module discovery in the breast cancer transcriptome
abstract
BACKGROUND: Although microarray-based studies have revealed global view of gene expression in cancer cells, we still have little knowledge about regulatory mechanisms underlying the transcriptome. Several computational methods applied to yeast data have recently succeeded in identifying expression modules, which is defined as co-expressed gene sets under common regulatory mechanisms. However, such module discovery methods are not applied cancer transcriptome data. RESULTS: In order to decode oncogenic regulatory programs in cancer cells, we developed a novel module discovery method termed EEM by extending a previously reported module discovery method, and applied it to breast cancer expression data. Starting from seed gene sets prepared based on cis-regulatory elements, ChIP-chip data, and gene locus information, EEM identified 10 principal expression modules in breast cancer based on their expression coherence. Moreover, EEM depicted their activity profiles, which predict regulatory programs in each subtypes of breast tumors. For example, our analysis revealed that the expression module regulated by the Polycomb repressive complex 2 (PRC2) is downregulated in triple negative breast cancers, suggesting similarity of transcriptional programs between stem cells and aggressive breast cancer cells. We also found that the activity of the PRC2 expression module is negatively correlated to the expression of EZH2, a component of PRC2 which belongs to the E2F expression module. E2F-driven EZH2 overexpression may be responsible for the repression of the PRC2 expression modules in triple negative tumors. Furthermore, our network analysis predicts regulatory circuits in breast cancer cells. CONCLUSION: These results demonstrate that the gene set-based module discovery approach is a powerful tool to decode regulatory programs in cancer cells.
Atsushi Niida, Andrew D. Smith, Seiya Imoto, Hiroyuki Aburatani, Michael Q. Zhang, Tetsu Akiyama
BMC Bioinform.3
2009 Network-Based Predictions and Simulations by Biological State Space Models: Search for Drug Mode of Action
Rui Yamaguchi, Seiya Imoto, Satoru Miyano
J. Comput. Sci. Technol.2
2008 Partial Order-Based Bayesian Network Learning Algorithm for Estimating Gene Networks
abstract
For learning Bayesian network structure from data, order-based algorithms such as K2 algorithm are widely used.In this paper, we consider a problem of constructing the order of nodes in such algorithms based on prior knowledge of gene networks. However, in many cases the prior knowledge is given as partial order of genes and we need to extend the order-based algorithm to partial order-based one. By extending our prior work we propose an efficient partial order-based algorithm for estimating gene networks based on Bayesian networks. The computational complexity of the proposed algorithm is shown.
Kazuyuki Numata, Seiya Imoto, Satoru Miyano
BIBM2
2008 Statistical inference of transcriptional module-based gene networks from time course gene expression profiles by using state space models
abstract
MOTIVATION: Statistical inference of gene networks by using time-course microarray gene expression profiles is an essential step towards understanding the temporal structure of gene regulatory mechanisms. Unfortunately, most of the current studies have been limited to analysing a small number of genes because the length of time-course gene expression profiles is fairly short. One promising approach to overcome such a limitation is to infer gene networks by exploring the potential transcriptional modules which are sets of genes sharing a common function or involved in the same pathway. RESULTS: In this article, we present a novel approach based on the state space model to identify the transcriptional modules and module-based gene networks simultaneously. The state space model has the potential to infer large-scale gene networks, e.g. of order 10(3), from time-course gene expression profiles. Particularly, we succeeded in the identification of a cell cycle system by using the gene expression profiles of Saccharomyces cerevisiae in which the length of the time-course and number of genes were 24 and 4382, respectively. However, when analysing shorter time-course data, e.g. of length 10 or less, the parameter estimations of the state space model often fail due to overfitting. To extend the applicability of the state space model, we provide an approach to use the technical replicates of gene expression profiles, which are often measured in duplicate or triplicate. The use of technical replicates is important for achieving highly-efficient inferences of gene networks with short time-course data. The potential of the proposed method has been demonstrated through the time-course analysis of the gene expression profiles of human umbilical vein endothelial cells (HUVECs) undergoing growth factor deprivation-induced apoptosis. AVAILABILITY: Supplementary Information and the software (TRANS-MNET) are available at http://daweb.ism.ac.jp/~yoshidar/software/ssm/.
Osamu Hirose, Ryo Yoshida, Seiya Imoto, Rui Yamaguchi, Tomoyuki Higuchi, Stephen D. Charnock-Jones, Cristin G. Print, Satoru Miyano
Bioinform.3
2008 Bayesian learning of biological pathways on genomic data assimilation
abstract
MOTIVATION: Mathematical modeling and simulation, based on biochemical rate equations, provide us a rigorous tool for unraveling complex mechanisms of biological pathways. To proceed to simulation experiments, it is an essential first step to find effective values of model parameters, which are difficult to measure from in vivo and in vitro experiments. Furthermore, once a set of hypothetical models has been created, any statistical criterion is needed to test the ability of the constructed models and to proceed to model revision. RESULTS: The aim of our research is to present a new statistical technology towards data-driven construction of in silico biological pathways. The method starts with a knowledge-based modeling with hybrid functional Petri net. It then proceeds to the Bayesian learning of model parameters for which experimental data are available. This process exploits quantitative measurements of evolving biochemical reactions, e.g. gene expression data. Another important issue that we consider is statistical evaluation and comparison of the constructed hypothetical pathways. For this purpose, we have developed a new Bayesian information-theoretic measure that assesses the predictability and the biological robustness of in silico pathways. AVAILABILITY: The FORTRAN source codes are available at the URL http://daweb.ism.ac.jpyoshidar/GDA/ SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ryo Yoshida, Masao Nagasaki, Rui Yamaguchi, Seiya Imoto, Satoru Miyano, Tomoyuki Higuchi
Bioinform.4
2008 Integrative bioinformatics analysis of transcriptional regulatory programs in breast cancer cells
abstract
BACKGROUND: Microarray technology has unveiled transcriptomic differences among tumors of various phenotypes, and, especially, brought great progress in molecular understanding of phenotypic diversity of breast tumors. However, compared with the massive knowledge about the transcriptome, we have surprisingly little knowledge about regulatory mechanisms underling transcriptomic diversity. RESULTS: To gain insights into the transcriptional programs that drive tumor progression, we integrated regulatory sequence data and expression profiles of breast cancer into a Bayesian Network, and searched for cis-regulatory motifs statistically associated with given histological grades and prognosis. Our analysis found that motifs bound by ELK1, E2F, NRF1 and NFY are potential regulatory motifs that positively correlate with malignant progression of breast cancer. CONCLUSION: The results suggest that these 4 motifs are principal regulatory motifs driving malignant progression of breast cancer. Our method offers a more concise description about transcriptome diversity among breast tumors with different clinical phenotypes.
Atsushi Niida, Andrew D. Smith, Seiya Imoto, Shuichi Tsutsumi, Hiroyuki Aburatani, Michael Q. Zhang, Tetsu Akiyama
BMC Bioinform.3
2008 ExonMiner: Web service for analysis of GeneChip Exon array data
abstract
BACKGROUND: Some splicing isoform-specific transcriptional regulations are related to disease. Therefore, detection of disease specific splice variations is the first step for finding disease specific transcriptional regulations. Affymetrix Human Exon 1.0 ST Array can measure exon-level expression profiles that are suitable to find differentially expressed exons in genome-wide scale. However, exon array produces massive datasets that are more than we can handle and analyze on personal computer. RESULTS: We have developed ExonMiner that is the first all-in-one web service for analysis of exon array data to detect transcripts that have significantly different splicing patterns in two cells, e.g. normal and cancer cells. ExonMiner can perform the following analyses: (1) data normalization, (2) statistical analysis based on two-way ANOVA, (3) finding transcripts with significantly different splice patterns, (4) efficient visualization based on heatmaps and barplots, and (5) meta-analysis to detect exon level biomarkers. We implemented ExonMiner on a supercomputer system in order to perform genome-wide analysis for more than 300,000 transcripts in exon array data, which has the potential to reveal the aberrant splice variations in cancer cells as exon level biomarkers. CONCLUSION: ExonMiner is well suited for analysis of exon array data and does not require any installation of software except for internet browsers. What all users need to do is to access the ExonMiner URL http://ae.hgc.jp/exonminer. Users can analyze full dataset of exon array data within hours by high-level statistical analysis with sound theoretical basis that finds aberrant splice variants as biomarkers.
Kazuyuki Numata, Ryo Yoshida, Masao Nagasaki, Ayumu Saito, Seiya Imoto, Satoru Miyano
BMC Bioinform.5
2007 A Structure Learning Algorithm for Inference of Gene Networks from Microarray Gene Expression Data Using Bayesian Networks
abstract
Estimation of gene networks based on microarray gene expression data is an important problem in systems biology. In this paper we use Bayesian networks as a mathematical model for reverse-engineering gene networks from microarray data. In such a case, structural learning of Bayesian networks is known as an NP-hard problem and we need to use heuristic algorithms to find better network structures. Recently, several algorithms have been proposed to estimate optimal Bayesian network structure, but the number of genes included in the network is limited less than 30 or so. In order to apply Bayesian network approach to drug target gene discovery, we need to consider gene networks with several hundreds of genes. Therefore we need to develop more efficient algorithms to learn Bayesian network structure based on observed data. In this paper we propose an efficient structural learning algorithm for Bayesian networks by extending K2 algorithm that is one of the standard learning algorithms in Bayesian networks. We conduct Monte Carlo simulations to examine the effectiveness of the proposed algorithm by comparing with greedy hill-climbing algorithm. We also show the application of yeast gene network estimation based on the proposed algorithm.
Kazuyuki Numata, Seiya Imoto, Satoru Miyano
BIBE2
2007 Computational Genome-Wide Discovery of Aberrant Splice Variations with Exon Expression Profiles
abstract
Alternative splicing plays a prominent role in eukaryotic gene regulations that allow a single gene to generate the multiple mRNA products. The recent advent of GeneChipregHuman Exon 1.0 ST Array enables us to measure the exon expression profiles of human cells on a genome-wide scale. With this advent, analysis of functional gene regulation could be extended to detect not only differentially expressed genes, but also specific splicing events that occur in target cells, but not in normal controls. We address some statistical issues for the identification of biomarker splice variations with exon expression data. The proposed method involves the following steps: (1) Whole transcript analysis with the nonparametric analysis of variance (ANOVA) to identify potential biomarkers that present specific splice variations. (2) Meta-analysis for discriminating non-specific splice variations that are caused by clinical heterogeneity in the collected samples. In the analysis of human cells, controlling non-specific splicing factors is essential for success in the detection of biomarker splice variations because splice patterns are possibly affected by inter-individual differences in the collected samples. We demonstrate its utility and perform a whole transcript analysis of exon expression profiles of colorectal carcinoma.
Ryo Yoshida, Kazuyuki Numata, Seiya Imoto, Masao Nagasaki, Atsushi Doi, Kazuko Ueno, Satoru Miyano
BIBE3
2007 Statistical Absolute Evaluation of Gene Ontology Terms with Gene Expression Data
Pramod K. Gupta, Ryo Yoshida, Seiya Imoto, Rui Yamaguchi, Satoru Miyano
ISBRA3
2006 ArrayCluster: an analytic tool for clustering, data visualization and module finder on gene expression profiles
abstract
SUMMARY: One of the significant challenges in gene expression analysis is to find unknown subtypes of several diseases at the molecular levels. This task can be addressed by grouping gene expression patterns of the collected samples on the basis of a large number of genes. Application of commonly used clustering methods to such a dataset however are likely to fail owing to over-learning, because the number of samples to be grouped is much smaller than the data dimension which is equal to the number of genes involved in the dataset. To overcome such difficulty, we developed a novel model-based clustering method, referred to as the mixed factors analysis. The ArrayCluster is a freely available software to perform the mixed factors analysis. It provides us some analytic tools for clustering DNA microarray experiments, data visualization and an automatic detector for module transcriptional of genes that are relevant to the calibrated molecular subtypes and so on.
Ryo Yoshida, Tomoyuki Higuchi, Seiya Imoto, Satoru Miyano
Bioinform.3
2005 Estimating Gene Networks from Expression Data and Binding Location Data via Boolean Networks
Osamu Hirose, Naoki Nariai, Yoshinori Tamada, Hideo Bannai, Seiya Imoto, Satoru Miyano
ICCSA (3)5
2005 A Penalized Likelihood Estimation on Transcriptional Module-Based Clustering
Ryo Yoshida, Seiya Imoto, Tomoyuki Higuchi
ICCSA (3)2
2004 Case-Control Study of Binary Disease Trait Considering Interactions between SNPs and Environmental Effects using Logistic Regression
abstract
In this paper, we propose a combination of logistic regression and genetic algorithm for the association study of the binary disease trait. We use a logistic regression model to describe the relation of multiple SNPs, environments and the target binary trait. The logistic regression model can capture the continuous effects of environments without categorization, which causes the loss of the information. To construct an accurate prediction rule for binary trait, we adopted Akaike information criterion (AIC) to find the most effective set of SNPs and environments. That is, the set of SNPs and environments that gives the smallest AIC is chosen as the optimal set. Since the number of combinations of SNPs and environments is usually huge, we propose the use of the genetic algorithm for choosing the optimal SNPs and environments in the sense of AIC. We show the effectiveness of the proposed method through the analysis of the case/control populations of diabetes patients. We succeeded in finding an efficient set to predict types of diabetes and some SNPs which have strong interactions to age while it is not significant as a single locus.
Reiichiro Nakamichi, Seiya Imoto, Satoru Miyano
BIBE2
2003 Inferring gene networks from time series microarray data using dynamic Bayesian networks
abstract
Dynamic Bayesian networks (DBNs) are considered as a promising model for inferring gene networks from time series microarray data. DBNs have overtaken Bayesian networks (BNs) as DBNs can construct cyclic regulations using time delay information. In this paper, a general framework for DBN modelling is outlined. Both discrete and continuous DBN models are constructed systematically and criteria for learning network structures are introduced from a Bayesian statistical viewpoint. This paper reviews the applications of DBNs over the past years. Real data applications for Saccharomyces cerevisiae time series gene expression data are also shown.
SunYong Kim, Seiya Imoto, Satoru Miyano
Briefings Bioinform.2
2002 Inferring Gene Regulatory Networks from Time-Ordered Gene Expression Data Using Differential Equations
Michiel J. L. de Hoon, Seiya Imoto, Satoru Miyano
Discovery Science2
2002 Statistical analysis of a small set of time-ordered gene expression data using linear splines
abstract
Abstract Motivation: Recently, the temporal response of genes to changes in their environment has been investigated using cDNA microarray technology by measuring the gene expression levels at a small number of time points. Conventional techniques for time series analysis are not suitable for such a short series of time-ordered data. The analysis of gene expression data has therefore usually been limited to a fold-change analysis, instead of a systematic statistical approach. Methods: We use the maximum likelihood method together with Akaike's Information Criterion to fit linear splines to a small set of time-ordered gene expression data in order to infer statistically meaningful information from the measurements. The significance of measured gene expression data is assessed using Student's t-test. Results: Previous gene expression measurements of the cyanobacterium Synechocystis sp. PCC6803 were reanalyzed using linear splines. The temporal response was identified of many genes that had been missed by a fold-change analysis. Based on our statistical analysis, we found that about four gene expression measurements or more are needed at each time point. Availability: An extension module for Python to calculate linear spline functions is available at http://bonsai.ims.u-tokyo.ac.jp/~mdehoon. This software package (with patent pending) is free of charge for academic use only. Contact: [email protected] * To whom correspondence should be addressed.
Michiel J. L. de Hoon, Seiya Imoto, Satoru Miyano
Bioinform.2