VLDB 2026 Research / reviewers in the wild / expert
Teppei Shimamura
dblp:75/7969
· DBLP profile ↗
19ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0003-2994-872XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | scSurv: a deep generative model for single-cell survival analysisabstractMOTIVATION: Single-cell omics analysis has unveiled the heterogeneity of various cell types within tumors. However, no methodology currently reveals how this heterogeneity influences cancer patient survival at single-cell resolution. Here, we introduce scSurv, combining a Cox proportional hazards model with a deep generative model of single-cell transcriptome, to estimate individual cellular contributions to clinical outcomes. RESULTS: The accuracy of scSurv was validated using both simulated and real datasets. This method identifies cells associated with favorable or adverse prognoses and extracts genes correlated with their contribution levels. In melanoma, scSurv reproduces known prognostic macrophage classifications and facilitates hazard mapping through spatial transcriptomics in renal cell carcinoma. We also identified genes consistently associated with prognosis across multiple cancers and demonstrated the applicability of this method to infectious diseases. scSurv is a novel framework for quantifying the heterogeneity of individual cellular effects on clinical outcomes. AVAILABILITY: The implementation of scSurv is available on GitHub (https://github.com/3254c/scSurv) and Zenodo (https://doi.org/10.5281/zenodo.17793054). Chikara Mizukoshi, Yasuhiro Kojima, Shuto Hayashi, Ko Abe, Daisuke Kasugai, Teppei Shimamura |
Bioinform. | 6 |
| 2024 | LineageVAE: reconstructing historical cell states and transcriptomes toward unobserved progenitorsabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) enables comprehensive characterization of the cell state. However, its destructive nature prohibits measuring gene expression changes during dynamic processes such as embryogenesis or cell state divergence due to injury or disease. Although recent studies integrating scRNA-seq with lineage tracing have provided clonal insights between progenitor and mature cells, challenges remain. Because of their experimental nature, observations are sparse, and cells observed in the early state are not the exact progenitors of cells observed at later time points. To overcome these limitations, we developed LineageVAE, a novel computational methodology that utilizes deep learning based on the property that cells sharing barcodes have identical progenitors. RESULTS: LineageVAE is a deep generative model that transforms scRNA-seq observations with identical lineage barcodes into sequential trajectories toward a common progenitor in a latent cell state space. This method enables the reconstruction of unobservable cell state transitions, historical transcriptomes, and regulatory dynamics at a single-cell resolution. Applied to hematopoiesis and reprogrammed fibroblast datasets, LineageVAE demonstrated its ability to restore backward cell state transitions and infer progenitor heterogeneity and transcription factor activity along differentiation trajectories. AVAILABILITY AND IMPLEMENTATION: The LineageVAE model was implemented in Python using the PyTorch deep learning library. The code is available on GitHub at https://github.com/LzrRacer/LineageVAE/. Koichiro Majima, Yasuhiro Kojima, Kodai Minoura, Ko Abe, Haruka Hirose, Teppei Shimamura |
Bioinform. | 6 |
| 2023 | UNMF: a unified nonnegative matrix factorization for multi-dimensional omics dataabstractFactor analysis, ranging from principal component analysis to nonnegative matrix factorization, represents a foremost approach in analyzing multi-dimensional data to extract valuable patterns, and is increasingly being applied in the context of multi-dimensional omics datasets represented in tensor form. However, traditional analytical methods are heavily dependent on the format and structure of the data itself, and if these change even slightly, the analyst must change their data analysis strategy and techniques and spend a considerable amount of time on data preprocessing. Additionally, many traditional methods cannot be applied as-is in the presence of missing values in the data. We present a new statistical framework, unified nonnegative matrix factorization (UNMF), for finding informative patterns in messy biological data sets. UNMF is designed for tidy data format and structure, making data analysis easier and simplifying the development of data analysis tools. UNMF can handle a wide range of data structures and formats, and works seamlessly with tensor data including missing observations and repeated measurements. The usefulness of UNMF is demonstrated through its application to several multi-dimensional omics data, offering user-friendly and unified features for analysis and integration. Its application holds great potential for the life science community. UNMF is implemented with R and is available from GitHub (https://github.com/abikoushi/moltenNMF). Ko Abe, Teppei Shimamura |
Briefings Bioinform. | 2 |
| 2021 | CYBERTRACK2.0: zero-inflated model-based cell clustering and population tracking method for longitudinal mass cytometry dataabstractSUMMARY: Recent advancements in high-dimensional single-cell technologies, such as mass cytometry, enable longitudinal experiments to track dynamics of cell populations and identify change points where the proportions vary significantly. However, current research is limited by the lack of tools specialized for analyzing longitudinal mass cytometry data. In order to infer cell population dynamics from such data, we developed a statistical framework named CYBERTRACK2.0. The framework's analytic performance was validated against synthetic and real data, showing that its results are consistent with previous research. AVAILABILITY AND IMPLEMENTATION: CYBERTRACK2.0 is available at https://github.com/kodaim1115/CYBERTRACK2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kodai Minoura, Ko Abe, Yuka Maeda, Hiroyoshi Nishikawa, Teppei Shimamura |
Bioinform. | 5 |
| 2020 | Model-based clustering for flow and mass cytometry data with clinical informationabstractBACKGROUND: High-dimensional flow cytometry and mass cytometry allow systemic-level characterization of more than 10 protein profiles at single-cell resolution and provide a much broader landscape in many biological applications, such as disease diagnosis and prediction of clinical outcome. When associating clinical information with cytometry data, traditional approaches require two distinct steps for identification of cell populations and statistical test to determine whether the difference between two population proportions is significant. These two-step approaches can lead to information loss and analysis bias. RESULTS: We propose a novel statistical framework, called LAMBDA (Latent Allocation Model with Bayesian Data Analysis), for simultaneous identification of unknown cell populations and discovery of associations between these populations and clinical information. LAMBDA uses specified probabilistic models designed for modeling the different distribution information for flow or mass cytometry data, respectively. We use a zero-inflated distribution for the mass cytometry data based the characteristics of the data. A simulation study confirms the usefulness of this model by evaluating the accuracy of the estimated parameters. We also demonstrate that LAMBDA can identify associations between cell populations and their clinical outcomes by analyzing real data. LAMBDA is implemented in R and is available from GitHub ( https://github.com/abikoushi/lambda ). Ko Abe, Kodai Minoura, Yuka Maeda, Hiroyoshi Nishikawa, Teppei Shimamura |
BMC Bioinform. | 5 |
| 2019 | A network of networks approach for modeling interconnected brain tissue-specific networksabstractMOTIVATION: Recent sequence-based analyses have identified a lot of gene variants that may contribute to neurogenetic disorders such as autism spectrum disorder and schizophrenia. Several state-of-the-art network-based analyses have been proposed for mechanical understanding of genetic variants in neurogenetic disorders. However, these methods were mainly designed for modeling and analyzing single networks that do not interact with or depend on other networks, and thus cannot capture the properties between interdependent systems in brain-specific tissues, circuits and regions which are connected each other and affect behavior and cognitive processes. RESULTS: We introduce a novel and efficient framework, called a 'Network of Networks' approach, to infer the interconnectivity structure between multiple networks where the response and the predictor variables are topological information matrices of given networks. We also propose Graph-Oriented SParsE Learning, a new sparse structural learning algorithm for network data to identify a subset of the topological information matrices of the predictors related to the response. We demonstrate on simulated data that propose Graph-Oriented SParsE Learning outperforms existing kernel-based algorithms in terms of F-measure. On real data from human brain region-specific functional networks associated with the autism risk genes, we show that the 'Network of Networks' model provides insights on the autism-associated interconnectivity structure between functional interaction networks and a comprehensive understanding of the genetic basis of autism across diverse regions of the brain. AVAILABILITY AND IMPLEMENTATION: Our software is available from https://github.com/infinite-point/GOSPEL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hideko Kawakubo, Yusuke Matsui 0002, Itaru Kushima, Norio Ozaki, Teppei Shimamura |
Bioinform. | 5 |
| 2019 | Model-based cell clustering and population tracking for time-series flow cytometry dataabstractBACKGROUND: Modern flow cytometry technology has enabled the simultaneous analysis of multiple cell markers at the single-cell level, and it is widely used in a broad field of research. The detection of cell populations in flow cytometry data has long been dependent on "manual gating" by visual inspection. Recently, numerous software have been developed for automatic, computationally guided detection of cell populations; however, they are not designed for time-series flow cytometry data. Time-series flow cytometry data are indispensable for investigating the dynamics of cell populations that could not be elucidated by static time-point analysis. Therefore, there is a great need for tools to systematically analyze time-series flow cytometry data. RESULTS: We propose a simple and efficient statistical framework, named CYBERTRACK (CYtometry-Based Estimation and Reasoning for TRACKing cell populations), to perform clustering and cell population tracking for time-series flow cytometry data. CYBERTRACK assumes that flow cytometry data are generated from a multivariate Gaussian mixture distribution with its mixture proportion at the current time dependent on that at a previous timepoint. Using simulation data, we evaluate the performance of CYBERTRACK when estimating parameters for a multivariate Gaussian mixture distribution, tracking time-dependent transitions of mixture proportions, and detecting change-points in the overall mixture proportion. The CYBERTRACK performance is validated using two real flow cytometry datasets, which demonstrate that the population dynamics detected by CYBERTRACK are consistent with our prior knowledge of lymphocyte behavior. CONCLUSIONS: Our results indicate that CYBERTRACK offers better understandings of time-dependent cell population dynamics to cytometry users by systematically analyzing time-series flow cytometry data. Kodai Minoura, Ko Abe, Yuka Maeda, Hiroyoshi Nishikawa, Teppei Shimamura |
BMC Bioinform. | 5 |
| 2018 | A latent allocation model for the analysis of microbial composition and diseaseabstractBACKGROUND: Establishing the relationship between microbiota and specific diseases is important but requires appropriate statistical methodology. A specialized feature of microbiome count data is the presence of a large number of zeros, which makes it difficult to analyze in case-control studies. Most existing approaches either add a small number called a pseudo-count or use probability models such as the multinomial and Dirichlet-multinomial distributions to explain the excess zero counts, which may produce unnecessary biases and impose a correlation structure taht is unsuitable for microbiome data. RESULTS: The purpose of this article is to develop a new probabilistic model, called BERnoulli and MUltinomial Distribution-based latent Allocation (BERMUDA), to address these problems. BERMUDA enables us to describe the differences in bacteria composition and a certain disease among samples. We also provide a simple and efficient learning procedure for the proposed model using an annealing EM algorithm. CONCLUSION: We illustrate the performance of the proposed method both through both the simulation and real data analysis. BERMUDA is implemented with R and is available from GitHub ( https://github.com/abikoushi/Bermuda ). Ko Abe, Masaaki Hirayama, Kinji Ohno, Teppei Shimamura |
BMC Bioinform. | 4 |
| 2017 | phyC: Clustering cancer evolutionary treesabstractMulti-regional sequencing provides new opportunities to investigate genetic heterogeneity within or between common tumors from an evolutionary perspective. Several state-of-the-art methods have been proposed for reconstructing cancer evolutionary trees based on multi-regional sequencing data to develop models of cancer evolution. However, there have been few studies on comparisons of a set of cancer evolutionary trees. We propose a clustering method (phyC) for cancer evolutionary trees, in which sub-groups of the trees are identified based on topology and edge length attributes. For interpretation, we also propose a method for evaluating the sub-clonal diversity of trees in the clusters, which provides insight into the acceleration of sub-clonal expansion. Simulation showed that the proposed method can detect true clusters with sufficient accuracy. Application of the method to actual multi-regional sequencing data of clear cell renal carcinoma and non-small cell lung cancer allowed for the detection of clusters related to cancer type or phenotype. phyC is implemented with R(≥3.2.2) and is available from https://github.com/ymatts/phyC. Yusuke Matsui 0002, Atsushi Niida, Ryutaro Uchi, Koshi Mimori, Satoru Miyano, Teppei Shimamura |
PLoS Comput. Biol. | 6 |
| 2016 | D3M: detection of differential distributions of methylation levelsabstractMOTIVATION: DNA methylation is an important epigenetic modification related to a variety of diseases including cancers. We focus on the methylation data from Illumina's Infinium HumanMethylation450 BeadChip. One of the key issues of methylation analysis is to detect the differential methylation sites between case and control groups. Previous approaches describe data with simple summary statistics or kernel function, and then use statistical tests to determine the difference. However, a summary statistics-based approach cannot capture complicated underlying structure, and a kernel function-based approach lacks interpretability of results. RESULTS: We propose a novel method D(3)M, for detection of differential distribution of methylation, based on distribution-valued data. Our method can detect the differences in high-order moments, such as shapes of underlying distributions in methylation profiles, based on the Wasserstein metric. We test the significance of the difference between case and control groups and provide an interpretable summary of the results. The simulation results show that the proposed method achieves promising accuracy and shows favorable results compared with previous methods. Glioblastoma multiforme and lower grade glioma data from The Cancer Genome Atlas show that our method supports recent biological advances and suggests new insights. AVAILABILITY AND IMPLEMENTATION: R implemented code is freely available from https://github.com/ymatts/D3M/ CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yusuke Matsui 0002, Masahiro Mizuta, Satoru Miyano, Teppei Shimamura |
Bioinform. | 5 |
| 2015 | High performance computing of a fusion gene detection pipeline on the K computerabstractRecently developed high-throughput sequencers can generate a huge amount of omics data, and TOP500-class supercomputers are required to analyze such large datasets. However, these supercomputers do not usually support grid engines, which are commonly used in bioinformatics, making it necessary to parallelize the software. Parallelization and optimization require domain and specific knowledge, which poses a challenge for most bioinformaticians. We here propose a simple methodology for the parallelization of pipeline software. To demonstrate the efficacy of the methodology, we employed the Genomon-fusion as a sample software pipeline and ported it onto the K computer by applying our method. Simultaneous analysis of a massive amount of samples was performed using the K computer in a very short period of time. Yuichi Shiraishi, Teppei Shimamura, Kenichi Chiba, Satoru Miyano |
BIBM | 3 |
| 2012 | Statistical model-based testing to evaluate the recurrence of genomic aberrationsabstractMOTIVATION: In cancer genomes, chromosomal regions harboring cancer genes are often subjected to genomic aberrations like copy number alteration and loss of heterozygosity. Given this, finding recurrent genomic aberrations is considered an apt approach for screening cancer genes. Although several permutation-based tests have been proposed for this purpose, none of them are designed to find recurrent aberrations from the genomic dataset without paired normal sample controls. Their application to unpaired genomic data may lead to false discoveries, because they retrieve pseudo-aberrations that exist in normal genomes as polymorphisms. RESULTS: We develop a new parametric method named parametric aberration recurrence test (PART) to test for the recurrence of genomic aberrations. The introduction of Poisson-binomial statistics allow us to compute small P-values more efficiently and precisely than the previously proposed permutation-based approach. Moreover, we extended PART to cover unpaired data (PART-up) so that there is a statistical basis for analyzing unpaired genomic data. PART-up uses information from unpaired normal sample controls to remove pseudo-aberrations in unpaired genomic data. Using PART-up, we successfully predict recurrent genomic aberrations in cancer cell line samples whose paired normal sample controls are unavailable. This article thus proposes a powerful statistical framework for the identification of driver aberrations, which would be applicable to ever-increasing amounts of cancer genomic data seen in the era of next generation sequencing. AVAILABILITY: Our implementations of PART and PART-up are available from http://www.hgc.jp/~niiyan/PART/manual.html. Atsushi Niida, Seiya Imoto, Teppei Shimamura, Satoru Miyano |
Bioinform. | 3 |
| 2012 | Identifying Gene Pathways Associated with Cancer Characteristics via Sparse Statistical MethodsabstractWe propose a statistical method for uncovering gene pathways that characterize cancer heterogeneity. To incorporate knowledge of the pathways into the model, we define a set of activities of pathways from microarray gene expression data based on the Sparse Probabilistic Principal Component Analysis (SPPCA). A pathway activity logistic regression model is then formulated for cancer phenotype. To select pathway activities related to binary cancer phenotypes, we use the elastic net for the parameter estimation and derive a model selection criterion for selecting tuning parameters included in the model estimation. Our proposed method can also reverse-engineer gene networks based on the identified multiple pathways that enables us to discover novel gene-gene associations relating with the cancer phenotypes. We illustrate the whole process of the proposed method through the analysis of breast cancer gene expression data. Shuichi Kawano, Teppei Shimamura, Atsushi Niida, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, Ryo Yoshida, Cristin G. Print, Satoru Miyano |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2011 | Estimating exogenous variables in data with more variables than observations
Yasuhiro Sogawa, Shohei Shimizu, Teppei Shimamura, Aapo Hyvärinen, Takashi Washio, Seiya Imoto |
Neural Networks | 3 |
| 2011 | Inferring Contagion in Regulatory NetworksabstractSeveral gene regulatory network models containing concepts of directionality at the edges have been proposed. However, only a few reports have an interpretable definition of directionality. Here, differently from the standard causality concept defined by Pearl, we introduce the concept of contagion in order to infer directionality at the edges, i.e., asymmetries in gene expression dependences of regulatory networks. Moreover, we present a bootstrap algorithm in order to test the contagion concept. This technique was applied in simulated data and, also, in an actual large sample of biological data. Literature review has confirmed some genes identified by contagion as actually belonging to the TP53 pathway. André Fujita, João R. Sato, Marcos Angelo Almeida Demasi, Rui Yamaguchi, Teppei Shimamura, Carlos Eduardo Ferreira, Mari Cleide Sogayar, Satoru Miyano |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2010 | Discovering functional gene pathways associated with cancer heterogeneity via sparse supervised learningabstractWe propose a statistical method for uncovering gene pathways that characterize cancer heterogeneity. To incorporate knowledge of the pathways into the model, we define a set of activities of pathways from microarray gene expression data based on the sparse probabilistic principal component analysis. A pathway activity logistic regression model is then formulated for cancer phenotype. To select pathway activities related to binary cancer phenotypes, we use the elastic net for the parameter estimation and derive a model selection criterion for selecting tuning parameters included in the model estimation. Our proposed method can also reverse-engineer gene networks based on the identified multiple pathways that enables us to discover novel gene-gene associations relating with the cancer phenotypes. We illustrate the whole process of the proposed method through the analysis of breast cancer gene expression data. Shuichi Kawano, Teppei Shimamura, Atsushi Niida, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, Ryo Yoshida, Cristin G. Print, Satoru Miyano |
BIBM | 2 |
| 2010 | Discovery of Exogenous Variables in Data with More Variables Than Observations
Yasuhiro Sogawa, Shohei Shimizu, Aapo Hyvärinen, Takashi Washio, Teppei Shimamura, Seiya Imoto |
ICANN (1) | 5 |
| 2010 | Model-free unsupervised gene set screening based on information enrichment in expression profilesabstractMOTIVATION: A number of unsupervised gene set screening methods have recently been developed for search of putative functional gene sets based on their expression profiles. Most of the methods statistically evaluate whether the expression profiles of each gene set are fit to assumed models: e.g. co-expression across all samples or a subgroup of samples. However, it is possible that they fail to capture informative gene sets whose expression profiles are not fit to the assumed models. RESULTS: To overcome this limitation, we propose a model-free unsupervised gene set screening method, Matrix Information Enrichment Analysis (MIEA). Without assuming any specific models, MIEA screens gene sets based on information richness of their expression profiles. We extensively compared the performance of MIEA to those of other unsupervised gene set screening methods, using various types of simulated and real data. The benchmark tests demonstrated that MIEA can detect singular expression profiles that the other methods fail to find, and performs broadly well for various types of input data. Taken together, this study introduces MIEA as a broadly applicable gene set screening tool for mining regulatory programs from transcriptome data. Atsushi Niida, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, André Fujita, Teppei Shimamura, Satoru Miyano |
Bioinform. | 6 |
| 2010 | Inferring dynamic gene networks under varying conditions for transcriptomic network comparisonabstractMOTIVATION: Elucidating the differences between cellular responses to various biological conditions or external stimuli is an important challenge in systems biology. Many approaches have been developed to reverse engineer a cellular system, called gene network, from time series microarray data in order to understand a transcriptomic response under a condition of interest. Comparative topological analysis has also been applied based on the gene networks inferred independently from each of the multiple time series datasets under varying conditions to find critical differences between these networks. However, these comparisons often lead to misleading results, because each network contains considerable noise due to the limited length of the time series. RESULTS: We propose an integrated approach for inferring multiple gene networks from time series expression data under varying conditions. To the best of our knowledge, our approach is the first reverse-engineering method that is intended for transcriptomic network comparison between varying conditions. Furthermore, we propose a state-of-the-art parameter estimation method, relevance-weighted recursive elastic net, for providing higher precision and recall than existing reverse-engineering methods. We analyze experimental data of MCF-7 human breast cancer cells stimulated by epidermal growth factor or heregulin with several doses and provide novel biological hypotheses through network comparison. AVAILABILITY: The software NETCOMP is available at http://bonsai.ims.u-tokyo.ac.jp/ approximately shima/NETCOMP/. Teppei Shimamura, Seiya Imoto, Rui Yamaguchi, Masao Nagasaki, Satoru Miyano |
Bioinform. | 1 |