VLDB 2026 Research / reviewers in the wild / expert
Osamu Maruyama
dblp:82/6722
· DBLP profile ↗
27ranked-venue papers
18as first author
4since 2021 · last 2025
0000-0002-5760-1507ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 13 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 5 first-authorTheory of computation · 6 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CDACHIE: chromatin domain annotation by integrating chromatin interaction and epigenomic data with contrastive learningabstractMOTIVATION: Chromatin domain annotation identifies functional genomic regions, such as active and inactive zones, based on epigenomic features like histone modifications, DNA methylation, and chromatin accessibility. While recent methods have utilized both chromatin interaction data (e.g. Hi-C) and epigenomic data, they often overlook the direct relationship between these data types. RESULTS: In this study, we introduce Chromatin Domain Annotation using Contrastive Learning for Hi-C and Epigenomic Data (CDACHIE), a method for identifying chromatin domains from Hi-C and epigenomic data. Our approach leverages contrastive learning to generate aligned representative vectors for both data types at each genomic bin. The concatenated vectors are then clustered using K-means to classify distinct chromatin domain types. CDACHIE achieves superior performance in Variance Explained, evaluated across gene expression, replication timing, and ChIA-PET data. This highlights its robust ability to integrate semantic associations between Hi-C and epigenomic features within the embedding space. AVAILABILITY AND IMPLEMENTATION: The source code is available at GitHub: https://github.com/maruyama-lab-design/CDACHIE. An archival snapshot of the code used in this study is available on Zenodo: https://doi.org/10.5281/zenodo.15751780. Asato Yoshinaga, Osamu Maruyama |
Bioinform. | 2 |
| 2025 | Class-balanced negative training sets for improving classifier model predictions of enhancer-promoter interactionsabstractBACKGROUND: Enhancers regulate gene expression by forming DNA loops, thereby bringing themselves in close proximity to the target gene promoter. The human genome contains hundreds of thousands of enhancers, vastly outnumbering its 20,000-25,000 protein-coding genes, highlighting the importance of enhancer-promoter interactions (EPIs) in gene regulation. Supervised learning models have been developed to predict EPIs, often using experimentally validated interacting enhancer-promoter pairs and artificially generated negative samples. However, the lack of reliable negative samples presents a challenge. Current methods randomly select pairs from unlabeled data, leading to class imbalance and reduced predictive performance. This imbalance, where enhancers and promoters are unevenly distributed between the positive and negative sets, hinders classifiers from learning meaningful patterns. Therefore, constructing more reliable negative samples is crucial for improving the accuracy of EPI predictions. RESULTS: We developed two methods to generate class-balanced negative training sets for EPI classifiers: one based on maximum flow and the other on Gibbs sampling. We evaluated these methods with the TargetFinder and TransEPI classifiers across five and six cell lines, respectively. The trained models were tested using a common negative test set. Our negative training sets significantly improved the prediction performance across several metrics, including precision, recall, and area under the receiver operating characteristic curve. CONCLUSIONS: Our findings demonstrate that carefully designed negative samples can enhance the performance of EPI classifiers. Further advanced methods in generating negative EPIs should further improve prediction accuracy. The source code is available at https://github.com/maruyama-lab-design/CBOEP2 . Osamu Maruyama, Tsukasa Koga |
BMC Bioinform. | 1 |
| 2022 | CMIC: predicting DNA methylation inheritance of CpG islands with embedding vectors of variable-length k-mersabstractBACKGROUND: Epigenetic modifications established in mammalian gametes are largely reprogrammed during early development, however, are partly inherited by the embryo to support its development. In this study, we examine CpG island (CGI) sequences to predict whether a mouse blastocyst CGI inherits oocyte-derived DNA methylation from the maternal genome. Recurrent neural networks (RNNs), including that based on gated recurrent units (GRUs), have recently been employed for variable-length inputs in classification and regression analyses. One advantage of this strategy is the ability of RNNs to automatically learn latent features embedded in inputs by learning their model parameters. However, the available CGI dataset applied for the prediction of oocyte-derived DNA methylation inheritance are not large enough to train the neural networks. RESULTS: We propose a GRU-based model called CMIC (CGI Methylation Inheritance Classifier) to augment CGI sequence by converting it into variable-length k-mers, where the length k is randomly selected from the range [Formula: see text] to [Formula: see text], N times, which were then used as neural network input. N was set to 1000 in the default setting. In addition, we proposed a new embedding vector generator for k-mers called splitDNA2vec. The randomness of this procedure was higher than the previous work, dna2vec. CONCLUSIONS: We found that CMIC can predict the inheritance of oocyte-derived DNA methylation at CGIs in the maternal genome of blastocysts with a high F-measure (0.93). We also show that the F-measure can be improved by increasing the parameter N, that is, the number of sequences of variable-length k-mers derived from a single CGI sequence. This implies the effectiveness of augmenting input data by converting a DNA sequence to N sequences of variable-length k-mers. This approach can be applied to different DNA sequence classification and regression analyses, particularly those involving a small amount of data. Osamu Maruyama, Yinuo Li, Hiroki Narita, Hidehiro Toh, Wan Kin Au Yeung, Hiroyuki Sasaki |
BMC Bioinform. | 1 |
| 2021 | A convolutional neural network-based regression model to infer the epigenetic crosstalk responsible for CG methylation patternsabstractBACKGROUND: Epigenetic modifications, including CG methylation (a major form of DNA methylation) and histone modifications, interact with each other to shape their genomic distribution patterns. However, the entire picture of the epigenetic crosstalk regulating the CG methylation pattern is unknown especially in cells that are available only in a limited number, such as mammalian oocytes. Most machine learning approaches developed so far aim at finding DNA sequences responsible for the CG methylation patterns and were not tailored for studying the epigenetic crosstalk. RESULTS: We built a machine learning model named epiNet to predict CG methylation patterns based on other epigenetic features, such as histone modifications, but not DNA sequence. Using epiNet, we identified biologically relevant epigenetic crosstalk between histone H3K36me3, H3K4me3, and CG methylation in mouse oocytes. This model also predicted the altered CG methylation pattern of mutant oocytes having perturbed histone modification, was applicable to cross-species prediction of the CG methylation pattern of human oocytes, and identified the epigenetic crosstalk potentially important in other cell types. CONCLUSIONS: Our findings provide insight into the epigenetic crosstalk regulating the CG methylation pattern in mammalian oocytes and other cells. The use of epiNet should help to design or complement biological experiments in epigenetics studies. Wan Kin Au Yeung, Osamu Maruyama, Hiroyuki Sasaki |
BMC Bioinform. | 2 |
| 2019 | DegSampler3: Pairwise Dependency Model in Degradation Motif Site Prediction of Substrate Protein SequencesabstractIn the ubiquitin-proteasome system, E3 ubiquitin ligase (E3s for short) selectively recognize and bind specific regions of their substrate proteins. Sequence motifs whose sites are bound by E3 ubiquitin ligases are called degrons. Because much remains unclear about the relationship between substrate proteins of E3s and their binding sites, there is a need to computationally identify such binding sites from the substrate proteins. For this motif identification problem, in our previous works, we have proposed a series of collapsed Gibbs sampling algorithms, called DegSampler1 and DegSampler2, both of which use position-specific prior information. In this work, we propose a new collapsed Gibbs sampling algorithm, called DegSampler3, by integrating intra-motif pair-wise dependency model into the posterior probability distribution of DegSampler2. In our preliminary experiments, we found that DegSampler3 has the ability of finding more various degron sites than DegSampler2 while keeping the prediction accuracy almost the same as that of the previous method, DegSampler2. Osamu Maruyama, Fumiko Matsuzaki |
BIBE | 1 |
| 2018 | DegSampler: Collapsed Gibbs Sampler for Detecting E3 Binding SitesabstractIn this paper, we address the problem of finding sequence motifs in substrate proteins specific to E3 ubiquitin ligases (E3s). We formulated a posterior probability distribution of sites by designing a likelihood function based on amino acid indexing and a prior distribution based on the disorderness of protein sequences. These designs are derived from known characteristics of E3 binding sites in substrate proteins. Then, we devise a collapsed Gibbs sampling algorithm for the posterior probability distribution called DegSampler. We performed computational experiments using 36 sets of substrate proteins specific to E3s and compared the performance of DegSampler with those of popular motif finders, MEME and GLAM2. The results showed that DegSampler was superior to the others in finding E3 binding motifs. Thus, DegSampler is a promising tool for finding E3 motifs in substrate proteins. Osamu Maruyama, Fumiko Matsuzaki |
BIBE | 1 |
| 2017 | RocSampler: regularizing overlapping protein complexes in protein-protein interaction networksabstractBACKGROUND: In recent years, protein-protein interaction (PPI) networks have been well recognized as important resources to elucidate various biological processes and cellular mechanisms. In this paper, we address the problem of predicting protein complexes from a PPI network. This problem has two difficulties. One is related to small complexes, which contains two or three components. It is relatively difficult to identify them due to their simpler internal structure, but unfortunately complexes of such sizes are dominant in major protein complex databases, such as CYC2008. Another difficulty is how to model overlaps between predicted complexes, that is, how to evaluate different predicted complexes sharing common proteins because CYC2008 and other databases include such protein complexes. Thus, it is critical how to model overlaps between predicted complexes to identify them simultaneously. RESULTS: In this paper, we propose a sampling-based protein complex prediction method, RocSampler (Regularizing Overlapping Complexes), which exploits, as part of the whole scoring function, a regularization term for the overlaps of predicted complexes and that for the distribution of sizes of predicted complexes. We have implemented RocSampler in MATLAB and its executable file for Windows is available at the site, http://imi.kyushu-u.ac.jp/~om/software/RocSampler/ . CONCLUSIONS: We have applied RocSampler to five yeast PPI networks and shown that it is superior to other existing methods. This implies that the design of scoring functions including regularization terms is an effective approach for protein complex prediction. Osamu Maruyama, Yuki Kuwahara |
BMC Bioinform. | 1 |
| 2015 | Regularizing predicted complexes by mutually exclusive protein-protein interactionsabstractProtein complexes are key entities in the cell responsible for various cellular mechanisms and biological processes. We propose here a method for predicting protein complexes from a protein-protein interaction (PPI) network, using information on mutually exclusive PPIs. If two interactions are mutually exclusive, they are not allowed to exist simultaneously in the same predicted complex. We introduce a new regularization term which checks whether predicted complexes are connected by mutually exclusive PPIs. This regularization term is added into the scoring function of our earlier protein complex prediction tool, PPSampler2. We show that PPSampler2 with mutually exclusive PPIs outperforms the original one. Furthermore, the performance is superior to well-known representative conventional protein complex prediction methods. Thus, it is is effective to use mutual exclusiveness of PPIs in protein complex prediction. Osamu Maruyama, Limsoon Wong |
ASONAM | 1 |
| 2014 | A scale-free structure prior for Bayesian inference of Gaussian graphical modelsabstractThe inference of gene association networks from gene expression profiles is an important approach to elucidate various cellular mechanisms. However, there exists a problematic issue that the number of samples is relatively small than that of genes. A promising approach to this problem will be to design regularization terms for characteristic network structures like sparsity and scale-freeness and optimize a scoring function including those regularization terms. The inference problem for gene association networks is often formulated as the problem of estimating the inverse covariance matrix of a Gaussian distribution from its samples. For this Bayesian inference problem, we propose a novel scale-free structure prior and devise a sampling method for optimizing a posterior probability including the prior. In a simulation study, scale-free graphs of 30 and 100 nodes are generated by the Barabási-Albert model, and the proposed method is shown to outperform another method which also use a scale-free regularization term. Our method is also applied to real gene expression profiles, and the resulting graph shows biologically meaningful features. Thus, we empirically conclude that our scale-free structure prior is effective in Bayesian inference of Gaussian graphical models. Osamu Maruyama, Shota Shikita |
BIBM | 1 |
| 2014 | Prediction of heterotrimeric protein complexes by two-phase learning using neighboring kernelsabstractBACKGROUND: Protein complexes play important roles in biological systems such as gene regulatory networks and metabolic pathways. Most methods for predicting protein complexes try to find protein complexes with size more than three. It, however, is known that protein complexes with smaller sizes occupy a large part of whole complexes for several species. In our previous work, we developed a method with several feature space mappings and the domain composition kernel for prediction of heterodimeric protein complexes, which outperforms existing methods. RESULTS: We propose methods for prediction of heterotrimeric protein complexes by extending techniques in the previous work on the basis of the idea that most heterotrimeric protein complexes are not likely to share the same protein with each other. We make use of the discriminant function in support vector machines (SVMs), and design novel feature space mappings for the second phase. As the second classifier, we examine SVMs and relevance vector machines (RVMs). We perform 10-fold cross-validation computational experiments. The results suggest that our proposed two-phase methods and SVM with the extended features outperform the existing method NWE, which was reported to outperform other existing methods such as MCL, MCODE, DPClus, CMC, COACH, RRW, and PPSampler for prediction of heterotrimeric protein complexes. CONCLUSIONS: We propose two-phase prediction methods with the extended features, the domain composition kernel, SVMs and RVMs. The two-phase method with the extended features and the domain composition kernel using SVM as the second classifier is particularly useful for prediction of heterotrimeric protein complexes. Peiying Ruan, Morihiro Hayashida, Osamu Maruyama, Tatsuya Akutsu |
BMC Bioinform. | 3 |
| 2013 | Heterodimeric protein complex identification by naïve Bayes classifiersabstractBACKGROUND: Protein complexes are basic cellular entities that carry out the functions of their components. It can be found that in databases of protein complexes of yeast like CYC2008, the major type of known protein complexes is heterodimeric complexes. Although a number of methods for trying to predict sets of proteins that form arbitrary types of protein complexes simultaneously have been proposed, it can be found that they often fail to predict heterodimeric complexes. RESULTS: In this paper, we have designed several features characterizing heterodimeric protein complexes based on genomic data sets, and proposed a supervised-learning method for the prediction of heterodimeric protein complexes. This method learns the parameters of the features, which are embedded in the naïve Bayes classifier. The log-likelihood ratio derived from the naïve Bayes classifier with the parameter values obtained by maximum likelihood estimation gives the score of a given pair of proteins to predict whether the pair is a heterodimeric complex or not. A five-fold cross-validation shows good performance on yeast. The trained classifiers also show higher predictability than various existing algorithms on yeast data sets with approximate and exact matching criteria. CONCLUSIONS: Heterodimeric protein complex prediction is a rather harder problem than heteromeric protein complex prediction because heterodimeric protein complex is topologically simpler. However, it turns out that by designing features specialized for heterodimeric protein complexes, predictability of them can be improved. Thus, the design of more sophisticate features for heterodimeric protein complexes as well as the accumulation of more accurate and useful genome-wide data sets will lead to higher predictability of heterodimeric protein complexes. Our tool can be downloaded from http://imi.kyushu-u.ac.jp/~om/. Osamu Maruyama |
BMC Bioinform. | 1 |
| 2010 | NWE: Node-weighted expansion for protein complex prediction using random walk distancesabstractProtein complexes are important entities to organize various biological systems. However, they are still limited in availability. Thus, it is a challenging problem to predict protein complexes computationally from existing genome-wide data sets, like protein-protein interaction (PPI) networks. In this paper, we propose an efficient algorithm for predicting protein complexes by random walking on a PPI network. The algorithm is designed based on the method of node-weighted expansion of a cluster, which simulates a random walk with restarts with the weighted nodes of the cluster. We have validated the biological significance of the results using curated complexes in the CYC2008 database. We have compared our method to a clustering-based method, MCL, and a repeated random walk-based method, RRW, and found that our algorithm outperforms the other algorithms. Osamu Maruyama, Ayaka Chihara |
BIBM | 1 |
| 2008 | Evaluating Protein Sequence Signatures Inferred from Protein-Protein Interaction Data by Gene Ontology AnnotationsabstractWe propose a systematic method to find sequence signatures of proteins which share a common interacting partner, using a protein similarity measure based on gene ontology (GO) annotations. In a computational experiment on our original human data set of protein-protein interactions determined by Y2H assays, we have succeeded in discovering convincing candidates for interacting sites and other functional regions. Osamu Maruyama, Hideki Hirakawa, Takao Iwayanagi, Yoshiko Ishida, Shizu Takeda, Jun Otomo, Satoru Kuhara |
BIBM | 1 |
| 2004 | Searching for Regulatory Elements of Alternative Splicing Events Using Phylogenetic Footprinting
Daichi Shigemizu, Osamu Maruyama |
WABI | 2 |
| 2003 | Identification of genetic networks by strategic gene disruptions and gene overexpressions under a boolean model
Tatsuya Akutsu, Satoru Kuhara, Osamu Maruyama, Satoru Miyano |
Theor. Comput. Sci. | 3 |
| 2002 | Toward Drawing an Atlas of Hypothesis Classes: Approximating a Hypothesis via Another Hypothesis Model
Osamu Maruyama, Takayoshi Shoudai, Satoru Miyano |
Discovery Science | 1 |
| 2002 | Extensive feature detection of N-terminal protein sorting signalsabstractMOTIVATION: The prediction of localization sites of various proteins is an important and challenging problem in the field of molecular biology. TargetP, by Emanuelsson et al. (J. Mol. Biol., 300, 1005-1016, 2000) is a neural network based system which is currently the best predictor in the literature for N-terminal sorting signals. One drawback of neural networks, however, is that it is generally difficult to understand and interpret how and why they make such predictions. In this paper, we aim to generate simple and interpretable rules as predictors, and still achieve a practical prediction accuracy. We adopt an approach which consists of an extensive search for simple rules and various attributes which is partially guided by human intuition. RESULTS: We have succeeded in finding rules whose prediction accuracies come close to that of TargetP, while still retaining a very simple and interpretable form. We also discuss and interpret the discovered rules. Hideo Bannai, Yoshinori Tamada, Osamu Maruyama, Kenta Nakai, Satoru Miyano |
Bioinform. | 3 |
| 2002 | Fast algorithm for extracting multiple unordered short motifs using bit operations
Osamu Maruyama, Hideo Bannai, Yoshinori Tamada, Satoru Kuhara, Satoru Miyano |
Inf. Sci. | 1 |
| 2001 | VML: A View Modeling Language for Computational Knowledge Discovery
Hideo Bannai, Yoshinori Tamada, Osamu Maruyama, Satoru Miyano |
Discovery Science | 3 |
| 2001 | Learning Conformation Rules
Osamu Maruyama, Takayoshi Shoudai, Emiko Furuichi, Satoru Kuhara, Satoru Miyano |
Discovery Science | 1 |
| 1999 | Designing Views in HypothesisCreator: System for Assisting in Discovery
Osamu Maruyama, Tomoyuki Uchida, Kim Lan Sim, Satoru Miyano |
Discovery Science | 1 |
| 1998 | Toward Genomic Hypothesis Creator: View Designer for Discovery
Osamu Maruyama, Tomoyuki Uchida, Takayoshi Shoudai, Satoru Miyano |
Discovery Science | 1 |
| 1998 | Identification of Gene Regulatory Networks by Strategic Gene Disruptions and Gene Overexpressions
Tatsuya Akutsu, Satoru Kuhara, Osamu Maruyama, Satoru Miyano |
SODA | 3 |
| 1996 | Extracting Best Consensus Motifs from Positive and Negative Examples
Erika Tateishi, Osamu Maruyama, Satoru Miyano |
STACS | 2 |
| 1996 | Inferring a Tree from Walks
Osamu Maruyama, Satoru Miyano |
Theor. Comput. Sci. | 1 |
| 1995 | Graph Inference from a Walk for TRees of Bounded Degree 3 is NP-Complete
Osamu Maruyama, Satoru Miyano |
MFCS | 1 |
| 1992 | Inferring a Tree from Walks
Osamu Maruyama, Satoru Miyano |
MFCS | 1 |