VLDB 2026 Research / reviewers in the wild / expert
Yasushi Okuno
dblp:79/6343
· DBLP profile ↗
11ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0003-3596-4208ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Diverse, Stable, and Efficient All-Atom Structure Sampling Based on Parallelized Coarse-Grained Elastic Network ModelabstractWe propose a methodology for efficiently obtaining diverse and stable all-atom (AA) protein structures on the basis of coarse-grained molecular dynamics (CG-MD) simulations and back-mapping to AA structures. In computational drug discovery, expanding the sampling space of MD simulations and identifying diverse conformational states of target proteins is crucial. The CG elastic network model (ENM) is suitable for expanding the sampling space due to its low computational complexity and ease of state transitions. However, detailed analysis of sampled conformations and their application to drug discovery requires reconstructing$\mathbf{C G}$structures into$\mathbf{A A}$structures. To obtain stable AA structures, we need to address two major challenges: (1) determining an appropriate parameter set for CG-ENM, and (2) back-mapping from sampled CG structures to stable AA structures. Since these operations consume a significant part of execution time, we developed a novel approach to leverage the abundant computing resources of a supercomputer by parallelizing both operations. Moreover, we applied dynamic load balancing to the parallelized implementation to mitigate the increase in wall time caused by load imbalances. We also optimized the implementation for the supercomputer Fugaku to maximize performance by minimizing memory and disk contentions. A detailed performance analysis shows that our implementation reduced node-hours by over 86.79% compared with replica-exchange MD simulations, demonstrating 2.156 to 4.062 times faster sampling speed. Shingo Okuno, Yuta Yoshimoto, Nana Takeda, Yusuke Nagasaka, Ryo Kanada, Atsushi Tokuhisa, Yasushi Okuno |
CCGrid | 7 |
| 2024 | Subgrouping Causal Networks of Disease Onset in Large-scale Health and Medical Data using Supercomputer FugakuabstractBayesian networks can deduce statistical causal relationships from observed data. When applied to a large-scale health and medical dataset, it becomes feasible to employ the deduced networks to identify potential factors related to disease onset. Factors contributing to the onset of lifestyle-related diseases, such as the social environment and habits, vary significantly among individuals. Thus, it can be hypothesized that networks illustrating disease onset mechanisms would also exhibit substantial diversity. However, typical statistical causal discovery methods challenge the analysis of relationships specific to the sub-groups in a dataset because they use the entire data. In response to this, we use a pattern mining technique for Iwaki Health Promotion Project Health Checkup data to derive subgroups exhibiting strong correlations with the target variables. We estimated the Bayesian networks for the characteristic subgroups out of those derived, and compared them with the Bayesian network estimated for the total (hereafter, base network). Our target was the onset of eight lifestyle-related diseases within three years, resulting in a total of 359 subgroups. By comparing the estimated subgroup networks with the base network, we confirmed the numerous relationships specific to the subgroup networks. These encompassed not only clinically known but also non-trivial relationships. Our approach, which uses target-wise correlation-based rule subgrouping and network estimation is beneficial for constructing hypotheses on the differences in disease onset causes among potential subgroups. Taisei Tosaki, Eiichiro Uchino, Yohei Harada, Minoru Sakuragi, Yusuke Koyanagi, Seiji Okajima, Hirofumi Suzuki, Kentaro Kanamori, Masahiro Asaoka, Kouji Kurihara, Takuya Takagi, Koji Maruhashi, Yoshinori Tamada, Tatsuya Mikami, Koichi Murashita, Shigeyuki Nakaji, Yasushi Okuno |
BIBM | 17 |
| 2023 | An Auto-Encoder to Reconstruct Structure with Cryo-EM Images via Theoretically Guaranteed Isometric Latent Space, and Its Application for Automatically Computing the Conformational Pathway
Kimihiro Yamazaki, Yuichiro Wada, Atsushi Tokuhisa, Mutsuyo Wada, Takashi Katoh, Yuhei Umeda, Yasushi Okuno, Akira Nakagawa |
MICCAI (1) | 7 |
| 2023 | Network-based prediction approach for cancer-specific driver missense mutations using a graph neural networkabstractBACKGROUND: In cancer genomic medicine, finding driver mutations involved in cancer development and tumor growth is crucial. Machine-learning methods to predict driver missense mutations have been developed because variants are frequently detected by genomic sequencing. However, even though the abnormalities in molecular networks are associated with cancer, many of these methods focus on individual variants and do not consider molecular networks. Here we propose a new network-based method, Net-DMPred, to predict driver missense mutations considering molecular networks. Net-DMPred consists of the graph part and the prediction part. In the graph part, molecular networks are learned by a graph neural network (GNN). The prediction part learns whether variants are driver variants using features of individual variants combined with the graph features learned in the graph part. RESULTS: Net-DMPred, which considers molecular networks, performed better than conventional methods. Furthermore, the prediction performance differed by the molecular network structure used in learning, suggesting that it is important to consider not only the local network related to cancer but also the large-scale network in living organisms. CONCLUSIONS: We propose a network-based machine learning method, Net-DMPred, for predicting cancer driver missense mutations. Our method enables us to consider the entire graph architecture representing the molecular network because it uses GNN. Net-DMPred is expected to detect driver mutations from a lot of missense mutations that are not known to be associated with cancer. Narumi Hatano, Mayumi Kamada, Ryosuke Kojima, Yasushi Okuno |
BMC Bioinform. | 4 |
| 2023 | Individual health-disease phase diagrams for disease prevention based on machine learningabstractEarly disease detection and prevention methods based on effective interventions are gaining attention worldwide. Progress in precision medicine has revealed that substantial heterogeneity exists in health data at the individual level and that complex health factors are involved in chronic disease development. Machine-learning techniques have enabled precise personal-level disease prediction by capturing individual differences in multivariate data. However, it is challenging to identify what aspects should be improved for disease prevention based on future disease-onset prediction because of the complex relationships among multiple biomarkers. Here, we present a health-disease phase diagram (HDPD) that represents an individual's health state by visualizing the future-onset boundary values of multiple biomarkers that fluctuate early in the disease progression process. In HDPDs, future-onset predictions are represented by perturbing multiple biomarker values while accounting for dependencies among variables. We constructed HDPDs for 11 diseases using longitudinal health checkup cohort data of 3,238 individuals, comprising 3,215 measurement items and genetic data. The improvement of biomarker values to the non-onset region in HDPD remarkably prevented future disease onset in 7 out of 11 diseases. HDPDs can represent individual physiological states in the onset process and be used as intervention goals for disease prevention. Eiichiro Uchino, Noriaki Sato, Ayano Araki, Kei Terayama, Ryosuke Kojima, Koichi Murashita, Ken Itoh, Tatsuya Mikami, Yoshinori Tamada, Yasushi Okuno |
J. Biomed. Informatics | 11 |
| 2022 | CBNplot: Bayesian network plots for enrichment analysisabstractSUMMARY: When investigating gene expression profiles, determining important directed edges between genes can provide valuable insights in addition to identifying differentially expressed genes. In the subsequent functional enrichment analysis (EA), understanding how enriched pathways or genes in the pathway interact with one another can help infer the gene regulatory network (GRN), important for studying the underlying molecular mechanisms. However, packages for easy inference of the GRN based on EA are scarce. Here, we developed an R package, CBNplot, which infers the Bayesian network (BN) from gene expression data, explicitly utilizing EA results obtained from curated biological pathway databases. The core features include convenient wrapping for structure learning, visualization of the BN from EA results, comparison with reference networks, and reflection of gene-related information on the plot. As an example, we demonstrate the analysis of bladder cancer-related datasets using CBNplot, including probabilistic reasoning, which is a unique aspect of BN analysis. We display the transformability of results obtained from one dataset to another, the validity of the analysis as assessed using established knowledge and literature, and the possibility of facilitating knowledge discovery from gene expression datasets. AVAILABILITY AND IMPLEMENTATION: The library, documentation and web server are available at https://github.com/noriakis/CBNplot. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Noriaki Sato, Yoshinori Tamada, Guangchuang Yu, Yasushi Okuno |
Bioinform. | 4 |
| 2018 | Machine learning accelerates MD-based binding pose prediction between ligands and proteinsabstractMotivation: Fast and accurate prediction of protein-ligand binding structures is indispensable for structure-based drug design and accurate estimation of binding free energy of drug candidate molecules in drug discovery. Recently, accurate pose prediction methods based on short Molecular Dynamics (MD) simulations, such as MM-PBSA and MM-GBSA, among generated docking poses have been used. Since molecular structures obtained from MD simulation depend on the initial condition, taking the average over different initial conditions leads to better accuracy. Prediction accuracy of protein-ligand binding poses can be improved with multiple runs at different initial velocity. Results: This paper shows that a machine learning method, called Best Arm Identification, can optimally control the number of MD runs for each binding pose. It allows us to identify a correct binding pose with a minimum number of total runs. Our experiment using three proteins and eight inhibitors showed that the computational cost can be reduced substantially without sacrificing accuracy. This method can be applied for controlling all kinds of molecular simulations to obtain best results under restricted computational resources. Availability and implementation: Code and data are available on GitHub at https://github.com/tsudalab/bpbi. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Kei Terayama, Hiroaki Iwata, Mitsugu Araki, Yasushi Okuno, Koji Tsuda |
Bioinform. | 4 |
| 2013 | On computational complexity of graph inference from counting
Szilárd Zsolt Fazekas, Hiro Ito, Yasushi Okuno, Shinnosuke Seki 0001, Kei Taneishi |
Nat. Comput. | 3 |
| 2012 | On the Behavior of Tile Assembly System at High Temperatures
Shinnosuke Seki 0001, Yasushi Okuno |
CiE | 2 |
| 2010 | A novel chemogenomics analysis of G protein-coupled receptors (GPCRs) and their ligands: a potential strategy for receptor de-orphanizationabstractBACKGROUND: G protein-coupled receptors (GPCRs) represent a family of well-characterized drug targets with significant therapeutic value. Phylogenetic classifications may help to understand the characteristics of individual GPCRs and their subtypes. Previous phylogenetic classifications were all based on the sequences of receptors, adding only minor information about the ligand binding properties of the receptors. In this work, we compare a sequence-based classification of receptors to a ligand-based classification of the same group of receptors, and evaluate the potential to use sequence relatedness as a predictor for ligand interactions thus aiding the quest for ligands of orphan receptors. RESULTS: We present a classification of GPCRs that is purely based on their ligands, complementing sequence-based phylogenetic classifications of these receptors. Targets were hierarchically classified into phylogenetic trees, for both sequence space and ligand (substructure) space. The overall organization of the sequence-based tree and substructure-based tree was similar; in particular, the adenosine receptors cluster together as well as most peptide receptor subtypes (e.g. opioid, somatostatin) and adrenoceptor subtypes. In ligand space, the prostanoid and cannabinoid receptors are more distant from the other targets, whereas the tachykinin receptors, the oxytocin receptor, and serotonin receptors are closer to the other targets, which is indicative for ligand promiscuity. In 93% of the receptors studied, de-orphanization of a simulated orphan receptor using the ligands of related receptors performed better than random (AUC > 0.5) and for 35% of receptors de-orphanization performance was good (AUC > 0.7). CONCLUSIONS: We constructed a phylogenetic classification of GPCRs that is solely based on the ligands of these receptors. The similarities and differences with traditional sequence-based classifications were investigated: our ligand-based classification uncovers relationships among GPCRs that are not apparent from the sequence-based classification. This will shed light on potential cross-reactivity of GPCR ligands and will aid the design of new ligands with the desired activity profiles. In addition, we linked the ligand-based classification with a ligand-focused sequence-based classification described in literature and proved the potential of this method for de-orphanization of GPCRs. Eelke van der Horst, Julio E. Peironcely, Adriaan P. IJzerman, Margot W. Beukers, Jonathan Robert Lane, Herman van Vlijmen, Michael T. M. Emmerich, Yasushi Okuno, Andreas Bender 0002 |
BMC Bioinform. | 8 |
| 2009 | Laplacian Linear Discriminant Analysis Approach to Unsupervised Feature SelectionabstractUntil recently, numerous feature selection techniques have been proposed and found wide applications in genomics and proteomics. For instance, feature/gene selection has proven to be useful for biomarker discovery from microarray and mass spectrometry data. While supervised feature selection has been explored extensively, there are only a few unsupervised methods that can be applied to exploratory data analysis. In this paper, we address the problem of unsupervised feature selection. First, we extend Laplacian linear discriminant analysis (LLDA) to unsupervised cases. Second, we propose a novel algorithm for computing LLDA, which is efficient in the case of high dimensionality and small sample size as in microarray data. Finally, an unsupervised feature selection method, called LLDA-based Recursive Feature Elimination (LLDA-RFE), is proposed. We apply LLDA-RFE to several public data sets of cancer microarrays and compare its performance with those of Laplacian score and SVD-entropy, two state-of-the-art unsupervised methods, and with that of Fisher score, a supervised filter method. Our results demonstrate that LLDA-RFE outperforms Laplacian score and shows favorable performance against SVD-entropy. It performs even better than Fisher score for some of the data sets, despite the fact that LLDA-RFE is fully unsupervised. Satoshi Niijima, Yasushi Okuno |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |