Hiroyuki Kurata

dblp:64/3014 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
6since 2021 · last 2022
0000-0003-4254-2214ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2022 iACVP: markedly enhanced identification of anti-coronavirus peptides using a dataset-specific word2vec model
abstract
The COVID-19 pandemic caused several million deaths worldwide. Development of anti-coronavirus drugs is thus urgent. Unlike conventional non-peptide drugs, antiviral peptide drugs are highly specific, easy to synthesize and modify, and not highly susceptible to drug resistance. To reduce the time and expense involved in screening thousands of peptides and assaying their antiviral activity, computational predictors for identifying anti-coronavirus peptides (ACVPs) are needed. However, few experimentally verified ACVP samples are available, even though a relatively large number of antiviral peptides (AVPs) have been discovered. In this study, we attempted to predict ACVPs using an AVP dataset and a small collection of ACVPs. Using conventional features, a binary profile and a word-embedding word2vec (W2V), we systematically explored five different machine learning methods: Transformer, Convolutional Neural Network, bidirectional Long Short-Term Memory, Random Forest (RF) and Support Vector Machine. Via exhaustive searches, we found that the RF classifier with W2V consistently achieved better performance on different datasets. The two main controlling factors were: (i) the dataset-specific W2V dictionary was generated from the training and independent test datasets instead of the widely used general UniProt proteome and (ii) a systematic search was conducted and determined the optimal k-mer value in W2V, which provides greater discrimination between positive and negative samples. Therefore, our proposed method, named iACVP, consistently provides better prediction performance compared with existing state-of-the-art methods. To assist experimentalists in identifying putative ACVPs, we implemented our model as a web server accessible via the following link: http://kurata35.bio.kyutech.ac.jp/iACVP.
Hiroyuki Kurata, Sho Tsukiyama, Balachandran Manavalan
Briefings Bioinform.1
2022 BERT6mA: prediction of DNA N6-methyladenine site using deep learning-based approaches
abstract
N6-methyladenine (6mA) is associated with important roles in DNA replication, DNA repair, transcription, regulation of gene expression. Several experimental methods were used to identify DNA modifications. However, these experimental methods are costly and time-consuming. To detect the 6mA and complement these shortcomings of experimental methods, we proposed a novel, deep leaning approach called BERT6mA. To compare the BERT6mA with other deep learning approaches, we used the benchmark datasets including 11 species. The BERT6mA presented the highest AUCs in eight species in independent tests. Furthermore, BERT6mA showed higher and comparable performance with the state-of-the-art models while the BERT6mA showed poor performances in a few species with a small sample size. To overcome this issue, pretraining and fine-tuning between two species were applied to the BERT6mA. The pretrained and fine-tuned models on specific species presented higher performances than other models even for the species with a small sample size. In addition to the prediction, we analyzed the attention weights generated by BERT6mA to reveal how the BERT6mA model extracts critical features responsible for the 6mA prediction. To facilitate biological sciences, the BERT6mA online web server and its source codes are freely accessible at https://github.com/kuratahiroyuki/BERT6mA.git, respectively.
Sho Tsukiyama, Md. Mehedi Hasan 0002, Hong-Wen Deng, Hiroyuki Kurata
Briefings Bioinform.4
2022 MLAGO: machine learning-aided global optimization for Michaelis constant estimation of kinetic modeling
abstract
Abstract Background Kinetic modeling is a powerful tool for understanding the dynamic behavior of biochemical systems. For kinetic modeling, determination of a number of kinetic parameters, such as the Michaelis constant (Km), is necessary, and global optimization algorithms have long been used for parameter estimation. However, the conventional global optimization approach has three problems: (i) It is computationally demanding. (ii) It often yields unrealistic parameter values because it simply seeks a better model fitting to experimentally observed behaviors. (iii) It has difficulty in identifying a unique solution because multiple parameter sets can allow a kinetic model to fit experimental data equally well (the non-identifiability problem). Results To solve these problems, we propose the Machine Learning-Aided Global Optimization (MLAGO) method for Km estimation of kinetic modeling. First, we use a machine learning-based Km predictor based only on three factors: EC number, KEGG Compound ID, and Organism ID, then conduct a constrained global optimization-based parameter estimation by using the machine learning-predicted Km values as the reference values. The machine learning model achieved relatively good prediction scores: RMSE = 0.795 and R2 = 0.536, making the subsequent global optimization easy and practical. The MLAGO approach reduced the error between simulation and experimental data while keeping Km values close to the machine learning-predicted values. As a result, the MLAGO approach successfully estimated Km values with less computational cost than the conventional method. Moreover, the MLAGO approach uniquely estimated Km values, which were close to the measured values. Conclusions MLAGO overcomes the major problems in parameter estimation, accelerates kinetic modeling, and thus ultimately leads to better understanding of complex cellular systems. The web application for our machine learning-based Km predictor is accessible at https://sites.google.com/view/kazuhiro-maeda/software-tools-web-apps , which helps modelers perform MLAGO on their own parameter estimation tasks.
Kazuhiro Maeda, Aoi Hatae, Yukie Sakai, Fred C. Boogerd, Hiroyuki Kurata
BMC Bioinform.5
2021 NeuroPred-FRL: an interpretable prediction model for identifying neuropeptide using feature representation learning
abstract
Neuropeptides (NPs) are the most versatile neurotransmitters in the immune systems that regulate various central anxious hormones. An efficient and effective bioinformatics tool for rapid and accurate large-scale identification of NPs is critical in immunoinformatics, which is indispensable for basic research and drug development. Although a few NP prediction tools have been developed, it is mandatory to improve their NPs' prediction performances. In this study, we have developed a machine learning-based meta-predictor called NeuroPred-FRL by employing the feature representation learning approach. First, we generated 66 optimal baseline models by employing 11 different encodings, six different classifiers and a two-step feature selection approach. The predicted probability scores of NPs based on the 66 baseline models were combined to be deemed as the input feature vector. Second, in order to enhance the feature representation ability, we applied the two-step feature selection approach to optimize the 66-D probability feature vector and then inputted the optimal one into a random forest classifier for the final meta-model (NeuroPred-FRL) construction. Benchmarking experiments based on both cross-validation and independent tests indicate that the NeuroPred-FRL achieves a superior prediction performance of NPs compared with the other state-of-the-art predictors. We believe that the proposed NeuroPred-FRL can serve as a powerful tool for large-scale identification of NPs, facilitating the characterization of their functional mechanisms and expediting their applications in clinical therapy. Moreover, we interpreted some model mechanisms of NeuroPred-FRL by leveraging the robust SHapley Additive exPlanation algorithm.
Md. Mehedi Hasan 0002, Md. Ashad Alam, Watshara Shoombuatong, Hong-Wen Deng, Balachandran Manavalan, Hiroyuki Kurata
Briefings Bioinform.6
2021 Meta-i6mA: an interspecies predictor for identifying DNA N6-methyladenine sites of plant genomes by exploiting informative features in an integrative machine-learning framework
abstract
DNA N6-methyladenine (6mA) represents important epigenetic modifications, which are responsible for various cellular processes. The accurate identification of 6mA sites is one of the challenging tasks in genome analysis, which leads to an understanding of their biological functions. To date, several species-specific machine learning (ML)-based models have been proposed, but majority of them did not test their model to other species. Hence, their practical application to other plant species is quite limited. In this study, we explored 10 different feature encoding schemes, with the goal of capturing key characteristics around 6mA sites. We selected five feature encoding schemes based on physicochemical and position-specific information that possesses high discriminative capability. The resultant feature sets were inputted to six commonly used ML methods (random forest, support vector machine, extremely randomized tree, logistic regression, naïve Bayes and AdaBoost). The Rosaceae genome was employed to train the above classifiers, which generated 30 baseline models. To integrate their individual strength, Meta-i6mA was proposed that combined the baseline models using the meta-predictor approach. In extensive independent test, Meta-i6mA showed high Matthews correlation coefficient values of 0.918, 0.827 and 0.635 on Rosaceae, rice and Arabidopsis thaliana, respectively and outperformed the existing predictors. We anticipate that the Meta-i6mA can be applied across different plant species. Furthermore, we developed an online user-friendly web server, which is available at http://kurata14.bio.kyutech.ac.jp/Meta-i6mA/.
Md. Mehedi Hasan 0002, Shaherin Basith, Mst. Shamima Khatun, Gwang Lee, Balachandran Manavalan, Hiroyuki Kurata
Briefings Bioinform.6
2021 LSTM-PHV: prediction of human-virus protein-protein interactions by LSTM with word2vec
abstract
Viral infection involves a large number of protein-protein interactions (PPIs) between human and virus. The PPIs range from the initial binding of viral coat proteins to host membrane receptors to the hijacking of host transcription machinery. However, few interspecies PPIs have been identified, because experimental methods including mass spectrometry are time-consuming and expensive, and molecular dynamic simulation is limited only to the proteins whose 3D structures are solved. Sequence-based machine learning methods are expected to overcome these problems. We have first developed the LSTM model with word2vec to predict PPIs between human and virus, named LSTM-PHV, by using amino acid sequences alone. The LSTM-PHV effectively learnt the training data with a highly imbalanced ratio of positive to negative samples and achieved AUCs of 0.976 and 0.973 and accuracies of 0.984 and 0.985 on the training and independent datasets, respectively. In predicting PPIs between human and unknown or new virus, the LSTM-PHV learned greatly outperformed the existing state-of-the-art PPI predictors. Interestingly, learning of only sequence contexts as words is sufficient for PPI prediction. Use of uniform manifold approximation and projection demonstrated that the LSTM-PHV clearly distinguished the positive PPI samples from the negative ones. We presented the LSTM-PHV online web server and support data that are freely available at http://kurata35.bio.kyutech.ac.jp/LSTM-PHV.
Sho Tsukiyama, Md. Mehedi Hasan 0002, Satoshi Fujii, Hiroyuki Kurata
Briefings Bioinform.4
2019 Improvement of the memory function of a mutual repression network in a stochastic environment by negative autoregulation
abstract
BACKGROUND: Cellular memory is a ubiquitous function of biological systems. By generating a sustained response to a transient inductive stimulus, often due to bistability, memory is central to the robust control of many important biological processes. However, our understanding of the origins of cellular memory remains incomplete. Stochastic fluctuations that are inherent to most biological systems have been shown to hamper memory function. Yet, how stochasticity changes the behavior of genetic circuits is generally not clear from a deterministic analysis of the network alone. Here, we apply deterministic rate equations, stochastic simulations, and theoretical analyses of Fokker-Planck equations to investigate how intrinsic noise affects the memory function in a mutual repression network. RESULTS: We find that the addition of negative autoregulation improves the persistence of memory in a small gene regulatory network by reducing stochastic fluctuations. Our theoretical analyses reveal that this improved memory function stems from an increased stability of the steady states of the system. Moreover, we show how the tuning of critical network parameters can further enhance memory. CONCLUSIONS: Our work illuminates the power of stochastic and theoretical approaches to understanding biological circuits, and the importance of considering stochasticity when designing synthetic circuits with memory function.
A. B. M. Shamim Ul Hasan, Hiroyuki Kurata, Sebastian Pechmann
BMC Bioinform.2
2018 iLMS, Computational Identification of Lysine-Malonylation Sites by Combining Multiple Sequence Features
abstract
Lysine malonylation is a newly discovered post-translational modification of proteins, which plays an important role in regulating many cellular functions. Several approaches are available to identify malonylation proteins and its malonylation sites, however; experimental identification of malonylation sites is often laborious and costly. Therefore, computational schemes are needed to identify potential malonylation sites prior to in vitro experimentation. In this paper, a novel computational scheme iLMS (Identification of Lysine-Malonylation Sites) has been developed by combining primary sequences and evolutionary features via a support vector machine classifier. The final iLMS scheme achieved a robust performance in cross-validation test in both human and mouse datasets. For the mouse data, the iLMS predictor outperformed other existing implementations. The iLMS is a promising computational scheme for the prediction of malonylation sites.
Md. Mehedi Hasan 0002, Hiroyuki Kurata
BIBE2
2018 SIPMA: A Systematic Identification of Protein-Protein Interactions in Zea mays Using Autocorrelation Features in a Machine-Learning Framework
abstract
Zea mays (maize) is one of the most vital crops which are grown widely in the world. To understand the molecular structures and functions of maize, the identification of protein-protein interaction (PPI) is very important. PPI identification by wet lab experiments is time-consuming, expensive and laborious. These days in silico methods that accurately predict potential PPIs based on protein sequence information are highly demanded. Research on PPI prediction in maize is currently very limited, and no dedicated bioinformatics schemes are available. In this work, we proposed a novel approach, termed SIPMA (Systematic Identification of PPI in Maize using Autocorrelation). A machine learning random forest classifier was trained with autocorrelation features to build the prediction model. The SIPMA, which was tested by the experimentally verified PPI dataset of maize, yielded a prediction accuracy of 0.899 when the specificity was 0.969 on the training set. The SIPMA achieved promising performances on the test datasets. Compared with different sequence-based encoding and statistical learning methods, the SIPMA was a powerful computational resource for identifying PPIs in maize.
Mst. Shamima Khatun, Md. Mehedi Hasan 0002, Md. Nurul Haque Mollah, Hiroyuki Kurata
BIBE4
2014 BioFNet: biological functional network database for analysis and synthesis of biological systems
abstract
In synthetic biology and systems biology, a bottom-up approach can be used to construct a complex, modular, hierarchical structure of biological networks. To analyze or design such networks, it is critical to understand the relationship between network structure and function, the mechanism through which biological parts or biomolecules are assembled into building blocks or functional networks. A functional network is defined as a subnetwork of biomolecules that performs a particular function. Understanding the mechanism of building functional networks would help develop a methodology for analyzing the structure of large-scale networks and design a robust biological circuit to perform a target function. We propose a biological functional network database, named BioFNet, which can cover the whole cell at the level of molecular interactions. The BioFNet takes an advantage in implementing the simulation program for the mathematical models of the functional networks, visualizing the simulated results. It presents a sound basis for rational design of biochemical networks and for understanding how functional networks are assembled to create complex high-level functions, which would reveal design principles underlying molecular architectures.
Hiroyuki Kurata, Kazuhiro Maeda, Toshikazu Onaka, Takenori Takata
Briefings Bioinform.1
2012 Biological Design Principles of Complex Feedback Modules in the E. coli Ammonia Assimilation System
abstract
To synthesize natural or artificial life, it is critically important to understand the design principles of how biochemical networks generate particular cellular functions and evolve complex systems in comparison with engineering systems. Cellular systems maintain their robustness in the face of perturbations arising from environmental and genetic variations. In analogy to control engineering architectures, the complexity of modular structures within a cell can be attributed to the necessity of achieving robustness. To reveal such biological design, the E. coli ammonia assimilation system is analyzed, which consists of complex but highly structured modules: the glutamine synthetase (GS) activity feedback control module with bifunctional enzyme cascades for catalyzing reversible reactions, and the GS synthesis feedback control module with positive and negative feedback loops. We develop a full-scale dynamic model that unifies the two modules, and we analyze its robustness and fine tuning with respect to internal and external perturbations. The GS activity control is added to the GS synthesis module to improve its transient response to ammonia depletion, compensating the tradeoffs of each module, but its robustness to internal perturbations is lost. These findings suggest some design principles necessary for the synthesis of life.
Koichi Masaki, Kazuhiro Maeda, Hiroyuki Kurata
Artif. Life3
2009 Genetic modification of flux for flux prediction of mutants
abstract
MOTIVATION: Gene deletion and overexpression are critical technologies for designing or improving the metabolic flux distribution of microbes. Some algorithms including flux balance analysis (FBA) and minimization of metabolic adjustment (MOMA) predict a flux distribution from a stoichiometric matrix in the mutants in which some metabolic genes are deleted or non-functional, but there are few algorithms that predict how a broad range of genetic modifications, such as over- and underexpression of metabolic genes, alters the phenotypes of the mutants at the metabolic flux level. RESULTS: To overcome such existing limitations, we develop a novel algorithm that predicts the flux distribution of the mutants with a broad range of genetic modification, based on elementary mode analysis. It is denoted as genetic modification of flux (GMF), which couples two algorithms that we have developed: modified control effective flux (mCEF) and enzyme control flux (ECF). mCEF is proposed based on CEF to estimate the gene expression patterns in genetically modified mutants in terms of specific biological functions. GMF is demonstrated to predict the flux distribution of not only gene deletion mutants, but also the mutants with underexpressed and overexpressed genes in Escherichia coli and Corynebacterium glutamicum. This achieves breakthrough in the a priori flux prediction of a broad range of genetically modified mutants. SUPPLEMENTARY INFORMATION: Supplementary file and programs are available at Bioinformatics online or http://www.cadlive.jp.
Quanyu Zhao, Hiroyuki Kurata
Bioinform.2
2006 Effective and Fast Optimization for a Dynamic Model of the Drosophila Circadian Oscillator
abstract
In the field of systems biology, biochemical networks are being reconstructed in computer to understand their dynamic features. In this study, we focused on the estimation problem for a dynamic model of the Drosophila circadian oscillator, which is defined by two evaluation functions that represent oscillatory features. However, since the search space is multimodal and logarithmically large, it is quite difficult for ordinary GAs to optimize this evaluation problem. Therefore, we had used two-step optimizing, a random search with GA. It successfully optimized the circadian oscillator, but required a long calculation time, causing a local search. On the other hand, to optimize two evaluation functions simultaneously, we proposed the survival ratio GA, where genes have lifetimes and a population holds search histories. By alternating two evaluation functions every generation, the population holds previously evaluated genes and children inherit each and mixed properties. The survival ratio GA exhibited wider space search and higher success ratio, thus we applied the one-step optimizing with the survival ratio GA to the circadian estimation problem, which shortens the total calculation times and finds out more local optimums than the previous two-step optimizing method.
Tanaka Shin, Hiroyuki Kurata, Takeshi Ohashi 0001
SMC2
2006 Module-Based Analysis of Robustness Tradeoffs in the Heat Shock Response System
abstract
Biological systems have evolved complex regulatory mechanisms, even in situations where much simpler designs seem to be sufficient for generating nominal functionality. Using module-based analysis coupled with rigorous mathematical comparisons, we propose that in analogy to control engineering architectures, the complexity of cellular systems and the presence of hierarchical modular structures can be attributed to the necessity of achieving robustness. We employ the Escherichia coli heat shock response system, a strongly conserved cellular mechanism, as an example to explore the design principles of such modular architectures. In the heat shock response system, the sigma-factor sigma32 is a central regulator that integrates multiple feedforward and feedback modules. Each of these modules provides a different type of robustness with its inherent tradeoffs in terms of transient response and efficiency. We demonstrate how the overall architecture of the system balances such tradeoffs. An extensive mathematical exploration nevertheless points to the existence of an array of alternative strategies for the existing heat shock response that could exhibit similar behavior. We therefore deduce that the evolutionary constraints facing the system might have steered its architecture toward one of many robustly functional solutions.
Hiroyuki Kurata, Hana El-Samad, Rei Iwasaki, Hisao Ohtake, John Doyle 0001, Irina Grigorova, Carol A. Gross, Mustafa Khammash
PLoS Comput. Biol.1
2005 Survival ratio GA for two-evaluation problem in parameter tuning of heat shock response in E. coli
abstract
In the field of bioinformatics, several studies have attempted to reconstruct biological reactions on a computer in order to understand their essence. In this study, we optimized the parameter tuning problem of the heat shock response in E. coli, one of the gene regulatory network simulations, which exhibits a transient peak by genetic algorithms (GAs). However, if the search area is too large, the optimization performance deteriorated significantly. To optimize this problem more efficiently, we defined two evaluation functions that represent two features of the simulation and used the survival ratio GA. This GA has a gene's lifetime as a new concept, that is, the population holds previous search histories. By alternating between two evaluation functions every generation, the population holds both previously evaluated genes and children inherit both properties. In the two-evaluation problem of the heat shock response, the survival ratio GA exhibited a considerably better optimization performance than traditional GA methods.
Tanaka Shin, Hiroyuki Kurata, Takeshi Ohashi 0001
SMC2
2005 A grid layout algorithm for automatic drawing of biochemical networks
abstract
MOTIVATION: Visualization is indispensable in the research of complex biochemical networks. Available graph layout algorithms are not adequate for satisfactorily drawing such networks. New methods are required to visualize automatically the topological architectures and facilitate the understanding of the functions of the networks. RESULTS: We propose a novel layout algorithm to draw complex biochemical networks. A network is modeled as a system of interacting nodes on squared grids. A discrete cost function between each node pair is designed based on the topological relation and the geometric positions of the two nodes. The layouts are produced by minimizing the total cost. We design a fast algorithm to minimize the discrete cost function, by which candidate layouts can be produced efficiently. A simulated annealing procedure is used to choose better candidates. Our algorithm demonstrates its ability to exhibit cluster structures clearly in relatively compact layout areas without any prior knowledge. We developed Windows software to implement the algorithm for CADLIVE. AVAILABILITY: All materials can be freely downloaded from http://kurata21.bio.kyutech.ac.jp/grid/grid_layout.htm; http://www.cadlive.jp/ SUPPLEMENTARY INFORMATION: http://kurata21.bio.kyutech.ac.jp/grid/grid_layout.htm; http://www.cadlive.jp/
Weijiang Li, Hiroyuki Kurata
Bioinform.2