Shigehiko Kanaya

dblp:48/1812 · DBLP profile ↗
← Back
27ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-2147-6900ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 26 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 An Approach to Identify Mechanism of Actions of Predicted Natural Antibiotics
abstract
Antibiotic resistance is a mounting global health challenge, demanding accelerated strategies for discovering new drugs and understanding their mechanism of action (MoA). In this study, we propose a computational framework to predict the MoA of natural antibiotics by leveraging Tanimoto molecular similarity metrics and network-based clustering. A curated dataset of known antibiotics with established MoA's was used to generate a similarity network, into which predicted natural antibiotics were integrated. Using the DPClus algorithm, we identified structurally coherent clusters that frequently aligned with known functional categories. The results suggest that predicted antibiotics sharing structural clusters with characterized compounds are likely to exhibit similar biological activities. Our approach provides an interpretable and scalable method for prioritizing novel antibiotic candidates for further validation, contributing to faster and more informed drug discovery pipelines.
Muhammad Hendrick Sedayu, Ahmad Kamal Nasution, Mahfujul Islam Rumman, Naoaki Ono, Shigehiko Kanaya, Md. Altaf-Ul-Amin
BIBM5
2025 Flexible variational information bottleneck: Achieving diverse compression with a single training
Sota Kudo, Naoaki Ono, Shigehiko Kanaya, Ming Huang 0002
Neurocomputing3
2023 Investigating Potential Natural Antibiotics Plants Based on Unani Formula Using Supervised Network Analysis and Machine Learning Approach
abstract
This study employs a multi-faceted approach to identify potential natural antibiotics derived from plants used in the Unani formula. The first approach involves a supervised network analysis that utilizes distance measurement to analyze the interconnectivity of the network data. This approach identifies 26 candidate plants with the potential for natural antibiotics. The second approach involves machine learning techniques, specifically deep learning methods, to extract relevant features from the Unani data. This method produces a list of 29 potential plants that exhibit characteristics of natural antibiotics. Notably, seven plants overlap between the approaches - Piper longum, Trachyspermum ammi carum copticum, Santalum album, Cyperus rotundus, Vitis vinifera, Matricaria chamomilla and Zingiber officinale - shown in the literature to exhibit antibacterial properties direct or indirect. The study employs two analytical methods to comprehensively identify potential plants with natural antibiotic properties from the Unani formula. These findings could have significant implications for developing novel antibiotics to combat antibiotic-resistant pathogens.
Ahmad Kamal Nasution, Naoaki Ono, Shigehiko Kanaya, Md. Altaf-Ul-Amin
BIBM3
2022 Prediction of Potential Natural Antibiotics based on Jamu Formula Using Machine Learning Approach
abstract
In order to address antibiotics resistance, multi-drug resistance, and superbugs phenomena, our research explored the utility of Jamu ingredients on the molecular level to predict new natural antibiotic candidates. Jamu is one of the popular traditional medicines from Indonesia, with different therapeutics usage including curing diseases caused by bacterial infection. We used three types of machine learning methods, such as Random Forest (RF), Support Vector Machine (SVM), Deep Learning (DL), to classify Jamu formulas according to their effectiveness against different types of bacterial diseases. The best accuracy for RF, SVM, and DL models are 89%, 84%, and 80%, respectively. We extracted the potential compounds based on the best model as candidate antibiotics corresponding to five groups of efficacies, e.g., digestive systems, respiratory systems, reproductive systems, skin and soft tissue, and urinary systems. Overall, we mined 111 compounds, and many of them could be validated by published literature, and considering structural similarities with known antibiotics.
Ahmad Kamal Nasution, Sony Hartono Wijaya, Ming Huang 0002, Naoaki Ono, Shigehiko Kanaya, Md. Altaf-Ul-Amin
BIBE5
2022 Hierarchical Categorical Generative Modeling for Multi-omics Cancer Subtyping
abstract
Identifying a specific cancer subtype from a variety of candidates is vital for precise and effective treatment. However, cancer subtyping is highly non-trivial as a result of cancer heterogeneity. While significant efforts have been put into understanding the mechanism of cancer subtypes via studying the omics data, existing methods run the risk of presenting biased analyses resulted from overfitting the high-dimensional and scarce omics data. In this paper, we propose a novel generative model that directly models the cancer data distribution by which downstream tasks can circumvent the curse of overfitting and achieve better performance. Unlike conventional generative modeling schemes, the proposed method underlines hierarchical categorical latent spaces to extract global features and local details respectively from transcriptomics and genomics profiles, which is the first to be considered in the cancer subtyping literature. By extensive experiments we verify that the proposed architecture achieves more clearly separated subtypes, as well as medically significant insights into real subtyping.
Ziwei Yang 0002, Lingwei Zhu, Chen Li 0027, Zheng Chen 0012, Naoki Ono, Md. Altaf-Ul-Amin, Shigehiko Kanaya
BIBM7
2021 An End-to-End Sleep Staging Simulator Based on Mixed Deep Neural Networks
abstract
Sleep screening is not only a major tool in the assessment of pathophysiology, but also a bridge between the central neuronal systems and behaviour/cognition. Automatic sleep staging is an alternative for the time-consuming gold standard manual scoring procedure. Most of the existing works designed such procedure by using deep neural networks without considering the medical criterion of the sleep staging task. We argue that capturing the stage-specific features which meet the criterion is of significant importance for the automatic sleep staging alternative. In this work we propose an end-to-end sleep staging simulator based on mixed neural networks, i.e., CNN, LSTM, and Transformer. The framework consists of two subnetworks: stage dependent feature mapping network which is constructed by the idea of physiological sleep nature, and an attention-based parallel staging network. Moreover, we adopt a mixed precision training strategy to quantize the model for exploring feasible usage in the clinical settings. Through an experiment with a large EEG database (Sleep Heart Health Study), the proposed method has a competitive stage scoring performance, especially in stages Wake, N2, and N3, with higher precision of 0.92, 0.85, and 0.86, respectively. Our study proves that the quantized model has potential capability for further application in the clinical staging task.
Zheng Chen 0012, Ziwei Yang 0002, Dong Wang 0044, Ming Huang 0002, Naoaki Ono, Md. Altaf-Ul-Amin, Shigehiko Kanaya
BIBM7
2021 An Integrated Multi-Omics Approach for AMR Phenotype Prediction of Gut Microbiota
abstract
The gut microbiota is crucial for human physiology and susceptibility to diseases. Knowing the AMR phenotype canfacilitate the understanding of the impact of antibiotics administration on the gut microbiota. Nowadays, whole-genome sequencing for antibiotic susceptibility testing (WGS-AST) is widely used in clinical microbiology to predict the AMR phenotype. To release the limitations of the genomic information and improve the WGS-AST prediction, we propose an integrated multi-omics approach, employing a deep generative neural network (VAE: variational auto-encoder). We evaluate the proposed approach by two machine learning techniques (i.e., K-means for clustering and Random Forest for classification). Our evaluation results show that the integrated multi-omics approach achieves relatively better performance than the conventional WGS-AST. Moreover, the integrated multi-omics approach is able to visually reveal AMR phenotype of the gutmicrobiota via antibacterial spectrum. Our work provides evidence that multi-omics information is useful to enhance the WGS-AST prediction.
Pei Gao, Zheng Chen 0012, Dong Wang 0044, Ming Huang 0002, Naoaki Ono, Md. Altaf-Ul-Amin, Shigehiko Kanaya
BIBM7
2021 Exploring Feasibility of Truth-Involved Automatic Sleep Staging Combined with Transformer
abstract
Recently, deep learning-based methods have been successfully proposed for electrophysiology signal-based sleep staging with promising results. Most existing methods use convolutional layers and recurrent-based architectures to implement a model structure from feature extraction to sequence signal classification. In this study, we propose a method of segmenting electroencephalogram (EEG) and electrooculogram (EOG) data according to frequency bands and construct a Transformer based automatic sleep classification model on top of it. The results show that the classifications of the stage Wake, N3, and REM outperform the state-of-art works, with the Fl-scores of 0.92, 0.85 and 0.91. Our work is the first attempt to explore the feasibility of a truth-involved Transformer-based model with a large-scale sleep database.
Ziwei Yang 0002, Dong Wang 0044, Zheng Chen 0012, Ming Huang 0002, Naoaki Ono, Md. Altaf-Ul-Amin, Shigehiko Kanaya
BIBM7
2020 iVAE: An Improved Deep Learning Structure for EEG Signal Characterization and Reconstruction
abstract
Due to the inherent variability such as inter-users anatomical variability and the inter-systems differences, the design of new EEG-based index and a reliable model for sleep stages classification is still the main topic in sleep science. The unsupervised deep learning framework-variational autoencoder (VAE) which can capture the major characteristics of the input by imposing a Gaussian prior distribution on the latent features is suitable in EEG characterization and reconstruction. Although vanilla VAE and convolutional autoencoder (CAE) have been tried, it has yet been discussed that whether a deep structure or a multi-scale structure is more appropriate. In this paper, we constructed a shallow iVAE model, which will capture the multi-scale features of the spectrogram of EEG by replacing the main structure in encoder and decoder with the inception-like structure. By comparing with the vanilla VAE and the CAE, a more accurate reconstruction and a better classification using the latent features of the iVAE can be confirmed.
Zheng Chen 0012, Naoaki Ono, Md. Altaf-Ul-Amin, Shigehiko Kanaya, Ming Huang 0002
BIBM4
2020 BiClusO: A Novel Biclustering Approach and Its Application to Species-VOC Relational Data
abstract
In this paper, we propose a novel biclustering approach called BiClusO. Biclustering can be applied to various types of bipartite data such as gene-condition or gene-disease relations. For example, we applied BiClusO to bipartite relations between species and volatile organic compounds (VOCs). VOCs, which are emitted by different species, have huge environmental and ecological impacts. The biosynthesis of VOCs depends on different metabolic pathways which can be used to categorize the species. A previous study related to the KNApSAcK VOC database classified microorganisms based on their VOC profiles, which confirmed the consistency between VOC-based and pathogenicity-based classifications. However, due to limited data, classification of all species in terms of VOC profiles was not performed. In this study, we enriched our database with additional data collected from different online sources and journals. Then, by applying BiClusO to species-VOC relational data, we determined that VOC-based classification is consistent with taxonomy-based classification of the species. We also assessed the diversity of VOC pathways across different kingdoms of species.
Mohammad Bozlul Karim, Ming Huang 0002, Naoaki Ono, Shigehiko Kanaya, Md. Altaf-Ul-Amin
IEEE ACM Trans. Comput. Biol. Bioinform.4
2019 Inter Disease Relations Based on Human Biomarkers by Network Analysis
abstract
A biomarker (short for biological marker) is a medical sign of a disease or condition which indicates a normal or abnormal state of a body. The biomarker is a key factor in the analysis of diseases and also for analyzing inter disease relations. In the previous study, we designed and developed a human biomarker (metabolites and proteins) database and the database is currently available online. This work was supported by the Ministry of Education, Japan and NAIST Big Data Project. We have used our previously developed database and collected 486 human biomarkers and their respective diseases. We determined the similarity among NCBI disease classes based on associated biomarker fingerprints. For this purpose, we collected biomarker PubChem IDs and using them downloaded the SDF files in a batch, then with those molecular description files determined their atom pair fingerprints using ChemmineR package. We constructed a network of biomarkers based on Tanimoto similarity between their fingerprints and applied DPclusO algorithm to find clusters consisting of biomarkers with similar chemical structures. We also conducted hierarchical clustering of the biomarkers. We categorized all the diseases in our data into 18 NCBI disease classes. Combining all information, we finally determined inter disease relations based on structural similarity between biomarkers.
Shaikh Farhad Hossain, Ming Huang 0002, Naoaki Ono, Shigehiko Kanaya, Md. Altaf-Ul-Amin
BIBE4
2019 Cardiotoxicity Prediction Based on Integreted hERG Database with Molecular Convolution Model
abstract
Cardiotoxicity caused by drug candidates and chemical compounds that block hERG channels may lead to malignant ventricular arrhythmias and even sudden cardiac death (SCD). Various in-silico models have been built to predict the cardiotoxicity during early stages of drug design. The largest public database of hERG-related compounds by integrating several major databases has been constructed recently, which made it possible to build more sophisticated machine learning models for accurate prediction of cardiotoxicity. Here we developed a novel molecular graph convolution neural network (MGCNN) model, based on the new integrated database. The MGCNN models were built by altering the number of graph convolutional layers (GC) from 1 to 5. A random forest (RF) model input with the extended-connectivity fingerprint (ECFP) of different maximal radii (1 ~ 5) was built to enable a direct comparison with the MCGNN models. We found that the MGCNN model with 2 GCs has the best performance in terms of the ROC-AUC-score (0.84), whereas the RF model input with ECFP has a stable performance (0.77 ~ 0.80) over the preset radii. The machine learning models promise a potential new approach for harnessing the big data to achieve accurate prediction of drug cardiotoxicity.
Jieying Hu, Ming Huang 0002, Naoaki Ono, Ye Chen-Izu, Leighton T. Izu, Shigehiko Kanaya
BIBM6
2019 Classification of alkaloids according to the starting substances of their biosynthetic pathways using graph convolutional neural networks
abstract
BACKGROUND: Alkaloids, a class of organic compounds that contain nitrogen bases, are mainly synthesized as secondary metabolites in plants and fungi, and they have a wide range of bioactivities. Although there are thousands of compounds in this class, few of their biosynthesis pathways are fully identified. In this study, we constructed a model to predict their precursors based on a novel kind of neural network called the molecular graph convolutional neural network. Molecular similarity is a crucial metric in the analysis of qualitative structure-activity relationships. However, it is sometimes difficult for current fingerprint representations to emphasize specific features for the target problems efficiently. It is advantageous to allow the model to select the appropriate features according to data-driven decisions for extracting more useful information, which influences a classification or regression problem substantially. RESULTS: In this study, we applied a neural network architecture for undirected graph representation of molecules. By encoding a molecule as an abstract graph and applying "convolution" on the graph and training the weight of the neural network framework, the neural network can optimize feature selection for the training problem. By incorporating the effects from adjacent atoms recursively, graph convolutional neural networks can extract the features of latent atoms that represent chemical features of a molecule efficiently. In order to investigate alkaloid biosynthesis, we trained the network to distinguish the precursors of 566 alkaloids, which are almost all of the alkaloids whose biosynthesis pathways are known, and showed that the model could predict starting substances with an averaged accuracy of 97.5%. CONCLUSION: We have showed that our model can predict more accurately compared to the random forest and general neural network when the variables and fingerprints are not selected, while the performance is comparable when we carefully select 507 variables from 18000 dimensions of descriptors. The prediction of pathways contributes to understanding of alkaloid synthesis mechanisms and the application of graph based neural network models to similar problems in bioinformatics would therefore be beneficial. We applied our model to evaluate the precursors of biosynthesis of 12000 alkaloids found in various organisms and found power-low-like distribution.
Ryohei Eguchi, Naoaki Ono, Aki Hirai, Tetsuo Katsuragi, Satoshi Nakamura 0001, Ming Huang 0002, Md. Altaf-Ul-Amin, Shigehiko Kanaya
BMC Bioinform.8
2018 Prediction of Plant-Disease Relations Based on Unani Formulas by Network Analysis
abstract
Various medicinal plants are available in Bangladesh and these plants are used as traditional medicines for healing and health maintenance. Unani is one of the traditional medicine systems popular among Bangladeshi people because of its high success rate. Disease phenotype is changing constantly. It is Challenging for researchers to get the right medicinal ingredients, for the right disease, within a reasonable time. So we need to analyze the right plants for the right disease based on the existing formulas and to find out the relationship between plant and disease. The predicted plant-disease relations will help the health researcher or pharmacist for finding new drugs for new diseases. In our datasets, we have 409 plants, which are used as ingredients of 609 Unani formulas. Based on 609 formulas, we enlisted and sorted the relationship between diseases and plants. We assigned 609 Unani formulas to 18 National Center for Biotechnology Information (NCBI) disease classes. We then constructed the network of Unani formulas based on their ingredient similarity and applied DPclusO algorithm to find clusters. Clusters are associated with dominant disease and dominant plants by voting thus we established relations between plants and diseases. We predicted associations between 12 diseases and 151 plants. We validated our prediction based on the global set of Unani formulas and obtained 85.57% accuracy
Shaikh Farhad Hossain, Sony Hartono Wijaya, Ming Huang 0002, Irmanida Batubara, Shigehiko Kanaya, Md. Altaf-Ul-Amin
BIBE5
2018 Feature extraction and Cluster analysis of Pancreatic Pathological Image Based on Unsupervised Convolutional Neural Network
Konosuke Asanou, Naoaki Ono, Chika Iwamoto, Kenoki Ohuchida, Koji Shindo, Shigehiko Kanaya
BIBM6
2018 UC2 search: using unique connectivity of uncharged compounds for metabolite annotation by database searching in mass spectrometry-based metabolomics
abstract
Summary: For metabolite annotation in metabolomics, variations in the registered states of compounds (charged molecules and multiple components, such as salts) and their redundancy among compound databases could be the cause of misannotations and hamper immediate recognition of the uniqueness of metabolites while searching by mass values measured using mass spectrometry. We developed a search system named UC2 (Unique Connectivity of Uncharged Compounds), where compounds are tentatively neutralized into uncharged states and stored on the basis of their unique connectivity of atoms after removing their stereochemical information using the first block in the hash of the IUPAC International Chemical Identifier, by which false-positive hits are remarkably reduced, both charged and uncharged compounds are properly searched in a single query and records having a unique connectivity are compiled in a single search result. Availability and implementation: The UC2 search tool is available free of charge as a REST web service (http://webs2.kazusa.or.jp/mfsearcher) and a Java-based GUI tool. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Nozomu Sakurai, Takafumi Narise, Joon-Soo Sim, Chang-Muk Lee, Chiaki Ikeda, Nayumi Akimoto, Shigehiko Kanaya
Bioinform.7
2018 An integrative network-based approach to identify novel disease genes and pathways: a case study in the context of inflammatory bowel disease
abstract
BACKGROUND: There are different and complicated associations between genes and diseases. Finding the causal associations between genes and specific diseases is still challenging. In this work we present a method to predict novel associations of genes and pathways with inflammatory bowel disease (IBD) by integrating information of differential gene expression, protein-protein interaction and known disease genes related to IBD. RESULTS: We downloaded IBD gene expression data from NCBI's Gene Expression Omnibus, performed statistical analysis to determine differentially expressed genes, collected known IBD genes from DisGeNet database, which were used to construct a IBD related PPI network with HIPPIE database. We adapted our graph-based clustering algorithm DPClusO to cluster the disease PPI network. We evaluated the statistical significance of the identified clusters in the context of determining the richness of IBD genes using Fisher's exact test and predicted novel genes related to IBD. We showed 93.8% of our predictions are correct in the context of other databases and published literatures related to IBD. CONCLUSIONS: Finding disease-causing genes is necessary for developing drugs with synergistic effect targeting many genes simultaneously. Here we present an approach to identify novel disease genes and pathways and discuss our approach in the context of IBD. The approach can be generalized to find disease-associated genes for other diseases.
Ryohei Eguchi, Mohammad Bozlul Karim, Pingzhao Hu, Tetsuo Sato, Naoaki Ono, Shigehiko Kanaya, Md. Altaf-Ul-Amin
BMC Bioinform.6
2017 A Wearable Thermometry for Core Body Temperature Measurement and Its Experimental Verification
abstract
A wearable thermometry for core body temperature (CBT) measurement has both healthcare and clinical applications. On the basis of the mechanism of bioheat transfer, we earlier designed and improved a wearable thermometry using the dual-heat-flux method for CBT measurement. In this study, this thermometry is examined experimentally. We studied a fast-changing CBT measurement (FCCM, 55 min, 12 subjects) inside a thermostatic chamber and performed long-term monitoring of CBT (LTM, 24 h, six subjects). When compared with a reference, the CoreTemp CM-210 by Terumo, FCCM shows 0.07 °C average difference and a 95% CI of [-0.27, 0.12] °C. LTM shows no significant difference in parameters for the inference of circadian rhythm. The FCCM and LTM both simulated scenarios in which this thermometry could be used for intensive monitoring and daily healthcare, respectively. The results suggest that because of its convenient design, this thermometry may be an ideal choice for conventional CBT measurements.
Ming Huang 0002, Toshiyo Tamura, Zunyi Tang, Shigehiko Kanaya
IEEE J. Biomed. Health Informatics5
2017 A Chair-Based Unobtrusive Cuffless Blood Pressure Monitoring System Based on Pulse Arrival Time
abstract
In this paper, we present an unobtrusive cuffless blood pressure (BP) monitoring system based on pulse arrival time (PAT) for facilitating long-term home BP monitoring. The proposed system consists of an electrocardiograph (ECG), a photoplethysmograph (PPG), and a control circuit with a Bluetooth module, all of which are mounted on a common armchair to measure ECG and PPG signals from users while sitting on the armchair in order to calculate continuous PAT. Considering the good linear correlation of systolic BP (SBP) and the nonlinear correlation of diastolic BP (DBP) with PAT, a new BP estimation method was proposed. Ten subjects underwent BP monitoring experiments involving stationary sitting on a chair, lying on a bed, and pedaling using an ergometer in order to assess the accuracy of the estimated BP. A cuff-type BP monitor was used as reference in the experiments. Results showed that the mean difference of the estimated SBP and DBP was within 0.2 ± 5.8 mmHg ( p < 0.00001) and 0.4 ± 5.7 mmHg ( p < 0.00001), respectively, and the mean absolute difference of the estimated SBP and DBP were 4.4 and 4.6 mmHg, respectively, compared to references. Additionally, five subjects participated in data collections consisting of sitting on a chair twice a day for one month. Compared to the reference, the difference did not obviously increase along with time, even though individualized calibration was executed only once at the beginning. These results suggest that the proposed system has quite the potential for long-term home BP monitoring.
Zunyi Tang, Toshiyo Tamura, Masaki Sekine, Ming Huang 0002, Masaki Yoshida, Kaoru Sakatani, Hiroshi Kobayashi, Shigehiko Kanaya
IEEE J. Biomed. Health Informatics9
2016 Finding an appropriate equation to measure similarity between binary vectors: case studies on Indonesian and Japanese herbal medicines
abstract
BACKGROUND: The binary similarity and dissimilarity measures have critical roles in the processing of data consisting of binary vectors in various fields including bioinformatics and chemometrics. These metrics express the similarity and dissimilarity values between two binary vectors in terms of the positive matches, absence mismatches or negative matches. To our knowledge, there is no published work presenting a systematic way of finding an appropriate equation to measure binary similarity that performs well for certain data type or application. A proper method to select a suitable binary similarity or dissimilarity measure is needed to obtain better classification results. RESULTS: In this study, we proposed a novel approach to select binary similarity and dissimilarity measures. We collected 79 binary similarity and dissimilarity equations by extensive literature search and implemented those equations as an R package called bmeasures. We applied these metrics to quantify the similarity and dissimilarity between herbal medicine formulas belonging to the Indonesian Jamu and Japanese Kampo separately. We assessed the capability of binary equations to classify herbal medicine pairs into match and mismatch efficacies based on their similarity or dissimilarity coefficients using the Receiver Operating Characteristic (ROC) curve analysis. According to the area under the ROC curve results, we found Indonesian Jamu and Japanese Kampo datasets obtained different ranking of binary similarity and dissimilarity measures. Out of all the equations, the Forbes-2 similarity and the Variant of Correlation similarity measures are recommended for studying the relationship between Jamu formulas and Kampo formulas, respectively. CONCLUSIONS: The selection of binary similarity and dissimilarity measures for multivariate analysis is data dependent. The proposed method can be used to find the most suitable binary similarity and dissimilarity equation wisely for a particular data. Our finding suggests that all four types of matching quantities in the Operational Taxonomic Unit (OTU) table are important to calculate the similarity and dissimilarity coefficients between herbal medicine formulas. Also, the binary similarity and dissimilarity measures that include the negative match quantity d achieve better capability to separate herbal medicine pairs compared to equations that exclude d.
Sony Hartono Wijaya, Farit Mochamad Afendi, Irmanida Batubara, Latifah K. Darusman, Md. Altaf-Ul-Amin, Shigehiko Kanaya
BMC Bioinform.6
2016 Integrated pathway-based transcription regulation network mining and visualization based on gene expression profiles
abstract
Conventionally, workflows examining transcription regulation networks from gene expression data involve distinct analytical steps. There is a need for pipelines that unify data mining and inference deduction into a singular framework to enhance interpretation and hypotheses generation. We propose a workflow that merges network construction with gene expression data mining focusing on regulation processes in the context of transcription factor driven gene regulation. The pipeline implements pathway-based modularization of expression profiles into functional units to improve biological interpretation. The integrated workflow was implemented as a web application software (TransReguloNet) with functions that enable pathway visualization and comparison of transcription factor activity between sample conditions defined in the experimental design. The pipeline merges differential expression, network construction, pathway-based abstraction, clustering and visualization. The framework was applied in analysis of actual expression datasets related to lung, breast and prostrate cancer.
Nelson Kibinge, Naoaki Ono, Masafumi Horie, Tetsuo Sato, Tadao Sugiura, Md. Altaf-Ul-Amin, Akira Saito, Shigehiko Kanaya
J. Biomed. Informatics8
2013 An application of a relational database system for high-throughput prediction of elemental compositions from accurate mass values
abstract
SUMMARY: High-accuracy mass values detected by high-resolution mass spectrometry analysis enable prediction of elemental compositions, and thus are used for metabolite annotations in metabolomic studies. Here, we report an application of a relational database to significantly improve the rate of elemental composition predictions. By searching a database of pre-calculated elemental compositions with fixed kinds and numbers of atoms, the approach eliminates redundant evaluations of the same formula that occur in repeated calculations with other tools. When our approach is compared with HR2, which is one of the fastest tools available, our database search times were at least 109 times shorter than those of HR2. When a solid-state drive (SSD) was applied, the search time was 488 times shorter at 5 ppm mass tolerance and 1833 times at 0.1 ppm. Even if the search by HR2 was performed with 8 threads in a high-spec Windows 7 PC, the database search times were at least 26 and 115 times shorter without and with the SSD. These improvements were enhanced in a low spec Windows XP PC. We constructed a web service 'MFSearcher' to query the database in a RESTful manner. AVAILABILITY AND IMPLEMENTATION: Available for free at http://webs2.kazusa.or.jp/mfsearcher. The web service is implemented in Java, MySQL, Apache and Tomcat, with all major browsers supported. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Nozomu Sakurai, Takeshi Ara, Shigehiko Kanaya, Yukiko Nakamura, Yoko Iijima, Mitsuo Enomoto, Takeshi Motegi, Koh Aoki, Hideyuki Suzuki, Daisuke Shibata
Bioinform.3
2011 AMDORAP: Non-targeted metabolic profiling based on high-resolution LC-MS
abstract
BACKGROUND: Liquid chromatography-mass spectrometry (LC-MS) utilizing the high-resolution power of an orbitrap is an important analytical technique for both metabolomics and proteomics. Most important feature of the orbitrap is excellent mass accuracy. Thus, it is necessary to convert raw data to accurate and reliable m/z values for metabolic fingerprinting by high-resolution LC-MS. RESULTS: In the present study, we developed a novel, easy-to-use and straightforward m/z detection method, AMDORAP. For assessing the performance, we used real biological samples, Bacillus subtilis strains 168 and MGB874, in the positive mode by LC-orbitrap. For 14 identified compounds by measuring the authentic compounds, we compared obtained m/z values with other LC-MS processing tools. The errors by AMDORAP were distributed within ±3 ppm and showed the best performance in m/z value accuracy. CONCLUSIONS: Our method can detect m/z values of biological samples much more accurately than other LC-MS analysis tools. AMDORAP allows us to address the relationships between biological effects and cellular metabolites based on accurate m/z values. Obtaining the accurate m/z values from raw data should be indispensable as a starting point for comparative LC-orbitrap analysis. AMDORAP is freely available under an open-source license at http://amdorap.sourceforge.net/.
Hiroki Takahashi, Takuya Morimoto, Naotake Ogasawara, Shigehiko Kanaya
BMC Bioinform.4
2007 Predicting state transitions in the transcriptome and metabolome using a linear dynamical system model
abstract
BACKGROUND: Modelling of time series data should not be an approximation of input data profiles, but rather be able to detect and evaluate dynamical changes in the time series data. Objective criteria that can be used to evaluate dynamical changes in data are therefore important to filter experimental noise and to enable extraction of unexpected, biologically important information. RESULTS: Here we demonstrate the effectiveness of a Markov model, named the Linear Dynamical System, to simulate the dynamics of a transcript or metabolite time series, and propose a probabilistic index that enables detection of time-sensitive changes. This method was applied to time series datasets from Bacillus subtilis and Arabidopsis thaliana grown under stress conditions; in the former, only gene expression was studied, whereas in the latter, both gene expression and metabolite accumulation. Our method not only identified well-known changes in gene expression and metabolite accumulation, but also detected novel changes that are likely to be responsible for each stress response condition. CONCLUSION: This general approach can be applied to any time-series data profile from which one wishes to identify elements responsible for state transitions, such as rapid environmental adaptation by an organism.
Ryouko Morioka, Shigehiko Kanaya, Masami Y. Hirai, Mitsuru Yano, Naotake Ogasawara, Kazuki Saito
BMC Bioinform.2
2006 Development and implementation of an algorithm for detection of protein complexes in large interaction networks
abstract
BACKGROUND: After complete sequencing of a number of genomes the focus has now turned to proteomics. Advanced proteomics technologies such as two-hybrid assay, mass spectrometry etc. are producing huge data sets of protein-protein interactions which can be portrayed as networks, and one of the burning issues is to find protein complexes in such networks. The enormous size of protein-protein interaction (PPI) networks warrants development of efficient computational methods for extraction of significant complexes. RESULTS: This paper presents an algorithm for detection of protein complexes in large interaction networks. In a PPI network, a node represents a protein and an edge represents an interaction. The input to the algorithm is the associated matrix of an interaction network and the outputs are protein complexes. The complexes are determined by way of finding clusters, i. e. the densely connected regions in the network. We also show and analyze some protein complexes generated by the proposed algorithm from typical PPI networks of Escherichia coli and Saccharomyces cerevisiae. A comparison between a PPI and a random network is also performed in the context of the proposed algorithm. CONCLUSION: The proposed algorithm makes it possible to detect clusters of proteins in PPI networks which mostly represent molecular biological functional units. Therefore, protein complexes determined solely based on interaction data can help us to predict the functions of proteins, and they are also useful to understand and explain certain biological processes.
Md. Altaf-Ul-Amin, Yoko Shinbo, Kenji Mihara, Ken Kurokawa, Shigehiko Kanaya
BMC Bioinform.5
2005 Accurate extraction of functional associations between proteins based on common interaction partners and common domains
abstract
MOTIVATION: Genomic and proteomic approaches have accumulated a huge amount of data which provide clues to protein function. However, interpreting single omic data for predicting uncharacterized protein functions has been a challenging task, because the data contain a lot of false positives. To overcome this problem, methods for integrating data from various omic approaches are needed for more accurate function prediction. RESULT: In this paper, we have developed a method which extracts functionally similar proteins with high confidence by integrating protein-protein interaction data and domain information. We used this method to analyze publicly available data from Saccharomyces cerevisiae. We identified 1042 functional associations, involving 765 proteins of which 98 (12.8%) had no previously ascribed function. Our method extracts functionally similar protein pairs more accurately than conventional methods, and predicting function for previously uncharacterized proteins can be achieved. Our method can of course be applied to protein-protein interaction data for any species.
Kinya Okada, Shigehiko Kanaya, Kiyoshi Asai
Bioinform.2
1996 Detection of genes in Escherichia coli sequences determined by genome projects and prediction of protein production levels, based on multivariate diversity in codon usage
abstract
We used principal component analysis to develop measures (called Z-parameters in this study) which reflect the diversity of codon usage in Escherichia coli genes. Protein production levels for 1500 CDSs (protein-coding sequences) identified by E.coli genome projects in Japan and the US were estimated from a correlation equation between Z1 and cellular protein content obtained through analysis of the genes experimentally characterized. Through the profile analysis of Z1 for E.coli sequences obtained by the Japanese Project, we predicted an additional 36 CDSs that had not been annotated in the International DNA Database. Thirty-one out of the 36 CDSs could be assigned to presumptive protein genes through a BLASTX search for recent protein databases in the Genome Net in Japan. Detailed examination of the Z1-parameter profile led us to assess sequencing errors which cause frame-shift.
Shigehiko Kanaya, Yoshihiro Kudo, Yasukazu Nakamura, Toshimichi Ikemura
Comput. Appl. Biosci.1