Tatsuhiko Tsunoda

dblp:13/5504 · DBLP profile ↗
← Back
24ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-5439-7918ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 VBayesMM: variational Bayesian neural network to prioritize important relationships of high-dimensional microbiome multiomics data
abstract
The analysis of high-dimensional microbiome multiomics datasets is crucial for understanding the complex interactions between microbial communities and host physiological states across health and disease conditions. Despite their importance, current methods, such as the microbe-metabolite vectors approach, often face challenges in predicting metabolite abundances from microbial data and identifying keystone species. This arises from the vast dimensionality of metagenomics data, which complicates the inference of significant relationships, particularly the estimation of co-occurrence probabilities between microbes and metabolites. Here we propose the variational Bayesian microbiome multiomics (VBayesMM) approach, which aims to improve the prediction of metabolite abundances from microbial metagenomics data by incorporating a spike-and-slab prior within a Bayesian neural network. This allows VBayesMM to rapidly and precisely identify crucial microbial species, leading to more accurate estimations of co-occurrence probabilities between microbes and metabolites, while also robustly managing the uncertainty inherent in high-dimensional data. Moreover, we have implemented variational inference to address computational bottlenecks, enabling scalable analysis across extensive multiomics datasets. Our large-scale comparative evaluations demonstrate that VBayesMM not only outperforms existing methods in predicting metabolite abundances but also provides a scalable solution for analyzing massive datasets. VBayesMM enhances the interpretability of the Bayesian neural network by identifying a core set of influential microbial species, thus facilitating a deeper understanding of their probabilistic relationships with the host.
Tung Dang, Artem Lysenko, Keith A. Boroevich, Tatsuhiko Tsunoda
Briefings Bioinform.4
2025 scHDeepInsight: a hierarchical deep learning framework for precise immune cell annotation in single-cell RNA-seq data
abstract
Accurate classification of immune cells is crucial for elucidating their diverse roles in health and disease. However, this task remains very challenging in single-cell RNA sequencing (scRNA-seq) data due to the complex and hierarchical relationships of immune cell types. To address this, we introduce scHDeepInsight, a deep learning framework that extends our previous scDeepInsight model by integrating a biologically-informed classification architecture with an adaptive hierarchical focal loss (AHFL). The framework builds on our established method of converting gene expression data into two-dimensional structured images, enabling convolutional neural networks to effectively capture both global and fine-grained transcriptomic features. This design utilizes hierarchical relationships among immune cell types to enhance the classification ability beyond the flat classification approaches. scHDeepInsight dynamically adjusts loss contributions to balance performance across the hierarchy levels. Comprehensive benchmarking across seven diverse tissue datasets shows scHDeepInsight achieves an average accuracy of 93.2%, surpassing contemporary methods by 5.1 percentage points. The model successfully distinguishes 50 distinct immune cell subtypes with high accuracy, demonstrating proficiency for identifying rare and closely related cell subtypes. Additionally, SHAP-based interpretability quantifies individual gene contributions to reveal the biological basis of classification decisions. These qualities make scHDeepInsight a robust tool for high-resolution cell subtype characterization, well-suited for detailed profiling in immunological studies and extensible to nonimmune cell types.
Shangru Jia, Artem Lysenko, Keith A. Boroevich, Alok Sharma, Tatsuhiko Tsunoda
Briefings Bioinform.5
2024 multi-GAT: Integrative Analysis of scRNA-seq and scATAC-seq Data Using Graph Attention Networks for Cell Annotation
Shangru Jia, Tatsuhiko Tsunoda, Alok Sharma
PRICAI (1)2
2023 scDeepInsight: a supervised cell-type identification method for scRNA-seq data with deep learning
abstract
Annotation of cell-types is a critical step in the analysis of single-cell RNA sequencing (scRNA-seq) data that allows the study of heterogeneity across multiple cell populations. Currently, this is most commonly done using unsupervised clustering algorithms, which project single-cell expression data into a lower dimensional space and then cluster cells based on their distances from each other. However, as these methods do not use reference datasets, they can only achieve a rough classification of cell-types, and it is difficult to improve the recognition accuracy further. To effectively solve this issue, we propose a novel supervised annotation method, scDeepInsight. The scDeepInsight method is capable of performing manifold assignments. It is competent in executing data integration through batch normalization, performing supervised training on the reference dataset, doing outlier detection and annotating cell-types on query datasets. Moreover, it can help identify active genes or marker genes related to cell-types. The training of the scDeepInsight model is performed in a unique way. Tabular scRNA-seq data are first converted to corresponding images through the DeepInsight methodology. DeepInsight can create a trainable image transformer to convert non-image RNA data to images by comprehensively comparing interrelationships among multiple genes. Subsequently, the converted images are fed into convolutional neural networks such as EfficientNet-b3. This enables automatic feature extraction to identify the cell-types of scRNA-seq samples. We benchmarked scDeepInsight with six other mainstream cell annotation methods. The average accuracy rate of scDeepInsight reached 87.5%, which is more than 7% higher compared with the state-of-the-art methods.
Shangru Jia, Artem Lysenko, Keith A. Boroevich, Alok Sharma, Tatsuhiko Tsunoda
Briefings Bioinform.5
2021 DeepFeature: feature selection in nonimage data using convolutional neural network
abstract
Artificial intelligence methods offer exciting new capabilities for the discovery of biological mechanisms from raw data because they are able to detect vastly more complex patterns of association that cannot be captured by classical statistical tests. Among these methods, deep neural networks are currently among the most advanced approaches and, in particular, convolutional neural networks (CNNs) have been shown to perform excellently for a variety of difficult tasks. Despite that applications of this type of networks to high-dimensional omics data and, most importantly, meaningful interpretation of the results returned from such models in a biomedical context remains an open problem. Here we present, an approach applying a CNN to nonimage data for feature selection. Our pipeline, DeepFeature, can both successfully transform omics data into a form that is optimal for fitting a CNN model and can also return sets of the most important genes used internally for computing predictions. Within the framework, the Snowfall compression algorithm is introduced to enable more elements in the fixed pixel framework, and region accumulation and element decoder is developed to find elements or genes from the class activation maps. In comparative tests for cancer type prediction task, DeepFeature simultaneously achieved superior predictive performance and better ability to discover key pathways and biological processes meaningful for this context. Capabilities offered by the proposed framework can enable the effective use of powerful deep learning methods to facilitate the discovery of causal mechanisms in high-dimensional biomedical data.
Alok Sharma, Artem Lysenko, Keith A. Boroevich, Edwin Vans, Tatsuhiko Tsunoda
Briefings Bioinform.5
2021 Forecasting the spread of COVID-19 using LSTM network
abstract
BACKGROUND: The novel coronavirus (COVID-19) is caused by severe acute respiratory syndrome coronavirus 2, and within a few months, it has become a global pandemic. This forced many affected countries to take stringent measures such as complete lockdown, shutting down businesses and trade, as well as travel restrictions, which has had a tremendous economic impact. Therefore, having knowledge and foresight about how a country might be able to contain the spread of COVID-19 will be of paramount importance to the government, policy makers, business partners and entrepreneurs. To help social and administrative decision making, a model that will be able to forecast when a country might be able to contain the spread of COVID-19 is needed. RESULTS: The results obtained using our long short-term memory (LSTM) network-based model are promising as we validate our prediction model using New Zealand's data since they have been able to contain the spread of COVID-19 and bring the daily new cases tally to zero. Our proposed forecasting model was able to correctly predict the dates within which New Zealand was able to contain the spread of COVID-19. Similarly, the proposed model has been used to forecast the dates when other countries would be able to contain the spread of COVID-19. CONCLUSION: The forecasted dates are only a prediction based on the existing situation. However, these forecasted dates can be used to guide actions and make informed decisions that will be practically beneficial in influencing the real future. The current forecasting trend shows that more stringent actions/restrictions need to be implemented for most of the countries as the forecasting model shows they will take over three months before they can possibly contain the spread of COVID-19.
Shiu Kumar, Ronesh Sharma, Tatsuhiko Tsunoda, Thirumananseri Kumarevel, Alok Sharma
BMC Bioinform.3
2021 SPECTRA: a tool for enhanced brain wave signal recognition
abstract
BACKGROUND: Brain wave signal recognition has gained increased attention in neuro-rehabilitation applications. This has driven the development of brain-computer interface (BCI) systems. Brain wave signals are acquired using electroencephalography (EEG) sensors, processed and decoded to identify the category to which the signal belongs. Once the signal category is determined, it can be used to control external devices. However, the success of such a system essentially relies on significant feature extraction and classification algorithms. One of the commonly used feature extraction technique for BCI systems is common spatial pattern (CSP). RESULTS: The performance of the proposed spatial-frequency-temporal feature extraction (SPECTRA) predictor is analysed using three public benchmark datasets. Our proposed predictor outperformed other competing methods achieving lowest average error rates of 8.55%, 17.90% and 20.26%, and highest average kappa coefficient values of 0.829, 0.643 and 0.595 for BCI Competition III dataset IVa, BCI Competition IV dataset I and BCI Competition IV dataset IIb, respectively. CONCLUSIONS: Our proposed SPECTRA predictor effectively finds features that are more separable and shows improvement in brain wave signal recognition that can be instrumental in developing improved real-time BCI systems that are computationally efficient.
Shiu Kumar, Tatsuhiko Tsunoda, Alok Sharma
BMC Bioinform.2
2019 Subject-Specific-Frequency-Band for Motor Imagery EEG Signal Recognition Based on Common Spatial Spectral Pattern
Shiu Kumar, Alok Sharma, Tatsuhiko Tsunoda
PRICAI (2)3
2019 Computational Prediction of Lysine Pupylation Sites in Prokaryotic Proteins Using Position Specific Scoring Matrix into Bigram for Feature Extraction
Alok Sharma, Abel Avitesh Chandra, Abdollah Dehzangi, Daichi Shigemizu, Tatsuhiko Tsunoda
PRICAI (3)6
2019 Clustering of Small-Sample Single-Cell RNA-Seq Data via Feature Clustering and Selection
Edwin Vans, Alok Sharma, Ashwini Patil, Daichi Shigemizu, Tatsuhiko Tsunoda
PRICAI (3)5
2019 Navigating the disease landscape: knowledge representations for contextualizing molecular signatures
abstract
Large amounts of data emerging from experiments in molecular medicine are leading to the identification of molecular signatures associated with disease subtypes. The contextualization of these patterns is important for obtaining mechanistic insight into the aberrant processes associated with a disease, and this typically involves the integration of multiple heterogeneous types of data. In this review, we discuss knowledge representations that can be useful to explore the biological context of molecular signatures, in particular three main approaches, namely, pathway mapping approaches, molecular network centric approaches and approaches that represent biological statements as knowledge graphs. We discuss the utility of each of these paradigms, illustrate how they can be leveraged with selected practical examples and identify ongoing challenges for this field of research.
Mansoor A. S. Saqi, Artem Lysenko, Yike Guo, Tatsuhiko Tsunoda, Charles Auffray
Briefings Bioinform.4
2019 GlyStruct: glycation prediction using structural properties of amino acid residues
abstract
BACKGROUND: Glycation is a one of the post-translational modifications (PTM) where sugar molecules and residues in protein sequences are covalently bonded. It has become one of the clinically important PTM in recent times attributed to many chronic and age related complications. Being a non-enzymatic reaction, it is a great challenge when it comes to its prediction due to the lack of significant bias in the sequence motifs. RESULTS: We developed a classifier, GlyStruct based on support vector machine, to predict glycated and non-glycated lysine residues using structural properties of amino acid residues. The features used were secondary structure, accessible surface area and the local backbone torsion angles. For this work, a benchmark dataset was extracted containing 235 glycated and 303 non-glycated lysine residues. GlyStruct demonstrated improved performance of approximately 10% in comparison to benchmark method of Gly-PseAAC. The performance for GlyStruct on the metrics, sensitivity, specificity, accuracy and Mathew's correlation coefficient were 0.7013, 0.7989, 0.7562, and 0.5065, respectively for 10-fold cross-validation. CONCLUSION: Glycation has emerged to be one of the clinically important PTM of proteins in recent times. Therefore, the development of computational tools become necessary to predict glycation, which could help medical professionals administer drugs and manage patients more effectively. The proposed predictor manages to classify glycated and non-glycated lysine residues with promising results consistently on various cross-validation schemes and outperforms other state of the art methods.
Hamendra Manhar Reddy, Alok Sharma, Abdollah Dehzangi, Daichi Shigemizu, Abel Avitesh Chandra, Tatsuhiko Tsunoda
BMC Bioinform.6
2019 Discovering MoRFs by trisecting intrinsically disordered protein sequence into terminals and middle regions
abstract
BACKGROUND: Molecular Recognition Features (MoRFs) are short protein regions present in intrinsically disordered protein (IDPs) sequences. MoRFs interact with structured partner protein and upon interaction, they undergo a disorder-to-order transition to perform various biological functions. Analyses of MoRFs are important towards understanding their function. RESULTS: Performance is reported using the MoRF dataset that has been previously used to compare the other existing MoRF predictors. The performance obtained in this study is equivalent to the benchmarked OPAL predictor, i.e., OPAL achieved AUC of 0.815, whereas the model in this study achieved AUC of 0.819 using TEST set. CONCLUSION: Achieving comparable performance, the proposed method can be used as an alternative approach for MoRF prediction.
Ronesh Sharma, Alok Sharma, Ashwini Patil, Tatsuhiko Tsunoda
BMC Bioinform.4
2018 OPAL: prediction of MoRF regions in intrinsically disordered protein sequences
abstract
Motivation: Intrinsically disordered proteins lack stable 3-dimensional structure and play a crucial role in performing various biological functions. Key to their biological function are the molecular recognition features (MoRFs) located within long disordered regions. Computationally identifying these MoRFs from disordered protein sequences is a challenging task. In this study, we present a new MoRF predictor, OPAL, to identify MoRFs in disordered protein sequences. OPAL utilizes two independent sources of information computed using different component predictors. The scores are processed and combined using common averaging method. The first score is computed using a component MoRF predictor which utilizes composition and sequence similarity of MoRF and non-MoRF regions to detect MoRFs. The second score is calculated using half-sphere exposure (HSE), solvent accessible surface area (ASA) and backbone angle information of the disordered protein sequence, using information from the amino acid properties of flanks surrounding the MoRFs to distinguish MoRF and non-MoRF residues. Results: OPAL is evaluated using test sets that were previously used to evaluate MoRF predictors, MoRFpred, MoRFchibi and MoRFchibi-web. The results demonstrate that OPAL outperforms all the available MoRF predictors and is the most accurate predictor available for MoRF prediction. It is available at http://www.alok-ai-lab.com/tools/opal/. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Ronesh Sharma, Gaurav Raicar, Tatsuhiko Tsunoda, Ashwini Patil, Alok Sharma
Bioinform.3
2017 An improved discriminative filter bank selection approach for motor imagery EEG signal classification using mutual information
abstract
BACKGROUND: Common spatial pattern (CSP) has been an effective technique for feature extraction in electroencephalography (EEG) based brain computer interfaces (BCIs). However, motor imagery EEG signal feature extraction using CSP generally depends on the selection of the frequency bands to a great extent. METHODS: In this study, we propose a mutual information based frequency band selection approach. The idea of the proposed method is to utilize the information from all the available channels for effectively selecting the most discriminative filter banks. CSP features are extracted from multiple overlapping sub-bands. An additional sub-band has been introduced that cover the wide frequency band (7-30 Hz) and two different types of features are extracted using CSP and common spatio-spectral pattern techniques, respectively. Mutual information is then computed from the extracted features of each of these bands and the top filter banks are selected for further processing. Linear discriminant analysis is applied to the features extracted from each of the filter banks. The scores are fused together, and classification is done using support vector machine. RESULTS: The proposed method is evaluated using BCI Competition III dataset IVa, BCI Competition IV dataset I and BCI Competition IV dataset IIb, and it outperformed all other competing methods achieving the lowest misclassification rate and the highest kappa coefficient on all three datasets. CONCLUSIONS: Introducing a wide sub-band and using mutual information for selecting the most discriminative sub-bands, the proposed method shows improvement in motor imagery EEG signal classification.
Shiu Kumar, Alok Sharma, Tatsuhiko Tsunoda
BMC Bioinform.3
2017 2D-EM clustering approach for high-dimensional data through folding feature vectors
abstract
BACKGROUND: Clustering methods are becoming widely utilized in biomedical research where the volume and complexity of data is rapidly increasing. Unsupervised clustering of patient information can reveal distinct phenotype groups with different underlying mechanism, risk prognosis and treatment response. However, biological datasets are usually characterized by a combination of low sample number and very high dimensionality, something that is not adequately addressed by current algorithms. While the performance of the methods is satisfactory for low dimensional data, increasing number of features results in either deterioration of accuracy or inability to cluster. To tackle these challenges, new methodologies designed specifically for such data are needed. RESULTS: We present 2D-EM, a clustering algorithm approach designed for small sample size and high-dimensional datasets. To employ information corresponding to data distribution and facilitate visualization, the sample is folded into its two-dimension (2D) matrix form (or feature matrix). The maximum likelihood estimate is then estimated using a modified expectation-maximization (EM) algorithm. The 2D-EM methodology was benchmarked against several existing clustering methods using 6 medically-relevant transcriptome datasets. The percentage improvement of Rand score and adjusted Rand index compared to the best performing alternative method is up to 21.9% and 155.6%, respectively. To present the general utility of the 2D-EM method we also employed 2 methylome datasets, again showing superior performance relative to established methods. CONCLUSIONS: The 2D-EM algorithm was able to reproduce the groups in transcriptome and methylome data with high accuracy. This build confidence in the methods ability to uncover novel disease subtypes in new datasets. The design of 2D-EM algorithm enables it to handle a diverse set of challenging biomedical dataset and cluster with higher accuracy than established methods. MATLAB implementation of the tool can be freely accessed online ( http://www.riken.jp/en/research/labs/ims/med_sci_math or http://www.alok-ai-lab.com /).
Alok Sharma, Piotr J. Kamola, Tatsuhiko Tsunoda
BMC Bioinform.3
2017 Divisive hierarchical maximum likelihood clustering
abstract
BACKGROUND: Biological data comprises various topologies or a mixture of forms, which makes its analysis extremely complicated. With this data increasing in a daily basis, the design and development of efficient and accurate statistical methods has become absolutely necessary. Specific analyses, such as those related to genome-wide association studies and multi-omics information, are often aimed at clustering sub-conditions of cancers and other diseases. Hierarchical clustering methods, which can be categorized into agglomerative and divisive, have been widely used in such situations. However, unlike agglomerative methods divisive clustering approaches have consistently proved to be computationally expensive. RESULTS: The proposed clustering algorithm (DRAGON) was verified on mutation and microarray data, and was gauged against standard clustering methods in the literature. Its validation included synthetic and significant biological data. When validated on mixed-lineage leukemia data, DRAGON achieved the highest clustering accuracy with data of four different dimensions. Consequently, DRAGON outperformed previous methods with 3-,4- and 5-dimensional acute leukemia data. When tested on mutation data, DRAGON achieved the best performance with 2-dimensional information. CONCLUSIONS: This work proposes a computationally efficient divisive hierarchical clustering method, which can compete equally with agglomerative approaches. The proposed method turned out to correctly cluster data with distinct topologies. A MATLAB implementation can be extraced from http://www.riken.jp/en/research/labs/ims/med_sci_math/ or http://www.alok-ai-lab.com.
Alok Sharma, Yosvany López, Tatsuhiko Tsunoda
BMC Bioinform.3
2016 Decimation filter with Common Spatial Pattern and Fishers Discriminant Analysis for motor imagery classification
abstract
Brain Computer Interface (BCI) system converts thoughts into commands for driving external device with Electroencephalography (EEG). This paper presents the use of decimation filters for filtering the EEG signal. Common Spatial Pattern (CSP) technique is used to transform the filtered signal to a new time series in order to have optimal variance for the discrimination of different tasks. Fishers Discriminant Analysis (FDA) is applied to the CSP features and the FDA scores are fed to a Support Vector Machine (SVM) classifier. The method is evaluated on BCI Competition III Dataset IVa and compared with other related state-of-the-art approaches. The results show that our method outperforms all other approaches in terms of average classification error rate. Compared to best performing method that uses only CSP features, the results obtained in this research offer on average a reduction of 1.07% in the classification error rate.
Shiu Kumar, Ronesh Sharma, Alok Sharma, Tatsuhiko Tsunoda
IJCNN4
2016 Predicting MoRFs in protein sequences using HMM profiles
abstract
BACKGROUND: Intrinsically Disordered Proteins (IDPs) lack an ordered three-dimensional structure and are enriched in various biological processes. The Molecular Recognition Features (MoRFs) are functional regions within IDPs that undergo a disorder-to-order transition on binding to a partner protein. Identifying MoRFs in IDPs using computational methods is a challenging task. METHODS: In this study, we introduce hidden Markov model (HMM) profiles to accurately identify the location of MoRFs in disordered protein sequences. Using windowing technique, HMM profiles are utilised to extract features from protein sequences and support vector machines (SVM) are used to calculate a propensity score for each residue. Two different SVM kernels with high noise tolerance are evaluated with a varying window size and the scores of the SVM models are combined to generate the final propensity score to predict MoRF residues. The SVM models are designed to extract maximal information between MoRF residues, its neighboring regions (Flanks) and the remainder of the sequence (Others). RESULTS: To evaluate the proposed method, its performance was compared to that of other MoRF predictors; MoRFpred and ANCHOR. The results show that the proposed method outperforms these two predictors. CONCLUSIONS: Using HMM profile as a source of feature extraction, the proposed method indicates improvement in predicting MoRFs in disordered protein sequences.
Ronesh Sharma, Shiu Kumar, Tatsuhiko Tsunoda, Ashwini Patil, Alok Sharma
BMC Bioinform.3
2016 Stepwise iterative maximum likelihood clustering approach
abstract
BACKGROUND: Biological/genetic data is a complex mix of various forms or topologies which makes it quite difficult to analyze. An abundance of such data in this modern era requires the development of sophisticated statistical methods to analyze it in a reasonable amount of time. In many biological/genetic analyses, such as genome-wide association study (GWAS) analysis or multi-omics data analysis, it is required to cluster the plethora of data into sub-categories to understand the subtypes of populations, cancers or any other diseases. Traditionally, the k-means clustering algorithm is a dominant clustering method. This is due to its simplicity and reasonable level of accuracy. Many other clustering methods, including support vector clustering, have been developed in the past, but do not perform well with the biological data, either due to computational reasons or failure to identify clusters. RESULTS: The proposed SIML clustering algorithm has been tested on microarray datasets and SNP datasets. It has been compared with a number of clustering algorithms. On MLL datasets, SIML achieved highest clustering accuracy and rand score on 4/9 cases; similarly on SRBCT dataset, it got for 3/5 cases; on ALL subtype it got highest clustering accuracy for 5/7 cases and highest rand score for 4/7 cases. In addition, SIML overall clustering accuracy on a 3 cluster problem using SNP data were 97.3, 94.7 and 100 %, respectively, for each of the clusters. CONCLUSIONS: In this paper, considering the nature of biological data, we proposed a maximum likelihood clustering approach using a stepwise iterative procedure. The advantage of this proposed method is that it not only uses the distance information, but also incorporate variance information for clustering. This method is able to cluster when data appeared in overlapping and complex forms. The experimental results illustrate its performance and usefulness over other clustering methods. A Matlab package of this method (SIML) is provided at the web-link http://www.riken.jp/en/research/labs/ims/med_sci_math/ .
Alok Sharma, Daichi Shigemizu, Keith A. Boroevich, Yosvany López, Yoichiro Kamatani, Michiaki Kubo, Tatsuhiko Tsunoda
BMC Bioinform.7
2008 MOCSphaser: a haplotype inference tool from a mixture of copy number variation and single nucleotide polymorphism data
abstract
Abstract Summary: Detailed analyses of the population-genetic nature of copy number variations (CNVs) and the linkage disequilibrium between CNV and single nucleotide polymorphism (SNP) loci from high-throughput experimental data require a computational tool to accurately infer alleles of CNVs and haplotypes composed of both CNV alleles and SNP alleles. Here we developed a new tool to infer population frequencies of such alleles and haplotypes from observed copy numbers and SNP genotypes, using the expectation–maximization algorithm. This tool can also handle copy numbers ambiguously determined, such as 2 or 3 copies, due to experimental noise. Availability: http://emu.src.riken.jp/MOCSphaser/MOCSphaser.zip Contact: [email protected] Supplementary information: Additional materials can be found at http://emu.src.riken.jp/MOCSphaser/SuppInfor.doc
Mamoru Kato, Yusuke Nakamura, Tatsuhiko Tsunoda
Bioinform.3
2007 MotifCombinator: a web-based tool to search for combinations of cis-regulatory motifs
abstract
BACKGROUND: A combination of multiple types of transcription factors and cis-regulatory elements is often required for gene expression in eukaryotes, and the combinatorial regulation confers specific gene expression to tissues or environments. To reveal the combinatorial regulation, computational methods are developed that efficiently infer combinations of cis-regulatory motifs that are important for gene expression as measured by DNA microarrays. One promising type of computational method is to utilize regression analysis between expression levels and scores of motifs in input sequences. This type takes full advantage of information on expression levels because it does not require that the expression level of each gene be dichotomized according to whether or not it reaches a certain threshold level. However, there is no web-based tool that employs regression methods to systematically search for motif combinations and that practically handles combinations of more than two or three motifs. RESULTS: We here introduced MotifCombinator, an online tool with a user-friendly interface, to systematically search for combinations composed of any number of motifs based on regression methods. The tool utilizes well-known regression methods (the multivariate linear regression, the multivariate adaptive regression spline or MARS, and the multivariate logistic regression method) for this purpose, and uses the genetic algorithm to search for combinations composed of any desired number of motifs. The visualization systems in this tool help users to intuitively grasp the process of the combination search, and the backup system allows users to easily stop and restart calculations that are expected to require large computational time. This tool also provides preparatory steps needed for systematic combination search--i.e., selecting single motifs to constitute combinations and cutting out redundant similar motifs based on clustering analysis. CONCLUSION: MotifCombinator helps users to systematically search for motif combinations that play an important role in gene expression as measured by microarrays.
Mamoru Kato, Tatsuhiko Tsunoda
BMC Bioinform.2
1999 Estimating transcription factor bindability on DNA
abstract
MOTIVATION: Precise analysis of the genetic network, gene function and transcription regulation requires accurate prediction of transcription factor (TF) bindability on DNA. For calculating the matching score between an input sequence and a set of known TF binding sites, we use positional weight matrices (PWMs) and Bucher's calculating method (Bucher, J. Mol. Biol., 212, 563-578, 1990). Since estimating TF binding sites requires cut-off values, we propose a robust cut-off value determining algorithm. RESULTS: We generalize the concept of local overrepresentation with statistics, and propose a new algorithm for determining the cut-off value using the background rate estimated on non-promoters. The algorithm iteratively determines parameters separating instances into phenomena-dependent and phenomena-independent subsets. Our system includes the method of re-estimating cut-off values of TFs that mis-recognize other TF preferred regions. Our data source comprised 433 non-redundant vertebrate promoters including viral promoters, from Eukaryotic Promoter Database (EPD) R.50. The method is applied to 205 vertebrate TFs that have frequency matrices in TRANSFAC Ver.3. 4 and the cut-off values of all of them can be determined. AVAILABILITY: The cut-off values and TF binding site predicting tool are available at http://www.hgc.ims.u-tokyo.ac. jp/service/tooldoc/TFBIND. We also provide the cut-off value estimating programs.
Tatsuhiko Tsunoda, Toshihisa Takagi
Bioinform.1
1994 Analysis of Scene Identification Ability of Associative Memory with Pictorial Dictionary
Tatsuhiko Tsunoda, Hidehiko Tanaka
COLING1