EDBT 2026 Demo / reviewers in the wild / expert
Daisuke Kihara
dblp:13/915
· DBLP profile ↗
52ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0003-4091-6614ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 36 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 7 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Unsupervised to Zero-Shot: 3D Domain Adaptation for Cryo-Electron Tomography SegmentationabstractAbstract With the refinement of 3D imaging modality, cryo-electron tomography (cryo-ET) has emerged as a powerful technique for the structural analysis of macromolecular complexes at near-atomic resolution. Recent advancements in volumetric segmentation methods applied to cryo-ET datasets have garnered significant attention within the biomedical sector. However, existing methods rely heavily on manually labeled data, which demands highly specialized expertise, making fully supervised approaches less feasible for cryo-ET images. To address this, a number of unsupervised domain adaptation (UDA) techniques have been developed to improve segmentation network performance using unlabeled data. Nevertheless, directly applying these methods to cryo-ET image segmentation presents two major challenges: 1) the source dataset, usually obtained through simulation, contains a fixed level of noise, while the target dataset, being directly collected from raw-data from the real-world scenario, has unpredictable noise levels; 2) the source data used for training typically consists of known macromolecules, in contrast, the target domain data are often unknown, causing the model to be biased towards those known macromolecules, leading to a domain shift problem. To address such challenges, in this paper, we introduce a voxel-wise unsupervised domain adaptation approach, termed Vox-UDA, specifically for cryo-ET subtomogram segmentation. Vox-UDA incorporates a noise generation module to simulate target-like noises in the source dataset for cross-noise level adaptation, and a denoised pseudo-labeling strategy based on the improved bilateral filter to alleviate the domain shift problem. Additionally, we further consider a scenario that is more in line with the real world, where the target (experimental) dataset might not be accessible during training or has a very small sample size, and we present a voxel-wise zero-shot domain adaptation (ZSDA) approach, named Vox-ZSDA. In Vox-ZSDA, we introduce a self-supervised graph learning strategy to eliminate any dependency on the target data, accompanied by a dynamic graph contrastive learning technique to enhance the model’s sensitivity to macromolecular structures to boost the segmentation performance. More importantly, we construct the first UDA and ZSDA cryo-ET subtomogram segmentation benchmark on three experimental datasets. Extensive experimental results on multiple benchmarks and newly curated real-world datasets demonstrate the superiority of our proposed approach compared to state-of-the-art UDA and ZSDA methods. Haoran Li 0024, Xingjian Li 0002, Jiahua Shi, Huaming Chen, Bo Du 0004, Johan Barthélemy, Daisuke Kihara, Jun Shen 0001, Min Xu 0009 |
Int. J. Comput. Vis. | 8 |
| 2025 | Vox-UDA: Voxel-wise Unsupervised Domain Adaptation for Cryo-Electron Subtomogram Segmentation with Denoised Pseudo-LabelingabstractCryo-Electron Tomography (cryo-ET) is a 3D imaging technology that facilitates the study of macromolecular structures at near-atomic resolution. Recent volumetric segmentation approaches on cryo-ET images have drawn widespread interest in the biological sector. However, existing methods heavily rely on manually labeled data, which requires highly professional skills, thereby hindering the adoption of fully-supervised approaches for cryo-ET images. Some unsupervised domain adaptation (UDA) approaches have been designed to enhance the segmentation network performance using unlabeled data. However, applying these methods directly to cryo-ET image segmentation tasks remains challenging due to two main issues: 1) the source dataset, usually obtained through simulation, contains a fixed level of noise, while the target dataset, directly collected from raw-data from the real-world scenario, have unpredictable noise levels. 2) the source data used for training typically consists of known macromoleculars. In contrast, the target domain data are often unknown, causing the model to be biased towards those known macromolecules, leading to a domain shift problem. To address such challenges, in this work, we introduce a voxel-wise unsupervised domain adaptation approach, termed Vox-UDA, specifically for cryo-ET subtomogram segmentation. Vox-UDA incorporates a noise generation module to simulate target-like noises in the source dataset for cross-noise level adaptation. Additionally, we propose a denoised pseudo-labeling strategy based on the improved Bilateral Filter to alleviate the domain shift problem. More importantly, we construct the first UDA cryo-ET subtomogram segmentation benchmark on three experimental datasets. Extensive experimental results on multiple benchmarks and newly curated real-world datasets demonstrate the superiority of our proposed approach compared to state-of-the-art UDA methods. Haoran Li 0024, Xingjian Li 0002, Jiahua Shi, Huaming Chen, Bo Du 0004, Daisuke Kihara, Johan Barthélemy, Jun Shen 0001, Min Xu 0009 |
AAAI | 6 |
| 2025 | Twenty years of advances in prediction of nucleic acid-binding residues in protein sequencesabstractComputational prediction of nucleic acid-binding residues in protein sequences is an active field of research, with over 80 methods that were released in the past 2 decades. We identify and discuss 87 sequence-based predictors that include dozens of recently published methods that are surveyed for the first time. We overview historical progress and examine multiple practical issues that include availability and impact of predictors, key features of their predictive models, and important aspects related to their training and assessment. We observe that the past decade has brought increased use of deep neural networks and protein language models, which contributed to substantial gains in the predictive performance. We also highlight advancements in vital and challenging issues that include cross-predictions between deoxyribonucleic acid (DNA)-binding and ribonucleic acid (RNA)-binding residues and targeting the two distinct sources of binding annotations, structure-based versus intrinsic disorder-based. The methods trained on the structure-annotated interactions tend to perform poorly on the disorder-annotated binding and vice versa, with only a few methods that target and perform well across both annotation types. The cross-predictions are a significant problem, with some predictors of DNA-binding or RNA-binding residues indiscriminately predicting interactions with both nucleic acid types. Moreover, we show that methods with web servers are cited substantially more than tools without implementation or with no longer working implementations, motivating the development and long-term maintenance of the web servers. We close by discussing future research directions that aim to drive further progress in this area. Sushmita Basu, Daisuke Kihara, Lukasz A. Kurgan |
Briefings Bioinform. | 3 |
| 2025 | SHREC 2025: Protein surface shape retrieval including electrostatic potentialabstractThis SHREC 2025 track dedicated to protein surface shape retrieval involved 9 participating teams. We evaluated the performance in retrieval of 15 proposed methods on a large dataset of 11,555 protein surfaces with calculated electrostatic potential (a key molecular surface descriptor). The performance in retrieval of the proposed methods was evaluated through different metrics (Accuracy, Balanced accuracy, F1 score, Precision and Recall). The best retrieval performance was achieved by the proposed methods that used the electrostatic potential complementary to molecular surface shape. This observation was also valid for classes with limited data which highlights the importance of taking into account additional molecular surface descriptors. Taher Yacoub, Camille Depenveiller, Atsushi Tatsuma, Tin Barisin, Eugen Rusakov, Udo Göbel, Yuxu Peng, Shiqiang Deng, Yuki Kagaya, Joon Hong Park, Daisuke Kihara, Marco Guerra, Giorgio Palmieri, Andrea Ranieri, Ulderico Fugacci, Silvia Biasotti, He Ruiwen, Halim Benhabiles, Adnane Cabani, Karim Hammoudi, Hao Huang 0003, Chunyan Li 0002, Alireza Tehrani, Fanwang Meng, Farnaz Heidar-Zadeh, Tuan-Anh Yang, Matthieu Montès |
Comput. Graph. | 11 |
| 2025 | MVGFormer: Multi-view perspective with graph-guided transformer for cryo-ET segmentationabstract• We propose MVGFormer, a multi-view fusion framework with a dual-stream encoder guided by a visual graph. • We design two decoder variants: MF for hierarchical feature fusion and P3DA for multi-scale representation. • We introduce a view-masked self-supervised strategy that reconstructs masked views from remaining ones. • MVGFormer achieves superior performance over existing state-of-the-art cryo-ET segmentation methods. Cryo-Electron Tomography (cryo-ET) is a cutting-edge 3D imaging technology that enables detailed examination of biological macromolecular structures at near-atomic resolution. Recent deep learning applications on cryo-ET, such as cryo-ET segmentation, have drawn widespread interest for their potential to improve particle alignment, classification, and other tasks. However, current methods heavily rely on convolutional architectures, which prioritize local information while neglecting the global structural information inherent in cryo-ET data. Transformer-based models, known for their large receptive field, have become the de-facto design for 2D vision tasks due to their ability to effectively capture global information. This approach is also well-suited for 3D tasks, given the complex nature of 3D objects. Based on this, we extend 2D vision transformers into 3D and propose a novel transformer-based framework for cryo-ET segmentation, named MVGFormer. MVGFormer introduces a multi-view perspective fusion transformer encoder, which captures rich global structural information from multiple perspectives using unique positional embeddings. To enhance contextual awareness, we design a parallel context encoder that builds a visual graph to guide attention. We further introduce two complementary 3D decoders: multi-level feature fusion (MF) and parallel atrous convolutions (P3DA), which together capture multi-scale structural cues for precise segmentation. Furthermore, we introduce a view-masked self-supervised learning strategy to reinforce the effectiveness of the multi-view design and improve the model’s representation capability. To our knowledge, MVGFormer is the first transformer-based model for cryo-ET segmentation. We empirically evaluate MVGFormer on six cryo-ET datasets across three different tasks. Extensive experimental results demonstrate its superiority over state-of-the-art 3D segmentation methods. Haoran Li 0024, Xingjian Li 0002, Jiahua Shi, Huaming Chen, Bo Du 0004, Johan Barthélemy, Daisuke Kihara, Jun Shen 0001, Min Xu 0009 |
Knowl. Based Syst. | 9 |
| 2023 | ACC-UNet: A Completely Convolutional UNet Model for the 2020s
Nabil Ibtehaz, Daisuke Kihara |
MICCAI (3) | 2 |
| 2023 | Enhancing cryo-EM maps with 3D deep generative networks for assisting protein structure modelingabstractMOTIVATION: The tertiary structures of an increasing number of biological macromolecules have been determined using cryo-electron microscopy (cryo-EM). However, there are still many cases where the resolution is not high enough to model the molecular structures with standard computational tools. If the resolution obtained is near the empirical borderline (3-4.5 Å), improvement in the map quality facilitates structure modeling. RESULTS: We report EM-GAN, a novel approach that modifies an input cryo-EM map to assist protein structure modeling. The method uses a 3D generative adversarial network (GAN) that has been trained on high- and low-resolution density maps to learn the density patterns, and modifies the input map to enhance its suitability for modeling. The method was tested extensively on a dataset of 65 EM maps in the resolution range of 3-6 Å and showed substantial improvements in structure modeling using popular protein structure modeling tools. AVAILABILITY AND IMPLEMENTATION: https://github.com/kiharalab/EM-GAN, Google Colab: https://tinyurl.com/3ccxpttx. Sai Raghavendra Maddhuri Venkata Subramaniya, Genki Terashi, Daisuke Kihara |
Bioinform. | 3 |
| 2022 | On the Importance of Asymmetry for Siamese Representation LearningabstractMany recent self-supervised frameworks for visual representation learning are based on certain forms of Siamese networks. Such networks are conceptually symmetric with two parallel encoders, but often practically asymmetric as numerous mechanisms are devised to break the symmetry. In this work, we conduct a formal study on the importance of asymmetry by explicitly distinguishing the two encoders within the network - one produces source encodings and the other targets. Our key insight is keeping a relatively lower variance in target than source generally benefits learning. This is empirically justified by our results from five case studies covering different variance-oriented designs, and is aligned with our preliminary theoretical analysis on the baseline. Moreover, we find the improvements from asymmetric designs generalize well to longer training schedules, multiple other frameworks and newer backbones. Finally, the combined effect of several asymmetric designs achieves a state-of-the-art accuracy on ImageNet linear probing and competitive results on downstream transfer. We hope our exploration will inspire more research in exploiting asymmetry for Siamese representation learning. Haoqi Fan 0001, Yuandong Tian, Daisuke Kihara, Xinlei Chen |
CVPR | 4 |
| 2022 | SHREC 2022: Protein-ligand binding site recognition
Luca Gagliardi, Andrea Raffo, Ulderico Fugacci, Silvia Biasotti, Walter Rocchia, Hao Huang 0003, Boulbaba Ben Amor, Yi Fang 0006, Charles Christoffer, Daisuke Kihara, Apostolos Axenopoulos, Stelios K. Mylonas, Petros Daras |
Comput. Graph. | 12 |
| 2021 | Protein contact map refinement for improving structure prediction using generative adversarial networksabstractMOTIVATION: Protein structure prediction remains as one of the most important problems in computational biology and biophysics. In the past few years, protein residue-residue contact prediction has undergone substantial improvement, which has made it a critical driving force for successful protein structure prediction. Boosting the accuracy of contact predictions has, therefore, become the forefront of protein structure prediction. RESULTS: We show a novel contact map refinement method, ContactGAN, which uses Generative Adversarial Networks (GAN). ContactGAN was able to make a significant improvement over predictions made by recent contact prediction methods when tested on three datasets including protein structure modeling targets in CASP13 and CASP14. We show improvement of precision in contact prediction, which translated into improvement in the accuracy of protein tertiary structure models. On the other hand, observed improvement over trRosetta was relatively small, reasons for which are discussed. ContactGAN will be a valuable addition in the structure prediction pipeline to achieve an extra gain in contact prediction accuracy. AVAILABILITY AND IMPLEMENTATION: https://github.com/kiharalab/ContactGAN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sai Raghavendra Maddhuri Venkata Subramaniya, Genki Terashi, Aashish Jain, Yuki Kagaya, Daisuke Kihara |
Bioinform. | 5 |
| 2021 | SHREC 2021: Retrieval and classification of protein surfaces equipped with physical and chemical properties
Andrea Raffo, Ulderico Fugacci, Silvia Biasotti, Walter Rocchia, Yonghuai Liu, Ekpo Otu, Reyer Zwiggelaar, David Hunter, Evangelia I. Zacharaki, Eleftheria Psatha, Dimitrios Laskos, Gerasimos Arvanitis, Konstantinos Moustakas, Tunde Aderinwale, Charles Christoffer, Woong-Hee Shin, Daisuke Kihara, Andrea Giachetti 0001, Huu-Nghia Nguyen, Tuan-Duy Nguyen, Vinh-Thuyen Nguyen-Truong, Danh Le-Thanh, Hai-Dang Nguyen, Minh-Triet Tran |
Comput. Graph. | 17 |
| 2021 | EnAET: A Self-Trained Framework for Semi-Supervised and Supervised Learning With Ensemble TransformationsabstractDeep neural networks have been successfully applied to many real-world applications. However, such successes rely heavily on large amounts of labeled data that is expensive to obtain. Recently, many methods for semi-supervised learning have been proposed and achieved excellent performance. In this study, we propose a new EnAET framework to further improve existing semi-supervised methods with self-supervised information. To our best knowledge, all current semi-supervised methods improve performance with prediction consistency and confidence ideas. We are the first to explore the role of self-supervised representations in semi-supervised learning under a rich family of transformations. Consequently, our framework can integrate the self-supervised information as a regularization term to further improve all current semi-supervised methods. In the experiments, we use MixMatch, which is the current state-of-the-art method on semi-supervised learning, as a baseline to test the proposed EnAET framework. Across different datasets, we adopt the same hyper-parameters, which greatly improves the generalization ability of the EnAET framework. Experiment results on different datasets demonstrate that the proposed EnAET framework greatly improves the performance of current semi-supervised algorithms. Moreover, this framework can also improve supervised learning by a large margin, including the extremely challenging scenarios with only 10 images per class. The code and experiment records are available in https://github.com/maple-research-lab/EnAET. Xiao Wang 0013, Daisuke Kihara, Jiebo Luo 0001, Guo-Jun Qi |
IEEE Trans. Image Process. | 2 |
| 2020 | A Simple But Effective Bert Model for Dialog State Tracking on Resource-Limited SystemsabstractIn a task-oriented dialog system, the goal of dialog state tracking (DST) is to monitor the state of the conversation from the dialog history. Recently, many deep learning based methods have been proposed for the task. Despite their impressive performance, current neural architectures for DST are typically heavily-engineered and conceptually complex, making it difficult to implement, debug, and maintain them in a production setting. In this work, we propose a simple but effective DST model based on BERT. In addition to its simplicity, our approach also has a number of other advantages: (a) the number of parameters does not grow with the ontology size (b) the model can operate in situations where the domain ontology may change dynamically. Experimental results demonstrate that our BERT-based model outperforms previous methods by a large margin, achieving new state-of-the-art results on the standard WoZ 2.0 dataset1. Finally, to make the model small and fast enough for resource-restricted systems, we apply the knowledge distillation method to compress our model. The final compressed model achieves comparable results with the original model while being 8x smaller and 7x faster. Tuan Manh Lai, Quan Hung Tran, Trung Bui, Daisuke Kihara |
ICASSP | 4 |
| 2020 | Protein docking model evaluation by 3D deep convolutional neural networksabstractMOTIVATION: Many important cellular processes involve physical interactions of proteins. Therefore, determining protein quaternary structures provide critical insights for understanding molecular mechanisms of functions of the complexes. To complement experimental methods, many computational methods have been developed to predict structures of protein complexes. One of the challenges in computational protein complex structure prediction is to identify near-native models from a large pool of generated models. RESULTS: We developed a convolutional deep neural network-based approach named DOcking decoy selection with Voxel-based deep neural nEtwork (DOVE) for evaluating protein docking models. To evaluate a protein docking model, DOVE scans the protein-protein interface of the model with a 3D voxel and considers atomic interaction types and their energetic contributions as input features applied to the neural network. The deep learning models were trained and validated on docking models available in the ZDock and DockGround databases. Among the different combinations of features tested, almost all outperformed existing scoring functions. AVAILABILITY AND IMPLEMENTATION: Codes available at http://github.com/kiharalab/DOVE, http://kiharalab.org/dove/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiao Wang 0013, Genki Terashi, Charles Christoffer, Daisuke Kihara |
Bioinform. | 5 |
| 2020 | SHREC 2020: Classification in cryo-electron tomograms
Ilja Gubins, Marten L. Chaillet, Gijs van der Schot, Remco C. Veltkamp, Friedrich Förster, Xiaohua Wan 0001, Xuefeng Cui, Fa Zhang 0001, Emmanuel Moebel, Xiao Wang 0004, Daisuke Kihara, Min Xu 0009, Nguyen P. Nguyen, Tommi A. White, Filiz Bunyak |
Comput. Graph. | 12 |
| 2020 | SHREC 2020: Multi-domain protein shape retrieval challenge
Florent Langenfeld, Yuxu Peng, Yukun Lai, Paul L. Rosin, Tunde Aderinwale, Genki Terashi, Charles Christoffer, Daisuke Kihara, Halim Benhabiles, Karim Hammoudi, Adnane Cabani, Féryal Windal, Mahmoud Melkemi, Andrea Giachetti 0001, Stelios K. Mylonas, Apostolos Axenopoulos, Petros Daras, Ekpo Otu, Matthieu Montès |
Comput. Graph. | 8 |
| 2019 | A Gated Self-attention Memory Network for Answer SelectionabstractTuan Lai, Quan Hung Tran, Trung Bui, Daisuke Kihara. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Tuan Manh Lai, Quan Hung Tran, Trung Bui, Daisuke Kihara |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Phylo-PFP: improved automated protein function prediction using phylogenetic distance of distantly related sequencesabstractMOTIVATION: Function annotation of proteins is fundamental in contemporary biology across fields including genomics, molecular biology, biochemistry, systems biology and bioinformatics. Function prediction is indispensable in providing clues for interpreting omics-scale data as well as in assisting biologists to build hypotheses for designing experiments. As sequencing genomes is now routine due to the rapid advancement of sequencing technologies, computational protein function prediction methods have become increasingly important. A conventional method of annotating a protein sequence is to transfer functions from top hits of a homology search; however, this approach has substantial short comings including a low coverage in genome annotation. RESULTS: Here we have developed Phylo-PFP, a new sequence-based protein function prediction method, which mines functional information from a broad range of similar sequences, including those with a low sequence similarity identified by a PSI-BLAST search. To evaluate functional similarity between identified sequences and the query protein more accurately, Phylo-PFP reranks retrieved sequences by considering their phylogenetic distance. Compared to the Phylo-PFP's predecessor, PFP, which was among the top ranked methods in the second round of the Critical Assessment of Functional Annotation (CAFA2), Phylo-PFP demonstrated substantial improvement in prediction accuracy. Phylo-PFP was further shown to outperform prediction programs to date that were ranked top in CAFA2. AVAILABILITY AND IMPLEMENTATION: Phylo-PFP web server is available for at http://kiharalab.org/phylo_pfp.php. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Aashish Jain, Daisuke Kihara |
Bioinform. | 2 |
| 2019 | Prediction of protein group function by iterative classification on functional relevance networkabstractMOTIVATION: Biological experiments including proteomics and transcriptomics approaches often reveal sets of proteins that are most likely to be involved in a disease/disorder. To understand the functional nature of a set of proteins, it is important to capture the function of the proteins as a group, even in cases where function of individual proteins is not known. In this work, we propose a model that takes groups of proteins found to work together in a certain biological context, integrates them into functional relevance networks, and subsequently employs an iterative inference on graphical models to identify group functions of the proteins, which are then extended to predict function of individual proteins. RESULTS: The proposed algorithm, iterative group function prediction (iGFP), depicts proteins as a graph that represents functional relevance of proteins considering their known functional, proteomics and transcriptional features. Proteins in the graph will be clustered into groups by their mutual functional relevance, which is iteratively updated using a probabilistic graphical model, the conditional random field. iGFP showed robust accuracy even when substantial amount of GO annotations were missing. The perspective of 'group' function annotation opens up novel approaches for understanding functional nature of proteins in biological systems.Availability and implementation: http://kiharalab.org/iGFP/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ishita K. Khan, Aashish Jain, Reda Rawi, Halima Bensmail, Daisuke Kihara |
Bioinform. | 5 |
| 2019 | A global map of the protein shape universeabstractProteins are involved in almost all functions in a living cell, and functions of proteins are realized by their tertiary structures. Obtaining a global perspective of the variety and distribution of protein structures lays a foundation for our understanding of the building principle of protein structures. In light of the rapid accumulation of low-resolution structure data from electron tomography and cryo-electron microscopy, here we map and classify three-dimensional (3D) surface shapes of proteins into a similarity space. Surface shapes of proteins were represented with 3D Zernike descriptors, mathematical moment-based invariants, which have previously been demonstrated effective for biomolecular structure similarity search. In addition to single chains of proteins, we have also analyzed the shape space occupied by protein complexes. From the mapping, we have obtained various new insights into the relationship between shapes, main-chain folds, and complex formation. The unique view obtained from shape mapping opens up new ways to understand design principles, functions, and evolution of proteins. Xusi Han, Atilla Sit, Charles Christoffer, Daisuke Kihara |
PLoS Comput. Biol. | 5 |
| 2019 | Three-dimensional Krawtchouk descriptors for protein local surface shape comparison
Atilla Sit, Woong-Hee Shin, Daisuke Kihara |
Pattern Recognit. | 3 |
| 2018 | Modeling the assembly order of multimeric heteroprotein complexesabstractProtein-protein interactions are the cornerstone of numerous biological processes. Although an increasing number of protein complex structures have been determined using experimental methods, relatively fewer studies have been performed to determine the assembly order of complexes. In addition to the insights into the molecular mechanisms of biological function provided by the structure of a complex, knowing the assembly order is important for understanding the process of complex formation. Assembly order is also practically useful for constructing subcomplexes as a step toward solving the entire complex experimentally, designing artificial protein complexes, and developing drugs that interrupt a critical step in the complex assembly. There are several experimental methods for determining the assembly order of complexes; however, these techniques are resource-intensive. Here, we present a computational method that predicts the assembly order of protein complexes by building the complex structure. The method, named Path-LzerD, uses a multimeric protein docking algorithm that assembles a protein complex structure from individual subunit structures and predicts assembly order by observing the simulated assembly process of the complex. Benchmarked on a dataset of complexes with experimental evidence of assembly order, Path-LZerD was successful in predicting the assembly pathway for the majority of the cases. Moreover, when compared with a simple approach that infers the assembly path from the buried surface area of subunits in the native complex, Path-LZerD has the strong advantage that it can be used for cases where the complex structure is not known. The path prediction accuracy decreased when starting from unbound monomers, particularly for larger complexes of five or more subunits, for which only a part of the assembly path was correctly identified. As the first method of its kind, Path-LZerD opens a new area of computational protein structure modeling and will be an indispensable approach for studying protein complexes. Lenna X. Peterson, Yoichiro Togawa, Juan Esquivel-Rodríguez, Genki Terashi, Charles Christoffer, Amitava Roy, Woong-Hee Shin, Daisuke Kihara |
PLoS Comput. Biol. | 8 |
| 2018 | IAS: Interaction Specific GO Term Associations for Predicting Protein-Protein Interaction NetworksabstractProteins carry out their function in a cell through interactions with other proteins. A large scale protein-protein interaction (PPI) network of an organism provides static yet an essential structure of interactions, which is valuable clue for understanding the functions of proteins and pathways. PPIs are determined primarily by experimental methods; however, computational PPI prediction methods can supplement or verify PPIs identified by experiment. Here, we developed a novel scoring method for predicting PPIs from Gene Ontology (GO) annotations of proteins. Unlike existing methods that consider functional similarity as an indication of interaction between proteins, the new score, named the protein-protein Interaction Association Score (IAS), was computed from GO term associations of known interacting protein pairs in 49 organisms. IAS was evaluated on PPI data of six organisms and found to outperform existing GO term-based scoring methods. Moreover, consensus scoring methods that combine different scores further improved performance of PPI prediction. Satwica Yerneni, Ishita K. Khan, Daisuke Kihara |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2017 | DextMP: deep dive into text for predicting moonlighting proteinsabstractMOTIVATION: Moonlighting proteins (MPs) are an important class of proteins that perform more than one independent cellular function. MPs are gaining more attention in recent years as they are found to play important roles in various systems including disease developments. MPs also have a significant impact in computational function prediction and annotation in databases. Currently MPs are not labeled as such in biological databases even in cases where multiple distinct functions are known for the proteins. In this work, we propose a novel method named DextMP, which predicts whether a protein is a MP or not based on its textual features extracted from scientific literature and the UniProt database. RESULTS: DextMP extracts three categories of textual information for a protein: titles, abstracts from literature, and function description in UniProt. Three language models were applied and compared: a state-of-the-art deep unsupervised learning algorithm along with two other language models of different types, Term Frequency-Inverse Document Frequency in the bag-of-words and Latent Dirichlet Allocation in the topic modeling category. Cross-validation results on a dataset of known MPs and non-MPs showed that DextMP successfully predicted MPs with over 91% accuracy with significant improvement over existing MP prediction methods. Lastly, we ran DextMP with the best performing language models and text-based feature combinations on three genomes, human, yeast and Xenopus laevis , and found that about 2.5-35% of the proteomes are potential MPs. AVAILABILITY AND IMPLEMENTATION: Code available at http://kiharalab.org/DextMP . CONTACT: [email protected]. Ishita K. Khan, Mansurul Bhuiyan, Daisuke Kihara |
Bioinform. | 3 |
| 2017 | NaviGO: interactive tool for visualization and functional similarity and coherence analysis with gene ontologyabstractBACKGROUND: The number of genomics and proteomics experiments is growing rapidly, producing an ever-increasing amount of data that are awaiting functional interpretation. A number of function prediction algorithms were developed and improved to enable fast and automatic function annotation. With the well-defined structure and manual curation, Gene Ontology (GO) is the most frequently used vocabulary for representing gene functions. To understand relationship and similarity between GO annotations of genes, it is important to have a convenient pipeline that quantifies and visualizes the GO function analyses in a systematic fashion. RESULTS: NaviGO is a web-based tool for interactive visualization, retrieval, and computation of functional similarity and associations of GO terms and genes. Similarity of GO terms and gene functions is quantified with six different scores including protein-protein interaction and context based association scores we have developed in our previous works. Interactive navigation of the GO function space provides intuitive and effective real-time visualization of functional groupings of GO terms and genes as well as statistical analysis of enriched functions. CONCLUSIONS: We developed NaviGO, which visualizes and analyses functional similarity and associations of GO terms and genes. The NaviGO webserver is freely available at: http://kiharalab.org/web/navigo . Ishita K. Khan, Ziyun Ding, Satwica Yerneni, Daisuke Kihara |
BMC Bioinform. | 5 |
| 2017 | A Study of the Boltzmann Sequence-Structure ChannelabstractWe rigorously study a channel that maps sequences from a finite alphabet to self-avoiding walks in the two-dimensional grid, inspired by a model of protein folding from statistical physics and studied empirically by biophysicists. This channel, which we call the Boltzmann sequence-structure channel, is characterized by a Boltzmann/Gibbs distribution with a free parameter corresponding to temperature. In our previous work, we verified empirically that the channel capacity appears to have a phase transition for small temperature and decays to zero for high temperature. In this paper, we make some progress toward theoretically explaining these phenomena. We first estimate the conditional entropy between the input sequence and the output fold, giving an upper bound which exhibits a phase transition with respect to temperature. Next, we formulate a class of parameter settings under which the dependence between walk energies is governed by their number of shared contacts. In this setting, we derive a lower bound on the conditional entropy. This lower bound allows us to conclude that the mutual information tends to zero in a nontrivial regime of high temperature, giving some support to the empirical fact regarding capacity. Finally, we construct an example setting of the parameters of the model for which the conditional entropy is exactly calculable and which does not exhibit a phase transition. Abram Magner, Daisuke Kihara, Wojciech Szpankowski |
Proc. IEEE | 2 |
| 2017 | Modeling disordered protein interactions from biophysical principlesabstractDisordered protein-protein interactions (PPIs), those involving a folded protein and an intrinsically disordered protein (IDP), are prevalent in the cell, including important signaling and regulatory pathways. IDPs do not adopt a single dominant structure in isolation but often become ordered upon binding. To aid understanding of the molecular mechanisms of disordered PPIs, it is crucial to obtain the tertiary structure of the PPIs. However, experimental methods have difficulty in solving disordered PPIs and existing protein-protein and protein-peptide docking methods are not able to model them. Here we present a novel computational method, IDP-LZerD, which models the conformation of a disordered PPI by considering the biophysical binding mechanism of an IDP to a structured protein, whereby a local segment of the IDP initiates the interaction and subsequently the remaining IDP regions explore and coalesce around the initial binding site. On a dataset of 22 disordered PPIs with IDPs up to 69 amino acids, successful predictions were made for 21 bound and 18 unbound receptors. The successful modeling provides additional support for biophysical principles. Moreover, the new technique significantly expands the capability of protein structure modeling and provides crucial insights into the molecular mechanisms of disordered PPIs. Lenna X. Peterson, Amitava Roy, Charles Christoffer, Genki Terashi, Daisuke Kihara |
PLoS Comput. Biol. | 5 |
| 2016 | The Boltzmann sequence-structure channelabstractWe rigorously study a channel that maps binary sequences to self-avoiding walks in the two-dimensional grid, inspired by a model of protein statistics. This channel, which we also call the Boltzmann sequence-structure channel, is characterized by a Boltzmann/Gibbs distribution with a free parameter corresponding to temperature. In our previous work, we verified experimentally that the channel capacity has a phase transition for small temperature and decays to zero for high temperature. In this paper, we make some progress towards explaining these phenomena. We first upper bound the conditional entropy between the input sequence and the output which exhibits a phase transition with respect to temperature. Then we derive a lower bound on the conditional entropy for some specific set of parameters. This lower bound allows us to conclude that the mutual information tends to zero for high temperature. Abram Magner, Daisuke Kihara, Wojciech Szpankowski |
ISIT | 2 |
| 2016 | Ensemble-based evaluation for protein structure modelsabstractMOTIVATION: Comparing protein tertiary structures is a fundamental procedure in structural biology and protein bioinformatics. Structure comparison is important particularly for evaluating computational protein structure models. Most of the model structure evaluation methods perform rigid body superimposition of a structure model to its crystal structure and measure the difference of the corresponding residue or atom positions between them. However, these methods neglect intrinsic flexibility of proteins by treating the native structure as a rigid molecule. Because different parts of proteins have different levels of flexibility, for example, exposed loop regions are usually more flexible than the core region of a protein structure, disagreement of a model to the native needs to be evaluated differently depending on the flexibility of residues in a protein. RESULTS: We propose a score named FlexScore for comparing protein structures that consider flexibility of each residue in the native state of proteins. Flexibility information may be extracted from experiments such as NMR or molecular dynamics simulation. FlexScore considers an ensemble of conformations of a protein described as a multivariate Gaussian distribution of atomic displacements and compares a query computational model with the ensemble. We compare FlexScore with other commonly used structure similarity scores over various examples. FlexScore agrees with experts' intuitive assessment of computational models and provides information of practical usefulness of models. AVAILABILITY AND IMPLEMENTATION: https://bitbucket.org/mjamroz/flexscore CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Michal Jamroz 0001, Andrzej Kolinski, Daisuke Kihara |
Bioinform. | 3 |
| 2016 | Genome-scale prediction of moonlighting proteins using diverse protein association informationabstractMOTIVATION: Moonlighting proteins (MPs) show multiple cellular functions within a single polypeptide chain. To understand the overall landscape of their functional diversity, it is important to establish a computational method that can identify MPs on a genome scale. Previously, we have systematically characterized MPs using functional and omics-scale information. In this work, we develop a computational prediction model for automatic identification of MPs using a diverse range of protein association information. RESULTS: We incorporated a diverse range of protein association information to extract characteristic features of MPs, which range from gene ontology (GO), protein-protein interactions, gene expression, phylogenetic profiles, genetic interactions and network-based graph properties to protein structural properties, i.e. intrinsically disordered regions in the protein chain. Then, we used machine learning classifiers using the broad feature space for predicting MPs. Because many known MPs lack some proteomic features, we developed an imputation technique to fill such missing features. Results on the control dataset show that MPs can be predicted with over 98% accuracy when GO terms are available. Furthermore, using only the omics-based features the method can still identify MPs with over 75% accuracy. Last, we applied the method on three genomes: Saccharomyces cerevisiae, Caenorhabditis elegans and Homo sapiens, and found that about 2-10% of proteins in the genomes are potential MPs. AVAILABILITY AND IMPLEMENTATION: Code available at http://kiharalab.org/MPprediction CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ishita K. Khan, Daisuke Kihara |
Bioinform. | 2 |
| 2015 | PFP/ESG: automated protein function prediction servers enhanced with Gene Ontology visualization toolabstractUNLABELLED: Protein function prediction (PFP) is an automated function prediction method that predicts Gene Ontology (GO) annotations for a protein sequence using distantly related sequences and contextual associations of GO terms. Extended similarity group (ESG) is another GO prediction algorithm that makes predictions based on iterative sequence database searches. Here, we provide interactive web servers for the PFP and ESG algorithms that are equipped with an effective visualization of the GO predictions in a hierarchical topology. AVAILABILITY: PFP/ESG servers are freely available at http://kiharalab.org/web/pfp.php and http://kiharalab.org/web/esg.php, or access both at http://kiharalab.org/pfp_esg.php. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ishita K. Khan, Meghana Chitale, Daisuke Kihara |
Bioinform. | 4 |
| 2015 | Large-scale binding ligand prediction by improved patch-based method Patch-Surfer2.0abstractMOTIVATION: Ligand binding is a key aspect of the function of many proteins. Thus, binding ligand prediction provides important insight in understanding the biological function of proteins. Binding ligand prediction is also useful for drug design and examining potential drug side effects. RESULTS: We present a computational method named Patch-Surfer2.0, which predicts binding ligands for a protein pocket. By representing and comparing pockets at the level of small local surface patches that characterize physicochemical properties of the local regions, the method can identify binding pockets of the same ligand even if they do not share globally similar shapes. Properties of local patches are represented by an efficient mathematical representation, 3D Zernike Descriptor. Patch-Surfer2.0 has significant technical improvements over our previous prototype, which includes a new feature that captures approximate patch position with a geodesic distance histogram. Moreover, we constructed a large comprehensive database of ligand binding pockets that will be searched against by a query. The benchmark shows better performance of Patch-Surfer2.0 over existing methods. AVAILABILITY AND IMPLEMENTATION: http://kiharalab.org/patchsurfer2.0/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaolei Zhu 0001, Yi Xiong 0002, Daisuke Kihara |
Bioinform. | 3 |
| 2015 | Navigating 3D electron microscopy maps with EM-SURFERabstractBACKGROUND: The Electron Microscopy DataBank (EMDB) is growing rapidly, accumulating biological structural data obtained mainly by electron microscopy and tomography, which are emerging techniques for determining large biomolecular complex and subcellular structures. Together with the Protein Data Bank (PDB), EMDB is becoming a fundamental resource of the tertiary structures of biological macromolecules. To take full advantage of this indispensable resource, the ability to search the database by structural similarity is essential. However, unlike high-resolution structures stored in PDB, methods for comparing low-resolution electron microscopy (EM) density maps in EMDB are not well established. RESULTS: We developed a computational method for efficiently searching low-resolution EM maps. The method uses a compact fingerprint representation of EM maps based on the 3D Zernike descriptor, which is derived from a mathematical series expansion for EM maps that are considered as 3D functions. The method is implemented in a web server named EM-SURFER, which allows users to search against the entire EMDB in real-time. EM-SURFER compares the global shapes of EM maps. Examples of search results from different types of query structures are discussed. CONCLUSIONS: We developed EM-SURFER, which retrieves structurally relevant matches for query EM maps from EMDB within seconds. The unique capability of EM-SURFER to detect 3D shape similarity of low-resolution EM maps should prove invaluable in structural biology. Juan Esquivel-Rodríguez, Yi Xiong 0002, Xusi Han, Shuomeng Guang, Charles Christoffer, Daisuke Kihara |
BMC Bioinform. | 6 |
| 2014 | The International Society of Computational Biology presents: the Great Lakes Bioinformatics Conference, May 16-18, 2014, Cincinnati, OhioabstractThe Great Lakes Bioinformatics Consortium (GLBC) is pleased to announce its ninth annual conference, the Great Lakes Bioinformatics Conference (GLBIO), to be held May 16–18, 2014 in Cincinnati, OH. GLBIO 2014 will be hosted by Cincinnati Children’s Hospital Medical Center in conjunction with the University of Cincinnati. The conference, an official conference of the International Society for Computational Biology (ISCB), provides an interdisciplinary forum for the discussion of research findings and methods, and development of long-term relationships and networking opportunities, for researchers within the region, as well as from around the world. The program will include oral presentations, poster presentations, invited keynote speakers and tutorials. From novice to expert, attendees partake in a variety of workshops, tutorials, presentations, posters, networking events and exhibits. Keynote speakers include Gary Bader, The Donnelly Centre at the University of Toronto; Tanya Y. Berger-Wolf, Department of Computer Science, University of Illinois; Charles Brooks, University of Michigan and Mike Hawrylycz, Allen Institute for Brain Science. Hands-on tutorials will be led by Sorin Draghici, Wayne State University (advanced gene network); Anil Jegga, Children’s Hospital Medical Center, University of Cincinnati (advanced gene network); Jarek Meller, Children’s Hospital Medical Center, University of Cincinnati (an introduction to bioinformatics); Isaac Neuhaus, University of CA at San Francisco (visualization web interface JavaScript for programmers); Larsson Omberg, Sage Bionetworks (an introduction to Sage Synapys platform) and more! Interested in being a part of this dynamic group of speakers? See the GLBIO 2014 Web site for details and links, at www.iscb.org/glbio. GLBIO 2014 will also feature professional and career development sessions sponsored by Federation of American Societies of Experimental Biology (FASEB) Minority Access to Research Careers (MARC). This unique program will include a career fair, wherein students will receive career guidance (including a curriculum vitae (CV) critique) and will have the opportunity to meet with recruiters. Conference educational session topics will include Algorithm Development and Machine Learning, Bioimage Analysis, Biological Networks, Chemical Biology, Disease Models and Molecular Medicine, Evolutionary, Comparative and Metagenomics and many more! GLBIO has established a strong reputation for building relationships among a nationally prominent bioscience research community, showcasing the North American Great Lakes region as a perfect place to conduct computer-aided research in the life sciences. Of the 250 attendees of GLBIO 2013, 70% of conference attendees said they would attend the 2014 conference and 96% stated they would recommend the GLBIO conference to a colleague. For more information, please go to www.iscb.org/glbio to learn more about this exciting and unique gathering of industry experts and collegiate giants. Mark your calendar today and join us in Cincinnati at GLBIO 2014! Interested in supporting GLBIO 2014 through sponsorship or exhibiting? Please contact Stacy Slagor, Director of Corporate Relations and Development, ISCB, at [email protected]. About the Great Lakes Bioinformatics Consortium: The Great Lakes Bioinformatics Consortium strives to enhance educational opportunities and research infrastructure throughout the region, to make the Great Lakes a world leader in bioinformatics and to facilitate new discoveries in data-intensive biological research. The annual research meeting (GLBIO) serves as an informal communication and networking forum for professional development. We believe that by bringing together the Great Lakes bioinformatics community on a regular basis, many new initiatives will be born. The GLBC foresees development of regional research center grant proposals to identify central strengths for research centers in the Great Lakes region and to create funded centers for bioinformatics research. Additionally, the GLBC envisions scholarship and training investments that are focused on developing talent within the Great Lakes region. About the International Society for Computational Biology : The ISCB (www.iscb.org) is the sole society representing computational biology and bioinformatics on a worldwide scale. ISCB serves a global community of >3000 scientists who are dedicated to advancing the scientific understanding of living systems through computation. It convenes the world’s experts and future leaders in top conferences, including the Intelligent Systems in Molecular Biology (ISMB) Conference, and features journals that promote discovery and expand access to computational biology and bioinformatics. It delivers valuable information about training, education, employment and other relevant news. ISCB also provides an influential voice on government policies and scientific policies that are important to its members and to the general public. Jim Cavalcoli, Lonnie R. Welch, Bruce J. Aronow, Sorin Draghici, Daisuke Kihara |
Bioinform. | 5 |
| 2014 | Comparison of Image Patches Using Local Moment InvariantsabstractWe propose a new set of moment invariants based on Krawtchouk polynomials for comparison of local patches in 2D images. Being computed from discrete functions, these moments do not carry the error due to discretization. Unlike many orthogonal moments, which usually capture global features, Krawtchouk moments can be used to compute local descriptors from a region-of-interest in an image. This can be achieved by changing two parameters, and hence shifting the center of interest region horizontally or vertically or both. This property enables comparison of two arbitrary local regions. We show that Krawtchouk moments can be written as a linear combination of geometric moments, so easily converted to rotation, size, and position independent invariants. We also construct local Hu-based invariants using Hu invariants and utilizing them on images localized by the weight function given in the definition of Krawtchouk polynomials. We give the formulation of local Krawtchouk-based and Hu-based invariants, and evaluate their discriminative performance on local comparison of artificially generated test images. Atilla Sit, Daisuke Kihara |
IEEE Trans. Image Process. | 2 |
| 2013 | Evaluation of Arterial Stiffness during the Flow-Mediated Dilation TestabstractThe paper discusses the arterial stiffness during the flow-mediated dilation (FMD) test The FMD test is a method of evaluating the vascular endothelial function and has been popular as it is non-invasive and readily performed by a skillful ultrasound technician. The FMD test, however, evaluates only the maximal increase in vascular diameter mediated by the increases in blood flow after the release of the occlusive cuff and does not evaluate the arterial viscoelastic properties. This paper thus estimates the log-linearlized stiffness, to evaluate the arterial stiffness properties using the arterial diameter and blood pressure measured in a beat-to-beat manner during the FMD test. To six healthy volunteers, we performed the FMD test to measure the arterial diameter and blood pressure with ultrasound diagnostic imaging equipment and non-invasive continuous arterial blood pressure monitor, respectively. As a result, the maximal vasodilatation ratio of FMD (FMD) was obtained after cuff occlusion. In comparison with the arterial stiffness before the FMD test, the stiffness of the arterial wall is temporarily decrease and increase. It was concluded the the arterial stiffness can be estimated on a beat-to-beat basis during the FMD test. Harutoyo Hirano, Daisuke Kihara, Hiroki Hirano, Yuichi Kurita, Teiji Ukawa, Tsuneo Takayanagi, Haruka Morimoto, Ryuji Nakamura, Noboru Saeki, Yukihito Higashi, Masashi Kawamoto, Masao Yoshizumi, Toshio Tsuji |
SMC | 2 |
| 2013 | In-depth performance evaluation of PFP and ESG sequence-based function prediction methods in CAFA 2011 experimentabstractBACKGROUND: Many Automatic Function Prediction (AFP) methods were developed to cope with an increasing growth of the number of gene sequences that are available from high throughput sequencing experiments. To support the development of AFP methods, it is essential to have community wide experiments for evaluating performance of existing AFP methods. Critical Assessment of Function Annotation (CAFA) is one such community experiment. The meeting of CAFA was held as a Special Interest Group (SIG) meeting at the Intelligent Systems in Molecular Biology (ISMB) conference in 2011. Here, we perform a detailed analysis of two sequence-based function prediction methods, PFP and ESG, which were developed in our lab, using the predictions submitted to CAFA. RESULTS: We evaluate PFP and ESG using four different measures in comparison with BLAST, Prior, and GOtcha. In addition to the predictions submitted to CAFA, we further investigate performance of a different scoring function to rank order predictions by PFP as well as PFP/ESG predictions enriched with Priors that simply adds frequently occurring Gene Ontology terms as a part of predictions. Prediction accuracies of each method were also evaluated separately for different functional categories. Successful and unsuccessful predictions by PFP and ESG are also discussed in comparison with BLAST. CONCLUSION: The in-depth analysis discussed here will complement the overall assessment by the CAFA organizers. Since PFP and ESG are based on sequence database search results, our analyses are not only useful for PFP and ESG users but will also shed light on the relationship of the sequence similarity space and functions that can be inferred from the sequences. Meghana Chitale, Ishita K. Khan, Daisuke Kihara |
BMC Bioinform. | 3 |
| 2012 | Protein domain recurrence and order can enhance prediction of protein functionsabstractMOTIVATION: Burgeoning sequencing technologies have generated massive amounts of genomic and proteomic data. Annotating the functions of proteins identified in this data has become a big and crucial problem. Various computational methods have been developed to infer the protein functions based on either the sequences or domains of proteins. The existing methods, however, ignore the recurrence and the order of the protein domains in this function inference. RESULTS: We developed two new methods to infer protein functions based on protein domain recurrence and domain order. Our first method, DRDO, calculates the posterior probability of the Gene Ontology terms based on domain recurrence and domain order information, whereas our second method, DRDO-NB, relies on the naïve Bayes methodology using the same domain architecture information. Our large-scale benchmark comparisons show strong improvements in the accuracy of the protein function inference achieved by our new methods, demonstrating that domain recurrence and order can provide important information for inference of protein functions. AVAILABILITY: The new models are provided as open source programs at http://sfb.kaust.edu.sa/Pages/Software.aspx. CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics Online. Mario Abdel Messih, Meghana Chitale, Vladimir B. Bajic, Daisuke Kihara, Xin Gao 0001 |
Bioinform. | 4 |
| 2012 | Evaluation of multiple protein docking structures using correctly predicted pairwise subunitsabstractBACKGROUND: Many functionally important proteins in a cell form complexes with multiple chains. Therefore, computational prediction of multiple protein complexes is an important task in bioinformatics. In the development of multiple protein docking methods, it is important to establish a metric for evaluating prediction results in a reasonable and practical fashion. However, since there are only few works done in developing methods for multiple protein docking, there is no study that investigates how accurate structural models of multiple protein complexes should be to allow scientists to gain biological insights. METHODS: We generated a series of predicted models (decoys) of various accuracies by our multiple protein docking pipeline, Multi-LZerD, for three multi-chain complexes with 3, 4, and 6 chains. We analyzed the decoys in terms of the number of correctly predicted pair conformations in the decoys. RESULTS AND CONCLUSION: We found that pairs of chains with the correct mutual orientation exist even in the decoys with a large overall root mean square deviation (RMSD) to the native. Therefore, in addition to a global structure similarity measure, such as the global RMSD, the quality of models for multiple chain complexes can be better evaluated by using the local measurement, the number of chain pairs with correct mutual orientation. We termed the fraction of correctly predicted pairs (RMSD at the interface of less than 4.0Å) as fpair and propose to use it for evaluation of the accuracy of multiple protein docking. Juan Esquivel-Rodríguez, Daisuke Kihara |
BMC Bioinform. | 2 |
| 2012 | Protein docking prediction using predicted protein-protein interfaceabstractBACKGROUND: Many important cellular processes are carried out by protein complexes. To provide physical pictures of interacting proteins, many computational protein-protein prediction methods have been developed in the past. However, it is still difficult to identify the correct docking complex structure within top ranks among alternative conformations. RESULTS: We present a novel protein docking algorithm that utilizes imperfect protein-protein binding interface prediction for guiding protein docking. Since the accuracy of protein binding site prediction varies depending on cases, the challenge is to develop a method which does not deteriorate but improves docking results by using a binding site prediction which may not be 100% accurate. The algorithm, named PI-LZerD (using Predicted Interface with Local 3D Zernike descriptor-based Docking algorithm), is based on a pair wise protein docking prediction algorithm, LZerD, which we have developed earlier. PI-LZerD starts from performing docking prediction using the provided protein-protein binding interface prediction as constraints, which is followed by the second round of docking with updated docking interface information to further improve docking conformation. Benchmark results on bound and unbound cases show that PI-LZerD consistently improves the docking prediction accuracy as compared with docking without using binding site prediction or using the binding site prediction as post-filtering. CONCLUSION: We have developed PI-LZerD, a pairwise docking algorithm, which uses imperfect protein-protein binding interface prediction to improve docking accuracy. PI-LZerD consistently showed better prediction accuracy over alternative methods in the series of benchmark experiments including docking using actual docking interface site predictions as well as unbound docking cases. Daisuke Kihara |
BMC Bioinform. | 2 |
| 2012 | Constructing patch-based ligand-binding pocket database for predicting function of proteinsabstractBACKGROUND: Many of solved tertiary structures of unknown functions do not have global sequence and structural similarities to proteins of known function. Often functional clues of unknown proteins can be obtained by predicting small ligand molecules that bind to the proteins. METHODS: In our previous work, we have developed an alignment free local surface-based pocket comparison method, named Patch-Surfer, which predicts ligand molecules that are likely to bind to a protein of interest. Given a query pocket in a protein, Patch-Surfer searches a database of known pockets and finds similar ones to the query. Here, we have extended the database of ligand binding pockets for Patch-Surfer to cover diverse types of binding ligands. RESULTS AND CONCLUSION: We selected 9393 representative pockets with 2707 different ligand types from the Protein Data Bank. We tested Patch-Surfer on the extended pocket database to predict binding ligand of 75 non-homologous proteins that bind one of seven different ligands. Patch-Surfer achieved the average enrichment factor at 0.1 percent of over 20.0. The results did not depend on the sequence similarity of the query protein to proteins in the database, indicating that Patch-Surfer can identify correct pockets even in the absence of known homologous structures in the database. Lee Sael, Daisuke Kihara |
BMC Bioinform. | 2 |
| 2012 | Effective inter-residue contact definitions for accurate protein fold recognitionabstractBACKGROUND: Effective encoding of residue contact information is crucial for protein structure prediction since it has a unique role to capture long-range residue interactions compared to other commonly used scoring terms. The residue contact information can be incorporated in structure prediction in several different ways: It can be incorporated as statistical potentials or it can be also used as constraints in ab initio structure prediction. To seek the most effective definition of residue contacts for template-based protein structure prediction, we evaluated 45 different contact definitions, varying bases of contacts and distance cutoffs, in terms of their ability to identify proteins of the same fold. RESULTS: We found that overall the residue contact pattern can distinguish protein folds best when contacts are defined for residue pairs whose Cβ atoms are at 7.0 Å or closer to each other. Lower fold recognition accuracy was observed when inaccurate threading alignments were used to identify common residue contacts between protein pairs. In the case of threading, alignment accuracy strongly influences the fraction of common contacts identified among proteins of the same fold, which eventually affects the fold recognition accuracy. The largest deterioration of the fold recognition was observed for β-class proteins when the threading methods were used because the average alignment accuracy was worst for this fold class. When results of fold recognition were examined for individual proteins, we found that the effective contact definition depends on the fold of the proteins. A larger distance cutoff is often advantageous for capturing spatial arrangement of the secondary structures which are not physically in contact. For capturing contacts between neighboring β strands, considering the distance between Cα atoms is better than the Cβ-based distance because the side-chain of interacting residues on β strands sometimes point to opposite directions. CONCLUSION: Residue contacts defined by Cβ-Cβ distance of 7.0 Å work best overall among tested to identify proteins of the same fold. We also found that effective contact definitions differ from fold to fold, suggesting that using different residue contact definition specific for each template will lead to improvement of the performance of threading. Daisuke Kihara |
BMC Bioinform. | 3 |
| 2011 | Quantification of Protein Group Coherence and Pathway Assignment Using Functional AssociationabstractBACKGROUND: Genomics and proteomics experiments produce a large amount of data that are awaiting functional elucidation. An important step in analyzing such data is to identify functional units, which consist of proteins that play coherent roles to carry out the function. Importantly, functional coherence is not identical with functional similarity. For example, proteins in the same pathway may not share the same Gene Ontology (GO) terms, but they work in a coordinated fashion so that the aimed function can be performed. Thus, simply applying existing functional similarity measures might not be the best solution to identify functional units in omics data. RESULTS: We have designed two scores for quantifying the functional coherence by considering association of GO terms observed in two biological contexts, co-occurrences in protein annotations and co-mentions in literature in the PubMed database. The counted co-occurrences of GO terms were normalized in a similar fashion as the statistical amino acid contact potential is computed in the protein structure prediction field. We demonstrate that the developed scores can identify functionally coherent protein sets, i.e. proteins in the same pathways, co-localized proteins, and protein complexes, with statistically significant score values showing a better accuracy than existing functional similarity scores. The scores are also capable of detecting protein pairs that interact with each other. It is further shown that the functional coherence scores can accurately assign proteins to their respective pathways. CONCLUSION: We have developed two scores which quantify the functional coherence of sets of proteins. The scores reflect the actual associations of GO terms observed either in protein annotations or in literature. It has been shown that they have the ability to accurately distinguish biologically relevant groups of proteins from random ones as well as a good discriminative power for detecting interacting pairs of proteins. The scores were further successfully applied for assigning proteins to pathways. Meghana Chitale, Shriphani Palakodety, Daisuke Kihara |
BMC Bioinform. | 3 |
| 2010 | Functional enrichment analyses and construction of functional similarity networks with high confidence function prediction by PFPabstractBACKGROUND: A new paradigm of biological investigation takes advantage of technologies that produce large high throughput datasets, including genome sequences, interactions of proteins, and gene expression. The ability of biologists to analyze and interpret such data relies on functional annotation of the included proteins, but even in highly characterized organisms many proteins can lack the functional evidence necessary to infer their biological relevance. RESULTS: Here we have applied high confidence function predictions from our automated prediction system, PFP, to three genome sequences, Escherichia coli, Saccharomyces cerevisiae, and Plasmodium falciparum (malaria). The number of annotated genes is increased by PFP to over 90% for all of the genomes. Using the large coverage of the function annotation, we introduced the functional similarity networks which represent the functional space of the proteomes. Four different functional similarity networks are constructed for each proteome, one each by considering similarity in a single Gene Ontology (GO) category, i.e. Biological Process, Cellular Component, and Molecular Function, and another one by considering overall similarity with the funSim score. The functional similarity networks are shown to have higher modularity than the protein-protein interaction network. Moreover, the funSim score network is distinct from the single GO-score networks by showing a higher clustering degree exponent value and thus has a higher tendency to be hierarchical. In addition, examining function assignments to the protein-protein interaction network and local regions of genomes has identified numerous cases where subnetworks or local regions have functionally coherent proteins. These results will help interpreting interactions of proteins and gene orders in a genome. Several examples of both analyses are highlighted. CONCLUSION: The analyses demonstrate that applying high confidence predictions from PFP can have a significant impact on a researchers' ability to interpret the immense biological data that are being generated today. The newly introduced functional similarity networks of the three organisms show different network properties as compared with the protein-protein interaction networks. Troy Hawkins, Meghana Chitale, Daisuke Kihara |
BMC Bioinform. | 3 |
| 2010 | Improved protein surface comparison and application to low-resolution protein structure dataabstractBACKGROUND: Recent advancements of experimental techniques for determining protein tertiary structures raise significant challenges for protein bioinformatics. With the number of known structures of unknown function expanding at a rapid pace, an urgent task is to provide reliable clues to their biological function on a large scale. Conventional approaches for structure comparison are not suitable for a real-time database search due to their slow speed. Moreover, a new challenge has arisen from recent techniques such as electron microscopy (EM), which provide low-resolution structure data. Previously, we have introduced a method for protein surface shape representation using the 3D Zernike descriptors (3DZDs). The 3DZD enables fast structure database searches, taking advantage of its rotation invariance and compact representation. The search results of protein surface represented with the 3DZD has showngood agreement with the existing structure classifications, but some discrepancies were also observed. RESULTS: The three new surface representations of backbone atoms, originally devised all-atom-surface representation, and the combination of all-atom surface with the backbone representation are examined. All representations are encoded with the 3DZD. Also, we have investigated the applicability of the 3DZD for searching protein EM density maps of varying resolutions. The surface representations are evaluated on structure retrieval using two existing classifications, SCOP and the CE-based classification. CONCLUSIONS: Overall, the 3DZDs representing backbone atoms show better retrieval performance than the original all-atom surface representation. The performance further improved when the two representations are combined. Moreover, we observed that the 3DZD is also powerful in comparing low-resolution structures obtained by electron microscopy. Lee Sael, Daisuke Kihara |
BMC Bioinform. | 2 |
| 2009 | ESG: extended similarity group method for automated protein function predictionabstractMOTIVATION: Importance of accurate automatic protein function prediction is ever increasing in the face of a large number of newly sequenced genomes and proteomics data that are awaiting biological interpretation. Conventional methods have focused on high sequence similarity-based annotation transfer which relies on the concept of homology. However, many cases have been reported that simple transfer of function from top hits of a homology search causes erroneous annotation. New methods are required to handle the sequence similarity in a more robust way to combine together signals from strongly and weakly similar proteins for effectively predicting function for unknown proteins with high reliability. RESULTS: We present the extended similarity group (ESG) method, which performs iterative sequence database searches and annotates a query sequence with Gene Ontology terms. Each annotation is assigned with probability based on its relative similarity score with the multiple-level neighbors in the protein similarity graph. We will depict how the statistical framework of ESG improves the prediction accuracy by iteratively taking into account the neighborhood of query protein in the sequence similarity space. ESG outperforms conventional PSI-BLAST and the protein function prediction (PFP) algorithm. It is found that the iterative search is effective in capturing multiple-domains in a query protein, enabling accurately predicting several functions which originate from different domains. AVAILABILITY: ESG web server is available for automated protein function prediction at http://dragon.bio.purdue.edu/ESG/. Meghana Chitale, Troy Hawkins, Changsoon Park, Daisuke Kihara |
Bioinform. | 4 |
| 2009 | 3D-SURFER: software for high-throughput protein surface comparison and analysisabstractSUMMARY: We present 3D-SURFER, a web-based tool designed to facilitate high-throughput comparison and characterization of proteins based on their surface shape. As each protein is effectively represented by a vector of 3D Zernike descriptors, comparison times for a query protein against the entire PDB take, on an average, only a couple of seconds. The web interface has been designed to be as interactive as possible with displays showing animated protein rotations, CATH codes and structural alignments using the CE program. In addition, geometrically interesting local features of the protein surface, such as pockets that often correspond to ligand binding sites as well as protrusions and flat regions can also be identified and visualized. AVAILABILITY: 3D-SURFER is a web application that can be freely accessed from: http://dragon.bio.purdue.edu/3d-surfer CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. David La, Juan Esquivel-Rodríguez, Vishwesh Venkatraman, Lee Sael, Stephen Ueng, Steven Ahrendt, Daisuke Kihara |
Bioinform. | 8 |
| 2009 | Protein-protein docking using region-based 3D Zernike descriptorsabstractBACKGROUND: Protein-protein interactions are a pivotal component of many biological processes and mediate a variety of functions. Knowing the tertiary structure of a protein complex is therefore essential for understanding the interaction mechanism. However, experimental techniques to solve the structure of the complex are often found to be difficult. To this end, computational protein-protein docking approaches can provide a useful alternative to address this issue. Prediction of docking conformations relies on methods that effectively capture shape features of the participating proteins while giving due consideration to conformational changes that may occur. RESULTS: We present a novel protein docking algorithm based on the use of 3D Zernike descriptors as regional features of molecular shape. The key motivation of using these descriptors is their invariance to transformation, in addition to a compact representation of local surface shape characteristics. Docking decoys are generated using geometric hashing, which are then ranked by a scoring function that incorporates a buried surface area and a novel geometric complementarity term based on normals associated with the 3D Zernike shape description. Our docking algorithm was tested on both bound and unbound cases in the ZDOCK benchmark 2.0 dataset. In 74% of the bound docking predictions, our method was able to find a near-native solution (interface C-alphaRMSD < or = 2.5 A) within the top 1000 ranks. For unbound docking, among the 60 complexes for which our algorithm returned at least one hit, 60% of the cases were ranked within the top 2000. Comparison with existing shape-based docking algorithms shows that our method has a better performance than the others in unbound docking while remaining competitive for bound docking cases. CONCLUSION: We show for the first time that the 3D Zernike descriptors are adept in capturing shape complementarity at the protein-protein interface and useful for protein docking prediction. Rigorous benchmark studies show that our docking approach has a superior performance compared to existing methods. Vishwesh Venkatraman, Yifeng D. Yang, Lee Sael, Daisuke Kihara |
BMC Bioinform. | 4 |
| 2008 | Combining gene sequence similarity and textual information for gene function annotation in the literature
Luo Si, Danni Yu, Daisuke Kihara, Yi Fang 0008 |
Inf. Retr. | 3 |
| 2007 | Tracing Lineage in Multi-version Scientific DatabasesabstractThe critical need for better tracing of lineage in scientific databases is well known. It is clear that performance is not an issue for most domain scientists - rather the functionality is more important. In this paper, we highlight the importance of maintaining multiple versions of data and tracing fine-grained lineage in support of these needs. We study alternatives for managing versions, and propose a model for the example application of protein annotations. We present query rewriting algorithms for SPJ and ASP J queries that piggy-back lineage computation with query evaluation. Our models are implemented using PostgreSQL and tested using a large, real dataset from Uniprot. We establish the validity of the approach in enabling relevant queries and study the space and time overheads. While these overheads can be high in some cases, the real gain for scientists is the novel functionality that can allow them to ascertain reliability of derived data, and foster data-driven research. To the best of our knowledge, this is the first work that can handle these types of queries for lineage tracing. Mingwu Zhang, Daisuke Kihara, Sunil Prabhakar 0001 |
BIBE | 2 |
| 2007 | Salient critical points for meshesabstractA novel method for extracting the salient critical points of meshes, possibly with noise, is presented by combining mesh saliency with Morse theory. In this paper, we use the idea of mesh saliency as a measure of regional importance for meshes. The proposed method defines the salient critical points in a scalar function space using a center-surround filter operator on Gaussian-weighted average of the scalar of vertices. Compared to using a purely geometric measure of shape, such as curvature, our method yields more satisfactory results with the lower number of critical points. We demonstrate the effectiveness of this approach by comparing our results with the results of the conventional approaches in a number of examples. Furthermore, this work has a variety of potential applications. We give a direct application to the hierarchical topological representation for meshes by combining the salient critical points with the Morse-Smale complex. Yu-Shen Liu, Min Liu 0018, Daisuke Kihara, Karthik Ramani |
Symposium on Solid and Physical Modeling | 3 |
| 2006 | EMD: an ensemble algorithm for discovering regulatory motifs in DNA sequencesabstractBACKGROUND: Understanding gene regulatory networks has become one of the central research problems in bioinformatics. More than thirty algorithms have been proposed to identify DNA regulatory sites during the past thirty years. However, the prediction accuracy of these algorithms is still quite low. Ensemble algorithms have emerged as an effective strategy in bioinformatics for improving the prediction accuracy by exploiting the synergetic prediction capability of multiple algorithms. RESULTS: We proposed a novel clustering-based ensemble algorithm named EMD for de novo motif discovery by combining multiple predictions from multiple runs of one or more base component algorithms. The ensemble approach is applied to the motif discovery problem for the first time. The algorithm is tested on a benchmark dataset generated from E. coli RegulonDB. The EMD algorithm has achieved 22.4% improvement in terms of the nucleotide level prediction accuracy over the best stand-alone component algorithm. The advantage of the EMD algorithm is more significant for shorter input sequences, but most importantly, it always outperforms or at least stays at the same performance level of the stand-alone component algorithms even for longer sequences. CONCLUSION: We proposed an ensemble approach for the motif discovery problem by taking advantage of the availability of a large number of motif discovery programs. We have shown that the ensemble approach is an effective strategy for improving both sensitivity and specificity, thus the accuracy of the prediction. The advantage of the EMD algorithm is its flexibility in the sense that a new powerful algorithm can be easily added to the system. Jianjun Hu, Yifeng D. Yang, Daisuke Kihara |
BMC Bioinform. | 3 |