EDBT 2026 Demo / reviewers in the wild / expert
Masahito Ohue
dblp:120/8312
· DBLP profile ↗
22ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-0120-1643ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 7 since 2021Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GraphBioisostere: general bioisostere prediction model with deep graph neural networkabstractAbstract Lead optimization to improve pharmacokinetics and toxicity while maintaining biological activity is an important and costly stage in the drug discovery process, requiring computational approaches for increased efficiency. We propose GraphBioisostere, a bioisostere prediction model that uses graph neural networks. The proposed model leverages a large-scale matched molecular pair dataset constructed from the ChEMBL database and directly learns bioisosterism without target information by considering entire chemical structures. Our evaluation shows that incorporating whole-molecule context improves bioisostere prediction compared to fragment/substituent-only inputs. Compared with a strong fingerprint-based LightGBM baseline, GraphBioisostere achieves competitive prediction performance, with the best GNN variant approaching the baseline ROC-AUC. Additionally, models pre-trained on target-independent bioisostere prediction improved transfer learning performance for potency change prediction against specific targets, particularly in low-data settings. This suggests that GraphBioisostere acquires useful representations of the relationship between chemical structure and activity. Our research provides a tool to evaluate the potential of structural changes in molecular pairs to maintain activity independently of targets, contributing to improved efficiency in the drug discovery process. Sho Masunaga, Kairi Furui, Apakorn Kengkanna, Masahito Ohue |
J. Supercomput. | 4 |
| 2026 | PBPredictor.net: GBDT-based model and web tool for prediction of blood-placental barrier permeability of small moleculesabstractAbstract The extent to which a drug administered to a mother reaches the fetus is determined by its ability to cross the blood–placental barrier. Accurate knowledge of blood–placental barrier permeability is not only crucial for the development of safe drugs but also provides essential guidance for pharmacotherapy in pregnant women, where safety concerns are paramount. However, experimental evaluation remains challenging because animal models do not adequately recapitulate the human placenta, and human-based approaches such as cord blood analysis or placental perfusion are ethically and technically constrained. In this study, we employ gradient boosting decision trees (GBDT) to construct predictive models of blood–placental barrier permeability with relatively low computational cost. Two endpoints derived from publicly available human data were modeled separately: (i) in vivo log-transformed fetal–maternal blood concentration ratios (logFM), and (ii) ex vivo clearance indices (CI) from placental perfusion experiments. In both cases, our LightGBM-based models achieved higher predictive accuracy and better generalization compared with previous approaches. To facilitate practical use, we implemented a freely accessible web application, PBPredictor ( https://pbpredictor.net ), which provides real-time predictions of logFM and CI from SMILES inputs, along with programmatic access via a REST API. By integrating reliable machine learning with an easy-to-use platform, PBPredictor offers a scalable tool to support safer drug development and evidence-based treatment strategies during pregnancy. Masahito Ohue, Kairi Furui |
J. Supercomput. | 1 |
| 2025 | Innovations in mathematical modeling, AI, and optimization techniquesabstractAbstract This special issue is dedicated to examining the rapidly evolving fields of artificial intelligence, mathematical modeling, and optimization, with particular emphasis on their growing importance in computational science. It features the most notable papers from the "Mathematical Modeling and Problem Solving" workshop at PDPTA'24, the 30th International Conference on Parallel and Distributed Processing Techniques and Applications. The issue showcases pioneering research in areas such as natural language processing, system optimization, and high-performance computing. The nine selected studies include novel AI-driven methods for chemical compound generation, historical text recognition, and music recommendation, along with advancements in hardware optimization through reconfigurable accelerators and vector register sharing. Additionally, evolutionary and hyper-heuristic algorithms are explored for sophisticated problem-solving in engineering design, and innovative techniques are introduced for high-speed numerical methods in large-scale systems. Collectively, these contributions demonstrate the significance of AI, supercomputing, and advanced algorithms in driving the next generation of scientific discovery. Masahito Ohue, Nobuaki Yasuo, Masami Takata |
J. Supercomput. | 1 |
| 2025 | NPGPT: natural product-like compound generation with GPT-based chemical language modelsabstractAbstract Natural products are substances produced by organisms in nature and often possess biological activity and structural diversity. Drug development based on natural products has been common for many years. However, the intricate structures of these compounds present challenges in terms of structure determination and synthesis, particularly compared to the efficiency of high-throughput screening of synthetic compounds. In recent years, deep learning-based methods have been applied to the generation of molecules. In this study, we trained chemical language models on a natural product dataset and generated natural product-like compounds and verified the performance of the generated compounds as a drug candidate library. The results showed that the distribution of the compounds generated was similar to that of natural products. We also evaluated the effectiveness of the generated compounds as drug candidates. Our method can be used to explore the vast chemical space and reduce the time and cost of drug discovery of natural products. Koh Sakano, Kairi Furui, Masahito Ohue |
J. Supercomput. | 3 |
| 2024 | REALM: Region-Empowered Antibody Language Model for Antibody Property PredictionabstractProtein language models (pLM) are beneficial to build antibody property prediction models. However, current pLMs lacks the ability to understand antibody properties because region and structure information is not effectively embedded. We propose the Region-Empowered Antibody Language Model (REALM), a pLM built by multi-task pretraining strategy of residue prediction and region prediction tasks in antibodies, to incorporate not only co-evolution but also region information of antibodies. We demonstrate that our REALM improves the understanding of antibody properties, including hydrophobicity and thermo-stability. Toru Nishino, Noriji Kato, Takuya Tsutaoka, Yuanzhong Li, Masahito Ohue |
BIBM | 5 |
| 2024 | Predicting Antibody Stability pH Values from Amino Acid Sequences: Leveraging Protein Language Models for Formulation OptimizationabstractMonoclonal antibodies (mAbs) offer significant therapeutic benefits; however, their formulation requires careful optimization to prevent instability. Standard practices for determining optimal formulation conditions rely on time-consuming and costly wet lab experiments. We developed a machine learning-based approach to predict the optimal pH value for stabilizing mAbs using only their amino acid sequences by leveraging a protein language model. Due to the absence of directly relevant methods, we established a baseline by comparing various combinations of elements. We also conducted feature engineering to enhance the predictive performance by incorporating structural information and descriptors. Our approach achieved a high Pearson correlation coefficient of 0.88 on the test folds from 10-fold cross-validation, highlighting its potential to complement wet lab experiments and increase the efficiency of mAb formulation. Takuya Tsutaoka, Noriji Kato, Toru Nishino, Yuanzhong Li, Masahito Ohue |
BIBM | 5 |
| 2024 | Improving Performance on Replica-Exchange Molecular Dynamics Simulations by Optimizing GPU Core UtilizationabstractWhile GPUs are the main players of the accelerating devices on high performance computing systems, their performance depends on how to utilize a numerous number of cores in parallel on each device. Typically, a loop structure with a number of iterations is assigned to a device to utilize their cores to map calculations in iterations so that there must be enough count of iterations to fill the thousands of GPU cores in the high-end GPUs. Taisuke Boku, Masatake Sugita, Ryohei Kobayashi 0001, Shinnosuke Furuya, Takuya Fujie, Masahito Ohue, Yutaka Akiyama |
ICPP | 6 |
| 2024 | Fastlomap: faster lead optimization mapper algorithm for large-scale relative free energy perturbationabstractAbstract In recent years, free energy perturbation calculations have garnered increasing attention as tools to support drug discovery. The lead optimization mapper (Lomap) was proposed as an algorithm to calculate the relative free energy between ligands efficiently. However, Lomap requires checking whether each edge in the FEP graph is removable, which necessitates checking the constraints for all edges. Consequently, conventional Lomap requires significant computation time, at least several hours for cases involving hundreds of compounds, and is impractical for cases with more than tens of thousands of edges. In this study, we aimed to reduce the computational cost of Lomap to enable the construction of FEP graphs for hundreds of compounds. We can reduce the overall number of constraint checks required from an amount dependent on the number of edges to one dependent on the number of nodes by using the chunk check process to check the constraints for as many edges as possible simultaneously. Based on the analysis of the execution profiles, we also improved the speed of cycle constraint and diameter constraint checks. Moreover, the output graph is the same as that obtained using the conventional Lomap, enabling direct replacement of the original one with our method. With our improvement, the execution was hundreds of times faster than that of the original Lomap. Kairi Furui, Masahito Ohue |
J. Supercomput. | 2 |
| 2024 | Mathematical modeling and problem solving: from fundamentals to applicationsabstractAbstract The rapidly advancing fields of machine learning and mathematical modeling, greatly enhanced by the recent growth in artificial intelligence, are the focus of this special issue. This issue compiles extensively revised and improved versions of the top papers from the workshop on Mathematical Modeling and Problem Solving at PDPTA'23, the 29th International Conference on Parallel and Distributed Processing Techniques and Applications. Covering fundamental research in matrix operations and heuristic searches to real-world applications in computer vision and drug discovery, the issue underscores the crucial role of supercomputing and parallel and distributed computing infrastructure in research. Featuring nine key studies, this issue pushes forward computational technologies in mathematical modeling, refines techniques for analyzing images and time-series data, and introduces new methods in pharmaceutical and materials science, making significant contributions to these areas. Masahito Ohue, Kotoyu Sasayama, Masami Takata |
J. Supercomput. | 1 |
| 2024 | Antibody complementarity-determining region design using AlphaFold2 and DDG predictorabstractAbstract The constraints imposed by natural antibody affinity maturation often culminate in antibodies with suboptimal binding affinities, thereby limiting their therapeutic efficacy. As such, the augmentation of antibody binding affinity is pivotal for the advancement of efficacious antibody-based therapies. Classical experimental paradigms for antibody engineering are financially and temporally prohibitive due to the extensive combinatorial space of sequence variations in the complementarity-determining regions (CDRs). The advent of computational techniques presents a more expeditious and economical avenue for the systematic design and optimization of antibodies. In this investigation, we assess the performance of AlphaFold2 coupled with the binder hallucination technique for the computational refinement of antibody sequences to elevate the binding affinity of pre-existing antigen-antibody complexes. These methodologies exhibit the capability to predict protein tertiary structures with remarkable fidelity, even in the absence of empirically derived data. Our results intimate that the proposed approach is adept at designing antibodies with improved affinities for antigen-antibody complexes unrepresented in AlphaFold2’s training dataset, underscoring its potential as a robust and scalable strategy for antibody engineering. Takafumi Ueki, Masahito Ohue |
J. Supercomput. | 2 |
| 2023 | Enhancing Model Learning and Interpretation using Multiple Molecular Graph Representations for Compound Property and Activity PredictionabstractGraph neural networks (GNNs) demonstrate great performance in compound property and activity prediction due to their capability to efficiently learn complex molecular graph structures. However, two main limitations persist including compound representation and model interpretability. While atom-level molecular graph representations are commonly used because of their ability to capture natural topology, they may not fully express important substructures or functional groups which significantly influence molecular properties. Consequently, recent research proposes alternative representations employing reduction techniques to integrate higher-level information and leverages both representations for model learning. However, there is still a lack of study about different molecular graph representations on model learning and interpretation. Interpretability is also crucial for drug discovery as it can offer chemical insights and inspiration for optimization. Numerous studies attempt to include model interpretation to explain the rationale behind predictions, but most of them focus solely on individual prediction with little analysis of the interpretation on different molecular graph representations. This research introduces multiple molecular graph representations that incorporate higher-level information and investigates their effects on model learning and interpretation from diverse perspectives. Several experiments are conducted across a broad range of datasets and an attention mechanism is applied to identify significant features. The results indicate that combining atom graph representation with reduced molecular graph representation can yield promising model performance. Furthermore, the interpretation results can provide significant features and potential substructures consistently aligning with background knowledge. These multiple molecular graph representations and interpretation analysis can bolster model comprehension and facilitate relevant applications in drug discovery. Apakorn Kengkanna, Masahito Ohue |
CIBCB | 2 |
| 2022 | Compound Virtual Screening by Learning-to-Rank with Gradient Boosting Decision Tree and Enrichment-based Cumulative GainabstractLearning-to-rank, a machine learning technique widely used in information retrieval, has recently been applied to the problem of ligand-based virtual screening to accelerate the early stages of new drug development. Ranking prediction models learn based on ordinal relationships, making them suitable for integrating assay data from various environments. Existing studies of rank prediction in compound screening have generally used a learning-to-rank method called RankSVM. However, they have not been compared with or validated against the gradient boosting decision tree (GBDT)-based learning-to-rank methods that have gained popularity recently. Furthermore, although the ranking metric called Normalized Discounted Cumulative Gain (NDCG) is widely used in information retrieval, it only determines whether the predictions are better than those of other models. In other words, NDCG cannot recognize when a prediction model produces worse than random results. Nevertheless, NDCG is still used in the performance evaluation of compound screening using learning-to-rank. This study used the GBDT model with ranking loss functions, called lambdarank and lambdaloss, for ligand-based virtual screening; results were compared with existing RankSVM methods and GBDT models using regression. We also proposed a new ranking metric, Normalized Enrichment Discounted Cumulative Gain (NEDCG), aiming to evaluate the goodness of ranking predictions properly. In addition, the results showed that the GBDT model with learning-to-rank outperformed existing regression methods using GBDT and RankSVM on diverse datasets. Finally, NEDCG showed that the predictions by regression were comparable to random predictions in multi-assay, multi-family datasets, demonstrating its usefulness for a more direct assessment of compound screening performance. Kairi Furui, Masahito Ohue |
CIBCB | 2 |
| 2022 | Plasma protein binding prediction focusing on residue-level features and circularity of cyclic peptides by deep learningabstractMOTIVATION: In recent years, cyclic peptide drugs have been receiving increasing attention because they can target proteins that are difficult to be tackled by conventional small-molecule drugs or antibody drugs. Plasma protein binding rate (%PPB) is a significant pharmacokinetic property of a compound in drug discovery and design. However, due to structural differences, previous computational prediction methods developed for small-molecule compounds cannot be successfully applied to cyclic peptides, and methods for predicting the PPB rate of cyclic peptides with high accuracy are not yet available. RESULTS: Cyclic peptides are larger than small molecules, and their local structures have a considerable impact on PPB; thus, molecular descriptors expressing residue-level local features of cyclic peptides, instead of those expressing the entire molecule, as well as the circularity of the cyclic peptides should be considered. Therefore, we developed a prediction method named CycPeptPPB using deep learning that considers both factors. First, the macrocycle ring of cyclic peptides was decomposed residue by residue. The residue-based descriptors were arranged according to the sequence information of the cyclic peptide. Furthermore, the circular data augmentation method was used, and the circular convolution method CyclicConv was devised to express the cyclic structure. CycPeptPPB exhibited excellent performance, with mean absolute error (MAE) of 4.79% and correlation coefficient (R) of 0.92 for the public drug dataset, compared to the prediction performance of the existing PPB rate prediction software (MAE=15.08%, R=0.63). AVAILABILITY AND IMPLEMENTATION: The data underlying this article are available in the online supplementary material. The source code of CycPeptPPB is available at https://github.com/akiyamalab/cycpeptppb. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Keisuke Yanagisawa, Yasushi Yoshikawa, Masahito Ohue, Yutaka Akiyama |
Bioinform. | 4 |
| 2021 | Quantitative Estimate of Protein-Protein Interaction Targeting Drug-likenessabstractThe quantification of drug-likeness is very useful for screening drug candidates. The quantitative estimate of drug-likeness (QED) is the most commonly used quantitative drug efficacy assessment method proposed by Bickerton et al. However, QED is not considered suitable for screening compounds that target protein-protein interactions (PPI), which have garnered significant interest in recent years. Therefore, we developed a method called the quantitative estimate of protein-protein interaction targeting drug-likeness (QEPPI), specifically for early-stage screening of PPI-targeting compounds. QEPPI is an extension of the QED method for PPI-targeting drugs and developed using the QED concept, involving modeling physicochemical properties based on the information available on the drug. QEPPI models the physicochemical properties of compounds that have been reported in the literature to act on PPIs. Compounds in iPPI-DB, which comprises PPI inhibitors and stabilizers, and FDA-approved drugs were evaluated using QEPPI. The results showed that QEPPI is more suitable for the early screening of PPI-targeting compounds than QED. QEPPI was also considered an extended concept of “Rule-of-Four” (RO4), a PPI inhibitor index proposed by Morelli et al. We have been able to turn a discrete value indicator into a continuous value indicator. To compare the discriminatory performance of QEPPI and RO4, we evaluated their discriminatory performance using the datasets of PPI-target compounds and FDA-approved drugs using F-score and other indices. Results of the F-score of RO4 and QEPPI were 0.446 and 0.499, respectively. QEPPI demonstrated better performance and enabled quantification of drug-likeness for early-stage PPI drug discovery. Hence, it could be used as an initial filter for efficient screening of PPI-targeting compounds, which has been difficult in the past. Takatsugi Kosugi, Masahito Ohue |
CIBCB | 2 |
| 2021 | Drug-target affinity prediction using applicability domain based on data densityabstractIn the pursuit of research and development of drug discovery, the computational prediction of the target affinity of a drug candidate is useful for screening compounds at an early stage and for verifying the binding potential to an unknown target. The chemogenomics-based method has attracted increased attention as it integrates information pertaining to the drug and target to predict drug-target affinity (DTA). However, the compound and target spaces are vast, and without sufficient training data, proper DTA prediction is not possible. If a DTA prediction is made in this situation, it will potentially lead to false predictions. In this study, we propose a DTA prediction method that can advise whether/when there are insufficient samples in the compound/target spaces based on the concept of the applicability domain (AD) and the data density of the training dataset. AD indicates a data region in which a machine learning model can make reliable predictions. By preclassifying the samples to be predicted by the constructed AD into those within (In-AD) and those outside the AD (Out-AD), we can determine whether a reasonable prediction can be made for these samples. The results of the evaluation experiments based on the use of three different public datasets showed that the AD constructed by the k-nearest neighbor (k-NN) method worked well, i.e., the prediction accuracy of the samples classified by the AD as Out-AD was low, while the prediction accuracy of the samples classified by the AD as In-AD was high. Shunya Sugita, Masahito Ohue |
CIBCB | 2 |
| 2019 | Parallelized Pipeline for Whole Genome Shotgun Metagenomics with GHOSTZ-GPU and MEGANabstractMetagenome techniques allow analyses of microorganisms and their genes present in a given environment without isolation and culture. Thus, metagenomics has become a broadly applied tool to study various environments and elucidate the relationship between diseases and the host microbiota. With continuous improvement in the performance of genome sequencers, the number of sequence reads generated has increased exponentially; thus, methods for the efficient processing of such large numbers of sequences are required. To this end, we developed the pipeline system, GHOSTMEGAN, to speed up the processing of large-scale whole genome shotgun metagenome analysis, which integrates the sequence homology search tool GHOSTZ-GPU and the analyzing tool MEGAN. Assuming a cluster-type computer with a job scheduling system, the multi-node parallel processing of GHOSTZ-GPU and MEGAN was pipelined. Performance evaluation of GHOSTMEGAN with a whole genome sequence dataset, the oral metagenome demonstrated that execution of 128 nodes in parallel, which required 15 h on a single node, could be completed in only 20 min, thereby achieving about 45 times faster calculation. This pipeline is expected to greatly accelerate the field of metagenomics and broaden its application potential. Masahito Ohue, Marina Yamasawa, Kazuki Izawa, Yutaka Akiyama |
BIBE | 1 |
| 2019 | A playful tool for predicting protein-protein dockingabstractIn this paper, we propose a playful tool to predict protein-protein docking. The tool uses human bodies to explore better protein-protein docking, where it offers natural interactions to dock protein molecules in a flexible way and playful interactions through two users' cooperative body movements. We present the design and implementation of the proposed tool and show potential opportunities and pitfalls of the current tool. Keren Jiang, Tsubasa Iino, Risa Kimura, Tatsuo Nakajima, Kana Shimizu, Masahito Ohue, Yutaka Akiyama |
MUM | 7 |
| 2018 | MEGADOCK-Web: an integrated database of high-throughput structure-based protein-protein interaction predictionsabstractBACKGROUND: Protein-protein interactions (PPIs) play several roles in living cells, and computational PPI prediction is a major focus of many researchers. The three-dimensional (3D) structure and binding surface are important for the design of PPI inhibitors. Therefore, rigid body protein-protein docking calculations for two protein structures are expected to allow elucidation of PPIs different from known complexes in terms of 3D structures because known PPI information is not explicitly required. We have developed rapid PPI prediction software based on protein-protein docking, called MEGADOCK. In order to fully utilize the benefits of computational PPI predictions, it is necessary to construct a comprehensive database to gather prediction results and their predicted 3D complex structures and to make them easily accessible. Although several databases exist that provide predicted PPIs, the previous databases do not contain a sufficient number of entries for the purpose of discovering novel PPIs. RESULTS: In this study, we constructed an integrated database of MEGADOCK PPI predictions, named MEGADOCK-Web. MEGADOCK-Web provides more than 10 times the number of PPI predictions than previous databases and enables users to conduct PPI predictions that cannot be found in conventional PPI prediction databases. In MEGADOCK-Web, there are 7528 protein chains and 28,331,628 predicted PPIs from all possible combinations of those proteins. Each protein structure is annotated with PDB ID, chain ID, UniProt AC, related KEGG pathway IDs, and known PPI pairs. Additionally, MEGADOCK-Web provides four powerful functions: 1) searching precalculated PPI predictions, 2) providing annotations for each predicted protein pair with an experimentally known PPI, 3) visualizing candidates that may interact with the query protein on biochemical pathways, and 4) visualizing predicted complex structures through a 3D molecular viewer. CONCLUSION: MEGADOCK-Web provides a huge amount of comprehensive PPI predictions based on docking calculations with biochemical pathways and enables users to easily and quickly assess PPI feasibilities by archiving PPI predictions. MEGADOCK-Web also promotes the discovery of new PPIs and protein functions and is freely available for use at http://www.bi.cs.titech.ac.jp/megadock-web/ . Yuri Matsuzaki, Keisuke Yanagisawa, Masahito Ohue, Yutaka Akiyama |
BMC Bioinform. | 4 |
| 2018 | Computational prediction of plasma protein binding of cyclic peptides from small molecule experimental data using sparse modeling techniquesabstractBACKGROUND: Cyclic peptide-based drug discovery is attracting increasing interest owing to its potential to avoid target protein depletion. In drug discovery, it is important to maintain the biostability of a drug within the proper range. Plasma protein binding (PPB) is the most important index of biostability, and developing a computational method to predict PPB of drug candidate compounds contributes to the acceleration of drug discovery research. PPB prediction of small molecule drug compounds using machine learning has been conducted thus far; however, no study has investigated cyclic peptides because experimental information of cyclic peptides is scarce. RESULTS: First, we adopted sparse modeling and small molecule information to construct a PPB prediction model for cyclic peptides. As cyclic peptide data are limited, applying multidimensional nonlinear models involves concerns regarding overfitting. However, models constructed by sparse modeling can avoid overfitting, offering high generalization performance and interpretability. More than 1000 PPB data of small molecules are available, and we used them to construct a prediction models with two enumeration methods: enumerating lasso solutions (ELS) and forward beam search (FBS). The accuracies of the prediction models constructed by ELS and FBS were equal to or better than those of conventional non-linear models (MAE = 0.167-0.174) on cross-validation of a small molecule compound dataset. Moreover, we showed that the prediction accuracies for cyclic peptides were close to those for small molecule compounds (MAE = 0.194-0.288). Such high accuracy could not be obtained by a simple method of learning from cyclic peptide data directly by lasso regression (MAE = 0.286-0.671) or ridge regression (MAE = 0.244-0.354). CONCLUSION: In this study, we proposed a machine learning techniques that uses low-dimensional sparse modeling to predict the PPB value of cyclic peptides computationally. The low-dimensional sparse model not only exhibits excellent generalization performance but also improves interpretation of the prediction model. This can provide common an noteworthy knowledge for future cyclic peptide drug discovery studies. Takashi Tajimi, Naoki Wakui, Keisuke Yanagisawa, Yasushi Yoshikawa, Masahito Ohue, Yutaka Akiyama |
BMC Bioinform. | 5 |
| 2017 | Link Mining for Kernel-Based Compound-Protein Interaction Predictions Using a Chemogenomics Approach
Masahito Ohue, Takuro Yamazaki, Tomohiro Ban, Yutaka Akiyama |
ICIC (2) | 1 |
| 2017 | Spresso: an ultrafast compound pre-screening method based on compound decompositionabstractMOTIVATION: Recently, the number of available protein tertiary structures and compounds has increased. However, structure-based virtual screening is computationally expensive owing to docking simulations. Thus, methods that filter out obviously unnecessary compounds prior to computationally expensive docking simulations have been proposed. However, the calculation speed of these methods is not fast enough to evaluate ≥ 10 million compounds. RESULTS: In this article, we propose a novel, docking-based pre-screening protocol named Spresso (Speedy PRE-Screening method with Segmented cOmpounds). Partial structures (fragments) are common among many compounds; therefore, the number of fragment variations needed for evaluation is smaller than that of compounds. Our method increases calculation speeds by ∼200-fold compared to conventional methods. AVAILABILITY AND IMPLEMENTATION: Spresso is written in C ++ and Python, and is available as an open-source code (http://www.bi.cs.titech.ac.jp/spresso/) under the GPLv3 license. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Keisuke Yanagisawa, Shunta Komine, Shogo D. Suzuki, Masahito Ohue, Takashi Ishida 0002, Yutaka Akiyama |
Bioinform. | 4 |
| 2014 | MEGADOCK 4.0: an ultra-high-performance protein-protein docking software for heterogeneous supercomputersabstractSUMMARY: The application of protein-protein docking in large-scale interactome analysis is a major challenge in structural bioinformatics and requires huge computing resources. In this work, we present MEGADOCK 4.0, an FFT-based docking software that makes extensive use of recent heterogeneous supercomputers and shows powerful, scalable performance of >97% strong scaling. AVAILABILITY AND IMPLEMENTATION: MEGADOCK 4.0 is written in C++ with OpenMPI and NVIDIA CUDA 5.0 (or later) and is freely available to all academic and non-profit users at: http://www.bi.cs.titech.ac.jp/megadock. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Masahito Ohue, Takehiro Shimoda, Shuji Suzuki, Yuri Matsuzaki, Takashi Ishida 0002, Yutaka Akiyama |
Bioinform. | 1 |