EDBT 2026 Demo / reviewers in the wild / expert
Kairi Furui
dblp:319/6616
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0003-1097-0003ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GraphBioisostere: general bioisostere prediction model with deep graph neural networkabstractAbstract Lead optimization to improve pharmacokinetics and toxicity while maintaining biological activity is an important and costly stage in the drug discovery process, requiring computational approaches for increased efficiency. We propose GraphBioisostere, a bioisostere prediction model that uses graph neural networks. The proposed model leverages a large-scale matched molecular pair dataset constructed from the ChEMBL database and directly learns bioisosterism without target information by considering entire chemical structures. Our evaluation shows that incorporating whole-molecule context improves bioisostere prediction compared to fragment/substituent-only inputs. Compared with a strong fingerprint-based LightGBM baseline, GraphBioisostere achieves competitive prediction performance, with the best GNN variant approaching the baseline ROC-AUC. Additionally, models pre-trained on target-independent bioisostere prediction improved transfer learning performance for potency change prediction against specific targets, particularly in low-data settings. This suggests that GraphBioisostere acquires useful representations of the relationship between chemical structure and activity. Our research provides a tool to evaluate the potential of structural changes in molecular pairs to maintain activity independently of targets, contributing to improved efficiency in the drug discovery process. Sho Masunaga, Kairi Furui, Apakorn Kengkanna, Masahito Ohue |
J. Supercomput. | 2 |
| 2026 | PBPredictor.net: GBDT-based model and web tool for prediction of blood-placental barrier permeability of small moleculesabstractAbstract The extent to which a drug administered to a mother reaches the fetus is determined by its ability to cross the blood–placental barrier. Accurate knowledge of blood–placental barrier permeability is not only crucial for the development of safe drugs but also provides essential guidance for pharmacotherapy in pregnant women, where safety concerns are paramount. However, experimental evaluation remains challenging because animal models do not adequately recapitulate the human placenta, and human-based approaches such as cord blood analysis or placental perfusion are ethically and technically constrained. In this study, we employ gradient boosting decision trees (GBDT) to construct predictive models of blood–placental barrier permeability with relatively low computational cost. Two endpoints derived from publicly available human data were modeled separately: (i) in vivo log-transformed fetal–maternal blood concentration ratios (logFM), and (ii) ex vivo clearance indices (CI) from placental perfusion experiments. In both cases, our LightGBM-based models achieved higher predictive accuracy and better generalization compared with previous approaches. To facilitate practical use, we implemented a freely accessible web application, PBPredictor ( https://pbpredictor.net ), which provides real-time predictions of logFM and CI from SMILES inputs, along with programmatic access via a REST API. By integrating reliable machine learning with an easy-to-use platform, PBPredictor offers a scalable tool to support safer drug development and evidence-based treatment strategies during pregnancy. Masahito Ohue, Kairi Furui |
J. Supercomput. | 2 |
| 2025 | NPGPT: natural product-like compound generation with GPT-based chemical language modelsabstractAbstract Natural products are substances produced by organisms in nature and often possess biological activity and structural diversity. Drug development based on natural products has been common for many years. However, the intricate structures of these compounds present challenges in terms of structure determination and synthesis, particularly compared to the efficiency of high-throughput screening of synthetic compounds. In recent years, deep learning-based methods have been applied to the generation of molecules. In this study, we trained chemical language models on a natural product dataset and generated natural product-like compounds and verified the performance of the generated compounds as a drug candidate library. The results showed that the distribution of the compounds generated was similar to that of natural products. We also evaluated the effectiveness of the generated compounds as drug candidates. Our method can be used to explore the vast chemical space and reduce the time and cost of drug discovery of natural products. Koh Sakano, Kairi Furui, Masahito Ohue |
J. Supercomput. | 2 |
| 2024 | Fastlomap: faster lead optimization mapper algorithm for large-scale relative free energy perturbationabstractAbstract In recent years, free energy perturbation calculations have garnered increasing attention as tools to support drug discovery. The lead optimization mapper (Lomap) was proposed as an algorithm to calculate the relative free energy between ligands efficiently. However, Lomap requires checking whether each edge in the FEP graph is removable, which necessitates checking the constraints for all edges. Consequently, conventional Lomap requires significant computation time, at least several hours for cases involving hundreds of compounds, and is impractical for cases with more than tens of thousands of edges. In this study, we aimed to reduce the computational cost of Lomap to enable the construction of FEP graphs for hundreds of compounds. We can reduce the overall number of constraint checks required from an amount dependent on the number of edges to one dependent on the number of nodes by using the chunk check process to check the constraints for as many edges as possible simultaneously. Based on the analysis of the execution profiles, we also improved the speed of cycle constraint and diameter constraint checks. Moreover, the output graph is the same as that obtained using the conventional Lomap, enabling direct replacement of the original one with our method. With our improvement, the execution was hundreds of times faster than that of the original Lomap. Kairi Furui, Masahito Ohue |
J. Supercomput. | 1 |
| 2022 | Compound Virtual Screening by Learning-to-Rank with Gradient Boosting Decision Tree and Enrichment-based Cumulative GainabstractLearning-to-rank, a machine learning technique widely used in information retrieval, has recently been applied to the problem of ligand-based virtual screening to accelerate the early stages of new drug development. Ranking prediction models learn based on ordinal relationships, making them suitable for integrating assay data from various environments. Existing studies of rank prediction in compound screening have generally used a learning-to-rank method called RankSVM. However, they have not been compared with or validated against the gradient boosting decision tree (GBDT)-based learning-to-rank methods that have gained popularity recently. Furthermore, although the ranking metric called Normalized Discounted Cumulative Gain (NDCG) is widely used in information retrieval, it only determines whether the predictions are better than those of other models. In other words, NDCG cannot recognize when a prediction model produces worse than random results. Nevertheless, NDCG is still used in the performance evaluation of compound screening using learning-to-rank. This study used the GBDT model with ranking loss functions, called lambdarank and lambdaloss, for ligand-based virtual screening; results were compared with existing RankSVM methods and GBDT models using regression. We also proposed a new ranking metric, Normalized Enrichment Discounted Cumulative Gain (NEDCG), aiming to evaluate the goodness of ranking predictions properly. In addition, the results showed that the GBDT model with learning-to-rank outperformed existing regression methods using GBDT and RankSVM on diverse datasets. Finally, NEDCG showed that the predictions by regression were comparable to random predictions in multi-assay, multi-family datasets, demonstrating its usefulness for a more direct assessment of compound screening performance. Kairi Furui, Masahito Ohue |
CIBCB | 1 |