Bo-Han Su

dblp:52/170 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021
YearPublicationVenuePosition
2025 A robust and interpretable graph neural network-based protocol for predicting p-glycoprotein substrates
abstract
P-glycoprotein (P-gp), a key member of the ATP-binding cassette (ABC) transporter family, plays a significant role in drug absorption and distribution by binding to diverse xenobiotics and actively transporting them out of cells. Given P-gp's widespread expression, including its critical presence at the blood-brain barrier, identifying whether a compound functions as a P-gp substrate or inhibitor is essential in drug development to evaluate its ability to penetrate the central nervous system. However, most studies on P-gp focus on inhibitor models rather than substrate models. This study presents a robust graph neural network approach to predict P-gp substrates, leveraging graph convolutional networks, AttentiveFP, and an ensemble model. Using a dataset of 1995 drug molecules (1202 substrates, 793 nonsubstrates), AttentiveFP outperformed traditional methods, achieving an ROC-AUC of 0.848 and an accuracy of 0.815. Integrated gradient analysis identified 20 key substructures associated with P-gp substrates. Most noteworthy is that the top four conferring a >70% probability of substrate classification which can be used a quick assessment in the future. This interpretable framework enhances P-gp prediction and broader drug development efforts.
Kuang-Cheng Hsu, Pei-Hua Wang, Bo-Han Su, Yufeng J. Tseng
Briefings Bioinform.3
2022 Comparative studies of AlphaFold, RoseTTAFold and Modeller: a case study involving the use of G-protein-coupled receptors
abstract
Neural network (NN)-based protein modeling methods have improved significantly in recent years. Although the overall accuracy of the two non-homology-based modeling methods, AlphaFold and RoseTTAFold, is outstanding, their performance for specific protein families has remained unexamined. G-protein-coupled receptor (GPCR) proteins are particularly interesting since they are involved in numerous pathways. This work directly compares the performance of these novel deep learning-based protein modeling methods for GPCRs with the most widely used template-based software-Modeller. We collected the experimentally determined structures of 73 GPCRs from the Protein Data Bank. The official AlphaFold repository and RoseTTAFold web service were used with default settings to predict five structures of each protein sequence. The predicted models were then aligned with the experimentally solved structures and evaluated by the root-mean-square deviation (RMSD) metric. If only looking at each program's top-scored structure, Modeller had the smallest average modeling RMSD of 2.17 Å, which is better than AlphaFold's 5.53 Å and RoseTTAFold's 6.28 Å, probably since Modeller already included many known structures as templates. However, the NN-based methods (AlphaFold and RoseTTAFold) outperformed Modeller in 21 and 15 out of the 73 cases with the top-scored model, respectively, where no good templates were available for Modeller. The larger RMSD values generated by the NN-based methods were primarily due to the differences in loop prediction compared to the crystal structures.
Chien Lee, Bo-Han Su, Yufeng J. Tseng
Briefings Bioinform.2
2022 Ensemble modeling with machine learning and deep learning to provide interpretable generalized rules for classifying CNS drugs with high prediction power
abstract
The trade-off between a machine learning (ML) and deep learning (DL) model's predictability and its interpretability has been a rising concern in central nervous system-related quantitative structure-activity relationship (CNS-QSAR) analysis. Many state-of-the-art predictive modeling failed to provide structural insights due to their black box-like nature. Lack of interpretability and further to provide easy simple rules would be challenging for CNS-QSAR models. To address these issues, we develop a protocol to combine the power of ML and DL to generate a set of simple rules that are easy to interpret with high prediction power. A data set of 940 market drugs (315 CNS-active, 625 CNS-inactive) with support vector machine and graph convolutional network algorithms were used. Individual ML/DL modeling methods were also constructed for comparison. The performance of these models was evaluated using an additional external dataset of 117 market drugs (42 CNS-active, 75 CNS-inactive). Fingerprint-split validation was adopted to ensure model stringency and generalizability. The resulting novel hybrid ensemble model outperformed other constituent traditional QSAR models with an accuracy of 0.96 and an F1 score of 0.95. With the power of the interpretability provided with this protocol, our model laid down a set of simple physicochemical rules to determine whether a compound can be a CNS drug using six sub-structural features. These rules displayed higher classification ability than classical guidelines, with higher specificity and more mechanistic insights than just for blood-brain barrier permeability. This hybrid protocol can potentially be used for other drug property predictions.
Tzu-Hui Yu, Bo-Han Su, Leo Chander Battalora, Sin Liu, Yufeng J. Tseng
Briefings Bioinform.2
2021 Current development of integrated web servers for preclinical safety and pharmacokinetics assessments in drug development
abstract
In drug development, preclinical safety and pharmacokinetics assessments of candidate drugs to ensure the safety profile are a must. While in vivo and in vitro tests are traditionally used, experimental determinations have disadvantages, as they are usually time-consuming and costly. In silico predictions of these preclinical endpoints have each been developed in the past decades. However, only a few web-based tools have integrated different models to provide a simple one-step platform to help researchers thoroughly evaluate potential drug candidates. To efficiently achieve this approach, a platform for preclinical evaluation must not only predict key ADMET (absorption, distribution, metabolism, excretion and toxicity) properties but also provide some guidance on structural modifications to improve the undesired properties. In this review, we organized and compared several existing integrated web servers that can be adopted in preclinical drug development projects to evaluate the subject of interest. We also introduced our new web server, Virtual Rat, as an alternative choice to profile the properties of drug candidates. In Virtual Rat, we provide not only predictions of important ADMET properties but also possible reasons as to why the model made those structural predictions. Multiple models were implemented into Virtual Rat, including models for predicting human ether-a-go-go-related gene (hERG) inhibition, cytochrome P450 (CYP) inhibition, mutagenicity (Ames test), blood-brain barrier penetration, cytotoxicity and Caco-2 permeability. Virtual Rat is free and has been made publicly available at https://virtualrat.cmdm.tw/.
Yi Hsiao, Bo-Han Su, Yufeng J. Tseng
Briefings Bioinform.2
2021 PanGPCR: predictions for multiple targets, repurposing and side effects
abstract
SUMMARY: Drug discovery targeting G protein-coupled receptors (GPCRs), the largest known class of therapeutic targets, is challenging. To facilitate the rapid discovery and development of GPCR drugs, we built a system, PanGPCR, to predict multiple potential GPCR targets and their expression locations in the tissues, side effects and possible repurposing of GPCR drugs. With PanGPCR, the compound of interest is docked to a library of 36 experimentally determined crystal structures comprising of 46 docking sites for human GPCRs, and a ranked list is generated from the docking studies to assess all GPCRs and their binding affinities. Users can determine a given compound's GPCR targets and its repurposing potential accordingly. Moreover, potential side effects collected from the SIDER (Side-Effect Resource) database and mapped to 45 tissues and organs are provided by linking predicted off-targets and their expressed sequence tag profiles. With PanGPCR, multiple targets, repurposing potential and side effects can be determined by simply uploading a small ligand. AVAILABILITY AND IMPLEMENTATION: PanGPCR is freely accessible at https://gpcrpanel.cmdm.tw/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lu-Chi Liu, Ming-Yang Ho, Bo-Han Su, San-Yuan Wang, Ming-Tsung Hsu, Yufeng J. Tseng
Bioinform.3
2015 CypRules: a rule-based P450 inhibition prediction server
abstract
UNLABELLED: Cytochrome P450 (CYPs) are the major enzymes involved in drug metabolism and bioactivation. Inhibition models were constructed for five of the most popular enzymes from the CYP superfamily in human liver. The five enzymes chosen for this study, namely CYP1A2, CYP2D6, CYP2C19, CYP2C9 and CYP3A4, account for 90% of the xenobiotic and drug metabolism in human body. CYP enzymes can be inhibited or induced by various drugs or chemical compounds. In this work, a rule-based CYP inhibition prediction online server, CypRules, was created based on predictive models generated by the rule-based C5.0 algorithm. CypRules can predict and provide structural rulesets for CYP inhibition for each compound uploaded to the server. Capable of fast execution performance, it can be used for virtual high-throughput screening (VHTS) of a large set of testing compounds. AVAILABILITY AND IMPLEMENTATION: CypRules is freely accessible at http://cyprules.cmdm.tw/ and models, descriptor and program files for all compounds are publically available at http://cyprules.cmdm.tw/sources/sources.rar.
Chi-Yu Shao, Bo-Han Su, Yi-shu Tu, Chieh Lin, Olivia A. Lin, Yufeng J. Tseng
Bioinform.2
2006 A reinforced merging methodology for mapping unique peptide motifs in members of protein families
abstract
BACKGROUND: Members of a protein family often have highly conserved sequences; most of these sequences carry identical biological functions and possess similar three-dimensional (3-D) structures. However, enzymes with high sequence identity may acquire differential functions other than the common catalytic ability. It is probable that each of their variable regions consists of a unique peptide motif (UPM), which selectively interacts with other cellular proteins, rendering additional biological activities. The ability to identify and localize such UPMs is paramount in recognizing the characteristic role of each member of a protein family. RESULTS: We have developed a reinforced merging algorithm (RMA) with which non-gapped UPMs were identified in a variety of query protein sequences including members of human ribonuclease A (RNaseA), epidermal growth factor receptor (EGFR), matrix metalloproteinase (MMP), and Sma-and-Mad related protein families (Smad). The UPMs generally occupy specific positions in the resolved 3-D structures, especially the loop regions on the structural surfaces. These motifs coincide with the recognition sites for antibodies, as the epitopes of four monoclonal antibodies and two polyclonal antibodies were shown to overlap with the UPMs. Most of the UPMs were found to correlate well with the potential antigenic regions predicted by PROTEAN. Furthermore, an accuracy of 70% can be achieved in terms of mapping a UPM to an epitope. CONCLUSION: Our study provides a bioinformatic approach for searching and predicting potential epitopes and interacting motifs that distinguish different members of a protein family.
Hao-Teng Chang, Tun-Wen Pai, Tan-Chi Fan, Bo-Han Su, Pei-Chih Wu, Chuan Yi Tang, Chun-Tien Chang, Shi-Hwei Liu, Margaret Dah-Tsyr Chang
BMC Bioinform.4
2005 Unique peptide prediction of RNase family sequences based on reinforced merging algorithms
Hao-Teng Chang, Tan-Chi Fan, Margaret Dah-Tsyr Chang, Tun-Wen Pai, Bo-Han Su, Pei-Chih Wu
APBC5