Yufeng J. Tseng

dblp:86/8711 · also Yufeng Jane Tseng · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0002-8461-6181ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A robust and interpretable graph neural network-based protocol for predicting p-glycoprotein substrates
abstract
P-glycoprotein (P-gp), a key member of the ATP-binding cassette (ABC) transporter family, plays a significant role in drug absorption and distribution by binding to diverse xenobiotics and actively transporting them out of cells. Given P-gp's widespread expression, including its critical presence at the blood-brain barrier, identifying whether a compound functions as a P-gp substrate or inhibitor is essential in drug development to evaluate its ability to penetrate the central nervous system. However, most studies on P-gp focus on inhibitor models rather than substrate models. This study presents a robust graph neural network approach to predict P-gp substrates, leveraging graph convolutional networks, AttentiveFP, and an ensemble model. Using a dataset of 1995 drug molecules (1202 substrates, 793 nonsubstrates), AttentiveFP outperformed traditional methods, achieving an ROC-AUC of 0.848 and an accuracy of 0.815. Integrated gradient analysis identified 20 key substructures associated with P-gp substrates. Most noteworthy is that the top four conferring a >70% probability of substrate classification which can be used a quick assessment in the future. This interpretable framework enhances P-gp prediction and broader drug development efforts.
Kuang-Cheng Hsu, Pei-Hua Wang, Bo-Han Su, Yufeng J. Tseng
Briefings Bioinform.4
2024 Every Pixel Has Its Moments: Ultra-High-Resolution Unpaired Image-to-Image Translation via Dense Normalization
Ming-Yang Ho, Che-Ming Wu, Min-Sheng Wu, Yufeng J. Tseng
ECCV (45)4
2024 Digital annealing optimization for natural product structure elucidation
abstract
The digital annealer (DA) leverages its computational capabilities of up to 100 000 bits to address the complex nondeterministic polynomial-time (NP)-complete challenge inherent in elucidating complex structures of natural products. Conventional computational methods often face limitations with complex mixtures, as they struggle to manage the high dimensionality and intertwined relationships typical in natural products, resulting in inefficiencies and inaccuracies. This study reformulates the challenge into a Quadratic Unconstrained Binary Optimization framework, thereby harnessing the quantum-inspired computing power of the DA. Utilizing mass spectrometry data from three distinct herb species and various potential scaffolds, the DA proficiently locates optimal sidechain combinations that adhere to predefined target molecular weights. This methodology enhances the probability of selecting appropriate sidechains and substituted positions and ensures the generation of solutions within a reasonable 5-min window. The findings underscore the transformative potential of the DA in the realms of analytical chemistry and drug discovery, markedly improving both the precision and practicality of natural product structure elucidation.
Chien Lee, Pei-Hua Wang, Yufeng J. Tseng
Briefings Bioinform.3
2024 Pathological Gait Analysis With an Open-Source Cloud-Enabled Platform Empowered by Semi-Supervised Learning-PathoOpenGait
abstract
We present PathoOpenGait, a cloud-based platform for comprehensive gait analysis. Gait assessment is crucial in neurodegenerative diseases such as Parkinson's and multiple system atrophy, yet current techniques are neither affordable nor efficient. PathoOpenGait utilizes 2D and 3D data from a binocular 3D camera for monitoring and analyzing gait parameters. Our algorithms, including a semi-supervised learning-boosted neural network model for turn time estimation and deterministic algorithms to estimate gait parameters, were rigorously validated on annotated gait records, demonstrating high precision and consistency. We further demonstrate PathoOpenGait's applicability in clinical settings by analyzing gait trials from Parkinson's patients and healthy controls. PathoOpenGait is the first open-source, cloud-based system for gait analysis, providing a user-friendly tool for continuous patient care and monitoring. It offers a cost-effective and accessible solution for both clinicians and patients, revolutionizing the field of gait assessment. PathoOpenGait is available at https://pathoopengait.cmdm.tw.
Ming-Yang Ho, Ming-Che Kuo, Ciao-Sin Chen, Ruey-Meei Wu, Ching-Chi Chuang, Chi-Sheng Shih 0001, Yufeng J. Tseng
IEEE J. Biomed. Health Informatics7
2022 A general optimization protocol for molecular property prediction using a deep learning network
abstract
The key to generating the best deep learning model for predicting molecular property is to test and apply various optimization methods. While individual optimization methods from different past works outside the pharmaceutical domain each succeeded in improving the model performance, better improvement may be achieved when specific combinations of these methods and practices are applied. In this work, three high-performance optimization methods in the literature that have been shown to dramatically improve model performance from other fields are used and discussed, eventually resulting in a general procedure for generating optimized CNN models on different properties of molecules. The three techniques are the dynamic batch size strategy for different enumeration ratios of the SMILES representation of compounds, Bayesian optimization for selecting the hyperparameters of a model and feature learning using chemical features obtained by a feedforward neural network, which are concatenated with the learned molecular feature vector. A total of seven different molecular properties (water solubility, lipophilicity, hydration energy, electronic properties, blood-brain barrier permeability and inhibition) are used. We demonstrate how each of the three techniques can affect the model and how the best model can generally benefit from using Bayesian optimization combined with dynamic batch size tuning.
Jen-Hao Chen, Yufeng J. Tseng
Briefings Bioinform.2
2022 Comparative studies of AlphaFold, RoseTTAFold and Modeller: a case study involving the use of G-protein-coupled receptors
abstract
Neural network (NN)-based protein modeling methods have improved significantly in recent years. Although the overall accuracy of the two non-homology-based modeling methods, AlphaFold and RoseTTAFold, is outstanding, their performance for specific protein families has remained unexamined. G-protein-coupled receptor (GPCR) proteins are particularly interesting since they are involved in numerous pathways. This work directly compares the performance of these novel deep learning-based protein modeling methods for GPCRs with the most widely used template-based software-Modeller. We collected the experimentally determined structures of 73 GPCRs from the Protein Data Bank. The official AlphaFold repository and RoseTTAFold web service were used with default settings to predict five structures of each protein sequence. The predicted models were then aligned with the experimentally solved structures and evaluated by the root-mean-square deviation (RMSD) metric. If only looking at each program's top-scored structure, Modeller had the smallest average modeling RMSD of 2.17 Å, which is better than AlphaFold's 5.53 Å and RoseTTAFold's 6.28 Å, probably since Modeller already included many known structures as templates. However, the NN-based methods (AlphaFold and RoseTTAFold) outperformed Modeller in 21 and 15 out of the 73 cases with the top-scored model, respectively, where no good templates were available for Modeller. The larger RMSD values generated by the NN-based methods were primarily due to the differences in loop prediction compared to the crystal structures.
Chien Lee, Bo-Han Su, Yufeng J. Tseng
Briefings Bioinform.3
2022 Ensemble modeling with machine learning and deep learning to provide interpretable generalized rules for classifying CNS drugs with high prediction power
abstract
The trade-off between a machine learning (ML) and deep learning (DL) model's predictability and its interpretability has been a rising concern in central nervous system-related quantitative structure-activity relationship (CNS-QSAR) analysis. Many state-of-the-art predictive modeling failed to provide structural insights due to their black box-like nature. Lack of interpretability and further to provide easy simple rules would be challenging for CNS-QSAR models. To address these issues, we develop a protocol to combine the power of ML and DL to generate a set of simple rules that are easy to interpret with high prediction power. A data set of 940 market drugs (315 CNS-active, 625 CNS-inactive) with support vector machine and graph convolutional network algorithms were used. Individual ML/DL modeling methods were also constructed for comparison. The performance of these models was evaluated using an additional external dataset of 117 market drugs (42 CNS-active, 75 CNS-inactive). Fingerprint-split validation was adopted to ensure model stringency and generalizability. The resulting novel hybrid ensemble model outperformed other constituent traditional QSAR models with an accuracy of 0.96 and an F1 score of 0.95. With the power of the interpretability provided with this protocol, our model laid down a set of simple physicochemical rules to determine whether a compound can be a CNS drug using six sub-structural features. These rules displayed higher classification ability than classical guidelines, with higher specificity and more mechanistic insights than just for blood-brain barrier permeability. This hybrid protocol can potentially be used for other drug property predictions.
Tzu-Hui Yu, Bo-Han Su, Leo Chander Battalora, Sin Liu, Yufeng J. Tseng
Briefings Bioinform.5
2021 Different molecular enumeration influences in deep learning: an example using aqueous solubility
abstract
Aqueous solubility is the key property driving many chemical and biological phenomena and impacts experimental and computational attempts to assess those phenomena. Accurate prediction of solubility is essential and challenging, even with modern computational algorithms. Fingerprint-based, feature-based and molecular graph-based representations have all been used with different deep learning methods for aqueous solubility prediction. It has been clearly demonstrated that different molecular representations impact the model prediction and explainability. In this work, we reviewed different representations and also focused on using graph and line notations for modeling. In general, one canonical chemical structure is used to represent one molecule when computing its properties. We carefully examined the commonly used simplified molecular-input line-entry specification (SMILES) notation representing a single molecule and proposed to use the full enumerations in SMILES to achieve better accuracy. A convolutional neural network (CNN) was used. The full enumeration of SMILES can improve the presentation of a molecule and describe the molecule with all possible angles. This CNN model can be very robust when dealing with large datasets since no additional explicit chemistry knowledge is necessary to predict the solubility. Also, traditionally it is hard to use a neural network to explain the contribution of chemical substructures to a single property. We demonstrated the use of attention in the decoding network to detect the part of a molecule that is relevant to solubility, which can be used to explain the contribution from the CNN.
Jen-Hao Chen, Yufeng J. Tseng
Briefings Bioinform.2
2021 Current development of integrated web servers for preclinical safety and pharmacokinetics assessments in drug development
abstract
In drug development, preclinical safety and pharmacokinetics assessments of candidate drugs to ensure the safety profile are a must. While in vivo and in vitro tests are traditionally used, experimental determinations have disadvantages, as they are usually time-consuming and costly. In silico predictions of these preclinical endpoints have each been developed in the past decades. However, only a few web-based tools have integrated different models to provide a simple one-step platform to help researchers thoroughly evaluate potential drug candidates. To efficiently achieve this approach, a platform for preclinical evaluation must not only predict key ADMET (absorption, distribution, metabolism, excretion and toxicity) properties but also provide some guidance on structural modifications to improve the undesired properties. In this review, we organized and compared several existing integrated web servers that can be adopted in preclinical drug development projects to evaluate the subject of interest. We also introduced our new web server, Virtual Rat, as an alternative choice to profile the properties of drug candidates. In Virtual Rat, we provide not only predictions of important ADMET properties but also possible reasons as to why the model made those structural predictions. Multiple models were implemented into Virtual Rat, including models for predicting human ether-a-go-go-related gene (hERG) inhibition, cytochrome P450 (CYP) inhibition, mutagenicity (Ames test), blood-brain barrier penetration, cytotoxicity and Caco-2 permeability. Virtual Rat is free and has been made publicly available at https://virtualrat.cmdm.tw/.
Yi Hsiao, Bo-Han Su, Yufeng J. Tseng
Briefings Bioinform.3
2021 PanGPCR: predictions for multiple targets, repurposing and side effects
abstract
SUMMARY: Drug discovery targeting G protein-coupled receptors (GPCRs), the largest known class of therapeutic targets, is challenging. To facilitate the rapid discovery and development of GPCR drugs, we built a system, PanGPCR, to predict multiple potential GPCR targets and their expression locations in the tissues, side effects and possible repurposing of GPCR drugs. With PanGPCR, the compound of interest is docked to a library of 36 experimentally determined crystal structures comprising of 46 docking sites for human GPCRs, and a ranked list is generated from the docking studies to assess all GPCRs and their binding affinities. Users can determine a given compound's GPCR targets and its repurposing potential accordingly. Moreover, potential side effects collected from the SIDER (Side-Effect Resource) database and mapped to 45 tissues and organs are provided by linking predicted off-targets and their expressed sequence tag profiles. With PanGPCR, multiple targets, repurposing potential and side effects can be determined by simply uploading a small ligand. AVAILABILITY AND IMPLEMENTATION: PanGPCR is freely accessible at https://gpcrpanel.cmdm.tw/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lu-Chi Liu, Ming-Yang Ho, Bo-Han Su, San-Yuan Wang, Ming-Tsung Hsu, Yufeng J. Tseng
Bioinform.6
2019 PgpRules: a decision tree based prediction server for P-glycoprotein substrates and inhibitors
abstract
SUMMARY: P-glycoprotein (P-gp) is a member of ABC transporter family that actively pumps xenobiotics out of cells to protect organisms from toxic compounds. P-gp substrates can be easily pumped out of the cells to reduce their absorption; conversely P-gp inhibitors can reduce such pumping activity. Hence, it is crucial to know if a drug is a P-gp substrate or inhibitor in view of pharmacokinetics. Here we present PgpRules, an online P-gp substrate and P-gp inhibitor prediction server with ruled-sets. The two models were built using classification and regression tree algorithm. For each compound uploaded, PgpRules not only predicts whether the compound is a P-gp substrate or a P-gp inhibitor, but also provides the rules containing chemical structural features for further structural optimization. AVAILABILITY AND IMPLEMENTATION: PgpRules is freely accessible at https://pgprules.cmdm.tw/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pei-Hua Wang, Yi-shu Tu, Yufeng J. Tseng
Bioinform.3
2019 PgpRules: a decision tree based prediction server for P-glycoprotein substrates and inhibitors
abstract
Bioinformatics (2019) doi: 10.1093/bioinformatics/btz213 In the above article the following funding statement was inadvertently omitted: ‘This work was supported by the NTU SPARK program, Ministry of Science and Technology, Taiwan (most 107-2823-8-002-003-), and Computational Molecular Design and Metabolomics Laboratory, Department of Computer Science and Information Engineering at National Taiwan University.’ This has now been corrected. The author apologises for the error.
Pei-Hua Wang, Yi-shu Tu, Yufeng J. Tseng
Bioinform.3
2018 LipidPedia: a comprehensive lipid knowledgebase
abstract
Motivation: Lipids are divided into fatty acyls, glycerolipids, glycerophospholipids, sphingolipids, saccharolipids, sterols, prenol lipids and polyketides. Fatty acyls and glycerolipids are commonly used as energy storage, whereas glycerophospholipids, sphingolipids, sterols and saccharolipids are common used as components of cell membranes. Lipids in fatty acyls, glycerophospholipids, sphingolipids and sterols classes play important roles in signaling. Although more than 36 million lipids can be identified or computationally generated, no single lipid database provides comprehensive information on lipids. Furthermore, the complex systematic or common names of lipids make the discovery of related information challenging. Results: Here, we present LipidPedia, a comprehensive lipid knowledgebase. The content of this database is derived from integrating annotation data with full-text mining of 3923 lipids and more than 400 000 annotations of associated diseases, pathways, functions and locations that are essential for interpreting lipid functions and mechanisms from over 1 400 000 scientific publications. Each lipid in LipidPedia also has its own entry containing a text summary curated from the most frequently cited diseases, pathways, genes, locations, functions, lipids and experimental models in the biomedical literature. LipidPedia aims to provide an overall synopsis of lipids to summarize lipid annotations and provide a detailed listing of references for understanding complex lipid functions and mechanisms. Availability and implementation: LipidPedia is available at http://lipidpedia.cmdm.tw. Supplementary information: Supplementary data are available at Bioinformatics online.
Tien-Chueh Kuo, Yufeng J. Tseng
Bioinform.2
2015 CypRules: a rule-based P450 inhibition prediction server
abstract
UNLABELLED: Cytochrome P450 (CYPs) are the major enzymes involved in drug metabolism and bioactivation. Inhibition models were constructed for five of the most popular enzymes from the CYP superfamily in human liver. The five enzymes chosen for this study, namely CYP1A2, CYP2D6, CYP2C19, CYP2C9 and CYP3A4, account for 90% of the xenobiotic and drug metabolism in human body. CYP enzymes can be inhibited or induced by various drugs or chemical compounds. In this work, a rule-based CYP inhibition prediction online server, CypRules, was created based on predictive models generated by the rule-based C5.0 algorithm. CypRules can predict and provide structural rulesets for CYP inhibition for each compound uploaded to the server. Capable of fast execution performance, it can be used for virtual high-throughput screening (VHTS) of a large set of testing compounds. AVAILABILITY AND IMPLEMENTATION: CypRules is freely accessible at http://cyprules.cmdm.tw/ and models, descriptor and program files for all compounds are publically available at http://cyprules.cmdm.tw/sources/sources.rar.
Chi-Yu Shao, Bo-Han Su, Yi-shu Tu, Chieh Lin, Olivia A. Lin, Yufeng J. Tseng
Bioinform.6
2010 Chromaligner: a web server for chromatogram alignment
abstract
UNLABELLED: Chromaligner is a tool for chromatogram alignment to align retention time for chromatographic methods coupled to spectrophotometers such as high performance liquid chromatography and capillary electrophoresis for metabolomics works. Chromaligner resolves peak shifts by a constrained chromatogram alignment. For a collection of chromatograms and a set of defined peaks, Chromaligner aligns the chromatograms on defined peaks using correlation warping (COW). Chromaligner is faster than the original COW algorithm by k(2) times, where k is the number of defined peaks in a chromatogram. It also provides alignments based on known component peaks to reach the best results for further chemometric analysis. AVAILABILITY: Chromaligner is freely accessible at http://cmdd.csie.ntu.edu.tw/~chromaligner.
San-Yuan Wang, Tsung-Jung Ho, Ching-Hua Kuo, Yufeng J. Tseng
Bioinform.4