Nguyen-Quoc-Khanh Le

dblp:183/8863 · also Nguyen Quoc Khanh Le · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0003-4896-7926ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 8 first-author · 15 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Graph-Theoretic Consistency for Robust and Topology-Aware Semi-Supervised Histopathology Segmentation (Student Abstract)
abstract
Semi-supervised semantic segmentation (SSSS) is vital in computational pathology, where dense annotations are costly and limited. Existing methods often rely on pixel-level consistency, which propagates noisy pseudo-labels and produces fragmented or topologically invalid masks. We propose Topology Graph Consistency (TGC), a framework that integrates graph-theoretic constraints by aligning Laplacian spectra, component counts, and adjacency statistics between prediction graphs and references. This enforces global topology and improves segmentation accuracy. Experiments on GlaS and CRAG demonstrate that TGC achieves state-of-the-art performance under 5–10% supervision and significantly narrows the gap to full supervision.
Ha-Hieu Pham, Minh Le, Han Huynh, Nguyen-Quoc-Khanh Le
AAAI4
2026 Toward trustworthy artificial intelligence in multi-omics: a review of reproducibility, stability, and interpretability
abstract
The integration of multi-omics data has become increasingly important in advancing precision medicine and systems biology. However, the reliability and trustworthiness of artificial intelligence (AI) models applied to such data remain critical concerns. This review examines the evolution and current landscape of reproducibility, stability, and interpretability in AI-driven multi-omics analysis. We explore these three pillars of trustworthiness in recent literature, with a particular focus on methodological innovations, benchmarking practices, and biological relevance. Drawing from key publications, including those featured in Briefings in Bioinformatics, we highlight emerging frameworks that aim to make multi-omics models more robust, transparent, and translationally meaningful. We advocate for routine adoption of TRUST-aligned evaluation practices, including structured stability assessments, multi-cohort benchmarking, and standardized model-card reporting, as default components of future multi-omics AI development. We conclude by outlining key challenges and future directions for developing trustworthy AI systems capable of supporting reproducible, interpretable, and clinically meaningful multi-omics research.
Thanh Hoa Vo, Nguyen-Quoc-Khanh Le
Briefings Bioinform.2
2026 3DICE: interpretable 3D cross-modal learning for drug-target interaction prediction and large-scale drug discovery
abstract
MOTIVATION: Drug-target interaction (DTI) prediction is a crucial step in modern drug discovery. Accurate and efficient predictions can substantially reduce costs and development time. Applications of deep learning methods for this purpose have been extensively studied in recent years, yielding instrumental contributions to this field. However, existing methods face issues pertaining to efficient learning of drug and target feature representations, which is detrimental to generalizability and performance in cold-start scenarios. Most approaches extract representations from SMILES strings for drugs and FASTA sequences for target proteins, which encode limited 3D structural information. Additionally, many models lack explainability, being black boxes that provide little physical insight into the underlying mechanisms behind such interactions. RESULTS: We propose 3DICE, a novel framework leveraging co-attention-based fusion and massively pre-trained 3D structural encoders for both drugs and proteins. Uni-Mol and ESM-IF1 are employed to generate high-fidelity, 3D structure-aware embeddings which enable richer geometric and chemical understanding. Cross-modal fusion modules further augment representations to model intermolecular binding relationships. Importantly, this mechanism also provides intrinsic interpretability, highlighting and enabling qualitative analysis of most influential atoms or residues. Experiments conducted on two canonical benchmark datasets display the competitiveness of our model in real-world scenarios. 3DICE outperformed state-of-the-art models across multiple metrics on the DrugBank and KIBA datasets. Additional experiments provide a more rigorous analysis of interpretability than is typically reported in prior DTI studies, and we find that attention consistently highlights decision-critical regions which is not intrinsically class-specific. AVAILABILITY: Our model and dataset are freely available at: https://github.com/austinatose/3DICE.
Austin Zi Rui Liu, Nguyen-Quoc-Khanh Le, Matthew Chua 0001
Bioinform.2
2026 Bayesian hyperparameter optimization improves scGPT fine-tuning for single-cell multi-omics integration
abstract
MOTIVATION: Foundation models such as scGPT have demonstrated strong potential for single-cell multi-omics integration; however, their downstream performance is highly sensitive to hyperparameter selection. Manual fine-tuning remains computationally expensive, dataset-dependent, and often irreproducible. Despite the increasing adoption of foundation models in single-cell analysis, systematic strategies for robust hyperparameter optimization remain underexplored. RESULTS: We developed a Bayesian optimization framework based on Tree-structured Parzen Estimators (TPE) for automated fine-tuning of scGPT and evaluated its performance on two benchmark bone marrow mononuclear cell (BMMC) multi-omics datasets, including CITE-seq and GSE194122 datasets. Across datasets, Bayesian optimization consistently improved biological conservation and batch integration metrics compared with default scGPT configurations. On the original BMMC benchmark, optimization improved AvgBIO from 0.59 to 0.67 and PCR from 0.33 to 0.52. On the GSE194122 dataset, the default configuration exhibited unstable convergence and weak biological preservation (AvgBIO = 0.19; ARI = 0.007), whereas Bayesian optimization substantially improved integration performance (AvgBIO = 0.60; ARI = 0.63) while reducing validation loss from 137 to 47.1. These findings demonstrate substantial dataset-specific sensitivity of scGPT fine-tuning and highlight the importance of automated optimization for stable deployment across heterogeneous multi-omics datasets. Our study demonstrates that Bayesian optimization provides an effective and reproducible strategy for stabilizing scGPT fine-tuning across diverse single-cell multi-omics datasets. Rather than introducing a new integration architecture, this work emphasizes the importance of systematic optimization for improving robustness and reproducibility of foundation-model applications in computational biology. AVAILABILITY AND IMPLEMENTATION: Our model and dataset are freely available at: https://github.com/daren642/scGPT_multiomic_tuning.
Darren Yu Jun Tay, Nguyen-Quoc-Khanh Le, Matthew Chua 0001
Bioinform.2
2025 Prediction of protein binding residues for metal ions, nucleic acids, and small molecules
abstract
Abstract Background Protein–ligand interactions are central to cellular regulation and therapeutic targeting. Experimental identification of binding residues remains costly and time-consuming, motivating the need for efficient computational approaches. We present an interpretable, XGBoost-based framework for residue-level prediction of protein binding sites specific to three ligand classes—metal ions, nucleic acids, and small molecules—emphasizing predictive accuracy and user accessibility. Methods A benchmark dataset [1] comprising 1314 annotated protein sequences was processed into peptide windows (±5, ±11, ±20 residues) centered on binding sites, maintaining a 1:3 ratio of positive to negative samples. Sequence-derived descriptors were generated using the iLearn toolkit, including amino acid composition (AAC), grouped AAC (GAAC), CTD composition, and BLOSUM62 substitution scores, yielding >8000 features subsequently reduced by feature-importance heuristics. More than 60 classifiers were evaluated; XGBoost consistently achieved superior performance after hyperparameter optimization. Results Independent models trained for each ligand type achieved high predictive power and biological interpretability. For small-molecule binding, the model attained an AUC of 0.99 and F1 of 0.91, with key features involving small residues (G, C, S, A) reflecting steric constraints. The nucleic-acid model reached AUC 0.99 and F1 0.95, highlighting glycine-rich and positively charged motifs typical of RNA/DNA interactions. The metal-ion model (AUC 0.92, F1 0.65) emphasized cysteine and histidine, consistent with metalloprotein binding patterns. Compared with reference CNN models, our framework achieved lower log-loss, greater interpretability, and stable cross-validation. Conclusion A standalone graphical user interface (GUI) implemented in Python allows users to input protein sequences and obtain residue-level binding probabilities, exportable as CSV files. This work demonstrates that well-engineered classical machine learning can rival deep learning in protein binding prediction, offering interpretable, high-performance tools to support biomedical and pharmaceutical research. References [1] Littmann M, Heinzinger M, Dallago C, Weissenow K, Rost B. Protein embeddings and deep learning predict binding residues for various ligand classes. Sci Rep. 2021 Dec 13;11(1):23916.
Nguyen-Quoc-Khanh Le
Briefings Bioinform.1
2025 ProtBert-BFD token classification improves intrinsically disordered protein region prediction
abstract
Abstract Background Intrinsically disordered regions (IDRs) lack stable tertiary structures yet are crucial to transcriptional regulation, signal transduction, and molecular recognition. Experimental annotation of IDRs remains limited—only ~25,000 disordered proteins are cataloged compared with hundreds of millions of known sequences—highlighting the need for scalable computational prediction. We present a systematic evaluation of classical, deep, and transformer-based models for IDR prediction and introduce a fine-tuned ProtBert-BFD token classification framework that achieves high accuracy and generalizability in data-limited settings. Methods Using manually curated annotations from the DisProt database, redundant sequences were filtered with CD-HIT (<30% similarity) and restricted to ≤526 residues, producing ~700 high-quality sequences. The dataset was partitioned (70:15:15) into training, validation, and test sets. We compared multiple architectures: (1) logistic regression and multilayer perceptron as classical baselines; (2) bidirectional and Seq2Seq LSTMs capturing sequential dependencies; and (3) ProtBert-BFD models leveraging pretrained protein-language embeddings for residue-level classification. Results Classical models showed limited predictive power (AUC 0.56–0.69), while LSTM variants improved recall but overfitted due to data imbalance. ProtBert-BFD fine-tuning substantially enhanced accuracy (75.6%), recall (68.4%), and F1-score (64.0%). The token classification variant achieved precision 0.816, recall 0.823, and F1 0.815—surpassing all internal baselines and approaching the state-of-the-art PROFbval model (recall 0.835). Average inference time was only 1.8 s per 530-residue sequence, enabling proteome-scale deployment. Conclusion These findings demonstrate that pretrained transformers effectively capture disorder-related sequence contexts and outperform conventional deep learning in low-data regimes. By integrating transfer learning with efficient inference, the ProtBert-BFD token classification model provides a robust framework for large-scale IDR annotation. Future work will expand datasets, incorporate physicochemical descriptors, and analyze error distributions to refine understanding of protein disorder in health and disease. References Hu, G., Katuwawala, A., Wang, K. et al. flDPnn: Accurate intrinsic disorder prediction with putative propensities of disorder functions. Nat Commun 12, 4438 (2021).
Nguyen-Quoc-Khanh Le
Briefings Bioinform.1
2025 RNA-ModX: a multilabel prediction and interpretation framework for RNA modifications
abstract
Accurate prediction of RNA modifications holds profound implications for elucidating RNA function and mechanism, with potential applications in drug development. Here, the RNA-ModX presents a highly precise predictive model designed to forecast post-transcriptional RNA modifications, complemented by a user-friendly web application tailored for seamless utilization by future researchers. To achieve exceptional accuracy, the RNA-ModX systematically explored a range of machine learning models, including Long Short-Term Memory (LSTM), Gated Recurrent Unit, and Transformer-based architectures. The model underwent rigorous testing using a dataset comprising RNA sequences containing the four fundamental nucleotides (A, C, G, U) and spanning 12 prevalent modification classes (m6A, m1A, m5C, m5U, m6Am, m7G, Ψ, I, Am, Cm, Gm, and Um), with sequences of length 1001 nucleotides. Notably, the LSTM model, augmented with 3-mer encoding, demonstrated the highest level of model accuracy. Furthermore, Local Interpretable Model-Agnostic Explanations were employed to facilitate result interpretation, enhancing the transparency and interpretability of the model's predictions. In conjunction with the model development, a user-friendly web application was meticulously crafted, featuring an intuitive interface for researchers to effortlessly upload RNA sequences. Upon submission, the model executes in the backend, generating predictions which are seamlessly presented to the user in a coherent manner. This integration of cutting-edge predictive modeling with a user-centric interface signifies a significant step forward in facilitating the exploration and utilization of RNA modification prediction technologies by the broader research community.
Chelsea Chen Yuge, Ee Soon Hang, Madasamy Ravi Nadar Mamtha, Shashikant Vishwakarma, Nguyen-Quoc-Khanh Le
Briefings Bioinform.7
2025 Deep Learning-Based Integrated System for Intraoperative Blood Loss Quantification in Surgical Sponges
abstract
Accurate quantification of intraoperative blood loss is crucial for enhancing patient safety and the success rate of surgeries. Traditional estimation techniques, mainly reliant on visual assessments, are prone to significant inaccuracies due to their subjective nature. This study introduces MDCare, an innovative deep learning-integrated system designed to substantially improve the precision of blood loss quantification using surgical sponges. By integrating advanced hardware components, including a mass sensor and webcam, with sophisticated algorithms like ResNet-18 and YOLOv4, MDCare achieves classification accuracy up to 96.2% and sponge detection accuracies above 91% for both synthetic and real blood scenarios. The system processes images at 7.4 frames per second, aligning with the exigent pace of surgical environments, thereby supporting surgeons with real-time, accurate blood loss data essential for timely and informed decision-making. The contributions of the paper are: (1) Demonstrating the application of advanced machine learning models in a critical clinical setting, achieving significantly higher accuracy in blood loss estimation compared to traditional methods; (2) Validating the system's efficacy in real-time surgical environments, thereby enhancing the decision-making process and potentially reducing postoperative complications; (3) Setting a new standard in surgical care by integrating a complex system into real-world clinical workflows, showcasing its adaptability and potential for widespread adoption. Future work will focus on expanding the dataset and refining the algorithms to ensure the MDCare system's robustness and adaptability across surgical settings. The findings underscore the potential of MDCare to automate and refine critical aspects of surgery, marking a significant advancement in surgical care.
Minh Huu Nhat Le, Trung Q. Le, Chinyere Charles-Okezie, Michael J. Diaz, Cameron Sabet, Hung The Dang, Hoan Nguyen, Nguyen-Quoc-Khanh Le, Aaron Muncey, Phat K. Huynh
IEEE J. Biomed. Health Informatics11
2023 Sequence-based prediction model of protein crystallization propensity using machine learning and two-level feature selection
abstract
Protein crystallization is crucial for biology, but the steps involved are complex and demanding in terms of external factors and internal structure. To save on experimental costs and time, the tendency of proteins to crystallize can be initially determined and screened by modeling. As a result, this study created a new pipeline aimed at using protein sequence to predict protein crystallization propensity in the protein material production stage, purification stage and production of crystal stage. The newly created pipeline proposed a new feature selection method, which involves combining Chi-square (${\chi }^{2}$) and recursive feature elimination together with the 12 selected features, followed by a linear discriminant analysisfor dimensionality reduction and finally, a support vector machine algorithm with hyperparameter tuning and 10-fold cross-validation is used to train the model and test the results. This new pipeline has been tested on three different datasets, and the accuracy rates are higher than the existing pipelines. In conclusion, our model provides a new solution to predict multistage protein crystallization propensity which is a big challenge in computational biology.
Nguyen-Quoc-Khanh Le, Wanru Li, Yanshuang Cao
Briefings Bioinform.1
2023 Prediction of anticancer peptides based on an ensemble model of deep learning and machine learning using ordinal positional encoding
abstract
Anticancer peptides (ACPs) are the types of peptides that have been demonstrated to have anticancer activities. Using ACPs to prevent cancer could be a viable alternative to conventional cancer treatments because they are safer and display higher selectivity. Due to ACP identification being highly lab-limited, expensive and lengthy, a computational method is proposed to predict ACPs from sequence information in this study. The process includes the input of the peptide sequences, feature extraction in terms of ordinal encoding with positional information and handcrafted features, and finally feature selection. The whole model comprises of two modules, including deep learning and machine learning algorithms. The deep learning module contained two channels: bidirectional long short-term memory (BiLSTM) and convolutional neural network (CNN). Light Gradient Boosting Machine (LightGBM) was used in the machine learning module. Finally, this study voted the three models' classification results for the three paths resulting in the model ensemble layer. This study provides insights into ACP prediction utilizing a novel method and presented a promising performance. It used a benchmark dataset for further exploration and improvement compared with previous studies. Our final model has an accuracy of 0.7895, sensitivity of 0.8153 and specificity of 0.7676, and it was increased by at least 2% compared with the state-of-the-art studies in all metrics. Hence, this paper presents a novel method that can potentially predict ACPs more effectively and efficiently. The work and source codes are made available to the community of researchers and developers at https://github.com/khanhlee/acp-ope/.
Qitong Yuan, Keyi Chen 0005, Yimin Yu, Nguyen-Quoc-Khanh Le, Matthew Chua 0001
Briefings Bioinform.4
2022 Intelligent wavelet fuzzy brain emotional controller using dual function-link network for uncertain nonlinear control systems
Tuan-Tu Huynh, Chih-Min Lin, Nguyen-Quoc-Khanh Le, Mai The Vu, Ngoc Phi Nguyen, Fei Chao 0001
Appl. Intell.3
2022 mCNN-ETC: identifying electron transporters and their functional families by using multiple windows scanning techniques in convolutional neural networks with evolutionary information of protein sequences
abstract
In the past decade, convolutional neural networks (CNNs) have been used as powerful tools by scientists to solve visual data tasks. However, many efforts of convolutional neural networks in solving protein function prediction and extracting useful information from protein sequences have certain limitations. In this research, we propose a new method to improve the weaknesses of the previous method. mCNN-ETC is a deep learning model which can transform the protein evolutionary information into image-like data composed of 20 channels, which correspond to the 20 amino acids in the protein sequence. We constructed CNN layers with different scanning windows in parallel to enhance the useful pattern detection ability of the proposed model. Then we filtered specific patterns through the 1-max pooling layer before inputting them into the prediction layer. This research attempts to solve a basic problem in biology in terms of application: predicting electron transporters and classifying their corresponding complexes. The performance result reached an accuracy of 97.41%, which was nearly 6% higher than its predecessor. We have also published a web server on http://bio219.bioinfo.yzu.edu.tw, which can be used for research purposes free of charge.
Quang-Thai Ho, Nguyen-Quoc-Khanh Le, Yu-Yen Ou
Briefings Bioinform.2
2022 Use Chou's 5-Steps Rule With Different Word Embedding Types to Boost Performance of Electron Transport Protein Prediction Model
abstract
Living organisms receive necessary energy substances directly from cellular respiration. The completion of electron storage and transportation requires the process of cellular respiration with the aid of electron transport chains. Therefore, the work of deciphering electron transport proteins is inevitably needed. The identification of these proteins with high performance has a prompt dependence on the choice of methods for feature extraction and machine learning algorithm. In this study, protein sequences served as natural language sentences comprising words. The nominated word embedding-based feature sets, hinged on the word embedding modulation and protein motif frequencies, were useful for feature choosing. Five word embedding types and a variety of conjoint features were examined for such feature selection. The support vector machine algorithm consequentially was employed to perform classification. The performance statistics within the 5-fold cross-validation including average accuracy, specificity, sensitivity, as well as MCC rates surpass 0.95. Such metrics in the independent test are 96.82, 97.16, 95.76 percent, and 0.9, respectively. Compared to state-of-the-art predictors, the proposed method can generate more preferable performance above all metrics indicating the effectiveness of the proposed method in determining electron transport proteins. Furthermore, this study reveals insights about the applicability of various word embeddings for understanding surveyed sequences.
Trinh-Trung-Duong Nguyen, Quang-Thai Ho, Nguyen-Quoc-Khanh Le, Van-Dinh Phan, Yu-Yen Ou
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 An Extensive Examination of Discovering 5-Methylcytosine Sites in Genome-Wide DNA Promoters Using Machine Learning Based Approaches
abstract
It is well-known that the major reason for the rapid proliferation of cancer cells are the hypomethylation of the whole cancer genome and the hypermethylation of the promoter of particular tumor suppressor genes. Locating 5-methylcytosine (5mC) sites in promoters is therefore a crucial step in further understanding of the relationship between promoter methylation and the regulation of mRNA gene expression. High throughput identification of DNA 5mC in wet lab is still time-consuming and labor-extensive. Thus, finding the 5mC site of genome-wide DNA promoters is still an important task. We compared the effectiveness of the most popular and strong machine learning techniques namely XGBoost, Random Forest, Deep Forest, and Deep Feedforward Neural Network in predicting the 5mC sites of genome-wide DNA promoters. A feature extraction method based on k-mers embeddings learned from a language model were also applied. Overall, the performance of all the surveyed models surpassed deep learning models of the latest studies on the same dataset employing other encoding scheme. Furthermore, the best model achieved AUC scores of 0.962 on both cross-validation and independent test data. We concluded that our approach was efficient for identifying 5mC sites of promoters with high performance.
Trinh-Trung-Duong Nguyen, The-Anh Tran, Nguyen-Quoc-Khanh Le, Dinh-Minh Pham, Yu-Yen Ou
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 Self-Organizing Double Function-Link Fuzzy Brain Emotional Control System Design for Uncertain Nonlinear Systems
abstract
This article aims to propose a more efficient control algorithm for uncertain nonlinear systems. An intelligent self-organizing double function-link fuzzy brain emotional control system is proposed which comprises a self-organizing double function-link fuzzy brain emotional controller (SDFLFBEC) and a compensation controller. The proposed SDFLFBEC consists of four substructures and a fuzzy inference system. The substructures are the prefrontal cortex, the amygdala, a double function-link network (FLN) and a self-organizing structure. The prefrontal cortex and the amygdala networks work as a mathematical form that presumes the judgment and emotion of a brain. Specifically, a new double FLN is designed to support the above networks for updating their weights. Next, a self-organizing structure can automatically add or prune the layers to achieve efficient network structure. In addition, the fuzzy inference rules are presented to explain the inference processes of the amygdala and orbitofrontal networks. From the above factors, the proposed SDFLFBEC can effectively reduce the tracking error and achieve favorable control performance. The parameters of the control system are adjusted online using the derived adaptation laws that are taken from a Lyapunov function so that the stability of the system is ensured. Simulation studies of a biped robot and the experimental results of a magnetic levitation system are employed to validate the effectiveness and superiority of the proposed SDFLFBEC.
Tuan-Tu Huynh, Chih-Min Lin, Tien-Loc Le, Nguyen-Quoc-Khanh Le, Van-Phong Vu, Fei Chao 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2021 Using deep neural networks and biological subwords to detect protein S-sulfenylation sites
abstract
Protein S-sulfenylation is one kind of crucial post-translational modifications (PTMs) in which the hydroxyl group covalently binds to the thiol of cysteine. Some recent studies have shown that this modification plays an important role in signaling transduction, transcriptional regulation and apoptosis. To date, the dynamic of sulfenic acids in proteins remains unclear because of its fleeting nature. Identifying S-sulfenylation sites, therefore, could be the key to decipher its mysterious structures and functions, which are important in cell biology and diseases. However, due to the lack of effective methods, scientists in this field tend to be limited in merely a handful of some wet lab techniques that are time-consuming and not cost-effective. Thus, this motivated us to develop an in silico model for detecting S-sulfenylation sites only from protein sequence information. In this study, protein sequences served as natural language sentences comprising biological subwords. The deep neural network was consequentially employed to perform classification. The performance statistics within the independent dataset including sensitivity, specificity, accuracy, Matthews correlation coefficient and area under the curve rates achieved 85.71%, 69.47%, 77.09%, 0.5554 and 0.833, respectively. Our results suggested that the proposed method (fastSulf-DNN) achieved excellent performance in predicting S-sulfenylation sites compared to other well-known tools on a benchmark dataset.
Duyen Thi Do, Trang Thanh Quynh Le, Nguyen-Quoc-Khanh Le
Briefings Bioinform.3
2021 A transformer architecture based on BERT and 2D convolutional neural network to identify DNA enhancers from sequence information
abstract
Recently, language representation models have drawn a lot of attention in the natural language processing field due to their remarkable results. Among them, bidirectional encoder representations from transformers (BERT) has proven to be a simple, yet powerful language model that achieved novel state-of-the-art performance. BERT adopted the concept of contextualized word embedding to capture the semantics and context of the words in which they appeared. In this study, we present a novel technique by incorporating BERT-based multilingual model in bioinformatics to represent the information of DNA sequences. We treated DNA sequences as natural sentences and then used BERT models to transform them into fixed-length numerical matrices. As a case study, we applied our method to DNA enhancer prediction, which is a well-known and challenging problem in this field. We then observed that our BERT-based features improved more than 5-10% in terms of sensitivity, specificity, accuracy and Matthews correlation coefficient compared to the current state-of-the-art features in bioinformatics. Moreover, advanced experiments show that deep learning (as represented by 2D convolutional neural networks; CNN) holds potential in learning BERT features better than other traditional machine learning techniques. In conclusion, we suggest that BERT and 2D CNNs could open a new avenue in biological modeling using sequence information.
Nguyen-Quoc-Khanh Le, Quang-Thai Ho, Trinh-Trung-Duong Nguyen, Yu-Yen Ou
Briefings Bioinform.1
2021 Prediction of FMN Binding Sites in Electron Transport Chains Based on 2-D CNN and PSSM Profiles
abstract
Flavin mono-nucleotides (FMNs) are cofactors that hold responsibility for carrying and transferring electrons in the electron transport chain stage of cellular respiration. Without being facilitated by FMNs, energy production is stagnant due to the interruption in most of the cellular processes. Investigation on FMN's functions, therefore, can gain holistic understanding about human diseases and molecular information on drug targets. We proposed a deep learning model using a two-dimensional convolutional neural network and position specific scoring matrices that could identify FMN interacting residues with the sensitivity of 83.7 percent, specificity of 99.2 percent, accuracy of 98.2 percent, and Matthews correlation coefficients of 0.85 for an independent dataset containing 141 FMN binding sites and 1,920 non-FMN binding sites. The proposed method outperformed other previous studies using similar evaluation metrics. Our positive outcome can also promote the utilization of deep learning in dealing with various problems in bioinformatics and computational biology.
Nguyen-Quoc-Khanh Le, Binh P. Nguyen
IEEE ACM Trans. Comput. Biol. Bioinform.1
2020 DeepETC: A deep convolutional neural network architecture for investigating and classifying electron transport chain's complexes
Nguyen-Quoc-Khanh Le, Quang-Thai Ho, Edward Kien Yee Yapp, Yu-Yen Ou, Hui-Yuan Yeh
Neurocomputing1
2019 ET-GRU: using multi-layer gated recurrent units to identify electron transport proteins
abstract
BACKGROUND: Electron transport chain is a series of protein complexes embedded in the process of cellular respiration, which is an important process to transfer electrons and other macromolecules throughout the cell. It is also the major process to extract energy via redox reactions in the case of oxidation of sugars. Many studies have determined that the electron transport protein has been implicated in a variety of human diseases, i.e. diabetes, Parkinson, Alzheimer's disease and so on. Few bioinformatics studies have been conducted to identify the electron transport proteins with high accuracy, however, their performance results require a lot of improvements. Here, we present a novel deep neural network architecture to address this problem. RESULTS: Most of the previous studies could not use the original position specific scoring matrix (PSSM) profiles to feed into neural networks, leading to a lack of information and the neural networks consequently could not achieve the best results. In this paper, we present a novel approach by using deep gated recurrent units (GRU) on full PSSMs to resolve this problem. Our approach can precisely predict the electron transporters with the cross-validation and independent test accuracy of 93.5 and 92.3%, respectively. Our approach demonstrates superior performance to all of the state-of-the-art predictors on electron transport proteins. CONCLUSIONS: Through the proposed study, we provide ET-GRU, a web server for discriminating electron transport proteins in particular and other protein functions in general. Also, our achievement could promote the use of GRU in computational biology, especially in protein function prediction.
Nguyen-Quoc-Khanh Le, Edward Kien Yee Yapp, Hui-Yuan Yeh
BMC Bioinform.1
2018 DeepEfflux: a 2D convolutional neural network model for identifying families of efflux proteins in transporters
abstract
Motivation: Efflux protein plays a key role in pumping xenobiotics out of the cells. The prediction of efflux family proteins involved in transport process of compounds is crucial for understanding family structures, functions and energy dependencies. Many methods have been proposed to classify efflux pump transporters without considerations of any pump specific of efflux protein families. In other words, efflux proteins protect cells from extrusion of foreign chemicals. Moreover, almost all efflux protein families have the same structure based on the analysis of significant motifs. The motif sequences consisting of the same amount of residues will have high degrees of residue similarity and thus will affect the classification process. Consequently, it is challenging but vital to recognize the structures and determine energy dependencies of efflux protein families. In order to efficiently identify efflux protein families with considering about pump specific, we developed a 2 D convolutional neural network (2 D CNN) model called DeepEfflux. DeepEfflux tried to capture the motifs of sequences around hidden target residues to use as hidden features of families. In addition, the 2 D CNN model uses a position-specific scoring matrix (PSSM) as an input. Three different datasets, each for one family of efflux protein, was fed into DeepEfflux, and then a 5-fold cross validation approach was used to evaluate the training performance. Results: The model evaluation results show that DeepEfflux outperforms traditional machine learning algorithms. Furthermore, the accuracy of 96.02%, 94.89% and 90.34% for classes A, B and C, respectively, in the independent test results show that our model can perform well and can be used as a reliable tool for identifying families of efflux proteins in transporters. Availability and implementation: The online version of deepefflux is available at http://deepefflux.irit.fr. The source code of deepefflux is available both on the deepefflux website and at http://140.138.155.216/deepefflux/. Supplementary information: Supplementary data are available at Bioinformatics online.
Semmy Wellem Taju, Trinh-Trung-Duong Nguyen, Nguyen-Quoc-Khanh Le, Rosdyana Mangir Irawan Kusuma, Yu-Yen Ou
Bioinform.3
2016 Using Deep Learning with Position Specific Scoring Matrices to Identify Efflux Proteins in Membrane and Transport Proteins
abstract
In several years, deep learning is a new area of machine learning field, which is the motivation of developing machine learning near to artificial intelligent. The neural networks belongs to deep learning are progressively important ideas in a variety of fields with great performance. Accordingly, utilization of deep learning in bioinformatics to enhance performance is very important. Convolutional neural networks is a network of deep learning which is claimed to be the best model to solve the problem of object recognition and detection utilizing GPU computing. In this study, we try to use CNN to identify efflux proteins in membrane and transport proteins, which is a famous problem in bioinformatics field. We construct the CNN from PSSM profiles with CUDA and Keras package based on Theano backend. Finally this approach achieved a significant improvement after we compare with the previous paper on efflux proteins. The proposed method can serve as an effective tool for identifying efflux proteins and can help biologists understand the functions of the efflux proteins. Moreover this study provides a basis for further research that can enrich a field of applying deep learning in bioinformatics.
Semmy Wellem Taju, Nguyen-Quoc-Khanh Le, Yu-Yen Ou
BIBE2
2016 Prediction of FAD binding sites in electron transport proteins according to efficient radial basis function networks and significant amino acid pairs
abstract
BACKGROUND: Cellular respiration is a catabolic pathway for producing adenosine triphosphate (ATP) and is the most efficient process through which cells harvest energy from consumed food. When cells undergo cellular respiration, they require a pathway to keep and transfer electrons (i.e., the electron transport chain). Due to oxidation-reduction reactions, the electron transport chain produces a transmembrane proton electrochemical gradient. In case protons flow back through this membrane, this mechanical energy is converted into chemical energy by ATP synthase. The convert process is involved in producing ATP which provides energy in a lot of cellular processes. In the electron transport chain process, flavin adenine dinucleotide (FAD) is one of the most vital molecules for carrying and transferring electrons. Therefore, predicting FAD binding sites in the electron transport chain is vital for helping biologists understand the electron transport chain process and energy production in cells. RESULTS: We used an independent data set to evaluate the performance of the proposed method, which had an accuracy of 69.84 %. We compared the performance of the proposed method in analyzing two newly discovered electron transport protein sequences with that of the general FAD binding predictor presented by Mishra and Raghava and determined that the accuracy of the proposed method improved by 9-45 % and its Matthew's correlation coefficient was 0.14-0.5. Furthermore, the proposed method enabled reducing the number of false positives significantly and can provide useful information for biologists. CONCLUSIONS: We developed a method that is based on PSSM profiles and SAAPs for identifying FAD binding sites in newly discovered electron transport protein sequences. This approach achieved a significant improvement after we added SAAPs to PSSM features to analyze FAD binding proteins in the electron transport chain. The proposed method can serve as an effective tool for predicting FAD binding sites in electron transport proteins and can help biologists understand the functions of the electron transport chain, particularly those of FAD binding sites. We also developed a web server which identifies FAD binding sites in electron transporters available for academics.
Nguyen-Quoc-Khanh Le, Yu-Yen Ou
BMC Bioinform.1
2016 Incorporating efficient radial basis function networks and significant amino acid pairs for predicting GTP binding sites in transport proteins
abstract
BACKGROUND: Guanonine-protein (G-protein) is known as molecular switches inside cells, and is very important in signals transmission from outside to inside cell. Especially in transport protein, most of G-proteins play an important role in membrane trafficking; necessary for transferring proteins and other molecules to a variety of destinations outside and inside of the cell. The function of membrane trafficking is controlled by G-proteins via Guanosine triphosphate (GTP) binding sites. The GTP binding sites active G-proteins initiated to membrane vesicles by interacting with specific effector proteins. Without the interaction from GTP binding sites, G-proteins could not be active in membrane trafficking and consequently cause many diseases, i.e., cancer, Parkinson… Thus it is very important to identify GTP binding sites in membrane trafficking, in particular, and in transport protein, in general. RESULTS: We developed the proposed model with a cross-validation and examined with an independent dataset. We achieved an accuracy of 95.6% for evaluating with cross-validation and 98.7% for examining the performance with the independent data set. For newly discovered transport protein sequences, our approach performed remarkably better than similar methods such as GTPBinder, NsitePred and TargetSOS. Moreover, a friendly web server was developed for identifying GTP binding sites in transport proteins available for all users. CONCLUSIONS: We approached a computational technique using PSSM profiles and SAAPs for identifying GTP binding residues in transport proteins. When we included SAAPs into PSSM profiles, the predictive performance achieved a significant improvement in all measurement metrics. Furthermore, the proposed method could be a power tool for determining new proteins that belongs into GTP binding sites in transport proteins and can provide useful information for biologists.
Nguyen-Quoc-Khanh Le, Yu-Yen Ou
BMC Bioinform.1