Watshara Shoombuatong

dblp:48/10029 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
11since 2021 · last 2025
0000-0002-3394-8709ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 AttBiomarker: unveiling preeclampsia biomarkers and molecular pathways through two-stage gene selection techniques and attention-based CNN with gene regulatory network analysis
abstract
Preeclampsia is a complex pregnancy disorder that poses significant health risks to both mother and fetus. Despite its clinical importance, the underlying molecular mechanisms remain poorly understood. In this study, we developed an integrative deep learning and bioinformatics approach to identify potential biomarkers for preeclampsia. Three microarray datasets related to preeclampsia were initially analyzed to select a preliminary gene subset based on $P$-values. Feature selection was then performed in two consecutive rounds: first, the Fisher score method was applied to extract significant genes, followed by the minimum Redundancy Maximum Relevance method to refine the subset further. These selected gene subsets were trained using our proposed Attention-based Convolutional Neural Network (AttCNN), which achieved the highest classification accuracy compared with other models. From the experiments, a set of 58 common genes was identified between differentially expressed genes and the final optimized subset. Here, Gene Ontology and KEGG pathway enrichment analyses highlighted key biological processes and pathways associated with preeclampsia. Subsequently, a protein-protein interaction network was constructed, identifying 10 hub genes: TSC22D1, IRF3, MME, SRSF10, SOD1, HK2, ERO1L, SH3BP5, UBC, and ZFAND5. Further analysis of gene regulatory networks, including transcription factor-gene, gene-microRNA, and drug-gene interactions, revealed that seven hub genes (HK2, SRSF10, SOD1, ERO1L, IRF3, MME, and SH3BP5) were strongly associated with preeclampsia. Molecular docking analysis showed that HK2, SH3BP5, and SOD1 exhibited significant binding affinities with two preeclampsia drugs. These findings suggest that the identified hub genes hold promise as biomarkers for early prognosis, diagnosis, and potential therapeutic targets for preeclampsia.
Sakib Sarker, S. M. Hasan Mahmud, Md. Faruk Hosen, Michael Kah Ong Goh, Watshara Shoombuatong
Briefings Bioinform.5
2025 M3S-GRPred: a novel ensemble learning approach for the interpretable prediction of glucocorticoid receptor antagonists using a multi-step stacking strategy
abstract
Accelerating drug discovery for glucocorticoid receptor (GR)-related disorders, including innovative machine learning (ML)-based approaches, holds promise in advancing therapeutic development, optimizing treatment efficacy, and mitigating adverse effects. While experimental methods can accurately identify GR antagonists, they are often not cost-effective for large-scale drug discovery. Thus, computational approaches leveraging SMILES information for precise in silico identification of GR antagonists are crucial, enabling efficient and scalable drug discovery. Here, we develop a new ensemble learning approach using a multi-step stacking strategy (M3S), termed M3S-GRPred, aimed at rapidly and accurately discovering novel GR antagonists. To the best of our knowledge, M3S-GRPred is the first SMILES-based predictor designed to identify GR antagonists without the use of 3D structural information. In M3S-GRPred, we first constructed different balanced subsets using an under-sampling approach. Using these balanced subsets, we explored and evaluated heterogeneous base-classifiers trained with a variety of SMILES-based feature descriptors coupled with popular ML algorithms. Finally, M3S-GRPred was constructed by integrating probabilistic feature from the selected base-classifiers derived from a two-step feature selection technique. Our comparative experiments demonstrate that M3S-GRPred can precisely identify GR antagonists and effectively address the imbalanced dataset. Compared to traditional ML classifiers, M3S-GRPred attained superior performance in terms of both the training and independent test datasets. Additionally, M3S-GRPred was applied to identify potential GR antagonists among FDA-approved drugs confirmed through molecular docking, followed by detailed MD simulation studies for drug repurposing in Cushing's syndrome. We anticipate that M3S-GRPred will serve as an efficient screening tool for discovering novel GR antagonists from vast libraries of unknown compounds in a cost-effective manner.
Nalini Schaduangrat, Hathaichanok Chuntakaruk, Thanyada Rungrotmongkol, Pakpoom Mookdarsanit, Watshara Shoombuatong
BMC Bioinform.5
2025 M3S-ALG: Improved and robust prediction of allergenicity of chemical compounds by using a novel multi-step stacking strategy
Phasit Charoenkwan, Nalini Schaduangrat, Le Thi Phan, Balachandran Manavalan, Watshara Shoombuatong
Future Gener. Comput. Syst.5
2025 DeepHDAC3i: Leveraging an Interpretable Deep Learning-Based Framework for the Accelerated Discovery of HDAC3 Inhibitors
abstract
Epigenetics encompasses dynamic and reversible modifications that regulate gene activity without altering the underlying DNA sequence. Epigenetic processes, including non-coding RNA interactions, and DNA methylation regulate patterns of gene expression by responding to cellular signaling, environmental stimuli, and developmental cues. The balance of histone acetylation is maintained by histone deacetylase (HDAC) and histone acetyltransferase (HAT) activities. Aberrant HDAC upregulation, often seen in cancer cells, disrupts this balance. HDAC inhibitors (HDACi) are thus used in cancer treatment. However, most synthetic HDACis are not specific to HDAC classes or individual members, highlighting the need for highly selective HDAC inhibitors. Machine learning (ML)-driven methods are now recognized as rapid and cost-efficient tools in drug discovery and development, capable of identifying inhibitors solely from SMILES notation, without requiring the 3D ligand structure. Here, we present a novel and interpretable deep learning-based framework, DeepHDAC3i, for accurate in silico identification of HDAC3i using only the SMILES notation. Firstly, we employed five molecular encoding methods, namely CDKExt, KR, KRC, Pubchem, and RDKIT, to extract the biological and structural information in HDAC3i. These molecular representations were then fused to generate multi-view features. Secondly, elastic net was employed to determine the optimal feature subset and enhance prediction performance. Thirdly, a one-dimensional convolutional neural network (1D-CNN) coupled with the optimal feature set was chosen for the construction of the final model. Finally, our framework leveraged the Shapley Additive exPlanation algorithm to disclose the most important features for identifying HDAC3i. On the independent test dataset, DeepHDAC3i achieved an accuracy of 0.965, MCC of 0.930, and AUC of 0.985, which were significantly higher than several conventional machine learning and deep learning models. In addition, upon comparison with the existing methods, DeepHDAC3i secured the best performance with improvements of approximately 4.80, 4.70, 6.50, and 9.50% in accuracy, F1, AUC, and MCC, respectively. Taken together, DeepHDAC3i is superior to other compared models and can be a useful tool for precisely identifying HDAC3i.
Nalini Schaduangrat, Ittipat Meewan, Watshara Shoombuatong
IEEE Trans. Comput. Biol. Bioinform.4
2025 iMRSA-Fuse: A Fast and Accurate Computational Approach for Predicting Anti-MRSA Peptides by Fusing Multi-View Information
abstract
Methicillin-resistant S. aureus (MRSA) has prominently emerged among the recognized causes of community-acquired and hospital infections. We proposed a novel computational approach, iMRSA-Fuse, based on a multi-view feature fusion strategy for fast and accurate anti-MRSA peptide identification. In iMRSA-Fuse, we explored and integrated 12 different sequence-based feature descriptors from multiple perspectives, in conjunction with 12 popular machine learning (ML) algorithms, to construct multi-view features that were able to fully capture the useful information of anti-MRSA peptides. Additionally, we applied our customized genetic algorithm to determine a set of multi-view features to enhance its discriminative ability. Based on a series of comparative results, our multi-view features exhibited the most discriminative ability compared to several conventional feature descriptors. Moreover, concerning the independent test dataset, iMRSA-Fuse achieved the best balanced accuracy (BACC) and Matthew's correlation coefficient (MCC) of 0.997 and 0.981, respectively with an increase of 3.93 and 7.78%, respectively. Finally, to facilitate the large-scale identification of candidate anti-MRSA peptides, a user-friendly web server of the iMRSA-Fuse model is constructed and is freely accessible at https://pmlabqsar.pythonanywhere.com/iMRSA-Fuse. We anticipate that this new computational approach will be effectively applied to screen and prioritize candidate peptides that might exhibit the great anti-MRSA activities.
Phasit Charoenkwan, Nalini Schaduangrat, Mohammad Ali Moni, Watshara Shoombuatong
IEEE Trans. Comput. Biol. Bioinform.4
2024 The role of ncRNA regulatory mechanisms in diseases - case on gestational diabetes
abstract
Non-coding RNAs (ncRNAs) are a class of RNA molecules that do not have the potential to encode proteins. Meanwhile, they can occupy a significant portion of the human genome and participate in gene expression regulation through various mechanisms. Gestational diabetes mellitus (GDM) is a pathologic condition of carbohydrate intolerance that begins or is first detected during pregnancy, making it one of the most common pregnancy complications. Although the exact pathogenesis of GDM remains unclear, several recent studies have shown that ncRNAs play a crucial regulatory role in GDM. Herein, we present a comprehensive review on the multiple mechanisms of ncRNAs in GDM along with their potential role as biomarkers. In addition, we investigate the contribution of deep learning-based models in discovering disease-specific ncRNA biomarkers and elucidate the underlying mechanisms of ncRNA. This might assist community-wide efforts to obtain insights into the regulatory mechanisms of ncRNAs in disease and guide a novel approach for early diagnosis and treatment of disease.
Liping Ren, Yu-Duo Hao, Nalini Schaduangrat, Xiao-Wei Liu, Shi-Shi Yuan, Watshara Shoombuatong
Briefings Bioinform.9
2023 TIPred: a novel stacked ensemble approach for the accelerated discovery of tyrosinase inhibitory peptides
abstract
BACKGROUND: Tyrosinase is an enzyme involved in melanin production in the skin. Several hyperpigmentation disorders involve the overproduction of melanin and instability of tyrosinase activity resulting in darker, discolored patches on the skin. Therefore, discovering tyrosinase inhibitory peptides (TIPs) is of great significance for basic research and clinical treatments. However, the identification of TIPs using experimental methods is generally cost-ineffective and time-consuming. RESULTS: Herein, a stacked ensemble learning approach, called TIPred, is proposed for the accurate and quick identification of TIPs by using sequence information. TIPred explored a comprehensive set of various baseline models derived from well-known machine learning (ML) algorithms and heterogeneous feature encoding schemes from multiple perspectives, such as chemical structure properties, physicochemical properties, and composition information. Subsequently, 130 baseline models were trained and optimized to create new probabilistic features. Finally, the feature selection approach was utilized to determine the optimal feature vector for developing TIPred. Both tenfold cross-validation and independent test methods were employed to assess the predictive capability of TIPred by using the stacking strategy. Experimental results showed that TIPred significantly outperformed the state-of-the-art method in terms of the independent test, with an accuracy of 0.923, MCC of 0.757 and an AUC of 0.977. CONCLUSIONS: The proposed TIPred approach could be a valuable tool for rapidly discovering novel TIPs and effectively identifying potential TIP candidates for follow-up experimental validation. Moreover, an online webserver of TIPred is publicly available at http://pmlabstack.pythonanywhere.com/TIPred .
Phasit Charoenkwan, Sasikarn Kongsompong, Nalini Schaduangrat, Pramote Chumnanpuen, Watshara Shoombuatong
BMC Bioinform.5
2023 StackTTCA: a stacking ensemble learning-based framework for accurate and high-throughput identification of tumor T cell antigens
abstract
BACKGROUND: The identification of tumor T cell antigens (TTCAs) is crucial for providing insights into their functional mechanisms and utilizing their potential in anticancer vaccines development. In this context, TTCAs are highly promising. Meanwhile, experimental technologies for discovering and characterizing new TTCAs are expensive and time-consuming. Although many machine learning (ML)-based models have been proposed for identifying new TTCAs, there is still a need to develop a robust model that can achieve higher rates of accuracy and precision. RESULTS: In this study, we propose a new stacking ensemble learning-based framework, termed StackTTCA, for accurate and large-scale identification of TTCAs. Firstly, we constructed 156 different baseline models by using 12 different feature encoding schemes and 13 popular ML algorithms. Secondly, these baseline models were trained and employed to create a new probabilistic feature vector. Finally, the optimal probabilistic feature vector was determined based the feature selection strategy and then used for the construction of our stacked model. Comparative benchmarking experiments indicated that StackTTCA clearly outperformed several ML classifiers and the existing methods in terms of the independent test, with an accuracy of 0.932 and Matthew's correlation coefficient of 0.866. CONCLUSIONS: In summary, the proposed stacking ensemble learning-based framework of StackTTCA could help to precisely and rapidly identify true TTCAs for follow-up experimental verification. In addition, we developed an online web server ( http://2pmlab.camt.cmu.ac.th/StackTTCA ) to maximize user convenience for high-throughput screening of novel TTCAs.
Phasit Charoenkwan, Nalini Schaduangrat, Watshara Shoombuatong
BMC Bioinform.3
2021 NeuroPred-FRL: an interpretable prediction model for identifying neuropeptide using feature representation learning
abstract
Neuropeptides (NPs) are the most versatile neurotransmitters in the immune systems that regulate various central anxious hormones. An efficient and effective bioinformatics tool for rapid and accurate large-scale identification of NPs is critical in immunoinformatics, which is indispensable for basic research and drug development. Although a few NP prediction tools have been developed, it is mandatory to improve their NPs' prediction performances. In this study, we have developed a machine learning-based meta-predictor called NeuroPred-FRL by employing the feature representation learning approach. First, we generated 66 optimal baseline models by employing 11 different encodings, six different classifiers and a two-step feature selection approach. The predicted probability scores of NPs based on the 66 baseline models were combined to be deemed as the input feature vector. Second, in order to enhance the feature representation ability, we applied the two-step feature selection approach to optimize the 66-D probability feature vector and then inputted the optimal one into a random forest classifier for the final meta-model (NeuroPred-FRL) construction. Benchmarking experiments based on both cross-validation and independent tests indicate that the NeuroPred-FRL achieves a superior prediction performance of NPs compared with the other state-of-the-art predictors. We believe that the proposed NeuroPred-FRL can serve as a powerful tool for large-scale identification of NPs, facilitating the characterization of their functional mechanisms and expediting their applications in clinical therapy. Moreover, we interpreted some model mechanisms of NeuroPred-FRL by leveraging the robust SHapley Additive exPlanation algorithm.
Md. Mehedi Hasan 0002, Md. Ashad Alam, Watshara Shoombuatong, Hong-Wen Deng, Balachandran Manavalan, Hiroyuki Kurata
Briefings Bioinform.3
2021 StackIL6: a stacking ensemble model for improving the prediction of IL-6 inducing peptides
abstract
The release of interleukin (IL)-6 is stimulated by antigenic peptides from pathogens as well as by immune cells for activating aggressive inflammation. IL-6 inducing peptides are derived from pathogens and can be used as diagnostic biomarkers for predicting various stages of disease severity as well as being used as IL-6 inhibitors for the suppression of aggressive multi-signaling immune responses. Thus, the accurate identification of IL-6 inducing peptides is of great importance for investigating their mechanism of action as well as for developing diagnostic and immunotherapeutic applications. This study proposes a novel stacking ensemble model (termed StackIL6) for accurately identifying IL-6 inducing peptides. More specifically, StackIL6 was constructed from twelve different feature descriptors derived from three major groups of features (composition-based features, composition-transition-distribution-based features and physicochemical properties-based features) and five popular machine learning algorithms (extremely randomized trees, logistic regression, multi-layer perceptron, support vector machine and random forest). To enhance the utility of baseline models, they were effectively and systematically integrated through a stacking strategy to build the final meta-based model. Extensive benchmarking experiments demonstrated that StackIL6 could achieve significantly better performance than the existing method (IL6PRED) and outperformed its constituent baseline models on both training and independent test datasets, which thereby support its excellent discrimination and generalization abilities. To facilitate easy access to the StackIL6 model, it was established as a freely available web server accessible at http://camt.pythonanywhere.com/StackIL6. It is anticipated that StackIL6 can help to facilitate rapid screening of promising IL-6 inducing peptides for the development of diagnostic and immunotherapeutic applications in the future.
Phasit Charoenkwan, Wararat Chiangjong, Chanin Nantasenamat, Md. Mehedi Hasan 0002, Balachandran Manavalan, Watshara Shoombuatong
Briefings Bioinform.6
2021 BERT4Bitter: a bidirectional encoder representations from transformers (BERT)-based model for improving the prediction of bitter peptides
abstract
MOTIVATION: The identification of bitter peptides through experimental approaches is an expensive and time-consuming endeavor. Due to the huge number of newly available peptide sequences in the post-genomic era, the development of automated computational models for the identification of novel bitter peptides is highly desirable. RESULTS: In this work, we present BERT4Bitter, a bidirectional encoder representation from transformers (BERT)-based model for predicting bitter peptides directly from their amino acid sequence without using any structural information. To the best of our knowledge, this is the first time a BERT-based model has been employed to identify bitter peptides. Compared to widely used machine learning models, BERT4Bitter achieved the best performance with an accuracy of 0.861 and 0.922 for cross-validation and independent tests, respectively. Furthermore, extensive empirical benchmarking experiments on the independent dataset demonstrated that BERT4Bitter clearly outperformed the existing method with improvements of 8.0% accuracy and 16.0% Matthews coefficient correlation, highlighting the effectiveness and robustness of BERT4Bitter. We believe that the BERT4Bitter method proposed herein will be a useful tool for rapidly screening and identifying novel bitter peptides for drug development and nutritional research. AVAILABILITYAND IMPLEMENTATION: The user-friendly web server of the proposed BERT4Bitter is freely accessible at http://pmlab.pythonanywhere.com/BERT4Bitter. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Phasit Charoenkwan, Chanin Nantasenamat, Md. Mehedi Hasan 0002, Balachandran Manavalan, Watshara Shoombuatong
Bioinform.5
2020 HLPpred-Fuse: improved and robust prediction of hemolytic peptide and its activity by fusing multiple feature representation
abstract
MOTIVATION: Therapeutic peptides failing at clinical trials could be attributed to their toxicity profiles like hemolytic activity, which hamper further progress of peptides as drug candidates. The accurate prediction of hemolytic peptides (HLPs) and its activity from the given peptides is one of the challenging tasks in immunoinformatics, which is essential for drug development and basic research. Although there are a few computational methods that have been proposed for this aspect, none of them are able to identify HLPs and their activities simultaneously. RESULTS: In this study, we proposed a two-layer prediction framework, called HLPpred-Fuse, that can accurately and automatically predict both hemolytic peptides (HLPs or non-HLPs) as well as HLPs activity (high and low). More specifically, feature representation learning scheme was utilized to generate 54 probabilistic features by integrating six different machine learning classifiers and nine different sequence-based encodings. Consequently, the 54 probabilistic features were fused to provide sufficiently converged sequence information which was used as an input to extremely randomized tree for the development of two final prediction models which independently identify HLP and its activity. Performance comparisons over empirical cross-validation analysis, independent test and case study against state-of-the-art methods demonstrate that HLPpred-Fuse consistently outperformed these methods in the identification of hemolytic activity. AVAILABILITY AND IMPLEMENTATION: For the convenience of experimental scientists, a web-based tool has been established at http://thegleelab.org/HLPpred-Fuse. CONTACT: [email protected] or [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Md. Mehedi Hasan 0002, Nalini Schaduangrat, Shaherin Basith, Gwang Lee, Watshara Shoombuatong, Balachandran Manavalan
Bioinform.5
2017 AnkPlex: algorithmic structure for refinement of near-native ankyrin-protein docking
abstract
BACKGROUND: Computational analysis of protein-protein interaction provided the crucial information to increase the binding affinity without a change in basic conformation. Several docking programs were used to predict the near-native poses of the protein-protein complex in 10 top-rankings. The universal criteria for discriminating the near-native pose are not available since there are several classes of recognition protein. Currently, the explicit criteria for identifying the near-native pose of ankyrin-protein complexes (APKs) have not been reported yet. RESULTS: In this study, we established an ensemble computational model for discriminating the near-native docking pose of APKs named "AnkPlex". A dataset of APKs was generated from seven X-ray APKs, which consisted of 3 internal domains, using the reliable docking tool ZDOCK. The dataset was composed of 669 and 44,334 near-native and non-near-native poses, respectively, and it was used to generate eleven informative features. Subsequently, a re-scoring rank was generated by AnkPlex using a combination of a decision tree algorithm and logistic regression. AnkPlex achieved superior efficiency with ≥1 near-native complexes in the 10 top-rankings for nine X-ray complexes compared to ZDOCK, which only obtained six X-ray complexes. In addition, feature analysis demonstrated that the van der Waals feature was the dominant near-native pose out of the potential ankyrin-protein docking poses. CONCLUSION: The AnkPlex model achieved a success at predicting near-native docking poses and led to the discovery of informative characteristics that could further improve our understanding of the ankyrin-protein complex. Our computational study could be useful for predicting the near-native poses of binding proteins and desired targets, especially for ankyrin-protein complexes. The AnkPlex web server is freely accessible at http://ankplex.ams.cmu.ac.th .
Tanchanok Wisitponchai, Watshara Shoombuatong, Vannajan Sanghiran Lee, Kuntida Kitidee, Chatchai Tayapiwatana
BMC Bioinform.2
2013 Predicting protein crystallization using a simple scoring card method
abstract
Many computational methods have been developed to predict protein crystallization. Most methods use amino acid and dipeptide compositions as part of the informative features. To advance the prediction accuracy, the support vector machine (SVM) based classifiers and ensemble approaches were effective and commonly-used techniques. However, these techniques suffer from the low interpretation ability of insight into crystallization. In this study, we utilize a newly-developed scoring card method (SCM) with a dipeptide composition feature to predict protein crystallization. This SCM classifier obtains prediction results 74%, 0.55 and 0.83 for accuracy, sensitivity and specificity, respectively, which is comparable to the SVM classifier using the same benchmarks. The experimental results show that the SCM classifier has advantages of simplicity, high interpretability, and high accuracy in predicting protein crystallization, compared with existing SVM-basedensemble classifiers.
Watshara Shoombuatong, Hui-Ling Huang, Jeerayut Chaijaruwanich, Phasit Charoenkwan, Hua-Chin Lee, Shinn-Ying Ho
CIBCB1