VLDB 2026 Research / reviewers in the wild / expert
Chia-Ru Chung
dblp:266/5015
· DBLP profile ↗
16ranked-venue papers
4as first author
14since 2021 · last 2025
0000-0002-4548-7620ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Self-Attention Enhanced Deep Learning Models for Immune Cell Deconvolution from Bulk RNA-SeqabstractAccurate immune cell composition profiling is crucial for understanding immunological dynamics and disease mechanisms. Bulk RNA sequencing (bulk RNA-seq) is widely employed due to its cost-effectiveness and scalability; however, it lacks the resolution to identify cell-specific gene expression. To address this limitation, we propose a self-attention enhanced deep learning model designed for precise immune cell deconvolution from bulk RNA-seq data. We systematically annotated immune cell types from four single-cell RNA-seq (scRNA-seq) peripheral blood mononuclear cell (PBMC) datasets and validated these annotations against established automated identification tools (SingleR, Seurat, scPred, ScType). Leveraging these annotations, we generated realistic pseudo-bulk RNA-seq training samples using Dirichlet-distribution-based composition sampling, significantly enhancing the model’s performance, particularly for rare cell populations. Comparative evaluations demonstrated that our self-attention enhanced deep learning model consistently outperformed existing approaches, including CIBERSORTx and Scaden, achieving lower prediction errors and higher correlations on benchmark PBMC datasets. Integrating multi-head self-attention allowed the model to dynamically capture intricate dependencies among gene expression features, substantially improving deconvolution accuracy for specific cell subsets. While demonstrating robust performance on PBMC datasets, we acknowledge that broader validation is essential due to potential limitations in generalizability across different tissue types and conditions. Our study highlights the potential of self-attention mechanisms and realistic training data generation strategies to enhance computational deconvolution techniques, providing valuable tools for clinical diagnostics and translational immunology research. Chia-Ru Chung, Yen-Lin Chen, Justin Bo-Kai Hsu, Li-Ching Wu, Tzong-Yi Lee, Jorng-Tzong Horng |
CIBCB | 1 |
| 2025 | Explainable AI-Enhanced Kinase Activity Profiling Through PhosphoproteomicsabstractKinases play a critical role in regulating fundamental cellular processes, including metabolism, signal transduction, and cell growth, primarily through phosphorylation. The dysregulation of kinase activity is implicated in various diseases, highlighting the urgent need for robust and interpretable methodologies to profile this activity. Current approaches frequently depend on overly complex or limited datasets, lack generalizability, or fail to provide meaningful biological insights into the mechanisms governing kinase activity. To address these challenges, we developed an explainable deep learning framework that leverages mass spectrometry-based phosphoproteomics data to profile kinase activity effectively. Our study systematically evaluated deep neural networks (DNNs) and convolutional neural networks (CNNs), incorporating a diverse set of feature inputs, including phosphorylation sites and kinase-substrate relationships. A notable finding was that a three-layer CNN, optimized through rigorous feature selection techniques, demonstrated superior performance, achieving substantial improvements in prediction accuracy and stability when compared to established methods such as kinase-substrate enrichment analysis (KSEA) and the kinase activity ranking pipeline (KARP). We integrated Shapley additive explanations (SHAP) values to enhance interpretability, illuminating biologically significant phosphorylation sites. For example, PAK2-related phosphorylation sites associated with the progression of colon adenocarcinoma and CAMK2D sites integral to adrenergic signaling were identified, thereby effectively linking computational predictions to established molecular pathways. This research illustrates the potential of explainable artificial intelligence in advancing kinase activity profiling by providing accurate and interpretable predictions. Our framework is valuable for elucidating disease mechanisms and identifying therapeutic targets, facilitating broader applications in precision medicine. Chia-Ru Chung, Ming-Feng Ho, Li-Ching Wu, Justin Bo-Kai Hsu, Tzong-Yi Lee, Jorng-Tzong Horng |
CIBCB | 1 |
| 2025 | Integrative Framework for Functional analysis of multiple Traditional Chinese Medicines based on transcriptome-driven systems biology and supervised learning strategyabstractTraditional Chinese Medicines (TCMs) contain a wide variety of ingredients and are rich in bioactive chemical sources. However, their complex and unknown effects on the human body prove a challenge in TCM research. Multiple genes are involved in a biological system and disease progression; therefore, transcriptome-driven systems biology approach directs our attention to perturbed pathways or enriched gene sets from whole gene expression profile, allowing more explanatory power and less analytical complexity than solely considering differentially expressed genes (DEGs). This framework integrated differential expression analysis (DEA), pathway-based expression analysis and supervised learning approaches to explore TCM functions and extract signature genes of high importance. Supervised learning approaches were employed to extract signature genes that effectively discriminate between four TCMs (92% accuracy). These signature genes were further examined through database annotations and literature survey, where some were found to be TCM compound targets, such as fatty acid synthase (FASN) and thioredoxin reductase 1 (TXNRD1). Based on network pharmacology concept, the proposed TCM-compound-target-pathway network can link TCM compounds to signature genes, DEGs and perturbed pathways, revealing the rationale behind TCM’s mechanism of action in treatment. Integrating a supervised learning strategy in identifying TCM-associated genes can offer a new perspective for discovering potential drug targets. Chia-Ru Chung, Hsi-Yuan Huang, Yang-Chi-Dung Lin, Hua-Li Zuo, Siyao Hu, Hsiao-Chin Hong, Jinrui Bai, Hsien-Da Huang |
CIBCB | 2 |
| 2025 | AI-Enhanced MALDI-TOF MS Analysis for Important Peaks on Predicting Ciprofloxacin Resistance across Different Gram-Negative BacteriaabstractRapid identification of antibiotic-resistant infections is crucial, as antimicrobial resistance is a global health crisis. Yet, conventional antibiotic susceptibility tests (AST) often require days to yield results. Matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS) has emerged as a rapid, cost-effective tool for bacterial identification and shows promise for resistance profiling by detecting spectral biomarkers. In this study, we harness MALDI-TOF MS with machine learning and deep learning to predict ciprofloxacin resistance across four Gram-negative bacteria, Escherichia coli, Klebsiella pneumoniae, Acinetobacter baumannii, and Acinetobacter nosocomialis, using a cross-species "basket-wise" approach. We extracted features from mass spectra using kernel density estimation-based peak detection and m/z binning, then trained a random forest (RF) classifier and a convolutional neural network (CNN) to distinguish ciprofloxacin-resistant and susceptible isolates. To interpret the models, we employed dual feature importance analyses: gradient-weighted class activation mapping (Grad-CAM) for the CNN to highlight critical m/z regions and an ensemble RF-based method to identify significant peak features. The CNN achieved higher overall accuracy than the RF, especially in three of four species, while the ensemble RF approach identified interpretable sets of around 20 important m/z peaks per organism. Several informative peaks overlapped between species, indicating some common resistance-associated spectral signatures. However, no single universal marker was found across all species. These findings demonstrate an AI-enhanced MALDI-TOF MS framework for rapid AMR detection, yielding accurate predictions and interpretable spectral markers. The approach highlights clinical potential to guide effective therapy and bolster antimicrobial stewardship, particularly for underrepresented pathogens such as A. nosocomialis. Hsin-Yao Wang, Chia-Ru Chung, Wen-Rui Zhang, Li-Ching Wu, Justin Bo-Kai Hsu, Jang-Jih Lu, Jorng-Tzong Horng |
CIBCB | 2 |
| 2025 | Mobile Virtual Assistant for Multi-Modal Depression-Level StratificationabstractDepression not only afflicts hundreds of millions of people but also contributes to a global disability and healthcare burden. The primary method of diagnosing depression relies on the judgment of medical professionals in clinical interviews with patients, which is subjective and time-consuming. Recent studies have demonstrated that text, audio, facial attributes, heart rate, and eye movement could be utilized for depression-level stratification. In this paper, we construct a virtual assistant for automatic depression-level stratification on mobile devices that can actively guide users through voice dialogue and change conversation content using emotion perception. During the conversation, features from text, audio, facial attributes, heart rate, and eye movement are extracted for multi-modal depression-level stratification. We utilize a feature-level fusion framework to integrate five modalities and the deep neural network to classify the varying levels of depression, which include healthy, mild, moderate, or severe depression, as well as bipolar disorder (formerly called manic depression). With outcome data from 168 subjects, experimental results reveal that the total accuracy of feature-level fusion with five modal features achieves the highest accuracy of 90.26 percent. Eric Hsiao-Kuang Wu, Ting-Yu Gao, Chia-Ru Chung, Chun-Chuan Chen, Chia-Fen Tsai, Shih-Ching Yeh |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | A two-stage computational framework for identifying antiviral peptides and their functional types based on contrastive learning and multi-feature fusion strategyabstractAntiviral peptides (AVPs) have shown potential in inhibiting viral attachment, preventing viral fusion with host cells and disrupting viral replication due to their unique action mechanisms. They have now become a broad-spectrum, promising antiviral therapy. However, identifying effective AVPs is traditionally slow and costly. This study proposed a new two-stage computational framework for AVP identification. The first stage identifies AVPs from a wide range of peptides, and the second stage recognizes AVPs targeting specific families or viruses. This method integrates contrastive learning and multi-feature fusion strategy, focusing on sequence information and peptide characteristics, significantly enhancing predictive ability and interpretability. The evaluation results of the model show excellent performance, with accuracy of 0.9240 and Matthews correlation coefficient (MCC) score of 0.8482 on the non-AVP independent dataset, and accuracy of 0.9934 and MCC score of 0.9869 on the non-AMP independent dataset. Furthermore, our model can predict antiviral activities of AVPs against six key viral families (Coronaviridae, Retroviridae, Herpesviridae, Paramyxoviridae, Orthomyxoviridae, Flaviviridae) and eight viruses (FIV, HCV, HIV, HPIV3, HSV1, INFVA, RSV, SARS-CoV). Finally, to facilitate user accessibility, we built a user-friendly web interface deployed at https://awi.cuhk.edu.cn/∼dbAMP/AVP/. Jiahui Guan, Lantian Yao, Peilin Xie, Chia-Ru Chung, Yixian Huang, Ying-Chih Chiang, Tzong-Yi Lee |
Briefings Bioinform. | 4 |
| 2024 | ACP-CapsPred: an explainable computational framework for identification and functional prediction of anticancer peptides based on capsule networkabstractCancer is a severe illness that significantly threatens human life and health. Anticancer peptides (ACPs) represent a promising therapeutic strategy for combating cancer. In silico methods enable rapid and accurate identification of ACPs without extensive human and material resources. This study proposes a two-stage computational framework called ACP-CapsPred, which can accurately identify ACPs and characterize their functional activities across different cancer types. ACP-CapsPred integrates a protein language model with evolutionary information and physicochemical properties of peptides, constructing a comprehensive profile of peptides. ACP-CapsPred employs a next-generation neural network, specifically capsule networks, to construct predictive models. Experimental results demonstrate that ACP-CapsPred exhibits satisfactory predictive capabilities in both stages, reaching state-of-the-art performance. In the first stage, ACP-CapsPred achieves accuracies of 80.25% and 95.71%, as well as F1-scores of 79.86% and 95.90%, on benchmark datasets Set 1 and Set 2, respectively. In the second stage, tasked with characterizing the functional activities of ACPs across five selected cancer types, ACP-CapsPred attains an average accuracy of 90.75% and an F1-score of 91.38%. Furthermore, ACP-CapsPred demonstrates excellent interpretability, revealing regions and residues associated with anticancer activity. Consequently, ACP-CapsPred presents a promising solution to expedite the development of ACPs and offers a novel perspective for other biological sequence analyses. Lantian Yao, Peilin Xie, Jiahui Guan, Chia-Ru Chung, Wenyang Zhang, Junyang Deng, Yixian Huang, Ying-Chih Chiang, Tzong-Yi Lee |
Briefings Bioinform. | 4 |
| 2024 | Qigong Master: A Qigong-Based Attention Training Game Using Action Recognition and Balance AnalysisabstractBaduanjin is a type of martial art that is aimed at the development and health of the physical, emotional and spiritual aspects. It emphasizes on gentle movements, relaxed yet disciplined, and training of mental concentration through relaxation. In order to make Baduanjin attention training more effective and easier, we proposed a novel Baduanjin-based attention training game that instructs the subject to practice Baduanjin using virtual reality (VR) and motion analysis. Through a virtual instructor who demonstrates a series of Baduanjin actions, the subject is asked to follow the instructor's movement at any time. Meanwhile, the 3D position of the subject's body joints is collected by a motion capture device for further analysis using a deep learning model to evaluate the order correctness and precision of the Baduanjin actions. In addition, transfer learning technique is used to solve the problem of small size of Baduanjin data. Preliminary tests with 20 normal individuals showed that the recognition accuracy of the Baduanjin actions reached nearly 97%, and the balance analysis reflected the position change of the body center of mass to a certain extent. We also compared the performance of the model with and without pre-training to demonstrate the importance of transfer learning. The result of these explorations shows the feasibility of this prototype system as an attention training system and its potential as an assistive treatment option. In conclusion, we not only created the virtual instructor that could provide accurate movement demonstration, but also collected objective action data for more accurate balance analysis to provide appropriate feedback. Chia-Ru Chung, Shih-Ching Yeh, Eric Hsiao-Kuang Wu, Sheng-Yang Lin |
IEEE Trans. Games | 1 |
| 2024 | Anti-Drugs Chatbot: Chinese BERT-Based Cognitive Intent AnalysisabstractDrug abuse has always been a severe issue, but the proportion of drug abuse and addiction is rising. According to research reports, youth are motivated to access drugs mainly due to curiosity and peer influence. Additionally, youth especially lack proper knowledge and education surrounding drug abuse. Analyzing whether potential addicts intend to access drugs is helpful in preventing drug abuse and addiction. We developed an Anti-drug Chatbot for young people on a popular online social platform. We can detect potential risks, obtain warnings from the user-entered query and provide these to professional consultants for help. In this article, we present a hierarchical system with bidirectional encoder representation from transformers (BERT) to efficiently recognize and classify a user’s intent. We use the Chinese BERT-based model to utilize contextual information to perform classification and recognition. We evaluate our proposed system on our conversational dataset. Jui-Hsuan Lee, Eric Hsiao-Kuang Wu, Yu-Yen Ou, Yueh-Che Lee, Cheng-Hsun Lee, Chia-Ru Chung |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2023 | Identification of species-specific RNA N6-methyladinosine modification sites from RNA sequencesabstractN6-methyladinosine (m6A) modification is the most abundant co-transcriptional modification in eukaryotic RNA and plays important roles in cellular regulation. Traditional high-throughput sequencing experiments used to explore functional mechanisms are time-consuming and labor-intensive, and most of the proposed methods focused on limited species types. To further understand the relevant biological mechanisms among different species with the same RNA modification, it is necessary to develop a computational scheme that can be applied to different species. To achieve this, we proposed an attention-based deep learning method, adaptive-m6A, which consists of convolutional neural network, bi-directional long short-term memory and an attention mechanism, to identify m6A sites in multiple species. In addition, three conventional machine learning (ML) methods, including support vector machine, random forest and logistic regression classifiers, were considered in this work. In addition to the performance of ML methods for multi-species prediction, the optimal performance of adaptive-m6A yielded an accuracy of 0.9832 and the area under the receiver operating characteristic curve of 0.98. Moreover, the motif analysis and cross-validation among different species were conducted to test the robustness of one model towards multiple species, which helped improve our understanding about the sequence characteristics and biological functions of RNA modifications in different species. Rulan Wang, Chia-Ru Chung, Hsien-Da Huang, Tzong-Yi Lee |
Briefings Bioinform. | 2 |
| 2023 | A risk assessment framework for multidrug-resistant Staphylococcus aureus using machine learning and mass spectrometry technologyabstractThe emergence of multidrug-resistant bacteria is a critical global crisis that poses a serious threat to public health, particularly with the rise of multidrug-resistant Staphylococcus aureus. Accurate assessment of drug resistance is essential for appropriate treatment and prevention of transmission of these deadly pathogens. Early detection of drug resistance in patients is critical for providing timely treatment and reducing the spread of multidrug-resistant bacteria. This study aims to develop a novel risk assessment framework for S. aureus that can accurately determine the resistance to multiple antibiotics. The comprehensive 7-year study involved ˃20 000 isolates with susceptibility testing profiles of six antibiotics. By incorporating mass spectrometry and machine learning, the study was able to predict the susceptibility to four different antibiotics with high accuracy. To validate the accuracy of our models, we externally tested on an independent cohort and achieved impressive results with an area under the receiver operating characteristic curve of 0. 94, 0.90, 0.86 and 0.91, and an area under the precision-recall curve of 0.93, 0.87, 0.87 and 0.81, respectively, for oxacillin, clindamycin, erythromycin and trimethoprim-sulfamethoxazole. In addition, the framework evaluated the level of multidrug resistance of the isolates by using the predicted drug resistance probabilities, interpreting them in the context of a multidrug resistance risk score and analyzing the performance contribution of different sample groups. The results of this study provide an efficient method for early antibiotic decision-making and a better understanding of the multidrug resistance risk of S. aureus. Yuxuan Pang, Chia-Ru Chung, Hsin-Yao Wang, Haiyan Cui, Ying-Chih Chiang, Jorng-Tzong Horng, Jang-Jih Lu, Tzong-Yi Lee |
Briefings Bioinform. | 3 |
| 2022 | Neuronal Abnormalities Induced by an Intelligent Virtual Reality System for Methamphetamine Use DisorderabstractMethamphetamine use disorder (MUD) is a brain disease that leads to altered regional neuronal activity. Virtual reality (VR) is used to induce the drug cue reactivity. Previous studies reported significant frequency-specific neuronal abnormalities in patients with MUD during VR induction of drug craving. However, whether those patients exhibit neuronal abnormalities after VR induction that could serve as the treatment target remains unclear. Here, we used an integrated VR system for inducing drug related changes and investigated the neuronal abnormalities after VR exposure in patients. Fifteen patients with MUD and ten healthy subjects were recruited and exposed to drug-related VR environments. Resting-state EEG were recorded for 5 minutes twice-before and after VR and transformed to obtain the frequency-specific data. Three self-reported scales for measurement of the anxiety levels and impulsivity of participants were obtained after VR task. Statistical tests and machine learning methods were employed to reveal the differences between patients and healthy subjects. The result showed that patients with MUD and healthy subjects significantly differed in Θ, α, and γ power changes after VR. These neuronal abnormalities in patients were associated with the self-reported behavioral scales, indicating impaired impulse control. Our findings of resting-state EEG abnormalities in patients with MUD after VR exposure have the translational value and can be used to develop the treatment strategies for methamphetamine use disorder. Chun-Chuan Chen, Meng-Chang Tsai, Eric Hsiao-Kuang Wu, Chia-Ru Chung, Yuchi Lee, Po-Ru Chiu, Po-Yi Tsai, Shao-Rong Sheng, Shih-Ching Yeh |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | A large-scale investigation and identification of methicillin-resistant Staphylococcus aureus based on peaks binning of matrix-assisted laser desorption ionization-time of flight MS spectraabstractRecent studies have demonstrated that the matrix-assisted laser desorption ionization-time of flight mass spectrometry (MALDI-TOF MS) could be used to detect superbugs, such as methicillin-resistant Staphylococcus aureus (MRSA). Due to an increasingly clinical need to classify between MRSA and methicillin-sensitive Staphylococcus aureus (MSSA) efficiently and effectively, we were motivated to develop a systematic pipeline based on a large-scale dataset of MS spectra. However, the shifting problem of peaks in MS spectra induced a low effectiveness in the classification between MRSA and MSSA isolates. Unlike previous works emphasizing on specific peaks, this study employs a binning method to cluster MS shifting ions into several representative peaks. A variety of bin sizes were evaluated to coalesce drifted or shifted MS peaks to a well-defined structured data. Then, various machine learning methods were performed to carry out the classification between MRSA and MSSA samples. Totally 4858 MS spectra of unique S. aureus isolates, including 2500 MRSA and 2358 MSSA instances, were collected by Chang Gung Memorial Hospitals, at Linkou and Kaohsiung branches, Taiwan. Based on the evaluation of Pearson correlation coefficients and the strategy of forward feature selection, a total of 200 peaks (with the bin size of 10 Da) were identified as the marker attributes for the construction of predictive models. These selected peaks, such as bins 2410-2419, 2450-2459 and 6590-6599 Da, have indicated remarkable differences between MRSA and MSSA, which were effective in the prediction of MRSA. The independent testing has revealed that the random forest model can provide a promising prediction with the area under the receiver operating characteristic curve (AUC) at 0.8450. When comparing to previous works conducted with hundreds of MS spectra, the proposed scheme demonstrates that incorporating machine learning method with a large-scale dataset of clinical MS spectra may be a feasible means for clinical physicians on the administration of correct antibiotics in shorter turn-around-time, which could reduce mortality, avoid drug resistance and shorten length of stay in hospital in the future. Hsin-Yao Wang, Chia-Ru Chung, Shangfu Li, Bo-Yu Chu, Jorng-Tzong Horng, Jang-Jih Lu, Tzong-Yi Lee |
Briefings Bioinform. | 2 |
| 2021 | Large-scale mass spectrometry data combined with demographics analysis rapidly predicts methicillin resistance in Staphylococcus aureusabstractBACKGROUND: A mass spectrometry-based assessment of methicillin resistance in Staphylococcus aureus would have huge potential in addressing fast and effective prediction of antibiotic resistance. Since delays in the traditional antibiotic susceptibility testing, methicillin-resistant S. aureus remains a serious threat to human health. RESULTS: Here, linking a 7 years of longitudinal study from two cohorts in the Taiwan area of over 20 000 individually resolved methicillin susceptibility testing results, we identify associations of methicillin resistance with the demographics and mass spectrometry data. When combined together, these connections allow for machine-learning-based predictions of methicillin resistance, with an area under the receiver operating characteristic curve of >0.85 in both the discovery [95% confidence interval (CI) 0.88-0.90] and replication (95% CI 0.84-0.86) populations. CONCLUSIONS: Our predictive model facilitates early detection for methicillin resistance of patients with S. aureus infection. The large-scale antibiotic resistance study has unbiasedly highlighted putative candidates that could improve trials of treatment efficiency and inform on prescriptions. Hsin-Yao Wang, Chia-Ru Chung, Jorng-Tzong Horng, Jang-Jih Lu, Tzong-Yi Lee |
Briefings Bioinform. | 3 |
| 2020 | Characterization and identification of antimicrobial peptides with different functional activitiesabstractIn recent years, antimicrobial peptides (AMPs) have become an emerging area of focus when developing therapeutics hot spot residues of proteins are dominant against infections. Importantly, AMPs are produced by virtually all known living organisms and are able to target a wide range of pathogenic microorganisms, including viruses, parasites, bacteria and fungi. Although several studies have proposed different machine learning methods to predict peptides as being AMPs, most do not consider the diversity of AMP activities. On this basis, we specifically investigated the sequence features of AMPs with a range of functional activities, including anti-parasitic, anti-viral, anti-cancer and anti-fungal activities and those that target mammals, Gram-positive and Gram-negative bacteria. A new scheme is proposed to systematically characterize and identify AMPs and their functional activities. The 1st stage of the proposed approach is to identify the AMPs, while the 2nd involves further characterization of their functional activities. Sequential forward selection was employed to extract potentially informative features that are possibly associated with the functional activities of the AMPs. These features include hydrophobicity, the normalized van der Waals volume, polarity, charge and solvent accessibility-all of which are essential attributes in classifying between AMPs and non-AMPs. The results revealed the 1st stage AMP classifier was able to achieve an area under the receiver operating characteristic curve (AUC) value of 0.9894. During the 2nd stage, we found pseudo amino acid composition to be an informative attribute when differentiating between AMPs in terms of their functional activities. The independent testing results demonstrated that the AUCs of the multi-class models were 0.7773, 0.9404, 0.8231, 0.8578, 0.8648, 0.8745 and 0.8672 for anti-parasitic, anti-viral, anti-cancer, anti-fungal AMPs and those that target mammals, Gram-positive and Gram-negative bacteria, respectively. The proposed scheme helps facilitate biological experiments related to the functional analysis of AMPs. Additionally, it was implemented as a user-friendly web server (AMPfun, http://fdblab.csie.ncu.edu.tw/AMPfun/index.html) that allows individuals to explore the antimicrobial functions of peptides of interest. Chia-Ru Chung, Ting-Rung Kuo, Li-Ching Wu, Tzong-Yi Lee, Jorng-Tzong Horng |
Briefings Bioinform. | 1 |
| 2019 | Rapid classification of group B Streptococcus serotypes based on matrix-assisted laser desorption ionization-time of flight mass spectrometry and machine learning techniquesabstractBACKGROUND: Group B streptococcus (GBS) is an important pathogen that is responsible for invasive infections, including sepsis and meningitis. GBS serotyping is an essential means for the investigation of possible infection outbreaks and can identify possible sources of infection. Although it is possible to determine GBS serotypes by either immuno-serotyping or geno-serotyping, both traditional methods are time-consuming and labor-intensive. In recent years, the matrix-assisted laser desorption ionization-time of flight mass spectrometry (MALDI-TOF MS) has been reported as an effective tool for the determination of GBS serotypes in a more rapid and accurate manner. Thus, this work aims to investigate GBS serotypes by incorporating machine learning techniques with MALDI-TOF MS to carry out the identification. RESULTS: In this study, a total of 787 GBS isolates, obtained from three research and teaching hospitals, were analyzed by MALDI-TOF MS, and the serotype of the GBS was determined by a geno-serotyping experiment. The peaks of mass-to-charge ratios were regarded as the attributes to characterize the various serotypes of GBS. Machine learning algorithms, such as support vector machine (SVM) and random forest (RF), were then used to construct predictive models for the five different serotypes (Types Ia, Ib, III, V, and VI). After optimization of feature selection and model generation based on training datasets, the accuracies of the selected models attained 54.9-87.1% for various serotypes based on independent testing data. Specifically, for the major serotypes, namely type III and type VI, the accuracies were 73.9 and 70.4%, respectively. CONCLUSION: The proposed models have been adopted to implement a web-based tool (GBSTyper), which is now freely accessible at http://csb.cse.yzu.edu.tw/GBSTyper/, for providing efficient and effective detection of GBS serotypes based on a MALDI-TOF MS spectrum. Overall, this work has demonstrated that the combination of MALDI-TOF MS and machine intelligence could provide a practical means of clinical pathogen testing. Hsin-Yao Wang, Wen-Chi Li, Kai-Yao Huang, Chia-Ru Chung, Jorng-Tzong Horng, Jen-Fu Hsu, Jang-Jih Lu, Tzong-Yi Lee |
BMC Bioinform. | 4 |