EDBT 2026 Demo / reviewers in the wild / expert
Justin Bo-Kai Hsu
dblp:52/6923
· DBLP profile ↗
9ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Self-Attention Enhanced Deep Learning Models for Immune Cell Deconvolution from Bulk RNA-SeqabstractAccurate immune cell composition profiling is crucial for understanding immunological dynamics and disease mechanisms. Bulk RNA sequencing (bulk RNA-seq) is widely employed due to its cost-effectiveness and scalability; however, it lacks the resolution to identify cell-specific gene expression. To address this limitation, we propose a self-attention enhanced deep learning model designed for precise immune cell deconvolution from bulk RNA-seq data. We systematically annotated immune cell types from four single-cell RNA-seq (scRNA-seq) peripheral blood mononuclear cell (PBMC) datasets and validated these annotations against established automated identification tools (SingleR, Seurat, scPred, ScType). Leveraging these annotations, we generated realistic pseudo-bulk RNA-seq training samples using Dirichlet-distribution-based composition sampling, significantly enhancing the model’s performance, particularly for rare cell populations. Comparative evaluations demonstrated that our self-attention enhanced deep learning model consistently outperformed existing approaches, including CIBERSORTx and Scaden, achieving lower prediction errors and higher correlations on benchmark PBMC datasets. Integrating multi-head self-attention allowed the model to dynamically capture intricate dependencies among gene expression features, substantially improving deconvolution accuracy for specific cell subsets. While demonstrating robust performance on PBMC datasets, we acknowledge that broader validation is essential due to potential limitations in generalizability across different tissue types and conditions. Our study highlights the potential of self-attention mechanisms and realistic training data generation strategies to enhance computational deconvolution techniques, providing valuable tools for clinical diagnostics and translational immunology research. Chia-Ru Chung, Yen-Lin Chen, Justin Bo-Kai Hsu, Li-Ching Wu, Tzong-Yi Lee, Jorng-Tzong Horng |
CIBCB | 4 |
| 2025 | Explainable AI-Enhanced Kinase Activity Profiling Through PhosphoproteomicsabstractKinases play a critical role in regulating fundamental cellular processes, including metabolism, signal transduction, and cell growth, primarily through phosphorylation. The dysregulation of kinase activity is implicated in various diseases, highlighting the urgent need for robust and interpretable methodologies to profile this activity. Current approaches frequently depend on overly complex or limited datasets, lack generalizability, or fail to provide meaningful biological insights into the mechanisms governing kinase activity. To address these challenges, we developed an explainable deep learning framework that leverages mass spectrometry-based phosphoproteomics data to profile kinase activity effectively. Our study systematically evaluated deep neural networks (DNNs) and convolutional neural networks (CNNs), incorporating a diverse set of feature inputs, including phosphorylation sites and kinase-substrate relationships. A notable finding was that a three-layer CNN, optimized through rigorous feature selection techniques, demonstrated superior performance, achieving substantial improvements in prediction accuracy and stability when compared to established methods such as kinase-substrate enrichment analysis (KSEA) and the kinase activity ranking pipeline (KARP). We integrated Shapley additive explanations (SHAP) values to enhance interpretability, illuminating biologically significant phosphorylation sites. For example, PAK2-related phosphorylation sites associated with the progression of colon adenocarcinoma and CAMK2D sites integral to adrenergic signaling were identified, thereby effectively linking computational predictions to established molecular pathways. This research illustrates the potential of explainable artificial intelligence in advancing kinase activity profiling by providing accurate and interpretable predictions. Our framework is valuable for elucidating disease mechanisms and identifying therapeutic targets, facilitating broader applications in precision medicine. Chia-Ru Chung, Ming-Feng Ho, Li-Ching Wu, Justin Bo-Kai Hsu, Tzong-Yi Lee, Jorng-Tzong Horng |
CIBCB | 5 |
| 2025 | Alzheimer's Disease Risk Prediction in the Elderly: A Machine Learning Approach Combining Clinical Characteristics and Polygenic Risk ScoresabstractMotivation: Alzheimer's disease (AD) is the most common type of dementia. Given the lack of a cure, early identification of high-risk populations is crucial for timely prevention. While several studies have focused on AD risk prediction, single feature (e.g., age) may dominate model performance, limiting the discovery of other potential risk factors. This study incorporates key features identified in previous research and applies propensity score matching for age and sex, aiming to improve the predictive performance of AD risk models for older adults.Methods: This study utilized data from the UK Biobank to integrate genetic and clinical data and developed 5-year and 10- year AD risk prediction models for older adults, respectively. The workflow included genome-wide association studies (GWAS) on 433,589 participants were conducted to identify significant Single nucleotide polymorphisms (SNPs) under three p-value thresholds, followed by polygenic risk score (PRS) calculation for 13,282 participants using PRSice-2 and Lassosum, and the integration of multiple features to construct prediction models. Clinical features, PRS, and significant SNPs were then incorporated into four machine learning models: Logistic Regression, LightGBM, XGBoost, and Multi-Layer Perceptron (MLP) for prediction and performance comparison.Results: For the 5-year risk prediction, the MLP model demonstrated the best performance, achieving an AUC of 0.88 based on 37 clinical features and 206 significant SNPs. For the 10-year risk prediction, the MLP model also demonstrated the best performance, achieving an AUC of 0.89 based on 37 clinical features, 206 significant SNPs, and PRS based on these SNPs. SHAP analysis revealed that key contributors across both models included ApoE genotype, urinary tract infection (N390), disorientation, depressive symptoms, and pairs matching time. The 5-year model emphasized immediate clinical and cognitive indicators such as reaction time and number of medications taken, whereas the 10-year model highlighted long-term risk factors including BMI, diabetes, and peak expiratory flow.Conclusion: This study demonstrates that integrating clinical features with PRS can effectively enhance the accuracy of AD risk prediction models for older adults. However, to further validate the utility of PRS, future research should involve collaborations across diverse populations and databases. Additionally, further exploration of other potential risk factors is needed to enhance the clinical applicability of these models. Justin Bo-Kai Hsu, Cheng-Yang Lee, Jia-Ruey Tsai, Vijesh Kumar Yadav, Tzu-Hao Chang |
CIBCB | 1 |
| 2025 | Establishing a Cross-Population Machine Learning Model for Alzheimer's Disease Risk Prediction in the Taiwanese PopulationabstractAlzheimer's disease (AD) is the most common form of dementia worldwide, but most existing risk prediction models are based on European populations and lack generalizability to other ethnic groups. This study combined genetic and clinical data from both the UK Biobank and a Taiwanese population to develop a model more broadly applicable across populations. By selecting SNPs with similar minor allele frequencies (MAF) between groups and performing genotype imputation, models were built using polygenic risk scores (PRS), genotype, and clinical data. Logistic regression, XGBoost, and multilayer perceptron (MLP) were used for comparison. All models performed well on the UKB validation set (AUC = 0.81), and the logistic regression and MLP models showed improved performance (AUC = 0.87) in the Taiwanese test set. Key predictive features included PRSice2_PRS, lassosum_PRS, and age. The results highlight the potential of integrating genetic and clinical data to improve risk prediction for AD across populations, offering insights into AD pathogenesis and aiding the development of precision medicine strategies. Justin Bo-Kai Hsu, Jia-Ruey Tsai, Shih-Han Hung, Cheng-Yang Lee, Chaur-Jong Hu, Tzu-Hao Chang |
CIBCB | 1 |
| 2025 | Integrative Multi-Omics Prognostic Modeling of Glioma Recurrence Using Variational Autoencoder and Similarity Network FusionabstractGliomas represent the most prevalent form of malignant brain tumors, with glioblastoma (GBM) classified as the most aggressive subtype and lower-grade gliomas (LGGs) showing a high recurrence rate despite better survival outcomes. Traditional prognostic methods rely on clinical and molecular markers, yet they often fail to leverage the predictive potential of integrative multiomics data. Hence, early and accurate identification of glioma recurrence is critical for optimizing treatment strategies and improving patient outcomes. To address this challenge, we developed a deep learning-based approach that integrates multiomics data with clinical features to enhance recurrence risk stratification. Multiomics data, including mRNA, miRNA, DNA methylation, and CNV profiles from 512 LGG and 350 GBM patients in The Cancer Genome Atlas (TCGA), were processed using the variational autoencoder for nonlinear feature extraction, followed by similarity network fusion to capture cross-omics relationships. Recurrence-guided feature selection identified a robust biomarker panel that effectively categorized patients into high-and low-risk recurrence groups, with Kaplan-Meier survival analysis demonstrating a notable significant difference with p-value less than 0.001. Integration of clinical factors (age, tumor grade, IDH mutation status) further improved predictive performance, yielding a C-index of 0.695 (p < 2e-16) for LGG and 0.619 (p = 3.346e-11) for GBM. Together, these findings establish a multiomics-based predictive model that enables refined glioma recurrence risk assessment, offering personalized treatment strategies for LGG and GBM patients. Phuong Lam Tran, Justin Bo-Kai Hsu, Tzong-Yi Lee |
CIBCB | 3 |
| 2025 | AI-Enhanced MALDI-TOF MS Analysis for Important Peaks on Predicting Ciprofloxacin Resistance across Different Gram-Negative BacteriaabstractRapid identification of antibiotic-resistant infections is crucial, as antimicrobial resistance is a global health crisis. Yet, conventional antibiotic susceptibility tests (AST) often require days to yield results. Matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS) has emerged as a rapid, cost-effective tool for bacterial identification and shows promise for resistance profiling by detecting spectral biomarkers. In this study, we harness MALDI-TOF MS with machine learning and deep learning to predict ciprofloxacin resistance across four Gram-negative bacteria, Escherichia coli, Klebsiella pneumoniae, Acinetobacter baumannii, and Acinetobacter nosocomialis, using a cross-species "basket-wise" approach. We extracted features from mass spectra using kernel density estimation-based peak detection and m/z binning, then trained a random forest (RF) classifier and a convolutional neural network (CNN) to distinguish ciprofloxacin-resistant and susceptible isolates. To interpret the models, we employed dual feature importance analyses: gradient-weighted class activation mapping (Grad-CAM) for the CNN to highlight critical m/z regions and an ensemble RF-based method to identify significant peak features. The CNN achieved higher overall accuracy than the RF, especially in three of four species, while the ensemble RF approach identified interpretable sets of around 20 important m/z peaks per organism. Several informative peaks overlapped between species, indicating some common resistance-associated spectral signatures. However, no single universal marker was found across all species. These findings demonstrate an AI-enhanced MALDI-TOF MS framework for rapid AMR detection, yielding accurate predictions and interpretable spectral markers. The approach highlights clinical potential to guide effective therapy and bolster antimicrobial stewardship, particularly for underrepresented pathogens such as A. nosocomialis. Hsin-Yao Wang, Chia-Ru Chung, Wen-Rui Zhang, Li-Ching Wu, Justin Bo-Kai Hsu, Jang-Jih Lu, Jorng-Tzong Horng |
CIBCB | 5 |
| 2019 | Characterization and identification of lysine glutarylation based on intrinsic interdependence between positions in the substrate sitesabstractBACKGROUND: Glutarylation, the addition of a glutaryl group (five carbons) to a lysine residue of a protein molecule, is an important post-translational modification and plays a regulatory role in a variety of physiological and biological processes. As the number of experimentally identified glutarylated peptides increases, it becomes imperative to investigate substrate motifs to enhance the study of protein glutarylation. We carried out a bioinformatics investigation of glutarylation sites based on amino acid composition using a public database containing information on 430 non-homologous glutarylation sites. RESULTS: The TwoSampleLogo analysis indicates that positively charged and polar amino acids surrounding glutarylated sites may be associated with the specificity in substrate site of protein glutarylation. Additionally, the chi-squared test was utilized to explore the intrinsic interdependence between two positions around glutarylation sites. Further, maximal dependence decomposition (MDD), which consists of partitioning a large-scale dataset into subgroups with statistically significant amino acid conservation, was used to capture motif signatures of glutarylation sites. We considered single features, such as amino acid composition (AAC), amino acid pair composition (AAPC), and composition of k-spaced amino acid pairs (CKSAAP), as well as the effectiveness of incorporating MDD-identified substrate motifs into an integrated prediction model. Evaluation by five-fold cross-validation showed that AAC was most effective in discriminating between glutarylation and non-glutarylation sites, according to support vector machine (SVM). CONCLUSIONS: The SVM model integrating MDD-identified substrate motifs performed well, with a sensitivity of 0.677, a specificity of 0.619, an accuracy of 0.638, and a Matthews Correlation Coefficient (MCC) value of 0.28. Using an independent testing dataset (46 glutarylated and 92 non-glutarylated sites) obtained from the literature, we demonstrated that the integrated SVM model could improve the predictive performance effectively, yielding a balanced sensitivity and specificity of 0.652 and 0.739, respectively. This integrated SVM model has been implemented as a web-based system (MDDGlutar), which is now freely available at http://csb.cse.yzu.edu.tw/MDDGlutar/ . Kai-Yao Huang, Hui-Ju Kao, Justin Bo-Kai Hsu, Shun-Long Weng, Tzong-Yi Lee |
BMC Bioinform. | 3 |
| 2013 | An enhanced computational platform for investigating the roles of regulatory RNA and for identifying functional RNA motifsabstractBACKGROUND: Functional RNA molecules participate in numerous biological processes, ranging from gene regulation to protein synthesis. Analysis of functional RNA motifs and elements in RNA sequences can obtain useful information for deciphering RNA regulatory mechanisms. Our previous work, RegRNA, is widely used in the identification of regulatory motifs, and this work extends it by incorporating more comprehensive and updated data sources and analytical approaches into a new platform. METHODS AND RESULTS: An integrated web-based system, RegRNA 2.0, has been developed for comprehensively identifying the functional RNA motifs and sites in an input RNA sequence. Numerous data sources and analytical approaches are integrated, and several types of functional RNA motifs and sites can be identified by RegRNA 2.0: (i) splicing donor/acceptor sites; (ii) splicing regulatory motifs; (iii) polyadenylation sites; (iv) ribosome binding sites; (v) rho-independent terminator; (vi) motifs in mRNA 5'-untranslated region (5'UTR) and 3'UTR; (vii) AU-rich elements; (viii) C-to-U editing sites; (ix) riboswitches; (x) RNA cis-regulatory elements; (xi) transcriptional regulatory motifs; (xii) user-defined motifs; (xiii) similar functional RNA sequences; (xiv) microRNA target sites; (xv) non-coding RNA hybridization sites; (xvi) long stems; (xvii) open reading frames; (xviii) related information of an RNA sequence. User can submit an RNA sequence and obtain the predictive results through RegRNA 2.0 web page. CONCLUSIONS: RegRNA 2.0 is an easy to use web server for identifying regulatory RNA motifs and functional sites. Through its integrated user-friendly interface, user is capable of using various analytical approaches and observing results with graphical visualization conveniently. RegRNA 2.0 is now available at http://regrna2.mbc.nctu.edu.tw. Tzu-Hao Chang, Hsi-Yuan Huang, Justin Bo-Kai Hsu, Shun-Long Weng, Jorng-Tzong Horng, Hsien-Da Huang |
BMC Bioinform. | 3 |
| 2011 | miRTar: an integrated system for identifying miRNA-target interactions in HumanabstractBACKGROUND: MicroRNAs (miRNAs) are small non-coding RNA molecules that are ~22-nt-long sequences capable of suppressing protein synthesis. Previous research has suggested that miRNAs regulate 30% or more of the human protein-coding genes. The aim of this work is to consider various analyzing scenarios in the identification of miRNA-target interactions, as well as to provide an integrated system that will aid in facilitating investigation on the influence of miRNA targets by alternative splicing and the biological function of miRNAs in biological pathways. RESULTS: This work presents an integrated system, miRTar, which adopts various analyzing scenarios to identify putative miRNA target sites of the gene transcripts and elucidates the biological functions of miRNAs toward their targets in biological pathways. The system has three major features. First, the prediction system is able to consider various analyzing scenarios (1 miRNA:1 gene, 1:N, N:1, N:M, all miRNAs:N genes, and N miRNAs: genes involved in a pathway) to easily identify the regulatory relationships between interesting miRNAs and their targets, in 3'UTR, 5'UTR and coding regions. Second, miRTar can analyze and highlight a group of miRNA-regulated genes that participate in particular KEGG pathways to elucidate the biological roles of miRNAs in biological pathways. Third, miRTar can provide further information for elucidating the miRNA regulation, i.e., miRNA-target interactions, affected by alternative splicing. CONCLUSIONS: In this work, we developed an integrated resource, miRTar, to enable biologists to easily identify the biological functions and regulatory relationships between a group of known/putative miRNAs and protein coding genes. miRTar is now available at http://miRTar.mbc.nctu.edu.tw/. Justin Bo-Kai Hsu, Chih-Min Chiu, Sheng-Da Hsu, Wei-Yun Huang, Chia-Hung Chien, Tzong-Yi Lee, Hsien-Da Huang |
BMC Bioinform. | 1 |