EDBT 2026 Demo / reviewers in the wild / expert
Zhi-an Huang
dblp:175/8315 · also Zhi-An Huang
· DBLP profile ↗
35ranked-venue papers
7as first author
30since 2021 · last 2026
0000-0001-9974-148XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 16 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-Channel Learning Framework for Zero-Shot CircRNA-miRNA Interaction Prediction via State Space ModelingabstractCircRNA-miRNA interaction (CMI) plays a pivotal role in disease therapeutics and drug discovery. However, existing methods face several challenges in modeling complex biological networks and zero-shot learning scenarios. Biological networks encapsulate rich biological information, yet current approaches often fail to fully exploit this depth. Moreover, zero-shot prediction requires models to identify new interactions without relying on previously observed samples, imposing stringent requirements on generalization capabilities. To address these limitations, we propose a dual-channel learning framework leveraging State space modeling for Zero-shot CMI prediction (ZeroStem). ZeroStem first enhances the biological relevance of node using prior knowledge, and employs a graph Transformer to extract macro-topological representations. Subsequently, it generates semantic subgraphs based on meta-paths to focus on specific biological relationships, utilizing the Mamba to extract micro-semantic representations via state space modeling. Finally, macro-topological and micro-semantic representations are seamlessly integrated through linear transformation and residual connections, enabling high-precision zero-shot CMI prediction. Extensive experiments on multiple benchmark datasets demonstrate that ZeroStem significantly outperforms existing methods, validating its efficiency and robust generalization in CMI prediction. Case studies further illustrate that ZeroStem offers novel insights into the molecular mechanisms underlying intricate disease-associated networks. Mengmeng Wei, Lei Wang 0121, Zhu-Hong You, Pengwei Hu 0001, Bo-Wei Zhao, Zhi-an Huang |
AAAI | 6 |
| 2026 | PEGNet-CDA: A Propagation-Enhanced Graph Network for CircRNA-Disease Association Prediction
Yue-Chao Li, Yao-Lu Li, Chen-Yv Yang, Mengmeng Wei, Xinfei Wang 0001, Lei Wang 0121, Zhi-an Huang, Zhu-Hong You |
ICIC (27) | 8 |
| 2026 | Spatial-spectral fusion enables drug repositioning by capturing indirect and long-range associations in biological networks
Lei Wang 0121, Runzhou Tang, Zhi-an Huang, Feng Tan 0002, Lun Hu, Zhu-Hong You, Pengwei Hu 0001 |
Bioinform. | 4 |
| 2026 | HiCMamba: Enhancing Hi-C resolution and identifying 3D genome structures with state space modelingabstractHi-C technology measures genome-wide interaction frequencies, providing a powerful tool for studying the 3D genomic structure within the nucleus. However, high sequencing costs and technical challenges often result in Hi-C data with limited coverage, leading to imprecise estimates of chromatin interaction frequencies. To address this issue, we present a novel deep learning-based method HiCMamba to enhance the resolution of Hi-C contact maps using a state space model. We adopt the UNet-based auto-encoder architecture to stack the proposed holistic scan block, enabling the perception of both global and local receptive fields at multiple scales. Experimental results demonstrate that HiCMamba outperforms state-of-the-art methods while significantly reducing computational resources. Furthermore, the 3D genome structures, including topologically associating domains (TADs) and loops, identified in the contact maps recovered by HiCMamba are validated through associated epigenomic features. Our work demonstrates the potential of a state space model as foundational frameworks in the field of Hi-C resolution enhancement. The data and source code used in this work are available at GitHub: https://github.com/myang998/HiCMamba. Zhi-an Huang, Zhihang Zheng, Yuqiao Liu 0004, Pengfei Zhang 0005, Hui Xiong 0001, Shaojun Tang |
PLoS Comput. Biol. | 2 |
| 2026 | scProGraph: A Cell Bagging Strategy for Cell Type Annotation With Gene Interaction-Aware Explainability
Yue-Chao Li, Hai-Ru You, Xuequn Shang 0001, Leon Wong, Zhi-an Huang, Zhu-Hong You |
IEEE Trans. Big Data | 6 |
| 2026 | scALGSL: Active Learning and Graph Structure Learning for Cell Type Annotation From Single-Cell RNA-Seq DataabstractThe breakthrough development of single-cell RNA sequencing technology enables tissue heterogeneity analysis at single-cell resolution, where accurate cell type annotation is crucial for unlocking its full potential. To address three key challenges in current annotation methods-scarce labeled data, suboptimal graph topology, and missing cell state information-we propose scALGSL, an innovative framework integrating dynamic graph optimization with active learning. Our core contributions are threefold: (1) A graph-guided active learning mechanism adaptively selects high-value training samples, significantly alleviating label scarcity; (2) A learnable graph structure optimization module dynamically refines adjacency matrices to eliminate spurious connections caused by data sparsity; (3) A novel cell state auxiliary pathway extracts critical functional features via pre-trained models to enhance type discrimination. The systematic review showed that the average accuracy and f1 of scALGSL on the cancer dataset were 0.896 and 0.771, respectively, and it showed good robustness in cross-platform tasks. Integration of cell state information substantially boosts performance, while ablation studies validate the necessity of node selection and edge optimization modules. This framework provides a scalable solution for precise cell annotation, facilitating tumor microenvironment analysis and precision medicine applications. Zhihua Du, Jia-Le Yi, Wei-Lin Hu, Jianqiang Li 0001, Hai-Ru You, Zhu-Hong You, Zhi-an Huang |
IEEE Trans. Comput. Biol. Bioinform. | 7 |
| 2026 | Lightweight Network Enhancing High-Resolution Feature Representation for Efficient Low Dose CT DenoisingabstractLow-dose computed tomography plays a crucial role in reducing radiation exposure in clinical imaging, however, the resultant noise significantly impacts image quality and diagnostic precision. Recent transformer-based models have demonstrated strong denoising capabilities but are often constrained by high computational complexity. To overcome these limitations, we propose AMFA-Net, an adaptive multi-order feature aggregation network that provides a lightweight architecture for enhancing high-resolution feature representation in low-dose CT imaging. AMFA-Net effectively integrates local and global contexts within high-resolution feature maps while learning discriminative representations through multi-order context aggregation. We introduce an agent-based self-attention cross-shaped window transformer block that efficiently captures global context in high-resolution feature maps, which is subsequently fused with backbone features to preserve critical structural information. Our approach employs multi-order gated aggregation to adaptively guide the network in capturing expressive interactions that may be overlooked in fused features, thereby producing robust representations for denoised image reconstruction. Experiments on two challenging public datasets with 25% and 10% full-dose CT image quality demonstrate that our method surpasses state-of-the-art approaches in denoising performance with low computational cost, highlighting its potential for real-time medical applications. Yakang Li, Fazhi Qi, Shengxiang Wang, Zhengde Zhang, Zhi-an Huang, Zitong Yu |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | scBIT: Integrating Single-Cell Transcriptomic Data Into fMRI-Based Prediction for Alzheimer's Disease DiagnosisabstractFunctional MRI (fMRI) and single-cell transcriptomics are pivotal in Alzheimer's disease (AD) research, each providing unique insights into neural function and molecular mechanisms. However, integrating these complementary modalities remains largely unexplored. Here, we introduce scBIT, a novel method for enhancing AD prediction by combining fMRI with single-nucleus RNA (snRNA). scBIT leverages snRNA as an auxiliary modality, significantly improving fMRI-based prediction models and providing comprehensive interpretability. It employs a sampling strategy to segment snRNA data into cell-type-specific gene networks and utilizes a self-explainable graph neural network to extract critical subgraphs. Additionally, we use demographic and genetic similarities to pair snRNA and fMRI data across individuals, enabling robust cross-modal learning. Extensive experiments validate scBIT's effectiveness in revealing intricate brain region-gene associations and enhancing diagnostic prediction accuracy. By advancing brain imaging transcriptomics to the single-cell level, scBIT sheds new light on biomarker discovery in AD research. Experimental results show that incorporating snRNA data into the scBIT model significantly boosts accuracy, improving binary classification by 3.39% and five-class classification by 26.59%. The codes were implemented in Python and have been released on GitHub (https://github.com/77YQ77/scBIT) and Zenodo (https://zenodo.org/records/11599030) with detailed instructions. Yao Hu 0001, Yue-Chao Li, Xiyue Cao, Kay Chen Tan, Zhu-Hong You, Zhi-an Huang |
IEEE Trans. Medical Imaging | 8 |
| 2025 | Multimodal 3D Genome Pre-trainingabstractDeep learning techniques have driven significant progress in various analytical tasks within 3D genomics in computational biology. However, a holistic understanding of 3D genomics knowledge remains underexplored. Here, we propose ***MIX-HIC***, the first multimodal foundation model of 3D genome that integrates both 3D genome structure and epigenomic tracks, which obtains unified and comprehensive semantics. For accurate heterogeneous semantic fusion, we design the cross-modal interaction and mapping blocks for robust unified representation, yielding the accurate aggregation of 3D genome knowledge. Besides, we introduce the first large-scale dataset comprising over ***1 million*** pairwise samples of Hi-C contact maps and epigenomic tracks for high-quality pre-training, enabling the exploration of functional implications in 3D genomics. Extensive experiments show that MIX-HIC significantly surpasses existing state-of-the-art methods in diverse downstream tasks. This work provides a valuable resource for advancing 3D genomics research. Pengteng Li, Qianyi Cai, Zhihang Zheng, Pengfei Zhang 0005, Zhi-an Huang, Hui Xiong 0001 |
NeurIPS | 8 |
| 2025 | scExGraph: Explainable graph neural network for predicting tumor environment components with single-cell sequencing data
Zhihua Du, Jiale Yi, Jianqiang Li 0001, Hai-Ru You, Zhu-Hong You, Zhi-an Huang |
Knowl. Based Syst. | 6 |
| 2025 | CausalMixNet: A mixed-attention framework for causal intervention in robust medical image diagnosis
Yao Hu 0001, Rui Liu 0038, Jibin Wu, Zhi-an Huang, Kay Chen Tan |
Medical Image Anal. | 6 |
| 2025 | Dynamic Graph Representation Learning for Spatio-Temporal Neuroimaging AnalysisabstractNeuroimaging analysis aims to reveal the information-processing mechanisms of the human brain in a noninvasive manner. In the past, graph neural networks (GNNs) have shown promise in capturing the non-Euclidean structure of brain networks. However, existing neuroimaging studies focused primarily on spatial functional connectivity, despite temporal dynamics in complex brain networks. To address this gap, we propose a spatio-temporal interactive graph representation framework (STIGR) for dynamic neuroimaging analysis that encompasses different aspects from classification and regression tasks to interpretation tasks. STIGR leverages a dynamic adaptive-neighbor graph convolution network to capture the interrelationships between spatial and temporal dynamics. To address the limited global scope in graph convolutions, a self-attention module based on Transformers is introduced to extract long-term dependencies. Contrastive learning is used to adaptively contrast similarities between adjacent scanning windows, modeling cross-temporal correlations in dynamic graphs. Extensive experiments on six public neuroimaging datasets demonstrate the competitive performance of STIGR across different platforms, achieving state-of-the-art results in classification and regression tasks. The proposed framework enables the detection of remarkable temporal association patterns between regions of interest based on sequential neuroimaging signals, offering medical professionals a versatile and interpretable tool for exploring task-specific neurological patterns. Our codes and models are available at https://github.com/77YQ77/STIGR/. Rui Liu 0038, Yao Hu 0001, Jibin Wu, Ka-Chun Wong, Zhi-an Huang, Kay Chen Tan |
IEEE Trans. Cybern. | 5 |
| 2025 | Toward Multilabel Classification for Multiple Disease Prediction Using Gut Microbiota ProfilesabstractAdvancements in high-throughput technologies have yielded large-scale human gut microbiota profiles, sparking considerable interest in exploring the relationship between the gut microbiome and complex human diseases. Through extracting and integrating knowledge from complex microbiome data, existing machine learning (ML)-based studies have demonstrated their effectiveness in the precise identification of high-risk individuals. However, these approaches struggle to address the heterogeneity and sparsity of microbial features and explore the intrinsic relatedness among human diseases. In this work, we reframe human gut microbiome-based disease detection as a multilabel classification (MLC) problem and integrate a range of innovative techniques within the proposed MLC framework, aptly named GutMLC. Specifically, the entity semantic similarity as priori knowledge is incorporated into multilabel feature selection and loss functions by capturing the shared attributes and inherent associations among diseases and microbes. To tackle the issue of label imbalance, both within and between labels, we adapt the focal loss (FL) function for MLC using debiased inverse weighting. Extensive experiment results consistently demonstrate the competitive performance of GutMLC in comparison with commonly used MLC and single-label classification (SLC) algorithms. This work seeks to unlock the potential of gut microbiota as robust biomarkers for multiple disease prediction. Zhi-an Huang, Pengwei Hu 0001, Lun Hu, Zhu-Hong You, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Anti-Confounding Hashing: Enhancing Radiological Image Retrieval via Debiased Weighting and Counterfactual ReasoningabstractContent-based medical image retrieval (CBMIR) enables physicians to make evidence-based diagnoses by retrieving similar medical images and recalling previous cases stored in databases. However, existing CBMIR models are prone to capturing superficial correlations due to confounding factors such as complex host organs and lesions, imaging discrepancies, artifacts, and inconsistent protocols. To address this issue, we propose a plug-and-play anti-confounding hashing (ACH) method, which uses debiased sample weighting and lesion counterfactual reasoning (LCR) to directly capture the natural direct effect (NDE) of lesions on query medical images without bias. The devised debiased weighting (DBW) loss adopts a backdoor adjustment to separate lesions from confounders. To effectively locate salient areas of lesions, we present a coarse-to-fine lesion positioning (C2F-LP) module by counterfactual reasoning. On two real-world radiological image datasets, ACH achieves 0.2%-9% improvement in mean average precision (mAP) over the six state-of-the-art methods, when using code lengths ranging from 8-bit to 32-bit. Its robustness to confounding factors is demonstrated through explainable visual analysis. Yao Hu 0001, Chengjun Cai, Zhi-an Huang, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Heterogeneous Structured Federated Learning with Graph Convolutional Aggregation for MRI-Based Mental Disorder DiagnosisabstractTo relieve the growing burden of mental disorders, deep learning techniques have emerged as a promising tool to aid clinicians by detecting abnormal patterns in neuroimaging data. However, the efficacy of such models is contingent upon access to vast pools of patient data, which is impractical for individual healthcare institutions. Moreover, the privacy-preserving policy regulations governing medical images further complicate the pooling of information necessary for training robust models. Federated Learning (FL) offers a solution to this dilemma by aggregating the local model updates without compromising patient privacy. However, current studies fail to adequately account for the need to personalize models according to the diverse structures of local data. In this work, an effective heterogeneous structured FL framework using graph convolutional aggregation dubbed GAHFL is proposed to diagnose mental disorders on functional magnetic resonance imaging data. In addition, we propose to perform the global model self-evaluation to enable the training to emphasize the samples that are difficult to classify. To solve the catastrophic forgetting problem, we build a historical logit pool to awaken the global model’s recognition ability by performing a server knowledge self-distillation. Empirical evaluations demonstrate that the proposed framework achieves averaged diagnosis AUC values of 69.01% and 69.04% with different sizes of public datasets of ABIDE-I and ADHD-200 datasets, respectively. The ablation studies and robustness validation test further demonstrate the superior performance of our framework. Yao Hu 0001, Rui Liu 0038, Jiaqi Zhang 0004, Zhi-an Huang, Linqi Song, Kay Chen Tan |
IJCNN | 4 |
| 2024 | Mixed Prototype Correction for Causal Inference in Medical Image ClassificationabstractThe heterogeneity of medical images poses significant challenges to accurate disease diagnosis. To tackle this issue, the impact of such heterogeneity on the causal relationship between image features and diagnostic labels should be incorporated into model design, which however remains underexplored. In this paper, we propose a mixed prototype correction for causal inference (MPCCI) method, aimed at mitigating the impact of unseen confounding factors on the causal relationships between medical images and disease labels, so as to enhance the diagnostic accuracy of deep learning models. The MPCCI comprises a causal inference component based on front-door adjustment and an adaptive training strategy. The causal inference component employs a multi-view feature extraction (MVFE) module to establish mediators, and a mixed prototype correction (MPC) module to execute causal interventions. Moreover, the adaptive training strategy incorporates both information purity and maturity metrics to maintain stable model training. Experimental evaluations on four medical image datasets, encompassing CT and ultrasound modalities, demonstrate the superior diagnostic accuracy and reliability of the proposed MPCCI. The code will be available at https://github.com/Yajie-Zhang/MPCCI. Zhi-an Huang, Zhiliang Hong 0002, Songsong Wu, Jibin Wu, Kay Chen Tan |
ACM Multimedia | 2 |
| 2024 | DPoSt: Dynamic Proof of Storage-timeabstractProof of storage-time (PoSt) enforces highly reliable data storage services to data owners by providing lightweight and continuous data availability or possession checks to the uploaded data. Despite being promising, current PoSt schemes mainly focus on static data that will not change after a PoSt process starts. This limitation, however, would largely undermine the practicality of PoSt schemes as they cannot be effectively used for important and emerging cloud storage services like document sharing, code collaboration, and website management that will modify the uploaded data whenever needed. In this paper, we propose DPoSt, a new PoSt system that provides continuous auditing protection to data owners while supporting efficient data updates. DPoSt is built with a tailored suite of cryptographic primitives to update the intermediate auditing proofs while reducing the changes we need to conduct for subsequent challenge and proof pairs in the PoSt process. We develop a prototype of DPoSt, and the evaluation results demonstrate the efficiency of DPoSt for supporting dynamic PoSt operations. Qingyuan Xie, Chengjun Cai, Zhi-an Huang, Xiaohua Jia |
MSN | 3 |
| 2024 | Evolutionary Multitasking for Costly Task Offloading in Mobile-Edge Computing NetworksabstractThe offloading of computation-intensive tasks to an edge server near resource-constrained mobile devices can provide improved application performance and user experience. However, with the rapid growth of mobile devices connected to the edge server, it is challenging to directly obtain an optimal task offloading scheme due to increasing computational cost and problem scale. In this study, we model the costly task offloading problem (CTOP) in mobile edge computing networks to achieve efficient joint optimization of energy consumption and processing latency for mobile devices. Inspired by the success of evolutionary multitasking in solving complex optimization problems by leveraging the experience of simple optimization problems, we develop a novel multitasking framework whose effectiveness is demonstrated in solving the CTOP. In this framework, auxiliary tasks are created to optimize the local processing overhead and the edge processing overhead of task offloading. On this basis, we propose an effective multitask evolutionary algorithm that includes segmented knowledge transfer and auxiliary task update. Specifically, source and extended decision variables are considered as different knowledge to be utilized, while the auxiliary tasks are allowed to be updated dynamically. Related knowledge that is learned from cheap and simple auxiliary tasks promotes the evolutionary search for CTOP. Experimental results verify the effectiveness of knowledge transfer. Compared to existing multitasking and single-tasking algorithms, the proposed algorithm shows competitive performance in CTOP instances and achieves better comprehensive performance in terms of energy consumption and processing latency. Chen Yang 0011, Qunjian Chen, Zexuan Zhu 0001, Zhi-an Huang, Shulin Lan, Liehuang Zhu |
IEEE Trans. Evol. Comput. | 4 |
| 2024 | Multitask Learning for Joint Diagnosis of Multiple Mental Disorders in Resting-State fMRIabstractFacing the increasing worldwide prevalence of mental disorders, the symptom-based diagnostic criteria struggle to address the urgent public health concern due to the global shortfall in well-qualified professionals. Thanks to the recent advances in neuroimaging techniques, functional magnetic resonance imaging (fMRI) has surfaced as a new solution to characterize neuropathological biomarkers for detecting functional connectivity (FC) anomalies in mental disorders. However, the existing computer-aided diagnosis models for fMRI analysis suffer from unstable performance on large datasets. To address this issue, we propose an efficient multitask learning (MTL) framework for joint diagnosis of multiple mental disorders using resting-state fMRI data. A novel multiobjective evolutionary clustering algorithm is presented to group regions of interests (ROIs) into different clusters for FC pattern analysis. On the optimal clustering solution, the multicluster multigate mixture-of-expert model is used for the final classification by capturing the highly consistent feature patterns among related diagnostic tasks. Extensive simulation experiments demonstrate that the performance of the proposed framework is superior to that of the other state-of-the-art methods. Moreover, the potential for practical application of the framework is also validated in terms of limited computational resources, real-time analysis, and insufficient training data. The proposed model can identify the remarkable interpretative biomarkers associated with specific mental disorders for clinical interpretation analysis. Zhi-an Huang, Rui Liu 0038, Zexuan Zhu 0001, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Attention-Like Multimodality Fusion With Data Augmentation for Diagnosis of Mental Disorders Using MRIabstractThe globally rising prevalence of mental disorders leads to shortfalls in timely diagnosis and therapy to reduce patients' suffering. Facing such an urgent public health problem, professional efforts based on symptom criteria are seriously overstretched. Recently, the successful applications of computer-aided diagnosis approaches have provided timely opportunities to relieve the tension in healthcare services. Particularly, multimodal representation learning gains increasing attention thanks to the high temporal and spatial resolution information extracted from neuroimaging fusion. In this work, we propose an efficient multimodality fusion framework to identify multiple mental disorders based on the combination of functional and structural magnetic resonance imaging. A multioutput conditional generative adversarial network (GAN) is developed to address the scarcity of multimodal data for augmentation. Based on the augmented training data, the multiheaded gating fusion model is proposed for classification by extracting the complementary features across different modalities. The experiments demonstrate that the proposed model can achieve robust accuracies of 75.1 ± 1.5 %, 72.9 ± 1.1 %, and 87.2 ± 1.5 % for autism spectrum disorder (ASD), attention deficit/hyperactivity disorder, and schizophrenia, respectively. In addition, the interpretability of our model is expected to enable the identification of remarkable neuropathology diagnostic biomarkers, leading to well-informed therapeutic decisions. Rui Liu 0038, Zhi-an Huang, Yao Hu 0001, Zexuan Zhu 0001, Ka-Chun Wong, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Spatial-Temporal Co-Attention Learning for Diagnosis of Mental Disorders From Resting-State fMRI DataabstractNeuroimaging techniques have been widely adopted to detect the neurological brain structures and functions of the nervous system. As an effective noninvasive neuroimaging technique, functional magnetic resonance imaging (fMRI) has been extensively used in computer-aided diagnosis (CAD) of mental disorders, e.g., autism spectrum disorder (ASD) and attention deficit/hyperactivity disorder (ADHD). In this study, we propose a spatial-temporal co-attention learning (STCAL) model for diagnosing ASD and ADHD from fMRI data. In particular, a guided co-attention (GCA) module is developed to model the intermodal interactions of spatial and temporal signal patterns. A novel sliding cluster attention module is designed to address global feature dependency of self-attention mechanism in fMRI time series. Comprehensive experimental results demonstrate that our STCAL model can achieve competitive accuracies of 73.0 ± 4.5%, 72.0 ± 3.8%, and 72.5 ± 4.2% on the ABIDE I, ABIDE II, and ADHD-200 datasets, respectively. Moreover, the potential for feature pruning based on the co-attention scores is validated by the simulation experiment. The clinical interpretation analysis of STCAL can allow medical professionals to concentrate on the discriminative regions of interest and key time frames from fMRI data. Rui Liu 0038, Zhi-an Huang, Yao Hu 0001, Zexuan Zhu 0001, Ka-Chun Wong, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Interpretable artificial intelligence model for accurate identification of medical conditions using immune repertoireabstractUnderlying medical conditions, such as cancer, kidney disease and heart failure, are associated with a higher risk for severe COVID-19. Accurate classification of COVID-19 patients with underlying medical conditions is critical for personalized treatment decision and prognosis estimation. In this study, we propose an interpretable artificial intelligence model termed VDJMiner to mine the underlying medical conditions and predict the prognosis of COVID-19 patients according to their immune repertoires. In a cohort of more than 1400 COVID-19 patients, VDJMiner accurately identifies multiple underlying medical conditions, including cancers, chronic kidney disease, autoimmune disease, diabetes, congestive heart failure, coronary artery disease, asthma and chronic obstructive pulmonary disease, with an average area under the receiver operating characteristic curve (AUC) of 0.961. Meanwhile, in this same cohort, VDJMiner achieves an AUC of 0.922 in predicting severe COVID-19. Moreover, VDJMiner achieves an accuracy of 0.857 in predicting the response of COVID-19 patients to tocilizumab treatment on the leave-one-out test. Additionally, VDJMiner interpretively mines and scores V(D)J gene segments of the T-cell receptors that are associated with the disease. The identified associations between single-cell V(D)J gene segments and COVID-19 are highly consistent with previous studies. The source code of VDJMiner is publicly accessible at https://github.com/TencentAILabHealthcare/VDJMiner. The web server of VDJMiner is available at https://gene.ai.tencent.com/VDJMiner/. Yu Zhao 0009, Yidan Zhang 0001, Zhi-an Huang, Fan Yang 0081, Liang Wang 0015, Lei Duan, Jiangning Song, Jianhua Yao 0001 |
Briefings Bioinform. | 6 |
| 2023 | MIX-TPI: a flexible prediction framework for TCR-pMHC interactions based on multimodal representationsabstractMOTIVATION: The interactions between T-cell receptors (TCR) and peptide-major histocompatibility complex (pMHC) are essential for the adaptive immune system. However, identifying these interactions can be challenging due to the limited availability of experimental data, sequence data heterogeneity, and high experimental validation costs. RESULTS: To address this issue, we develop a novel computational framework, named MIX-TPI, to predict TCR-pMHC interactions using amino acid sequences and physicochemical properties. Based on convolutional neural networks, MIX-TPI incorporates sequence-based and physicochemical-based extractors to refine the representations of TCR-pMHC interactions. Each modality is projected into modality-invariant and modality-specific representations to capture the uniformity and diversities between different features. A self-attention fusion layer is then adopted to form the classification module. Experimental results demonstrate the effectiveness of MIX-TPI in comparison with other state-of-the-art methods. MIX-TPI also shows good generalization capability on mutual exclusive evaluation datasets and a paired TCR dataset. AVAILABILITY AND IMPLEMENTATION: The source code of MIX-TPI and the test data are available at: https://github.com/Wolverinerine/MIX-TPI. Zhi-an Huang, Wei Zhou 0001, Junkai Ji, Jun Zhang 0078, Shan He 0001, Zexuan Zhu 0001 |
Bioinform. | 2 |
| 2023 | Source Free Semi-Supervised Transfer Learning for Diagnosis of Mental Disorders on fMRI ScansabstractThe high prevalence of mental disorders gradually poses a huge pressure on the public healthcare services. Deep learning-based computer-aided diagnosis (CAD) has emerged to relieve the tension in healthcare institutions by detecting abnormal neuroimaging-derived phenotypes. However, training deep learning models relies on sufficient annotated datasets, which can be costly and laborious. Semi-supervised learning (SSL) and transfer learning (TL) can mitigate this challenge by leveraging unlabeled data within the same institution and advantageous information from source domain, respectively. This work is the first attempt to propose an effective semi-supervised transfer learning (SSTL) framework dubbed S3TL for CAD of mental disorders on fMRI data. Within S3TL, a secure cross-domain feature alignment method is developed to generate target-related source model in SSL. Subsequently, we propose an enhanced dual-stage pseudo-labeling approach to assign pseudo-labels for unlabeled samples in target domain. Finally, an advantageous knowledge transfer method is conducted to improve the generalization capability of the target model. Comprehensive experimental results demonstrate that S3TL achieves competitive accuracies of 69.14%, 69.65%, and 72.62% on ABIDE-I, ABIDE-II, and ADHD-200 datasets, respectively. Furthermore, the simulation experiments also demonstrate the application potential of S3TL through model interpretation analysis and federated learning extension. Yao Hu 0001, Zhi-an Huang, Rui Liu 0038, Xiaoming Xue 0001, Xiaoyan Sun 0002, Linqi Song, Kay Chen Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | A Dual-Stage Pseudo-Labeling Method for the Diagnosis of Mental Disorder on MRI ScansabstractThe high prevalence of mental disorders gradually poses a huge pressure on the public healthcare services. Recently, deep learning-based computer-aided diagnosis has been introduced to relieve the tension in healthcare institutions by automatically detecting abnormal neuroimaging-derived pheno-types in patients. However, the training of deep learning models relies on sufficiently large annotated datasets, which can be costly, time-consuming, and laborious. Semi-supervised learning (SSL) can mitigate this challenge by leveraging both labeled and unlabeled samples. In this work, an effective dual-stage pseudo-labeling based classification framework dubbed DSPL is proposed to diagnose mental disorders on functional magnetic resonance imaging data. A bicriteria-based pseudo-labels selection method is developed to filter out inferior pseudo-labeled samples. Subsequently, we further propose a self-mutual learning enhanced pseudo-labeling generation approach to mitigate the adverse effects bought by the noisy pseudo-labeled samples. On real-world datasets, the proposed method achieves diagnosis accuracies of 68.09%, 67.94%, and 68.13% on ABIDE-I, ABIDE-II, and ADHD-200, respectively. Ablation study suggests that each component in DSPL makes a great contribution to performance improvement. Yao Hu 0001, Zhi-an Huang, Rui Liu 0038, Xiaoming Xue 0001, Linqi Song, Kay Chen Tan |
IJCNN | 2 |
| 2022 | Prediction of biomarker-disease associations based on graph attention network and text representationabstractMOTIVATION: The associations between biomarkers and human diseases play a key role in understanding complex pathology and developing targeted therapies. Wet lab experiments for biomarker discovery are costly, laborious and time-consuming. Computational prediction methods can be used to greatly expedite the identification of candidate biomarkers. RESULTS: Here, we present a novel computational model named GTGenie for predicting the biomarker-disease associations based on graph and text features. In GTGenie, a graph attention network is utilized to characterize diverse similarities of biomarkers and diseases from heterogeneous information resources. Meanwhile, a pretrained BERT-based model is applied to learn the text-based representation of biomarker-disease relation from biomedical literature. The captured graph and text features are then integrated in a bimodal fusion network to model the hybrid entity representation. Finally, inductive matrix completion is adopted to infer the missing entries for reconstructing relation matrix, with which the unknown biomarker-disease associations are predicted. Experimental results on HMDD, HMDAD and LncRNADisease data sets showed that GTGenie can obtain competitive prediction performance with other state-of-the-art methods. AVAILABILITY: The source code of GTGenie and the test data are available at: https://github.com/Wolverinerine/GTGenie. Zhi-an Huang, Wenhao Gu, Wenying Pan, Xiao Yang 0019, Zexuan Zhu 0001 |
Briefings Bioinform. | 2 |
| 2021 | Mining the Associations between V(D)J Gene Segments and COVID-19 Disease CharacteristicsabstractThe emerging COVID-19 variants lead to a new wave of infections, spreading more rapidly with more severe illnesses. The adaptive immune system plays an essential role in the control and clearance of viral infection and influences clinical outcomes. However, the understanding of the adaptive immune responses to COVID-19 is not sufficient, which impedes the development progress of treatments and vaccines. To address this issue, we proposed a machine-learning-based method (termed as VDJ-Seg-Miner) to mine the underlying associations between the V(D)J gene segments of the T cell receptor in personalized immune repertoires and COVID-19 disease characteristics for immune system analysis. Our VDJ-Seg-Miner can interpretively reveal multiple associations between the V(D)J gene segments and COVID-19 disease characteristics and assign confidence scores to indicate its confidence in each revealed association. Furthermore, experimental results based on the real-world dataset suggested that the identified associations were highly consistent with those reported in previous work. Yu Zhao 0009, Yidan Zhang 0001, Zhi-an Huang, Fan Yang 0081, Lei Duan, Jianhua Yao 0001 |
BIBM | 3 |
| 2021 | Predicting microRNA-disease associations from lncRNA-microRNA interactions via Multiview Multitask LearningabstractMOTIVATION: Identifying microRNAs that are associated with different diseases as biomarkers is a problem of great medical significance. Existing computational methods for uncovering such microRNA-diseases associations (MDAs) are mostly developed under the assumption that similar microRNAs tend to associate with similar diseases. Since such an assumption is not always valid, these methods may not always be applicable to all kinds of MDAs. Considering that the relationship between long noncoding RNA (lncRNA) and different diseases and the co-regulation relationships between the biological functions of lncRNA and microRNA have been established, we propose here a multiview multitask method to make use of the known lncRNA-microRNA interaction to predict MDAs on a large scale. The investigation is performed in the absence of complete information of microRNAs and any similarity measurement for it and to the best knowledge, the work represents the first ever attempt to discover MDAs based on lncRNA-microRNA interactions. RESULTS: In this paper, we propose to develop a deep learning model called MVMTMDA that can create a multiview representation of microRNAs. The model is trained based on an end-to-end multitasking approach to machine learning so that, based on it, missing data in the side information can be determined automatically. Experimental results show that the proposed model yields an average area under ROC curve of 0.8410+/-0.018, 0.8512+/-0.012 and 0.8521+/-0.008 when k is set to 2, 5 and 10, respectively. In addition, we also propose here a statistical approach to predicting lncRNA-disease associations based on these associations and the MDA discovered using MVMTMDA. AVAILABILITY: Python code and the datasets used in our studies are made available at https://github.com/yahuang1991polyu/MVMTMDA/. Keith C. C. Chan, Zhu-Hong You, Pengwei Hu 0001, Lei Wang 0121, Zhi-an Huang |
Briefings Bioinform. | 6 |
| 2021 | Identifying Autism Spectrum Disorder From Resting-State fMRI Using Deep Belief NetworkabstractWith the increasing prevalence of autism spectrum disorder (ASD), it is important to identify ASD patients for effective treatment and intervention, especially in early childhood. Neuroimaging techniques have been used to characterize the complex biomarkers based on the functional connectivity anomalies in the ASD. However, the diagnosis of ASD still adopts the symptom-based criteria by clinical observation. The existing computational models tend to achieve unreliable diagnostic classification on the large-scale aggregated data sets. In this work, we propose a novel graph-based classification model using the deep belief network (DBN) and the Autism Brain Imaging Data Exchange (ABIDE) database, which is a worldwide multisite functional and structural brain imaging data aggregation. The remarkable connectivity features are selected through a graph extension of K -nearest neighbors and then refined by a restricted path-based depth-first search algorithm. Thanks to the feature reduction, lower computational complexity could contribute to the shortening of the training time. The automatic hyperparameter-tuning technique is introduced to optimize the hyperparameters of the DBN by exploring the potential parameter space. The simulation experiments demonstrate the superior performance of our model, which is 6.4% higher than the best result reported on the ABIDE database. We also propose to use the data augmentation and the oversampling technique to identify further the possible subtypes within the ASD. The interpretability of our model enables the identification of the most remarkable autistic neural correlation patterns from the data-driven outcomes. Zhi-an Huang, Zexuan Zhu 0001, Chuen Heung Yau, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Identification of Autistic Risk Candidate Genes and Toxic Chemicals via Multilabel LearningabstractAs a group of complex neurodevelopmental disorders, autism spectrum disorder (ASD) has been reported to have a high overall prevalence, showing an unprecedented spurt since 2000. Due to the unclear pathomechanism of ASD, it is challenging to diagnose individuals with ASD merely based on clinical observations. Without additional support of biochemical markers, the difficulty of diagnosis could impact therapeutic decisions and, therefore, lead to delayed treatments. Recently, accumulating evidence have shown that both genetic abnormalities and chemical toxicants play important roles in the onset of ASD. In this work, a new multilabel classification (MLC) model is proposed to identify the autistic risk genes and toxic chemicals on a large-scale data set. We first construct the feature matrices and partially labeled networks for autistic risk genes and toxic chemicals from multiple heterogeneous biological databases. Based on both global and local measure metrics, the simulation experiments demonstrate that the proposed model achieves superior classification performance in comparison with the other state-of-the-art MLC methods. Through manual validation with existing studies, 60% and 50% out of the top-20 predicted risk genes are confirmed to have associations with ASD and autistic disorder, respectively. To the best of our knowledge, this is the first computational tool to identify ASD-related risk genes and toxic chemicals, which could lead to better therapeutic decisions of ASD. Zhi-an Huang, Jia Zhang 0019, Zexuan Zhu 0001, Qi Wu 0003, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Identification of Autistic Risk Genes Using Developmental Brain Gene Expression Data
Zhi-an Huang, Zhu-Hong You, Shanwen Zhang, Wenzhun Huang |
ICIC (2) | 1 |
| 2020 | Multi-Task Learning for Efficient Diagnosis of ASD and ADHD using Resting-State fMRI DataabstractIncreasing mental disorders have emerged as an urgent public health concern such as autism spectrum disorder (ASD) and attention deficit hyperactivity disorder (ADHD). Related mental disorders may share high overlap in clinical symptoms. Therefore, their diagnosis can be challenging to merely rely on the observation of cognitive phenotypes and behavioral manifestations. Unfortunately, there is no additional support of biochemical markers, laboratory tests, or neuroimaging analysis, which can be used as a diagnostic gold standard currently. Over the past decades, resting-state functional magnetic resonance imaging (rs-fMRI) has been considered as one of the most promising modality to capture the intrinsic neural activation patterns between regions in the brain. In this work, we focus on ASD and ADHD due to their high prevalence and relevance with the aim to exploit the multi-task learning (MTL) paradigm for their diagnosis. To the best of our knowledge, this is the first time to make use of the disease-specific heterogeneities for the MTL classification of ASD and ADHD via rs-fMRI signal. We propose a novel graph-based feature selection method to filter out irrelevant functional connectivity features. Then an efficient structure of multi-gate mixture-of-experts (MMoE) is applied to the MTL classification framework. Finally, the experiment results demonstrate that the proposed model can achieve a reliable classification performance in a short term, yielding the mean accuracies of 0.687±0.005 and 0.650±0.014 in ASD and ADHD datasets, respectively. The graph-based feature selection method and MMoE model are demonstrated to make great contribution to performance improvement. Zhi-an Huang, Rui Liu 0038, Kay Chen Tan |
IJCNN | 1 |
| 2019 | Precise Prediction of Pathogenic Microorganisms Using 16S rRNA Gene Sequences
Zhi-an Huang, Zhu-Hong You, Pengwei Hu 0001, Liping Li 0003, Zhengwei Li 0001, Lei Wang 0121 |
ICIC (2) | 2 |
| 2017 | LW-FQZip 2: a parallelized reference-based compression of FASTQ filesabstractBACKGROUND: The rapid progress of high-throughput DNA sequencing techniques has dramatically reduced the costs of whole genome sequencing, which leads to revolutionary advances in gene industry. The explosively increasing volume of raw data outpaces the decreasing disk cost and the storage of huge sequencing data has become a bottleneck of downstream analyses. Data compression is considered as a solution to reduce the dependency on storage. Efficient sequencing data compression methods are highly demanded. RESULTS: In this article, we present a lossless reference-based compression method namely LW-FQZip 2 targeted at FASTQ files. LW-FQZip 2 is improved from LW-FQZip 1 by introducing more efficient coding scheme and parallelism. Particularly, LW-FQZip 2 is equipped with a light-weight mapping model, bitwise prediction by partial matching model, arithmetic coding, and multi-threading parallelism. LW-FQZip 2 is evaluated on both short-read and long-read data generated from various sequencing platforms. The experimental results show that LW-FQZip 2 is able to obtain promising compression ratios at reasonable time and memory space costs. CONCLUSIONS: The competence enables LW-FQZip 2 to serve as a candidate tool for archival or space-sensitive applications of high-throughput DNA sequencing data. LW-FQZip 2 is freely available at http://csse.szu.edu.cn/staff/zhuzx/LWFQZip2 and https://github.com/Zhuzxlab/LW-FQZip2 . Zhi-an Huang, Zhenkun Wen, Qingjin Deng, Zexuan Zhu 0001 |
BMC Bioinform. | 1 |
| 2017 | PBMDA: A novel and effective path-based computational model for miRNA-disease association predictionabstractIn the recent few years, an increasing number of studies have shown that microRNAs (miRNAs) play critical roles in many fundamental and important biological processes. As one of pathogenetic factors, the molecular mechanisms underlying human complex diseases still have not been completely understood from the perspective of miRNA. Predicting potential miRNA-disease associations makes important contributions to understanding the pathogenesis of diseases, developing new drugs, and formulating individualized diagnosis and treatment for diverse human complex diseases. Instead of only depending on expensive and time-consuming biological experiments, computational prediction models are effective by predicting potential miRNA-disease associations, prioritizing candidate miRNAs for the investigated diseases, and selecting those miRNAs with higher association probabilities for further experimental validation. In this study, Path-Based MiRNA-Disease Association (PBMDA) prediction model was proposed by integrating known human miRNA-disease associations, miRNA functional similarity, disease semantic similarity, and Gaussian interaction profile kernel similarity for miRNAs and diseases. This model constructed a heterogeneous graph consisting of three interlinked sub-graphs and further adopted depth-first search algorithm to infer potential miRNA-disease associations. As a result, PBMDA achieved reliable performance in the frameworks of both local and global LOOCV (AUCs of 0.8341 and 0.9169, respectively) and 5-fold cross validation (average AUC of 0.9172). In the cases studies of three important human diseases, 88% (Esophageal Neoplasms), 88% (Kidney Neoplasms) and 90% (Colon Neoplasms) of top-50 predicted miRNAs have been manually confirmed by previous experimental reports from literatures. Through the comparison performance between PBMDA and other previous models in case studies, the reliable performance also demonstrates that PBMDA could serve as a powerful computational tool to accelerate the identification of disease-miRNA associations. Zhu-Hong You, Zhi-an Huang, Zexuan Zhu 0001, Guiying Yan, Zhengwei Li 0001, Zhenkun Wen, Xing Chen 0001 |
PLoS Comput. Biol. | 2 |