EDBT 2026 Demo / reviewers in the wild / expert
Hao Zhang 0064
dblp:55/2270-64
· DBLP profile ↗
17ranked-venue papers
1as first author
16since 2021 · last 2026
0000-0002-2058-7123ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 13 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Annotating Spatial Multi-Omics Spot-Level Niche Types Using Bi-View Retrieval-Augmented Generation With SpotTypeLLMabstractSpatial multi-omics techniques generate extensive spot-level profiles without accompanying spot-type labels, forcing biologists into labor-intensive manual annotation. Although large language models (LLMs) promise automated annotation, they are poorly equipped to handle high-dimensional numeric inputs, struggle to convert complex spatial-omics structures into interpretable text, and lack the specialized biological knowledge needed to avoid hallucinations. Moreover, the intrinsic sparsity and heterogeneity of spatial omics data undermine robust feature extraction and accurate spot-level niche label assignment. To address these challenges, we propose SpotTypeLLM, a framework for annotating spatial multi-omics spot-level niche types using a Bi-view Retrieval-Augmented Generation (BiRAG) tailored for LLMs. Specifically, SpotTypeLLM encodes spatial multi-omics data through scLLM-based embeddings, graph convolutional networks (GCNs), and multi-view attention module, and then selects spot-specific genes to construct the prompts for LLMs. We validate the accuracy and practical utility of our framework by comparing it with ten state-of-the-art methods on labeled spatial multi-omics datasets. The results demonstrate the effectiveness of SpotTypeLLM in overcoming the limitations of current annotation methods and enhancing spatial multi-omics analysis. Longyi Li, Liyan Dong, Decheng Li, Hanbo Liu, Yuheng Zhu, Xiyuan Mei, Hao Zhang 0064, Dong Xu 0002 |
IEEE Trans. Comput. Biol. Bioinform. | 8 |
| 2026 | GALA: Integrating Weighted Graph Walks and Latent-Space Adversarial Training for Single-Cell Batch AlignmentabstractSingle-cell RNA sequencing (scRNA-seq) enables unprecedented exploration of cellular heterogeneity, yet technical variations across datasets introduce pervasive batch effects that severely compromise integrative analysis. We present Graph-based Adversarial Latent Alignment (GALA), a novel batch correction framework that synergistically integrates weighted graph random walks with latent space adversarial training to robustly align scRNA-seq data while preserving critical biological signals. GALA employs a Weighted Graph Mutual Nearest Neighbor (WGMNN) module that achieves up to 125% improvement in cross-batch cell pairing diversity and 48% increase in coverage compared to conventional approaches, substantially enhancing the detection of biologically meaningful correspondences between batches. These optimized cell pairs guide adversarial training within a carefully designed low-dimensional latent space, generating batch-agnostic representations that simultaneously eliminate technical artifacts while faithfully preserving biological variability. Evaluated across five benchmark datasets representing diverse real-world scenarios-including identical cell types, non-identical cell types, and complex multi-batch settings-GALA demonstrates robust performance in batch correction and biological signal preservation, achieving high F1 scores of 0.91, 0.82, 0.88, 0.82, and 0.76, respectively. GALA outperforms established methods including Seurat v4, Harmony, and Scanorama, with notable advantages in challenging datasets with heterogeneous cell populations. Aggregating performance across all datasets, GALA achieves the highest overall integration score and best average rank while maintaining high computational efficiency, demonstrating consistent superiority and strong robustness across diverse scRNA-seq integration scenarios. Liyan Dong, Longyi Li, Hao Zhang 0064 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | Tokenizing RNA Structures: A Discrete Generative Approach to Represent the RNA Folding LandscapesabstractRNA's diverse biological functions, including gene regulation, enzymatic catalysis, signal transduction, and even therapeutic activity, ultimately depend on its three-dimensional (3D) structure. Departing from the classical task of inferring structure directly from sequence, this study investigates how to efficiently encode and generate RNA 3D structures, thereby enabling systematic exploration of the molecules' vast folding landscape. We propose a generative framework, RNA-FrameEncoder (RNA-FE). First, a Vector-Quantised Variational AutoEncoder compresses continuous RNA backbones into sequences of discrete structural tokens, “verbalizing” the geometry while suppressing high-frequency noise. Next, a Transformer language model is trained on these token sequences to capture their joint distribution, and a decoder subsequently projects the tokens back into Cartesian coordinates. This design sidesteps the complexity of explicit equivariance modelling, lowers computational cost, and markedly enhances the representational capacity. Comprehensive benchmarks show that RNA-FE can unconditionally generate realistic, novel, and designable RNA structures without any sequence or structural priors. Overall, RNA-FE provides a scalable approach for efficiently charting RNA 3D structural space and opens a data-efficient avenue for a range of downstream applications. The code and pretrained models are publicly available at: https://github.com/users/Hanbo24-bit/RNA-FE. Hanbo Liu, Hao Zhang 0064, Longyi Li, Xiyuan Mei, Enshuang Zhao, Yuheng Zhu, Yinfei Dai, Jingxun Cao, Dong Xu 0002 |
BIBM | 2 |
| 2025 | RBTN: A Spatio Temporal Graph Neural Network for Integrating Single Cell and Spatial Transcriptomics During Magnaporthe Oryzae Infection in RiceabstractRice blast, caused by the fungal pathogen Magnaporthe oryzae (M. oryzae), remains one of the most devastating threats to global rice production, capable of reducing yields by up to 30%. Current integration methods for single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics (ST) treat tissues as static snapshots, overlooking the temporal dynamics crucial for host-pathogen analysis. Here, we introduce Rice Blast Temporal Graph Neural Network (RBTN), the first computational framework specifically designed to integrate scRNA-seq and ST data across the three-dimensional spot-cell-time manifold during rice blast infection. RBTN constructs a unified graph whose nodes represent either individual cell transcriptomes or spatial spots sampled at three infection time points (0 h,$\text{1 2 h}, \text{2 4 h}$). It learns stage-specific mappings via bidirectional cosine loss and enforces temporal smoothness through a cosinebased consistency term. Alternating self-attention and temporal-attention layers capture local spatial organization and cross-time transitions. When applied to a public rice blast dataset, RBTN consistently delivers high-fidelity spatial reconstructions across different sets of highly variable marker genes (HVGs) and produces interpretable embeddings that jointly encode spatial, cellular, and temporal relationships. RBTN therefore establishes a robust, temporally-aware computational platform for elucidating the spatiotemporal dynamics of host-pathogen interactions during rice blast infection. All code in this paper is available at https://github.com/zhuyuheng111/RBTN. Yuheng Zhu, Enshuang Zhao, Yinfei Dai, Hanbo Liu, Longyi Li, Qiyuan Mei, Hao Zhang 0064 |
BIBM | 10 |
| 2025 | spaLLM: enhancing spatial domain analysis in multi-omics data through large language model integrationabstractSpatial multi-omics technologies provide valuable data on gene expression from various omics in the same tissue section while preserving spatial information. However, deciphering spatial domains within spatial omics data remains challenging due to the sparse gene expression. We propose spaLLM, the first multi-omics spatial domain analysis method that integrates large language models to enhance data representation. Our method combines a pre-trained single-cell language model (scGPT) with graph neural networks and multi-view attention mechanisms to compensate for limited gene expression information in spatial omics while improving sensitivity and resolution within modalities. SpaLLM processes multiple spatial modalities, including RNA, chromatin, and protein data, potentially adapting to emerging technologies and accommodating additional modalities. Benchmarking against eight state-of-the-art methods across four different datasets and platforms demonstrates that our model consistently outperforms other advanced methods across multiple supervised evaluation metrics. The source code for spaLLM is freely available at https://github.com/liiilongyi/spaLLM. Longyi Li, Liyan Dong, Hao Zhang 0064, Dong Xu 0002 |
Briefings Bioinform. | 3 |
| 2025 | TVAE-RNA: ensemble-based RNA secondary structure prediction via transformer variational autoencodersabstractMOTIVATION: Accurate prediction of RNA secondary structure remains challenging due to the presence of pseudoknots, long-range dependencies, and limited labeled data. RESULTS: We propose TVAE, a novel framework that integrates a Transformer encoder with a Variational Autoencoder (VAE). The Transformer captures global dependencies in the sequence, while the VAE models structural variability by learning a probabilistic latent space. Unlike deterministic models, TVAE generates diverse and biologically plausible secondary structures, enabling more comprehensive structure discovery. To obtain discrete predictions, we introduce GHA-Pairing, a fast and biologically constrained base-pairing algorithm. TVAE demonstrates strong generalization across different RNA families and achieves state-of-the-art performance on benchmark datasets, reaching an F1 score of 0.89 and 83% accuracy, surpassing existing methods by 10%. These results highlight the advantage of probabilistic modeling for RNA structure prediction and its potential to enhance biological insights. AVAILABILITY AND IMPLEMENTATION: Code and pretrained models are available at https://github.com/mei-rna/TVAE-RNA. The released version of the dataset and models can also be accessed via DOI: 10.5281/zenodo.16946114. Xiyuan Mei, Hanbo Liu, Yuheng Zhu, Enshuang Zhao, Longyi Li, Hao Zhang 0064 |
Bioinform. | 6 |
| 2025 | TFS-Net: Temporal first simulation network for video saliency prediction
Longyi Li, Liyan Dong, Hao Zhang 0064, Zhengtai Zhang, Minghui Sun 0001 |
Expert Syst. Appl. | 3 |
| 2024 | SurvConvMixer: robust and interpretable cancer survival prediction based on ConvMixer using pathway-level gene expression imagesabstractCancer is one of the leading causes of deaths worldwide. Survival analysis and prediction of cancer patients is of great significance for their precision medicine. The robustness and interpretability of the survival prediction models are important, where robustness tells whether a model has learned the knowledge, and interpretability means if a model can show human what it has learned. In this paper, we propose a robust and interpretable model SurvConvMixer, which uses pathways customized gene expression images and ConvMixer for cancer short-term, mid-term and long-term overall survival prediction. With ConvMixer, the representation of each pathway can be learned respectively. We show the robustness of our model by testing the trained model on absolutely untrained external datasets. The interpretability of SurvConvMixer depends on gradient-weighted class activation mapping (Grad-Cam), by which we can obtain the pathway-level activation heat map. Then wilcoxon rank-sum tests are conducted to obtain the statistically significant pathways, thereby revealing which pathways the model focuses on more. SurvConvMixer achieves remarkable performance on the short-term, mid-term and long-term overall survival of lung adenocarcinoma, lung squamous cell carcinoma and skin cutaneous melanoma, and the external validation tests show that SurvConvMixer can generalize to external datasets so that it is robust. Finally, we investigate the activation maps generated by Grad-Cam, after wilcoxon rank-sum test and Kaplan-Meier estimation, we find that some survival-related pathways play important role in SurvConvMixer. Yuanning Liu, Hao Zhang 0064 |
BMC Bioinform. | 3 |
| 2024 | GERWR: Identifying the Key Pathogenicity- Associated sRNAs of Magnaporthe Oryzae Infection in Rice Based on Graph Embedding and Random Walk With RestartabstractRice blast, caused by Magnaporthe oryzae(M.oryzae), is a destructive rice disease that reduces rice yield by 10% to 30% annually. It also affects other cereal crops such as barley, wheat, rye, millet, sorghum, and maize. Small RNAs (sRNAs) play an essential regulatory role in fungus-plant interaction during the fungal invasion, but studies on pathogenic sRNAs during the fungal invasion of plants based on multi-omics data integration are rare. This paper proposes a novel approach called Graph Embedding combined with Random Walk with Restart (GERWR) to identify pathogenic sRNAs based on multi-omics data integration during M.oryzae invasion. By constructing a multi-omics network (MRMO), we identified 29 pathogenic sRNAs of rice blast fungus. Further analysis revealed that these sRNAs regulate rice genes in a many-to-many relationship, playing a significant regulatory role in the pathogenesis of rice blast disease. This paper explores the pathogenic factors of rice blast disease from the perspective of multi-omics data analysis, revealing the inherent connection between pathogenic factors of different omics. It has essential scientific significance for studying the pathogenic mechanism of rice blast fungus, the rice blast fungus-rice model system, and the pathogen-host interaction in related fields. Hao Zhang 0064, Tianheng Zhao, Enshuang Zhao, Lanhui Li, Guihua Li, Borui Zhang, Qing-Ming Qin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2023 | An enhanced attentive implicit relation embedding for social recommendation
Xintao Ma, Liyan Dong, Yuequn Wang, Hao Zhang 0064 |
Data Knowl. Eng. | 6 |
| 2023 | AKUPP: attention-enhanced joint propagation of knowledge and user preference for recommendation systems
Xintao Ma, Liyan Dong, Yuequn Wang, Hao Zhang 0064 |
Knowl. Inf. Syst. | 5 |
| 2022 | LTPConstraint: a transfer learning based end-to-end method for RNA secondary structure predictionabstractBACKGROUND: RNA secondary structure is very important for deciphering cell's activity and disease occurrence. The first method which was used by the academics to predict this structure is biological experiment, But this method is too expensive, causing the promotion to be affected. Then, computing methods emerged, which has good efficiency and low cost. However, the accuracy of computing methods are not satisfactory. Many machine learning methods have also been applied to this area, but the accuracy has not improved significantly. Deep learning has matured and achieves great success in many areas such as computer vision and natural language processing. It uses neural network which is a kind of structure that has good functionality and versatility, but its effect is highly correlated with the quantity and quality of the data. At present, there is no model with high accuracy, low data dependence and high convenience in predicting RNA secondary structure. RESULTS: This paper designs a neural network called LTPConstraint to predict RNA secondary structure. The network is based on many network structure such as Bidirectional LSTM, Transformer and generator. It also uses transfer learning to train modelso that the data dependence can be reduced. CONCLUSIONS: LTPConstraint has achieved high accuracy in RNA secondary structure prediction. Compared with the previous methods, the accuracy improves obviously both in predicting the structure with pseudoknot and the structure without pseudoknot. At the same time, LTPConstraint is easy to operate and can achieve result very quickly. Yinchao Fei, Hao Zhang 0064, Yili Wang 0004, Yuanning Liu |
BMC Bioinform. | 2 |
| 2022 | SMRI: A New Method for siRNA Design for COVID-19 Therapy
Meng-Xin Chen, Xiaodong Zhu 0001, Hao Zhang 0064, Yuanning Liu |
J. Comput. Sci. Technol. | 3 |
| 2021 | A novel end-to-end method to predict RNA secondary structure profile based on bidirectional LSTM and residual neural networkabstractBACKGROUND: Studies have shown that RNA secondary structure, a planar structure formed by paired bases, plays diverse vital roles in fundamental life activities and complex diseases. RNA secondary structure profile can record whether each base is paired with others. Hence, accurate prediction of secondary structure profile can help to deduce the secondary structure and binding site of RNA. RNA secondary structure profile can be obtained through biological experiment and calculation methods. Of them, the biological experiment method involves two ways: chemical reagent and biological crystallization. The chemical reagent method can obtain a large number of prediction data, but its cost is high and always associated with high noise, making it difficult to get results of all bases on RNA due to the limited of sequencing coverage. By contrast, the biological crystallization method can lead to accurate results, yet heavy experimental work and high costs are required. On the other hand, the calculation method is CROSS, which comprises a three-layer fully connected neural network. However, CROSS can not completely learn the features of RNA secondary structure profile since its poor network structure, leading to its low performance. RESULTS: In this paper, a novel end-to-end method, named as "RPRes, was proposed to predict RNA secondary structure profile based on Bidirectional LSTM and Residual Neural Network. CONCLUSIONS: RPRes utilizes data sets generated by multiple biological experiment methods as the training, validation, and test sets to predict profile, which can compatible with numerous prediction requirements. Compared with the biological experiment method, RPRes has reduced the costs and improved the prediction efficiency. Compared with the state-of-the-art calculation method CROSS, RPRes has significantly improved performance. Xiaodan Zhong, Hao Zhang 0064, Yuanning Liu |
BMC Bioinform. | 4 |
| 2021 | GAEBic: A Novel Biclustering Analysis Method for miRNA-Targeted Gene Data Based on Graph Autoencoder
Hao Zhang 0064, Haowu Chang, Qing-Ming Qin, Borui Zhang, Xue-Qing Li, Tianheng Zhao |
J. Comput. Sci. Technol. | 2 |
| 2021 | ncRFP: A Novel end-to-end Method for Non-Coding RNAs Family Prediction Based on Deep LearningabstractEvidence has accumulated enough to prove non-coding RNAs (ncRNAs) play important roles in cellular biological processes and disease pathogenesis. High throughput techniques have produced a large number of ncRNAs whose function remains unknown. Since the accurate identification of ncRNAs family is helpful to the research of their function, it is of necessity and urgency to predict the family of each ncRNAs. Although several traditional excellent methods are applicable to predict the family of ncRNAs, their complex procedures or inaccurate performance remain major problems confronting us. The main idea of those methods is first to predict the secondary structure, and then identify ncRNAs family according to properties of the secondary structure. Unfortunately, the multi-step error superposition, especially the imperfection of RNA secondary structure prediction tools, maybe the cause of low accuracy. In this paper, a novel end-to-end method 'ncRFP' was proposed to complete the prediction task based on Deep Learning. Instead of predicting the secondary structure, ncRFP predicts the ncRNAs family by automatically extracting features from ncRNAs sequences. Compared with other methods, ncRFP not only simplifies the process but also improves accuracy. The source code of ncRFP can be available at https://github.com/linyuwangPHD/ncRFP. Shaoge Zheng, Hao Zhang 0064, Zhiyang Qiu, Xiaodan Zhong, Yuanning Liu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2016 | A New Method to Predict RNA Secondary Structure Based on RNA Folding SimulationabstractRNA plays an important role in various biological processes; hence, it is essential when determining the functions of RNA to research its secondary structures. So far, the accuracy of RNA secondary structure prediction remains an area in need of improvement. This paper presents a novel method for predicting RNA secondary structure based on an RNA folding simulation model. This model assumes that the process of RNA folding from the random coil state to full structure is staged and in every stage of folding, the final state of an RNA is determined by the optimal combination of helical regions, which are urgently essential to dynamics of RNA formation. This paper proposes the First Large Free Energy Difference (FLED) in order to find the helical regions most urgently needed for optimal final state formation among all the possible helical regions. Tests on the datasets with known structures from public databases demonstrate that our method can outperform other current RNA secondary structure prediction methods in terms of prediction accuracy. Yuanning Liu, Qi Zhao 0008, Hao Zhang 0064, Liyan Wei |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |