Haozhou Li

dblp:259/7263 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
12since 2021 · last 2026
0009-0001-0955-470XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Multimodal learning on heterogeneous subgraphs and LLMs representation for MHC-peptide binding affinity prediction
abstract
Accurate prediction of MHC-peptide binding affinity remains a challenge for immunotherapeutic development. Existing methods struggle to jointly model functional semantics of polymorphic residues, evolutionary conservation constraints, and structural dynamic. We propose the Contrast learning-based Multi-feature Heterogeneous Subgraph model (CMHS) with sequence and structural representation. For sequence representation, we introduce LoRA fine-tuning to obtain the MHC-exclusive sequence representation from ESM2, then jointly BLOSUM50 to capture long-range functional dependencies and evolutionarily conserved residues. For structural representation, we use the biophysics-guided heterogeneous graph network. Constructing an MHC-peptide graph with a novel trainable Gaussian noise layer guided by crystallographic B-factors to dynamically simulate electron density uncertainty, coupled with a three-stage message-passing framework with subgraph aggregation, subgraph extraction and heterogeneous. Finally, to align sequence and graph representation spaces, we use contrastive learning to obtain a more comprehensive representation and to enhance the ability of model prediction. Evaluations on 16 HLA allele benchmarks show average SRCC improvements of 8.7%, with improvements of average AUC of 7.6%. This work establishes a new paradigm for predicting hypervariable immune interactions. The corresponding code can be founded in github.
Ruimeng Li, Haozhou Li, Biyi Zhou, Qinke Peng
BMC Bioinform.3
2026 SEHLP: A summary-enhanced large language model for financial report sentiment analysis via hybrid LoRA and dynamic prefix tuning
Haozhou Li, Qinke Peng, Xu Mou, Zeyuan Zeng, Ruimeng Li, Jinzhi Wang, Wentong Sun
Inf. Process. Manag.1
2025 Building a Chinese Medical Dialogue System: Integrating Large-scale Corpora and Novel Models
abstract
The global COVID-19 pandemic underscored major deficiencies in traditional healthcare systems, hastening the advancement of online medical services, especially in medical triage and consultation. However, existing studies face two main challenges. First, the scarcity of large-scale, publicly available, domain-specific medical datasets due to privacy concerns, with current datasets being small and limited to a few diseases, limiting the effectiveness of triage methods based on Pre-trained Language Models (PLMs). Second, existing methods lack medical knowledge and struggle to accurately understand professional terms and expressions in patient-doctor consultations. To overcome these obstacles, we construct the Large-scale Chinese Medical Dialogue Corpora (LCMDC), thereby addressing the data shortage in this field. Moreover, we further propose a novel triage system that combines BERT-based supervised learning with prompt learning, as well as a GPT-based medical consultation model. To enhance domain knowledge acquisition, we pre-trained PLMs using our self-constructed background corpus. Experimental results on the LCMDC demonstrate the efficacy of our proposed systems. We further implement a voice-based medical dialogue system that enables real-time interaction between users and our models. Our corpora and codes are available at this link.
Xinyuan Wang 0011, Haozhou Li, Dingfang Zheng, Qinke Peng
IJCNN2
2025 CSTE: A Context-enhanced Speaker-aware Triple Encoding model for intelligent triage and diagnosis in medical dialogue
Haozhou Li, Qinke Peng, Xinyuan Wang 0011, Wentong Sun, Defu Li, Ruimeng Li
Inf. Process. Manag.1
2025 Compound Interaction Presentation Learning for MHC-Peptide Binding Affinity Prediction
abstract
The interaction between peptides and Major Histocompatibility Complex Class I (MHC-I) molecules plays a critical role in adaptive immune recognition. Although computational prediction algorithms have advanced over traditional experimental methods, challenges still remain. There is a scarcity of standardized datasets that provide comprehensive profiles of MHC-peptide structure. The polymorphism of MHC molecules introduces diverse binding patterns, complicating the characterization of specific amino acid interaction pairs. To address these issues, we introduce GSM, a novel deep learning model that combines a Graph Attention Neural Network with a Self-Attention Convolutional Neural Network to predict MHC-peptide binding affinities. By integrating self-attention mechanisms to capture global peptide-MHC interactions and graph-based modeling to represent local amino acid pairwise interactions, GSM provides a comprehensive understanding of binding mode. Compared to existing algorithms, GSM shows superior performance and greater stability across diverse allele datasets, as demonstrated on the benchmarks. Furthermore, by leveraging real 3D structural data and attention visualization, GSM is capacity of selecting interaction sites, offering valuable insights for vaccine design and advancing immunological research.
Ruimeng Li, Qinke Peng, Haozhou Li, Zeyuan Zeng
IEEE Trans. Comput. Biol. Bioinform.3
2025 Dynamic Graph Attention Meets Pretrained Language Models: Adaptive K-Mer Decomposition for LncRNA-Protein Interaction Prediction
abstract
Protein-RNA complexes, particularly those involving RNA-binding proteins and long non-coding RNAs (lncRNA), are commonly found to influence gene expression and mediate fundamental cellular processes. Despite significant advances in representations for these biological sequences, sequence decomposition based on k-mer generally results in fix-length substrings, failing to detect the information of variable-length biological functional regions. In this paper, we develop a concept of expressiveness for k-mer decompositions as a theoretical underpinning for traversing all k-mer decompositions. Based on this concept, we propose an advanced approach, BERTDGA-LPI, to detect the information of variable-length biological functional regions utilizing dynamic graph attention and to capture the influence of RNA and protein context leveraging pretrained language models. The experimental results demonstrate the outperformance of BERTDGA-LPI over state-of-the-art methods across two homo sapiens datasets, one plant species dataset, and two species-unspecific datasets. Furthermore, BERTDGA-LPI is validated as effective in predicting unknown RNA-protein interactions (RPI) with 100% prediction accuracy in six independent validation sets from different species. This study lays a theoretical underpinning for traversing all k-mer decompositions and innovatively offers a broadly applicable and efficient tool for LPI prediction and RPI prediction based only on sequences.
Zeyuan Zeng, Jingxian Zeng, Defu Li, Qinke Peng, Haozhou Li, Ruimeng Li, Wentong Sun, Jinzhi Wang
IEEE Trans. Comput. Biol. Bioinform.5
2025 A Finetuning Deep Learning Framework for Pan-Species Promoters Identification With Pseudo Time Series Analysis on Time and Frequency Space
abstract
Promoters are genomic sequences harbouring specific motifs, such as the TATA- box for eukaryotes and the Pribnow box for prokaryotes, which are known as regulatory elements. Accurate identification of these regulatory elements is essential for deciphering transcriptional regulation mechanisms. However, the heterogeneity of promoters across different species poses a significant challenge in this task. In our study, we introduce two deep learning methods, ProTriCNN and TransPro, designed for promoter identification. Based on promoter representation, ProTriCNN treats promoters as pseudo-time series, utilizing this approach to capture the intricate heterogeneity of promoter elements. TransPro is a ProTriCNN-based Fine-tuning framework to improve identification performance across different species. TransPro lies in utilizes elements and species evolutionary trees to represent the locality difference between source and target species across various levels and time-frequency space, respectively. With systematic experiments using real datasets, we demonstrate that ProTriCNN outperfroms state-of-the-art methods across all species, achieving an average accuracy improvement of 2.1% and a 20% enhancement in the Matthews coefficient. TransPro further attains accuracy improvement of the highest 8% and a 25% enhancement in the Matthews coefficient compared to ProTriCNN.
Ruimeng Li, Qinke Peng, Haozhou Li, Wentong Sun
IEEE J. Biomed. Health Informatics3
2024 SADE: A Speaker-Aware Dual Encoding Model Based on Diagbert for Medical Triage and Pre-Diagnosis
abstract
Medical triage and diagnosis systems are crucial in alleviating the severe burden on healthcare systems, providing great assistance and convenience to patients with limited medical knowledge. Most current studies, however, struggle to effectively learn speaker-aware information in patient-doctor dialogues. Besides, the professional terms and disease-related expressions make it difficult for traditional methods to understand medical knowledge. To solve these challenges, we propose a novel Speaker-Aware Dual Encoding (SADE) model, which employs two encoders of identical structure to separately represent patient consultations and doctor diagnoses. To grasp medical knowledge, we introduce DiagBERT, pre-trained on numerous domain-specific texts to represent input utterances. We then integrate BiLSTM and Dendrite Network into each encoder to capture sequential and interactional semantics. Moreover, we have built two large-scale Triage and Diagnosis datasets to overcome the scarcity of public medical corpora. Extensive experimental results on these datasets indicate that SADE outperforms other baseline models.
Haozhou Li, Xinyuan Wang 0011, Hongkai Du, Wentong Sun, Qinke Peng
ICASSP1
2024 Multi-document influence on readers: augmenting social emotion prediction by learning document interactions
Xu Mou, Qinke Peng, Muhammad Fiaz Bashir, Haozhou Li
Neural Comput. Appl.5
2024 SEHF: A Summary-Enhanced Hierarchical Framework for Financial Report Sentiment Analysis
abstract
Financial reports serve as crucial resources for investors and researchers, providing analysts’ assessments of stocks that play a vital role in stock market applications. However, detecting analysts’ opinions and sentiments in financial reports is challenging. First, the formal and professional language used in these reports makes it difficult for previous methods to comprehend domain-specific knowledge. Second, financial reports often adopt lengthy and elaborate expressions to convey rich semantics, which exposes the existing methods to contextual information loss, especially on long-term dependencies. To address these problems, we propose a summary-enhanced hierarchical framework (SEHF), which leverages summary information to enhance financial report sentiment analysis. Our framework incorporates financial bidirectional and auto-regressive transformer (FinBART), equipped with extended position encoding to summarize lengthy report articles and capture long-range interactions. To mitigate information loss, we initially divide each report into segments and then propose the hierarchical analyst sentiment representation network (ASRN), which utilizes financial bidirectional encoder representation from transformer (FinBERT), bidirectional long short-term memory (BiLSTM)-Attention, and dendrite (DD) network to fuse information in the generated summary and report segments. Notably, FinBART and FinBERT are pretrained on large-scale financial corpora to effectively understand professional expressions. Furthermore, we construct a new dataset large-scale Chinese financial report (LCFR) for the lack of supervised datasets. Experimental results on LCFR and a benchmark dataset show that SEHF significantly outperforms state-of-the-art (SOTA) baselines, and the ablation study highlights the effectiveness of aggregating sentiment information in the summary and report segments.
Haozhou Li, Qinke Peng, Xinyuan Wang 0011, Xu Mou, Yonghao Wang
IEEE Trans. Comput. Soc. Syst.1
2023 Abstractive Financial News Summarization via Transformer-BiLSTM Encoder and Graph Attention-Based Decoder
abstract
Financial news summarization (FNS) has been an attractive research problem in recent years, which aims to generate a shorter highlight of the news article while preserving key factual aspects, emotions, and opinions, providing significant assistance in stock trading and investment decision-making. However, FNS faces two challenges compared to the common domain. Firstly, financial news involves professional qualitative and quantitative information and salient content always scatters across long-range interactions. Secondly, financial news contains latent causal relationships, where historical information in the early generated sequence can significantly affect the subsequent decoding process. To address these difficulties, we propose an enhanced Seq2Seq model named TLGA, where the hierarchical Transformer-BiLSTM encoder can capture long-range interactions and sequential semantics while the Graph Attention-based decoder can fully utilize the historical information of decoded tokens and capture key causal relations. Moreover, we propose history-enhanced attention to concentrate on salient input content based on history semantics, guiding our decoder to generate the summary around the corresponding contents. It is also the first attempt to reuse history information of previously generated summary sequences in FNS using the idea of the Graph Attention Mechanism. Additionally, we construct the LCFNS dataset with 430,820 news-summary pairs for the lack of large-scale high-quality datasets in FNS. Experimental results on two financial datasets and two benchmark datasets indicate that our model outperforms other baselines.
Haozhou Li, Qinke Peng, Xu Mou, Ying Wang 0065, Zeyuan Zeng, Muhammad Fiaz Bashir
IEEE ACM Trans. Audio Speech Lang. Process.1
2022 A successful hybrid deep learning model aiming at promoter identification
abstract
BACKGROUND: The zone adjacent to a transcription start site (TSS), namely, the promoter, is primarily involved in the process of DNA transcription initiation and regulation. As a result, proper promoter identification is critical for further understanding the mechanism of the networks controlling genomic regulation. A number of methodologies for the identification of promoters have been proposed. Nonetheless, due to the great heterogeneity existing in promoters, the results of these procedures are still unsatisfactory. In order to establish additional discriminative characteristics and properly recognize promoters, we developed the hybrid model for promoter identification (HMPI), a hybrid deep learning model that can characterize both the native sequences of promoters and the morphological outline of promoters at the same time. We developed the HMPI to combine a method called the PSFN (promoter sequence features network), which characterizes native promoter sequences and deduces sequence features, with a technique referred to as the DSPN (deep structural profiles network), which is specially structured to model the promoters in terms of their structural profile and to deduce their structural attributes. RESULTS: The HMPI was applied to human, plant and Escherichia coli K-12 strain datasets, and the findings showed that the HMPI was successful at extracting the features of the promoter while greatly enhancing the promoter identification performance. In addition, after the improvements of synthetic sampling, transfer learning and label smoothing regularization, the improved HMPI models achieved good results in identifying subtypes of promoters on prokaryotic promoter datasets. CONCLUSIONS: The results showed that the HMPI was successful at extracting the features of promoters while greatly enhancing the performance of identifying promoters on both eukaryotic and prokaryotic datasets, and the improved HMPI models are good at identifying subtypes of promoters on prokaryotic promoter datasets. The HMPI is additionally adaptable to different biological functional sequences, allowing for the addition of new features or models.
Ying Wang 0065, Qinke Peng, Xu Mou, Xinyuan Wang 0011, Haozhou Li, Tian Han 0008, Xiao Wang 0007
BMC Bioinform.5
2019 A Deep Leaming Model with Multi-Scale Skip Connections for Solar Flare Prediction Combined with Prior Information
abstract
Solar flare prediction has been drawing increasing attention due to its impact on the global environment. However, its prediction remains computationally challenging. To address this, we develop a solar flare prediction model using Long Short-Term Memory(LSTM) and Deep Neural Network (DNN) with multi-scale skip connections, referred to as LSTM-DNN Flare Net (LDFN). Different from the existing models, LDFN is constructed based on the physical meanings of variables, which is in accordance with the human cognitive process and thus enhances the interpretability of the model. In order to improve training efficiency, we integrate both long and short skip connections in the construction of the DNN subnet. The experimental results demonstrate that LDFN can outperform traditional machine learning models and typical deep learning models in terms of internationally recognized standard metrics. Furthermore, the trained LDFN is used to rank the importance of top 10 variables affecting solar flares. Results reveal that total unsigned vertical current, fraction of area with shear > 45 and total unsigned flux around high gradient polarity inversion lines are the most informative variables for predicting solar flare, which is consistent with the results of current research.
Tian Han 0008, Qinke Peng, Yiqing Shen 0001, Haozhou Li
IEEE BigData4