VLDB 2026 Research / reviewers in the wild / expert
Qinke Peng
dblp:08/2357
· DBLP profile ↗
49ranked-venue papers
0as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 15 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Software engineering, systems software and programming languages · 2Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal learning on heterogeneous subgraphs and LLMs representation for MHC-peptide binding affinity predictionabstractAccurate prediction of MHC-peptide binding affinity remains a challenge for immunotherapeutic development. Existing methods struggle to jointly model functional semantics of polymorphic residues, evolutionary conservation constraints, and structural dynamic. We propose the Contrast learning-based Multi-feature Heterogeneous Subgraph model (CMHS) with sequence and structural representation. For sequence representation, we introduce LoRA fine-tuning to obtain the MHC-exclusive sequence representation from ESM2, then jointly BLOSUM50 to capture long-range functional dependencies and evolutionarily conserved residues. For structural representation, we use the biophysics-guided heterogeneous graph network. Constructing an MHC-peptide graph with a novel trainable Gaussian noise layer guided by crystallographic B-factors to dynamically simulate electron density uncertainty, coupled with a three-stage message-passing framework with subgraph aggregation, subgraph extraction and heterogeneous. Finally, to align sequence and graph representation spaces, we use contrastive learning to obtain a more comprehensive representation and to enhance the ability of model prediction. Evaluations on 16 HLA allele benchmarks show average SRCC improvements of 8.7%, with improvements of average AUC of 7.6%. This work establishes a new paradigm for predicting hypervariable immune interactions. The corresponding code can be founded in github. Ruimeng Li, Haozhou Li, Biyi Zhou, Qinke Peng |
BMC Bioinform. | 5 |
| 2026 | SEHLP: A summary-enhanced large language model for financial report sentiment analysis via hybrid LoRA and dynamic prefix tuning
Haozhou Li, Qinke Peng, Xu Mou, Zeyuan Zeng, Ruimeng Li, Jinzhi Wang, Wentong Sun |
Inf. Process. Manag. | 2 |
| 2026 | MACC: Masked Adversarial Convex Combination Against Word-Level Adversarial AttacksabstractRobustness againstword substitution attacks iscrucial for text classifiers and fundamental to broader NLP robustness. Such attacks use semantically similar word substitutions. Existing certified defenses often compute loose outer bounds for the convex hull of word embeddings, including irrelevant words, degrading performance on clean and adversarial tests. Additionally, convex hull-based methods also struggle to emulate worst-case scenarios accurately. This paper proposes theMasked Adversarial Convex Combination(MACC) method, which models the solution space as a convex hull of word vectors and uses Variational Information Bottleneck theory to mask unnecessary words. We also introduce an empirical masking method based on the volume of the convex hull to enhance the performance. By reducing the number of words explored within the convex hull, MACC enables more precise mimicry of worst-case attacks. Experiments across models and datasets show MACC outperforms existing methods in clean and adversarial accuracy against word-level attacks. Xu Mou, Bo Zeng 0001, Zhi-Hong Mao, Qinke Peng |
IEEE Signal Process. Lett. | 4 |
| 2025 | Building a Chinese Medical Dialogue System: Integrating Large-scale Corpora and Novel ModelsabstractThe global COVID-19 pandemic underscored major deficiencies in traditional healthcare systems, hastening the advancement of online medical services, especially in medical triage and consultation. However, existing studies face two main challenges. First, the scarcity of large-scale, publicly available, domain-specific medical datasets due to privacy concerns, with current datasets being small and limited to a few diseases, limiting the effectiveness of triage methods based on Pre-trained Language Models (PLMs). Second, existing methods lack medical knowledge and struggle to accurately understand professional terms and expressions in patient-doctor consultations. To overcome these obstacles, we construct the Large-scale Chinese Medical Dialogue Corpora (LCMDC), thereby addressing the data shortage in this field. Moreover, we further propose a novel triage system that combines BERT-based supervised learning with prompt learning, as well as a GPT-based medical consultation model. To enhance domain knowledge acquisition, we pre-trained PLMs using our self-constructed background corpus. Experimental results on the LCMDC demonstrate the efficacy of our proposed systems. We further implement a voice-based medical dialogue system that enables real-time interaction between users and our models. Our corpora and codes are available at this link. Xinyuan Wang 0011, Haozhou Li, Dingfang Zheng, Qinke Peng |
IJCNN | 4 |
| 2025 | CSTE: A Context-enhanced Speaker-aware Triple Encoding model for intelligent triage and diagnosis in medical dialogue
Haozhou Li, Qinke Peng, Xinyuan Wang 0011, Wentong Sun, Defu Li, Ruimeng Li |
Inf. Process. Manag. | 2 |
| 2025 | Compound Interaction Presentation Learning for MHC-Peptide Binding Affinity PredictionabstractThe interaction between peptides and Major Histocompatibility Complex Class I (MHC-I) molecules plays a critical role in adaptive immune recognition. Although computational prediction algorithms have advanced over traditional experimental methods, challenges still remain. There is a scarcity of standardized datasets that provide comprehensive profiles of MHC-peptide structure. The polymorphism of MHC molecules introduces diverse binding patterns, complicating the characterization of specific amino acid interaction pairs. To address these issues, we introduce GSM, a novel deep learning model that combines a Graph Attention Neural Network with a Self-Attention Convolutional Neural Network to predict MHC-peptide binding affinities. By integrating self-attention mechanisms to capture global peptide-MHC interactions and graph-based modeling to represent local amino acid pairwise interactions, GSM provides a comprehensive understanding of binding mode. Compared to existing algorithms, GSM shows superior performance and greater stability across diverse allele datasets, as demonstrated on the benchmarks. Furthermore, by leveraging real 3D structural data and attention visualization, GSM is capacity of selecting interaction sites, offering valuable insights for vaccine design and advancing immunological research. Ruimeng Li, Qinke Peng, Haozhou Li, Zeyuan Zeng |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | Dynamic Graph Attention Meets Pretrained Language Models: Adaptive K-Mer Decomposition for LncRNA-Protein Interaction PredictionabstractProtein-RNA complexes, particularly those involving RNA-binding proteins and long non-coding RNAs (lncRNA), are commonly found to influence gene expression and mediate fundamental cellular processes. Despite significant advances in representations for these biological sequences, sequence decomposition based on k-mer generally results in fix-length substrings, failing to detect the information of variable-length biological functional regions. In this paper, we develop a concept of expressiveness for k-mer decompositions as a theoretical underpinning for traversing all k-mer decompositions. Based on this concept, we propose an advanced approach, BERTDGA-LPI, to detect the information of variable-length biological functional regions utilizing dynamic graph attention and to capture the influence of RNA and protein context leveraging pretrained language models. The experimental results demonstrate the outperformance of BERTDGA-LPI over state-of-the-art methods across two homo sapiens datasets, one plant species dataset, and two species-unspecific datasets. Furthermore, BERTDGA-LPI is validated as effective in predicting unknown RNA-protein interactions (RPI) with 100% prediction accuracy in six independent validation sets from different species. This study lays a theoretical underpinning for traversing all k-mer decompositions and innovatively offers a broadly applicable and efficient tool for LPI prediction and RPI prediction based only on sequences. Zeyuan Zeng, Jingxian Zeng, Defu Li, Qinke Peng, Haozhou Li, Ruimeng Li, Wentong Sun, Jinzhi Wang |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | A Finetuning Deep Learning Framework for Pan-Species Promoters Identification With Pseudo Time Series Analysis on Time and Frequency SpaceabstractPromoters are genomic sequences harbouring specific motifs, such as the TATA- box for eukaryotes and the Pribnow box for prokaryotes, which are known as regulatory elements. Accurate identification of these regulatory elements is essential for deciphering transcriptional regulation mechanisms. However, the heterogeneity of promoters across different species poses a significant challenge in this task. In our study, we introduce two deep learning methods, ProTriCNN and TransPro, designed for promoter identification. Based on promoter representation, ProTriCNN treats promoters as pseudo-time series, utilizing this approach to capture the intricate heterogeneity of promoter elements. TransPro is a ProTriCNN-based Fine-tuning framework to improve identification performance across different species. TransPro lies in utilizes elements and species evolutionary trees to represent the locality difference between source and target species across various levels and time-frequency space, respectively. With systematic experiments using real datasets, we demonstrate that ProTriCNN outperfroms state-of-the-art methods across all species, achieving an average accuracy improvement of 2.1% and a 20% enhancement in the Matthews coefficient. TransPro further attains accuracy improvement of the highest 8% and a 25% enhancement in the Matthews coefficient compared to ProTriCNN. Ruimeng Li, Qinke Peng, Haozhou Li, Wentong Sun |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | SADE: A Speaker-Aware Dual Encoding Model Based on Diagbert for Medical Triage and Pre-DiagnosisabstractMedical triage and diagnosis systems are crucial in alleviating the severe burden on healthcare systems, providing great assistance and convenience to patients with limited medical knowledge. Most current studies, however, struggle to effectively learn speaker-aware information in patient-doctor dialogues. Besides, the professional terms and disease-related expressions make it difficult for traditional methods to understand medical knowledge. To solve these challenges, we propose a novel Speaker-Aware Dual Encoding (SADE) model, which employs two encoders of identical structure to separately represent patient consultations and doctor diagnoses. To grasp medical knowledge, we introduce DiagBERT, pre-trained on numerous domain-specific texts to represent input utterances. We then integrate BiLSTM and Dendrite Network into each encoder to capture sequential and interactional semantics. Moreover, we have built two large-scale Triage and Diagnosis datasets to overcome the scarcity of public medical corpora. Extensive experimental results on these datasets indicate that SADE outperforms other baseline models. Haozhou Li, Xinyuan Wang 0011, Hongkai Du, Wentong Sun, Qinke Peng |
ICASSP | 5 |
| 2024 | CGN: A Simple Yet Effective Multi-Channel Gated Network for Long-Term Time Series ForecastingabstractTransformers have gained widespread attention in the field of time series forecasting due to their exceptional capability to capture intricate interactions within sequences. However, as the sequence length increases, Transformer-based models face disadvantages such as high memory consumption, blurred long-range dependency, and susceptibility to overfitting. To address this, we introduce CGN, a lightweight multi-channel gated network. CGN employs a channel gate to identify and filter out noisy inputs, focusing on the most relevant variables. Unlike traditional Transformer-based models, our model is exceptionally lightweight, making it suitable for tasks requiring a longer historical window to enhance the accuracy of long-term forecasting. Experimental results demonstrate that our model outperforms the state-of-the-art Transformer-based and MLP-based models across six real-world LTSF benchmarks. Specifically, CGN surpasses the latest strongest Patch-Transformer in 90% cases with a significant reduction in the average number of trainable parameters, maximum memory consumption, running time, and inference time. Yulong Pei, Defu Li, Qinke Peng |
ICASSP | 4 |
| 2024 | ACLNDA: an asymmetric graph contrastive learning framework for predicting noncoding RNA-disease associations in heterogeneous graphsabstractNoncoding RNAs (ncRNAs), including long noncoding RNAs (lncRNAs) and microRNAs (miRNAs), play crucial roles in gene expression regulation and are significant in disease associations and medical research. Accurate ncRNA-disease association prediction is essential for understanding disease mechanisms and developing treatments. Existing methods often focus on single tasks like lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs), or lncRNA-miRNA interactions (LMIs), and fail to exploit heterogeneous graph characteristics. We propose ACLNDA, an asymmetric graph contrastive learning framework for analyzing heterophilic ncRNA-disease associations. It constructs inter-layer adjacency matrices from the original lncRNA, miRNA, and disease associations, and uses a Top-K intra-layer similarity edges construction approach to form a triple-layer heterogeneous graph. Unlike traditional works, to account for both node attribute features (ncRNA/disease) and node preference features (association), ACLNDA employs an asymmetric yet simple graph contrastive learning framework to maximize one-hop neighborhood context and two-hop similarity, extracting ncRNA-disease features without relying on graph augmentations or homophily assumptions, reducing computational cost while preserving data integrity. Our framework is capable of being applied to a universal range of potential LDA, MDA, and LMI association predictions. Further experimental results demonstrate superior performance to other existing state-of-the-art baseline methods, which shows its potential for providing insights into disease diagnosis and therapeutic target identification. The source code and data of ACLNDA is publicly available at https://github.com/AI4Bread/ACLNDA. Laiyi Fu, Yangyi Zhou, Qinke Peng, Hongqiang Lyu |
Briefings Bioinform. | 4 |
| 2024 | findGSEP: estimating genome size of polyploid species usingk-mer frequenciesabstractSUMMARY: Estimating genome size using k-mer frequencies, which plays a fundamental role in designing genome sequencing and analysis projects, has remained challenging for polyploid species, i.e., ploidy p > 2. To address this, we introduce "findGSEP," which is designed based on iterative curve fitting of k-mer frequencies. Precisely, it first disentangles up to p normal distributions by analyzing k-mer frequencies in whole genome sequencing of the focal species. Second, it computes the sizes of genomic regions related to 1∼p (homologous) chromosome(s) using each respective curve fitting, from which it infers the full polyploid and average haploid genome size. "findGSEP" can handle any level of ploidy p, and infer more accurate genome size than other well-known tools, as shown by tests using simulated and real genomic sequencing data of various species including octoploids. AVAILABILITY AND IMPLEMENTATION: "findGSEP" was implemented as a web server, which is freely available at http://146.56.237.198:3838/findGSEP/. Also, "findGSEP" was implemented as an R package for parallel processing of multiple samples. Source code and tutorial on its installation and usage is available at https://github.com/sperfu/findGSEP. Laiyi Fu, Yanxin Xie, ShunKang Ling, Ying Wang 0065, Binzhong Wang, Hejun Du, Qinke Peng, Hequan Sun |
Bioinform. | 7 |
| 2024 | Multi-document influence on readers: augmenting social emotion prediction by learning document interactions
Xu Mou, Qinke Peng, Muhammad Fiaz Bashir, Haozhou Li |
Neural Comput. Appl. | 2 |
| 2024 | KGRACDA: A Model Based on Knowledge Graph from Recursion and Attention Aggregation for CircRNA-Disease Association PredictionabstractCircRNA is closely related to human disease, so it is important to predict circRNA-disease association (CDA). However, the traditional biological detection methods have high difficulty and low accuracy, and computational methods represented by deep learning ignore the ability of the model to explicitly extract local depth information of the CDA. We propose a model based on knowledge graph from recursion and attention aggregation for circRNA-disease association prediction (KGRACDA). This model combines explicit structural features and implicit embedding information of graphs, optimizing graph embedding vectors. First, we built large-scale, multi-source heterogeneous datasets and construct a knowledge graph of multiple RNAs and diseases. After that, we use a recursive method to build multi-hop subgraphs and optimize graph attention mechanism by gating mechanism, mining local depth information. At the same time, the model uses multi-head attention mechanism to balance global and local depth features of graphs, and generate CDA prediction scores. KGRACDA surpasses other methods by capturing local and global depth information related to CDA. We update an interactive web platform HNRBase v2.0, which visualizes circRNA data, and allows users to download data and predict CDA using model. Ying Wang 0065, Maoyuan Ma, Yanxin Xie, Qinke Peng, Hongqiang Lyu, Hequan Sun, Laiyi Fu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | SEHF: A Summary-Enhanced Hierarchical Framework for Financial Report Sentiment AnalysisabstractFinancial reports serve as crucial resources for investors and researchers, providing analysts’ assessments of stocks that play a vital role in stock market applications. However, detecting analysts’ opinions and sentiments in financial reports is challenging. First, the formal and professional language used in these reports makes it difficult for previous methods to comprehend domain-specific knowledge. Second, financial reports often adopt lengthy and elaborate expressions to convey rich semantics, which exposes the existing methods to contextual information loss, especially on long-term dependencies. To address these problems, we propose a summary-enhanced hierarchical framework (SEHF), which leverages summary information to enhance financial report sentiment analysis. Our framework incorporates financial bidirectional and auto-regressive transformer (FinBART), equipped with extended position encoding to summarize lengthy report articles and capture long-range interactions. To mitigate information loss, we initially divide each report into segments and then propose the hierarchical analyst sentiment representation network (ASRN), which utilizes financial bidirectional encoder representation from transformer (FinBERT), bidirectional long short-term memory (BiLSTM)-Attention, and dendrite (DD) network to fuse information in the generated summary and report segments. Notably, FinBART and FinBERT are pretrained on large-scale financial corpora to effectively understand professional expressions. Furthermore, we construct a new dataset large-scale Chinese financial report (LCFR) for the lack of supervised datasets. Experimental results on LCFR and a benchmark dataset show that SEHF significantly outperforms state-of-the-art (SOTA) baselines, and the ablation study highlights the effectiveness of aggregating sentiment information in the summary and report segments. Haozhou Li, Qinke Peng, Xinyuan Wang 0011, Xu Mou, Yonghao Wang |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | Graph Neural Networks With High-Order Polynomial Spectral FiltersabstractGeneral graph neural networks (GNNs) implement convolution operations on graphs based on polynomial spectral filters. Existing filters with high-order polynomial approximations can detect more structural information when reaching high-order neighborhoods but produce indistinguishable representations of nodes, which indicates their inefficiency of processing information in high-order neighborhoods, resulting in performance degradation. In this article, we theoretically identify the feasibility of avoiding this problem and attribute it to overfitting polynomial coefficients. To cope with it, the coefficients are restricted in two steps, dimensionality reduction of the coefficients' domain and sequential assignment of the forgetting factor. We transform the optimization of coefficients to the tuning of a hyperparameter and propose a flexible spectral-domain graph filter, which significantly reduces the memory demand and the adverse impacts on message transmission under large receptive fields. Utilizing our filter, the performance of GNNs is improved significantly in large receptive fields and the receptive fields of GNNs are multiplied as well. Meanwhile, the superiority of applying a high-order approximation is verified across various datasets, notably in strongly hyperbolic datasets. Codes are publicly available at: https://github.com/cengzeyuan/TNNLS-FFKSF. Zeyuan Zeng, Qinke Peng, Xu Mou, Ying Wang 0065, Ruimeng Li |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Utilizing enhanced membership functions to improve the accuracy of a multi-inputs and single-output fuzzy system
Salah-ud-din Khokhar, Qinke Peng |
Appl. Intell. | 2 |
| 2023 | KGETCDA: an efficient representation learning framework based on knowledge graph encoder from transformer for predicting circRNA-disease associationsabstractRecent studies have demonstrated the significant role that circRNA plays in the progression of human diseases. Identifying circRNA-disease associations (CDA) in an efficient manner can offer crucial insights into disease diagnosis. While traditional biological experiments can be time-consuming and labor-intensive, computational methods have emerged as a viable alternative in recent years. However, these methods are often limited by data sparsity and their inability to explore high-order information. In this paper, we introduce a novel method named Knowledge Graph Encoder from Transformer for predicting CDA (KGETCDA). Specifically, KGETCDA first integrates more than 10 databases to construct a large heterogeneous non-coding RNA dataset, which contains multiple relationships between circRNA, miRNA, lncRNA and disease. Then, a biological knowledge graph is created based on this dataset and Transformer-based knowledge representation learning and attentive propagation layers are applied to obtain high-quality embeddings with accurately captured high-order interaction information. Finally, multilayer perceptron is utilized to predict the matching scores of CDA based on their embeddings. Our empirical results demonstrate that KGETCDA significantly outperforms other state-of-the-art models. To enhance user experience, we have developed an interactive web-based platform named HNRBase that allows users to visualize, download data and make predictions using KGETCDA with ease. The code and datasets are publicly available at https://github.com/jinyangwu/KGETCDA. Jinyang Wu 0001, Zhiwei Ning, Yidong Ding, Ying Wang 0065, Qinke Peng, Laiyi Fu |
Briefings Bioinform. | 5 |
| 2023 | Generic and scalable periodicity adaptation framework for time-series anomaly detection
Qinke Peng, Xu Mou, Muhammad Fiaz Bashir |
Multim. Tools Appl. | 2 |
| 2023 | Abstractive Financial News Summarization via Transformer-BiLSTM Encoder and Graph Attention-Based DecoderabstractFinancial news summarization (FNS) has been an attractive research problem in recent years, which aims to generate a shorter highlight of the news article while preserving key factual aspects, emotions, and opinions, providing significant assistance in stock trading and investment decision-making. However, FNS faces two challenges compared to the common domain. Firstly, financial news involves professional qualitative and quantitative information and salient content always scatters across long-range interactions. Secondly, financial news contains latent causal relationships, where historical information in the early generated sequence can significantly affect the subsequent decoding process. To address these difficulties, we propose an enhanced Seq2Seq model named TLGA, where the hierarchical Transformer-BiLSTM encoder can capture long-range interactions and sequential semantics while the Graph Attention-based decoder can fully utilize the historical information of decoded tokens and capture key causal relations. Moreover, we propose history-enhanced attention to concentrate on salient input content based on history semantics, guiding our decoder to generate the summary around the corresponding contents. It is also the first attempt to reuse history information of previously generated summary sequences in FNS using the idea of the Graph Attention Mechanism. Additionally, we construct the LCFNS dataset with 430,820 news-summary pairs for the lack of large-scale high-quality datasets in FNS. Experimental results on two financial datasets and two benchmark datasets indicate that our model outperforms other baselines. Haozhou Li, Qinke Peng, Xu Mou, Ying Wang 0065, Zeyuan Zeng, Muhammad Fiaz Bashir |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | A Deep Learning Framework for News Readers' Emotion Prediction Based on Features From News Article and Pseudo CommentsabstractWith the rapid development of the Internet, readers tend to share their views and emotions about news events. Predicting these emotions provides a vital role in social media applications (e.g., sentiment retrieval, opinion summary, and election prediction). However, news articles usually consist of objective texts that lack emotion words, making emotion prediction challenging. From prior studies, we know that comments that come directly from readers are full of emotions. Therefore, in this article, we propose a deep learning framework that first merges article and comment information to predict readers' emotions. At the same time, in the prediction process, we design a pseudo comment representation for unpublished news articles by the comments of published news. In addition, a better model is required to encode articles that contain implicit emotions. To solve this problem, we propose a block emotion attention network (BEAN) to encode news articles better. It includes an emotion attention mechanism and a hierarchical structure to capture emotion words and generate structural information during encoding. Experiments performed on three public datasets show that BEAN achieves the state-of-the-art average Pearson (AP) and accuracy (Acc@1). Moreover, results on four self-collected datasets show that both the introduction of emotional comments and BEAN in our framework improve the ability to predict readers' emotions. Xu Mou, Qinke Peng, Ying Wang 0065, Muhammad Fiaz Bashir |
IEEE Trans. Cybern. | 2 |
| 2023 | BertNDA: A Model Based on Graph-Bert and Multi-Scale Information Fusion for ncRNA-Disease Association PredictionabstractNon-coding RNAs (ncRNAs) are a class of RNA molecules that lack the ability to encode proteins in human cells, but play crucial roles in various biological process. Understanding the interactions between different ncRNAs and their impact on diseases can significantly contribute to diagnosis, prevention, and treatment of diseases. However, predicting tertiary interactions between ncRNAs and diseases based on structural information in multiple scales remains a challenging task. To address this challenge, we propose a method called BertNDA, aiming to predict potential relationships between miRNAs, lncRNAs, and diseases. The framework identifies the local information through connectionless subgraph, which aggregate neighbor nodes' feature. And global information is extracted by leveraging Laplace transform of graph structures and WL (Weisfeiler-Lehman) absolute role coding. Additionally, an EMLP (Element-wise MLP) structure is designed to fuse pairwise global information. The transformer-encoder is employed as the backbone of our approach, followed by a prediction-layer to output the final correlation score. Extensive experiments demonstrate that BertNDA outperforms state-of-the-art methods in prediction assignment and exhibits significant potential for various biological applications. Moreover, we develop an online prediction platform that incorporates the prediction model, providing users with an intuitive and interactive experience. Overall, our model offers an efficient, accurate, and comprehensive tool for predicting tertiary associations between ncRNAs and diseases. Zhiwei Ning, Jinyang Wu 0001, Yidong Ding, Ying Wang 0065, Qinke Peng, Laiyi Fu |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | LncDLSM: Identification of Long Non-Coding RNAs With Deep Learning-Based Sequence ModelabstractLong non-coding RNAs (LncRNAs) serve a vital role in regulating gene expressions and other biological processes. Differentiation of lncRNAs from protein-coding transcripts helps researchers dig into the mechanism of lncRNA formation and its downstream regulations related to various diseases. Previous works have been proposed to identify lncRNAs, including traditional bio-sequencing and machine learning approaches. Considering the tedious work of biological characteristic-based feature extraction procedures and inevitable artifacts during bio-sequencing processes, those lncRNA detection methods are not always satisfactory. Hence, in this work, we presented lncDLSM, a deep learning-based framework differentiating lncRNA from other protein-coding transcripts without dependencies on prior biological knowledge. lncDLSM is a helpful tool for identifying lncRNAs compared with other biological feature-based machine learning methods and can be applied to other species by transfer learning achieving satisfactory results. Further experiments showed that different species display distinct boundaries among distributions corresponding to the homology and the specificity among species, respectively. Ying Wang 0065, Hongkai Du, Yingxin Cao, Qinke Peng, Laiyi Fu |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | The regulatory network analysis of single-cell RNA-seq and single-cell ATAC-seq data using differential equation models with perturbationsabstractA recent study presented a new differential equation model, MultiVelo, which simultaneously performs on chromatin accessibility and gene expression from single-cell data to reveals regulatory network of epigenomic and transcriptomic changes. However, it is unclear how regulatory network changes with perturbations. Here, we extend the ordinal differential equation models to time-delay differential equation models with controller to investigate the stability of regulatory network with perturbations. Hang Mi, Qinke Peng |
BIBM | 3 |
| 2022 | A successful hybrid deep learning model aiming at promoter identificationabstractBACKGROUND: The zone adjacent to a transcription start site (TSS), namely, the promoter, is primarily involved in the process of DNA transcription initiation and regulation. As a result, proper promoter identification is critical for further understanding the mechanism of the networks controlling genomic regulation. A number of methodologies for the identification of promoters have been proposed. Nonetheless, due to the great heterogeneity existing in promoters, the results of these procedures are still unsatisfactory. In order to establish additional discriminative characteristics and properly recognize promoters, we developed the hybrid model for promoter identification (HMPI), a hybrid deep learning model that can characterize both the native sequences of promoters and the morphological outline of promoters at the same time. We developed the HMPI to combine a method called the PSFN (promoter sequence features network), which characterizes native promoter sequences and deduces sequence features, with a technique referred to as the DSPN (deep structural profiles network), which is specially structured to model the promoters in terms of their structural profile and to deduce their structural attributes. RESULTS: The HMPI was applied to human, plant and Escherichia coli K-12 strain datasets, and the findings showed that the HMPI was successful at extracting the features of the promoter while greatly enhancing the promoter identification performance. In addition, after the improvements of synthetic sampling, transfer learning and label smoothing regularization, the improved HMPI models achieved good results in identifying subtypes of promoters on prokaryotic promoter datasets. CONCLUSIONS: The results showed that the HMPI was successful at extracting the features of promoters while greatly enhancing the performance of identifying promoters on both eukaryotic and prokaryotic datasets, and the improved HMPI models are good at identifying subtypes of promoters on prokaryotic promoter datasets. The HMPI is additionally adaptable to different biological functional sequences, allowing for the addition of new features or models. Ying Wang 0065, Qinke Peng, Xu Mou, Xinyuan Wang 0011, Haozhou Li, Tian Han 0008, Xiao Wang 0007 |
BMC Bioinform. | 2 |
| 2022 | Dynamical Failures Driven by False Load Injection Attacks Against Smart GridabstractExtensive studies have revealed that smart grid is vulnerable to the cyber-physical attacks. However, these strategies only focus on the cascading initiation phase to induce single-stage failures with multiple branch tripping, lacking of exploring the attack effectiveness in the propagation phase so that the deeply hidden cascading failures are underestimated. In this paper, we propose a novel false load injection attack strategy that can intentionally penetrate into the cascading propagation phase to drive the multi-stage dynamical failures with a cascading process. Specifically, we formulate a bi-level optimization problem to model the adversarial game between the operator and the attacker. The former is in charge of security-constrained economic dispatching to minimize the generation cost, and the latter aims to maximize the cumulative number of tripped branches. Further, we reformulate this NP-hard bi-level problem as a mixed integer linear program for tractable computation. Finally, we perform numerical simulations on different-scale IEEE test systems to validate our strategy in driving the dynamical failures. Datian Peng, Jianmin Dong 0001, Qinke Peng |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | SAILER: scalable and accurate invariant representation learning for single-cell ATAC-seq processing and integrationabstractMOTIVATION: Single-cell sequencing assay for transposase-accessible chromatin (scATAC-seq) provides new opportunities to dissect epigenomic heterogeneity and elucidate transcriptional regulatory mechanisms. However, computational modeling of scATAC-seq data is challenging due to its high dimension, extreme sparsity, complex dependencies and high sensitivity to confounding factors from various sources. RESULTS: Here, we propose a new deep generative model framework, named SAILER, for analyzing scATAC-seq data. SAILER aims to learn a low-dimensional nonlinear latent representation of each cell that defines its intrinsic chromatin state, invariant to extrinsic confounding factors like read depth and batch effects. SAILER adopts the conventional encoder-decoder framework to learn the latent representation but imposes additional constraints to ensure the independence of the learned representations from the confounding factors. Experimental results on both simulated and real scATAC-seq datasets demonstrate that SAILER learns better and biologically more meaningful representations of cells than other methods. Its noise-free cell embeddings bring in significant benefits in downstream analyses: clustering and imputation based on SAILER result in 6.9% and 18.5% improvements over existing methods, respectively. Moreover, because no matrix factorization is involved, SAILER can easily scale to process millions of cells. We implemented SAILER into a software package, freely available to all for large-scale scATAC-seq data analysis. AVAILABILITY AND IMPLEMENTATION: The software is publicly available at https://github.com/uci-cbcl/SAILER. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yingxin Cao, Laiyi Fu, Qinke Peng, Qing Nie, Xiaohui Xie |
Bioinform. | 4 |
| 2021 | Efficient and effective control of confounding in eQTL mapping studies through joint differential expression and Mendelian randomization analysesabstractMOTIVATION: Identifying cis-acting genetic variants associated with gene expression levels-an analysis commonly referred to as expression quantitative trait loci (eQTLs) mapping-is an important first step toward understanding the genetic determinant of gene expression variation. Successful eQTL mapping requires effective control of confounding factors. A common method for confounding effects control in eQTL mapping studies is the probabilistic estimation of expression residual (PEER) analysis. PEER analysis extracts PEER factors to serve as surrogates for confounding factors, which is further included in the subsequent eQTL mapping analysis. However, it is computationally challenging to determine the optimal number of PEER factors used for eQTL mapping. In particular, the standard approach to determine the optimal number of PEER factors examines one number at a time and chooses a number that optimizes eQTLs discovery. Unfortunately, this standard approach involves multiple repetitive eQTL mapping procedures that are computationally expensive, restricting its use in large-scale eQTL mapping studies that being collected today. RESULTS: Here, we present a simple and computationally scalable alternative, Effect size Correlation for COnfounding determination (ECCO), to determine the optimal number of PEER factors used for eQTL mapping studies. Instead of performing repetitive eQTL mapping, ECCO jointly applies differential expression analysis and Mendelian randomization analysis, leading to substantial computational savings. In simulations and real data applications, we show that ECCO identifies a similar number of PEER factors required for eQTL mapping analysis as the standard approach but is two orders of magnitude faster. The computational scalability of ECCO allows for optimized eQTL discovery across 48 GTEx tissues for the first time, yielding an overall 5.89% power gain on the number of eQTL harboring genes (eGenes) discovered as compared to the previous GTEx recommendation that does not attempt to determine tissue-specific optimal number of PEER factors. AVAILABILITYAND IMPLEMENTATION: Our method is implemented in the ECCO software, which, along with its GTEx mapping results, is freely available at www.xzlab.org/software.html. All R scripts used in this study are also available at this site. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Huanhuan Zhu, Yanyi Song, Qinke Peng |
Bioinform. | 4 |
| 2020 | An Effective Hybrid Deep Learning Model for Eukaryotic Promoter IdentificationabstractThe promoter is a region located near the transcription start site (TSS) and responsible for the initiation and regulation of DNA transcription. Hence, accurate identification of promoters is essential for further building and understanding the mechanism of genetic regulatory networks. Numerous approaches for eukaryotic promoter identification were proposed. Nevertheless, the performances of these approaches are still unsatisfactory due to the variety nature of promoters. To extract more discriminative features and accurately identify eukaryotic promoters, here, we develop an effective hybrid deep learning model HDLMepi, which is able to characterize the original promoter sequences and the structural profiles of promoters simultaneously. We integrate the method we name PromoterClCce which characterizes the original promoter sequences and extracts sequence features, with an approach DSPN, which we design to model the structural profile of promoters and extract structure features, in HDLMepi for precisely eukaryotic promoter identification. We apply HDLMepi on both human and plants datasets and the experimental results demonstrate it is effective in promoter features extraction and can improve the performance of promoter identification significantly. HDLMepi is also open to add new features or new models and can be applied to other biology functional sequences. Ying Wang 0065, Qinke Peng, Xu Mou, Tian Han 0008, Xiao Wang 0007 |
BIBM | 2 |
| 2020 | Overloaded Branch Chains Induced by False Data Injection Attack in Smart GridabstractIncreasing integration of the information communication technology introduces security vulnerabilities into smart grid. False data injection attack can exploit the cyber vulnerabilities to compromise the physical grid without being detected. However, the existing attack strategies only consider the electrical characteristic of power flow calculation to induce multiple overloaded branches. On this basis, our attack strategy absorbs the topological adjacency of chain subgraph to induce the overloaded branch chains, which can better characterize the sequential propagation of cascading failures. To this end, a game-theoretical bilevel optimization problem is formulated to model the decision-making interaction between the upper-level attacker and the lower-level security constrained economic dispatching system. The simulation results in IEEE test systems verify that our attack strategy can induce the overloaded branch chains and justify that the overloaded branch chains are more effective than the overloaded branches in triggering the cascading failures. Datian Peng, Jianmin Dong 0001, Qinke Peng |
IEEE Signal Process. Lett. | 3 |
| 2020 | Predicting DNA Methylation States with Hybrid Information Based Deep-Learning ModelabstractDNA methylation plays an important role in the regulation of some biological processes. Up to now, with the development of machine learning models, there are several sequence-based deep learning models designed to predict DNA methylation states, which gain better performance than traditional methods like random forest and SVM. However, convolutional network based deep learning models that use one-hot encoding DNA sequence as input may discover limited information and cause unsatisfactory prediction performance, so more data and model structures of diverse angles should be considered. In this work, we proposed a hybrid sequence-based deep learning model with both MeDIP-seq data and Histone information to predict DNA methylated CpG states (MHCpG). We combined both MeDIP-seq data and histone modification data with sequence information and implemented convolutional network to discover sequence patterns. In addition, we used statistical data gained from previous three input data and adopted a 3-layer feedforward neuron network to extract more high-level features. We compared our method with traditional predicting methods using random forest and other previous methods like CpGenie and DeepCpG, the result showed that MHCpG exceeded the other approaches and gained more satisfactory performance. Laiyi Fu, Qinke Peng, Ling Chai |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2019 | A Deep Leaming Model with Multi-Scale Skip Connections for Solar Flare Prediction Combined with Prior InformationabstractSolar flare prediction has been drawing increasing attention due to its impact on the global environment. However, its prediction remains computationally challenging. To address this, we develop a solar flare prediction model using Long Short-Term Memory(LSTM) and Deep Neural Network (DNN) with multi-scale skip connections, referred to as LSTM-DNN Flare Net (LDFN). Different from the existing models, LDFN is constructed based on the physical meanings of variables, which is in accordance with the human cognitive process and thus enhances the interpretability of the model. In order to improve training efficiency, we integrate both long and short skip connections in the construction of the DNN subnet. The experimental results demonstrate that LDFN can outperform traditional machine learning models and typical deep learning models in terms of internationally recognized standard metrics. Furthermore, the trained LDFN is used to rank the importance of top 10 variables affecting solar flares. Results reveal that total unsigned vertical current, fraction of area with shear > 45 and total unsigned flux around high gradient polarity inversion lines are the most informative variables for predicting solar flare, which is consistent with the results of current research. Tian Han 0008, Qinke Peng, Yiqing Shen 0001, Haozhou Li |
IEEE BigData | 2 |
| 2019 | Predicting Social Emotions from Readers' PerspectiveabstractDue to the rapid development of Web, large numbers of documents assigned by readers' emotions have been generated through new portals. Comparing to the previous studies which focused on author's perspective, our research focuses on readers' emotions invoked by news articles. Our research provides meaningful assistance in social media application such as sentiment retrieval, opinion summarization and election prediction. In this paper, we predict the readers' emotion of news based on the social opinion network. More specifically, we construct the opinion network based on the semantic distance. The communities in the news network indicate specific events which are related to the emotions. Therefore, the opinion network serves as the lexicon between events and corresponding emotions. We leverage neighbor relationship in network to predict readers' emotions. As a result, our methods obtain better result than the state-of-the-art methods. Moreover, we developed a growing strategy to prune the network for practical application. The experiment verifies the rationality of the reduction for application. Qinke Peng, Ling Chai, Ying Wang 0065 |
IEEE Trans. Affect. Comput. | 2 |
| 2019 | Attentive Aspect Modeling for Review-Aware RecommendationabstractIn recent years, many studies extract aspects from user reviews and integrate them with ratings for improving the recommendation performance. The common aspects mentioned in a user’s reviews and a product’s reviews indicate indirect connections between the user and product. However, these aspect-based methods suffer from two problems. First, the common aspects are usually very sparse, which is caused by the sparsity of user-product interactions and the diversity of individual users’ vocabularies. Second, a user’s interests on aspects could be different with respect to different products, which are usually assumed to be static in existing methods. In this article, we propose an Attentive Aspect-based Recommendation Model (AARM) to tackle these challenges. For the first problem, to enrich the aspect connections between user and product, besides common aspects, AARM also models the interactions between synonymous and similar aspects. For the second problem, a neural attention network which simultaneously considers user, product, and aspect information is constructed to capture a user’s attention toward aspects when examining different products. Extensive quantitative and qualitative experiments show that AARM can effectively alleviate the two aforementioned problems and significantly outperforms several state-of-the-art recommendation methods on the top-N recommendation task. Zhiyong Cheng 0001, Xiangnan He 0001, Yongfeng Zhang 0003, Zhibo Zhu, Qinke Peng, Tat-Seng Chua |
ACM Trans. Inf. Syst. | 6 |
| 2017 | A novel clustering oriented closeness measure based on neighborhood chainabstractCloseness measures are crucial to clustering methods. In most traditional clustering methods, the closeness between data points or clusters is measured by the geometric distances alone. These metrics quantify the closeness only based on the concerned data points' positions in the feature space, and they might cause problems when dealing with clustering tasks with arbitrary clusters shapes and different clusters scales (varying clusters densities). In this paper, a novel Closeness Measure between data points based on Neighborhood Chain (CMNC) is proposed. Instead of using geometric distances alone, CMNC measures the closeness between data points by quantifying the difficulty for one data point to reach another through a chain of neighbors. Experimental results show that by substituting the geometric-distances-based closeness measures with CMNC, modified versions of the traditional clustering algorithms (e.g. k-means, single-link and CURE) perform much better than their original versions, especially when dealing with clustering tasks with clusters having arbitrary shapes and different scales. Shaoyi Liang, Deqiang Han, Qinke Peng |
IJCNN | 4 |
| 2017 | A high-order representation and classification method for transcription factor binding sites recognition in Escherichia coli
Shiquan Sun, Xiongpan Zhang, Qinke Peng |
Artif. Intell. Medicine | 3 |
| 2016 | Inferring gene regulatory networks based on spline regression and Bayesian group lassoabstractWe propose a fully Bayesian model based on B-splines and Bayesian group lasso with a spike and slab prior, to inference gene regulatory networks from time series data, where the regulation relationships are nonlinear. The results show that our method simulated in DREAM4 time series data outperforms other linear models and can recognize regulation relationship while the linear models cannot in real data. This method is well suited for the inference of gene regulatory networks with nonlinear relationship. Qinke Peng |
SNPD | 2 |
| 2016 | A multi-objective heuristic algorithm for gene expression microarray data classification
Jia Lv, Qinke Peng |
Expert Syst. Appl. | 2 |
| 2016 | A prediction model of post subjects based on information lifecycle in forum
Qinke Peng, Jia Lv |
Inf. Sci. | 2 |
| 2016 | Global feature selection from microarray data using Lagrange multipliers
Shiquan Sun, Qinke Peng |
Knowl. Based Syst. | 2 |
| 2015 | GT-kernelPLS: Game theory based hybrid gene selection method for microarray data classificationabstractGene selection prior to classification has been an important topic in bioinformatics, since last decade. Small sample size and high dimensionality in microarray data pose great challenges for performing efficient classification. In this paper we propose efficient hybrid method (GTkernelPLS) with a combination of wrapper like technique coalitional game theory and kernel partial least square (kernelPLS) filter method. Experimental results on ten microarray data sets ensure that GTkernelPLS achieve higher accuracy than several state of the art feature selection methods, and it exhibits a reasonable execution time, even for the data sets having more than twenty thousand genes. Adnan Shakoor, Qinke Peng, Shiquan Sun, Xiao Wang 0007, Jia Lv |
SNPD | 2 |
| 2014 | A neighbor selection method based on network community detection for collaborative filteringabstractThe neighbor selection that determines which users are exploited to estimate a target user's ratings has an important influence on the accuracy of recommendations of collaborative filtering based recommender system. Two kinds of ways for neighbor selection: KNN and cluster-based, are lack of specificity which refers to selecting different appropriate neighbors for different given target users, and thus limit the accuracy of recommendation. Therefore, in this paper, firstly, we propose a method that employs the evolutionary algorithm to optimize neighbors for all target users. Secondly, overcoming the high time complexity of the first one, we present another approach in which community detection algorithm is utilized as a preprocessing, and then the evolution algorithm is employed to optimize the neighborhood size for every community. We present experiments on a standard benchmark data-set, and the results show that the two methods both realize the specificity in neighbor selection, and accordingly lead to a higher accuracy of recommendations. Besides, the second one makes a good compromise between the specificity and time complexity. Qinke Peng |
ICIS | 2 |
| 2014 | A Hopfield neural network based algorithm for haplotype assembly from low-quality dataabstractThe objective of the haplotype assembly problem is to conclude a pair of haplotypes from a set of aligned single nucleotide polymorphism (SNP) fragments from a single individual. Errors in the SNP fragments, which are inevitable in the real-world application, severely increase the difficulty of the problem. As a result, most methods could not get accurate haplotypes on the data with high error rate. In this paper, we introduce a Hopfield neural network based method, named HNHap, to solve the haplotype assembly problem. Hopfield neural network is a very promising and effective approach to solve the combinatorial optimization problem. The stochastic optimal competitive Hopfield network model that has the mechanism to escape from the local optimum is a great improvement for the original model. Thus we map the haplotype assembly problem onto the stochastic optimal competitive Hopfield network model, in which a group of neurons correspond to an SNP fragment and the states of neurons denote the classification of the fragment. We also design a proper energy function based on the minimum error correction model for the haplotype assembly problem. We compare HNHap with other algorithms and the experiment results show that HNHap is an effective method to solve the haplotype assembly problem, especially on data with high error rate. Qinke Peng, LiBin Han, Xiao Wang 0007 |
IJCNN | 2 |
| 2014 | Analysis of disease association and susceptibility for SNP data using emotional neural networksabstractThe risk of some complex diseases are likely related to single nucleotide polymorphisms (SNPs), which are the most common form of DNA variations. Rapidly developing bioinformatics have made it possible to recognize a group of SNPs as the risk/protective factors of a specific disease, which are related to the possibility of the sample be infected. However, a particular algorithm to consider this kind of tendency information together is still in need. In this paper, inspired form the process that human beings to make a decision, we regard the risk/protect factor in the gene variations as the emotional of our nervous system. In this way, we regard these SNP combination factor as prior knowledge and use the emotional neural networks (ENN) to analysis the disease susceptibility. By sending this kind of information to ENN and using particle swarm optimization with hierarchical structure (PSO_HS) to train the parameters, we get a better result of susceptibility classification. The experimental results about real dataset shows that consider the risk/protect factor by emotional neural networks improve the performance of disease susceptibility analysis. Xiao Wang 0007, Qinke Peng |
IJCNN | 2 |
| 2012 | Identifying the semantic orientation of terms using S-HAL for sentiment analysis
Qinke Peng, Yinzhao Cheng |
Knowl. Based Syst. | 2 |
| 2010 | Splice sites prediction of Human genome using length-variable Markov model and feature selection
Quanwei Zhang, Qinke Peng, Qi Zhang 0011, Yanhua Yan, Kankan Li, Jing Li 0024 |
Expert Syst. Appl. | 2 |
| 2006 | Concaving Space Algorithm for Neural NetworkabstractThe global optimum solution, convergence rate are much more concerned for neural network training process. The efficient use of information resources works noticeably in the design of training methods. In this paper, not only gradient information but also Hessian matrix resource is applied for improving the learning efficiency and stability of neural network The necessary and sufficient condition of semi-positive definite Hessian matrix is gained .So the connecting weight matrix W can be revised within concave domains. Hence, the stable distribution of weights can be reached. The algorithm guides the training process to developing towards optimum goal. Compared with standard gradient method, the oscillating divergent phenomenon is avoided. The convergence of the algorithm is accelerated. Qinke Peng, YongXuan Huang |
SMC | 2 |
| 2005 | Remote Controller Design of Networked Control Systems Based on Self-constructing Fuzzy Neural Network
Qinke Peng |
ISNN (3) | 2 |
| 2004 | An adjustable algorithm for color quantization
Qinke Peng |
Pattern Recognit. Lett. | 3 |