Ying Wang 0065

dblp:94/3104-65 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
9since 2021 · last 2024
0009-0006-8959-8345ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021
YearPublicationVenuePosition
2024 findGSEP: estimating genome size of polyploid species usingk-mer frequencies
abstract
SUMMARY: Estimating genome size using k-mer frequencies, which plays a fundamental role in designing genome sequencing and analysis projects, has remained challenging for polyploid species, i.e., ploidy p > 2. To address this, we introduce "findGSEP," which is designed based on iterative curve fitting of k-mer frequencies. Precisely, it first disentangles up to p normal distributions by analyzing k-mer frequencies in whole genome sequencing of the focal species. Second, it computes the sizes of genomic regions related to 1∼p (homologous) chromosome(s) using each respective curve fitting, from which it infers the full polyploid and average haploid genome size. "findGSEP" can handle any level of ploidy p, and infer more accurate genome size than other well-known tools, as shown by tests using simulated and real genomic sequencing data of various species including octoploids. AVAILABILITY AND IMPLEMENTATION: "findGSEP" was implemented as a web server, which is freely available at http://146.56.237.198:3838/findGSEP/. Also, "findGSEP" was implemented as an R package for parallel processing of multiple samples. Source code and tutorial on its installation and usage is available at https://github.com/sperfu/findGSEP.
Laiyi Fu, Yanxin Xie, ShunKang Ling, Ying Wang 0065, Binzhong Wang, Hejun Du, Qinke Peng, Hequan Sun
Bioinform.4
2024 KGRACDA: A Model Based on Knowledge Graph from Recursion and Attention Aggregation for CircRNA-Disease Association Prediction
abstract
CircRNA is closely related to human disease, so it is important to predict circRNA-disease association (CDA). However, the traditional biological detection methods have high difficulty and low accuracy, and computational methods represented by deep learning ignore the ability of the model to explicitly extract local depth information of the CDA. We propose a model based on knowledge graph from recursion and attention aggregation for circRNA-disease association prediction (KGRACDA). This model combines explicit structural features and implicit embedding information of graphs, optimizing graph embedding vectors. First, we built large-scale, multi-source heterogeneous datasets and construct a knowledge graph of multiple RNAs and diseases. After that, we use a recursive method to build multi-hop subgraphs and optimize graph attention mechanism by gating mechanism, mining local depth information. At the same time, the model uses multi-head attention mechanism to balance global and local depth features of graphs, and generate CDA prediction scores. KGRACDA surpasses other methods by capturing local and global depth information related to CDA. We update an interactive web platform HNRBase v2.0, which visualizes circRNA data, and allows users to download data and predict CDA using model.
Ying Wang 0065, Maoyuan Ma, Yanxin Xie, Qinke Peng, Hongqiang Lyu, Hequan Sun, Laiyi Fu
IEEE ACM Trans. Comput. Biol. Bioinform.1
2024 Graph Neural Networks With High-Order Polynomial Spectral Filters
abstract
General graph neural networks (GNNs) implement convolution operations on graphs based on polynomial spectral filters. Existing filters with high-order polynomial approximations can detect more structural information when reaching high-order neighborhoods but produce indistinguishable representations of nodes, which indicates their inefficiency of processing information in high-order neighborhoods, resulting in performance degradation. In this article, we theoretically identify the feasibility of avoiding this problem and attribute it to overfitting polynomial coefficients. To cope with it, the coefficients are restricted in two steps, dimensionality reduction of the coefficients' domain and sequential assignment of the forgetting factor. We transform the optimization of coefficients to the tuning of a hyperparameter and propose a flexible spectral-domain graph filter, which significantly reduces the memory demand and the adverse impacts on message transmission under large receptive fields. Utilizing our filter, the performance of GNNs is improved significantly in large receptive fields and the receptive fields of GNNs are multiplied as well. Meanwhile, the superiority of applying a high-order approximation is verified across various datasets, notably in strongly hyperbolic datasets. Codes are publicly available at: https://github.com/cengzeyuan/TNNLS-FFKSF.
Zeyuan Zeng, Qinke Peng, Xu Mou, Ying Wang 0065, Ruimeng Li
IEEE Trans. Neural Networks Learn. Syst.4
2023 KGETCDA: an efficient representation learning framework based on knowledge graph encoder from transformer for predicting circRNA-disease associations
abstract
Recent studies have demonstrated the significant role that circRNA plays in the progression of human diseases. Identifying circRNA-disease associations (CDA) in an efficient manner can offer crucial insights into disease diagnosis. While traditional biological experiments can be time-consuming and labor-intensive, computational methods have emerged as a viable alternative in recent years. However, these methods are often limited by data sparsity and their inability to explore high-order information. In this paper, we introduce a novel method named Knowledge Graph Encoder from Transformer for predicting CDA (KGETCDA). Specifically, KGETCDA first integrates more than 10 databases to construct a large heterogeneous non-coding RNA dataset, which contains multiple relationships between circRNA, miRNA, lncRNA and disease. Then, a biological knowledge graph is created based on this dataset and Transformer-based knowledge representation learning and attentive propagation layers are applied to obtain high-quality embeddings with accurately captured high-order interaction information. Finally, multilayer perceptron is utilized to predict the matching scores of CDA based on their embeddings. Our empirical results demonstrate that KGETCDA significantly outperforms other state-of-the-art models. To enhance user experience, we have developed an interactive web-based platform named HNRBase that allows users to visualize, download data and make predictions using KGETCDA with ease. The code and datasets are publicly available at https://github.com/jinyangwu/KGETCDA.
Jinyang Wu 0001, Zhiwei Ning, Yidong Ding, Ying Wang 0065, Qinke Peng, Laiyi Fu
Briefings Bioinform.4
2023 Abstractive Financial News Summarization via Transformer-BiLSTM Encoder and Graph Attention-Based Decoder
abstract
Financial news summarization (FNS) has been an attractive research problem in recent years, which aims to generate a shorter highlight of the news article while preserving key factual aspects, emotions, and opinions, providing significant assistance in stock trading and investment decision-making. However, FNS faces two challenges compared to the common domain. Firstly, financial news involves professional qualitative and quantitative information and salient content always scatters across long-range interactions. Secondly, financial news contains latent causal relationships, where historical information in the early generated sequence can significantly affect the subsequent decoding process. To address these difficulties, we propose an enhanced Seq2Seq model named TLGA, where the hierarchical Transformer-BiLSTM encoder can capture long-range interactions and sequential semantics while the Graph Attention-based decoder can fully utilize the historical information of decoded tokens and capture key causal relations. Moreover, we propose history-enhanced attention to concentrate on salient input content based on history semantics, guiding our decoder to generate the summary around the corresponding contents. It is also the first attempt to reuse history information of previously generated summary sequences in FNS using the idea of the Graph Attention Mechanism. Additionally, we construct the LCFNS dataset with 430,820 news-summary pairs for the lack of large-scale high-quality datasets in FNS. Experimental results on two financial datasets and two benchmark datasets indicate that our model outperforms other baselines.
Haozhou Li, Qinke Peng, Xu Mou, Ying Wang 0065, Zeyuan Zeng, Muhammad Fiaz Bashir
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 A Deep Learning Framework for News Readers' Emotion Prediction Based on Features From News Article and Pseudo Comments
abstract
With the rapid development of the Internet, readers tend to share their views and emotions about news events. Predicting these emotions provides a vital role in social media applications (e.g., sentiment retrieval, opinion summary, and election prediction). However, news articles usually consist of objective texts that lack emotion words, making emotion prediction challenging. From prior studies, we know that comments that come directly from readers are full of emotions. Therefore, in this article, we propose a deep learning framework that first merges article and comment information to predict readers' emotions. At the same time, in the prediction process, we design a pseudo comment representation for unpublished news articles by the comments of published news. In addition, a better model is required to encode articles that contain implicit emotions. To solve this problem, we propose a block emotion attention network (BEAN) to encode news articles better. It includes an emotion attention mechanism and a hierarchical structure to capture emotion words and generate structural information during encoding. Experiments performed on three public datasets show that BEAN achieves the state-of-the-art average Pearson (AP) and accuracy (Acc@1). Moreover, results on four self-collected datasets show that both the introduction of emotional comments and BEAN in our framework improve the ability to predict readers' emotions.
Xu Mou, Qinke Peng, Ying Wang 0065, Muhammad Fiaz Bashir
IEEE Trans. Cybern.4
2023 BertNDA: A Model Based on Graph-Bert and Multi-Scale Information Fusion for ncRNA-Disease Association Prediction
abstract
Non-coding RNAs (ncRNAs) are a class of RNA molecules that lack the ability to encode proteins in human cells, but play crucial roles in various biological process. Understanding the interactions between different ncRNAs and their impact on diseases can significantly contribute to diagnosis, prevention, and treatment of diseases. However, predicting tertiary interactions between ncRNAs and diseases based on structural information in multiple scales remains a challenging task. To address this challenge, we propose a method called BertNDA, aiming to predict potential relationships between miRNAs, lncRNAs, and diseases. The framework identifies the local information through connectionless subgraph, which aggregate neighbor nodes' feature. And global information is extracted by leveraging Laplace transform of graph structures and WL (Weisfeiler-Lehman) absolute role coding. Additionally, an EMLP (Element-wise MLP) structure is designed to fuse pairwise global information. The transformer-encoder is employed as the backbone of our approach, followed by a prediction-layer to output the final correlation score. Extensive experiments demonstrate that BertNDA outperforms state-of-the-art methods in prediction assignment and exhibits significant potential for various biological applications. Moreover, we develop an online prediction platform that incorporates the prediction model, providing users with an intuitive and interactive experience. Overall, our model offers an efficient, accurate, and comprehensive tool for predicting tertiary associations between ncRNAs and diseases.
Zhiwei Ning, Jinyang Wu 0001, Yidong Ding, Ying Wang 0065, Qinke Peng, Laiyi Fu
IEEE J. Biomed. Health Informatics4
2023 LncDLSM: Identification of Long Non-Coding RNAs With Deep Learning-Based Sequence Model
abstract
Long non-coding RNAs (LncRNAs) serve a vital role in regulating gene expressions and other biological processes. Differentiation of lncRNAs from protein-coding transcripts helps researchers dig into the mechanism of lncRNA formation and its downstream regulations related to various diseases. Previous works have been proposed to identify lncRNAs, including traditional bio-sequencing and machine learning approaches. Considering the tedious work of biological characteristic-based feature extraction procedures and inevitable artifacts during bio-sequencing processes, those lncRNA detection methods are not always satisfactory. Hence, in this work, we presented lncDLSM, a deep learning-based framework differentiating lncRNA from other protein-coding transcripts without dependencies on prior biological knowledge. lncDLSM is a helpful tool for identifying lncRNAs compared with other biological feature-based machine learning methods and can be applied to other species by transfer learning achieving satisfactory results. Further experiments showed that different species display distinct boundaries among distributions corresponding to the homology and the specificity among species, respectively.
Ying Wang 0065, Hongkai Du, Yingxin Cao, Qinke Peng, Laiyi Fu
IEEE J. Biomed. Health Informatics1
2022 A successful hybrid deep learning model aiming at promoter identification
abstract
BACKGROUND: The zone adjacent to a transcription start site (TSS), namely, the promoter, is primarily involved in the process of DNA transcription initiation and regulation. As a result, proper promoter identification is critical for further understanding the mechanism of the networks controlling genomic regulation. A number of methodologies for the identification of promoters have been proposed. Nonetheless, due to the great heterogeneity existing in promoters, the results of these procedures are still unsatisfactory. In order to establish additional discriminative characteristics and properly recognize promoters, we developed the hybrid model for promoter identification (HMPI), a hybrid deep learning model that can characterize both the native sequences of promoters and the morphological outline of promoters at the same time. We developed the HMPI to combine a method called the PSFN (promoter sequence features network), which characterizes native promoter sequences and deduces sequence features, with a technique referred to as the DSPN (deep structural profiles network), which is specially structured to model the promoters in terms of their structural profile and to deduce their structural attributes. RESULTS: The HMPI was applied to human, plant and Escherichia coli K-12 strain datasets, and the findings showed that the HMPI was successful at extracting the features of the promoter while greatly enhancing the promoter identification performance. In addition, after the improvements of synthetic sampling, transfer learning and label smoothing regularization, the improved HMPI models achieved good results in identifying subtypes of promoters on prokaryotic promoter datasets. CONCLUSIONS: The results showed that the HMPI was successful at extracting the features of promoters while greatly enhancing the performance of identifying promoters on both eukaryotic and prokaryotic datasets, and the improved HMPI models are good at identifying subtypes of promoters on prokaryotic promoter datasets. The HMPI is additionally adaptable to different biological functional sequences, allowing for the addition of new features or models.
Ying Wang 0065, Qinke Peng, Xu Mou, Xinyuan Wang 0011, Haozhou Li, Tian Han 0008, Xiao Wang 0007
BMC Bioinform.1
2020 An Effective Hybrid Deep Learning Model for Eukaryotic Promoter Identification
abstract
The promoter is a region located near the transcription start site (TSS) and responsible for the initiation and regulation of DNA transcription. Hence, accurate identification of promoters is essential for further building and understanding the mechanism of genetic regulatory networks. Numerous approaches for eukaryotic promoter identification were proposed. Nevertheless, the performances of these approaches are still unsatisfactory due to the variety nature of promoters. To extract more discriminative features and accurately identify eukaryotic promoters, here, we develop an effective hybrid deep learning model HDLMepi, which is able to characterize the original promoter sequences and the structural profiles of promoters simultaneously. We integrate the method we name PromoterClCce which characterizes the original promoter sequences and extracts sequence features, with an approach DSPN, which we design to model the structural profile of promoters and extract structure features, in HDLMepi for precisely eukaryotic promoter identification. We apply HDLMepi on both human and plants datasets and the experimental results demonstrate it is effective in promoter features extraction and can improve the performance of promoter identification significantly. HDLMepi is also open to add new features or new models and can be applied to other biology functional sequences.
Ying Wang 0065, Qinke Peng, Xu Mou, Tian Han 0008, Xiao Wang 0007
BIBM1
2019 Predicting Social Emotions from Readers' Perspective
abstract
Due to the rapid development of Web, large numbers of documents assigned by readers' emotions have been generated through new portals. Comparing to the previous studies which focused on author's perspective, our research focuses on readers' emotions invoked by news articles. Our research provides meaningful assistance in social media application such as sentiment retrieval, opinion summarization and election prediction. In this paper, we predict the readers' emotion of news based on the social opinion network. More specifically, we construct the opinion network based on the semantic distance. The communities in the news network indicate specific events which are related to the emotions. Therefore, the opinion network serves as the lexicon between events and corresponding emotions. We leverage neighbor relationship in network to predict readers' emotions. As a result, our methods obtain better result than the state-of-the-art methods. Moreover, we developed a growing strategy to prune the network for practical application. The experiment verifies the rationality of the reduction for application.
Qinke Peng, Ling Chai, Ying Wang 0065
IEEE Trans. Affect. Comput.5