Laiyi Fu

dblp:276/6319 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-9086-3982ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 5 first-author · 11 since 2021
YearPublicationVenuePosition
2026 Biased multi-view contrastive learning with attentive masking for spatial transcriptomic analysis
abstract
Spatial transcriptomics (ST) enables the simultaneous measurement of gene expression and spatial context, offering unprecedented insights into tissue architecture and cellular communication. However, existing approaches often fail to jointly capture spatial topology and transcriptional heterogeneity, leading to suboptimal representations and limited biological interpretability. To address this limitation, we propose stCAMBL, a biased multi-view contrastive framework that integrates spatial graph structure modeling with attentive feature masking and partial contrastive regularization. Built upon a variational graph autoencoder backbone, stCAMBL learns biologically informed and noise-robust embeddings by adaptively emphasizing informative molecular features while mitigating confounding patterns across spatial domains. Comprehensive evaluations on multiple 10$\times$ Visium datasets demonstrate that stCAMBL substantially improves clustering accuracy, gene ontology enrichment, and signal restoration, demonstrating strong generalizability for high-fidelity ST analysis.
Laiyi Fu, Wenkai Cui, Danyang Wu, Hequan Sun
Briefings Bioinform.1
2025 Dual balanced augmented topological noncoding RNA disease triplet association in heterogeneous graphs
abstract
Noncoding RNAs (ncRNAs), including long noncoding RNAs (lncRNAs) and microRNAs (miRNAs), play pivotal roles in various human diseases. Predicting associations such as lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs), and lncRNA-miRNA interactions (LMIs) is crucial for understanding disease mechanisms and identifying therapeutic targets. However, existing models face significant challenges in handling extreme data imbalance and often treat multiple ncRNA-disease and ncRNA-ncRNA interactions collectively, lacking the ability to provide precise, differentiated predictions for specific types of ncRNAs. This limitation reduces their practical applicability. To address these issues, we propose the Dual Balanced Augmented Topological Noncoding RNA Disease triplet Association (DBATNDA) model. DBATNDA constructs an Interaction Dual Graph with LDAs, MDAs, and LMIs as nodes and introduces an efficient graph-based balanced topological augmentation mechanism to enhance node structural representation and adaptability to imbalanced data. This innovative approach enables fast and accurate predictions of ncRNA-disease and ncRNA-ncRNA triplet associations through node classification view. To the best of our knowledge, no existing method employs such a dual-representation strategy to provide simultaneously differentiated predictions for the associations between diverse ncRNAs and diseases while also enhancing target specificity. Experimental results demonstrate DBATNDA's superior performance compared to state-of-the-art models, while case studies confirm its practical significance in these triple association prediction. The code and datasets are publicly available at https://github.com/AI4Bread/DBATNDA.
Laiyi Fu, Yangyi Zhou, Hongqiang Lyu, Hequan Sun
Briefings Bioinform.1
2025 DeepExDC interprets genomic compartmentalization changes in single-cell Hi-C data
abstract
Single-cell Hi-C (scHi-C) technology enables probing of higher-order chromatin structures in individual cells. It provides an opportunity to get a deeper insight into genomic compartmentalization changes of single cells across different conditions, paving the way to a common understanding of the interplay among compartmental organization, genome functions, and cellular phenotypes. Unfortunately, there are only a few methods currently available for the differential analysis of A/B compartments on Hi-C data at the bulk level; the computational analysis of compartmentalization changes at the single-cell level is a field in its infancy. Herein, we propose DeepExDC, an interpretable 1D convolutional neural network for differential analysis of A/B compartments in scHi-C data on a genome-wide scale. It accepts Hi-C contact matrices at the single-cell level, runs without any distribution assumption and differential pattern limitation, and interprets genomic compartmentalization changes across multiple conditions. The results on simulated and experimental scHi-C data show that our DeepExDC has higher accuracies in detecting different types of compartmentalization changes, and the interpretation values are demonstrated to be able to reflect compartment changes across cell types. It is also observed that the differential compartments given by DeepExDC agree well with those by state-of-the-art methods at the bulk level, help to characterize heterogeneity of single cells, and exhibit a reasonable biological relevance in multiple regards. In addition, considering that DeepExDC is free of distribution assumptions and differential patterns, we attempted to transfer it onto scRNA-seq and scATAC-seq data; it is interesting that our method also presents considerable power compared with the competing methods.
Hongqiang Lyu, Wenyao Long, Xiaoran Yin, Shengjun Xu, Laiyi Fu
Briefings Bioinform.6
2024 ACLNDA: an asymmetric graph contrastive learning framework for predicting noncoding RNA-disease associations in heterogeneous graphs
abstract
Noncoding RNAs (ncRNAs), including long noncoding RNAs (lncRNAs) and microRNAs (miRNAs), play crucial roles in gene expression regulation and are significant in disease associations and medical research. Accurate ncRNA-disease association prediction is essential for understanding disease mechanisms and developing treatments. Existing methods often focus on single tasks like lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs), or lncRNA-miRNA interactions (LMIs), and fail to exploit heterogeneous graph characteristics. We propose ACLNDA, an asymmetric graph contrastive learning framework for analyzing heterophilic ncRNA-disease associations. It constructs inter-layer adjacency matrices from the original lncRNA, miRNA, and disease associations, and uses a Top-K intra-layer similarity edges construction approach to form a triple-layer heterogeneous graph. Unlike traditional works, to account for both node attribute features (ncRNA/disease) and node preference features (association), ACLNDA employs an asymmetric yet simple graph contrastive learning framework to maximize one-hop neighborhood context and two-hop similarity, extracting ncRNA-disease features without relying on graph augmentations or homophily assumptions, reducing computational cost while preserving data integrity. Our framework is capable of being applied to a universal range of potential LDA, MDA, and LMI association predictions. Further experimental results demonstrate superior performance to other existing state-of-the-art baseline methods, which shows its potential for providing insights into disease diagnosis and therapeutic target identification. The source code and data of ACLNDA is publicly available at https://github.com/AI4Bread/ACLNDA.
Laiyi Fu, Yangyi Zhou, Qinke Peng, Hongqiang Lyu
Briefings Bioinform.1
2024 findGSEP: estimating genome size of polyploid species usingk-mer frequencies
abstract
SUMMARY: Estimating genome size using k-mer frequencies, which plays a fundamental role in designing genome sequencing and analysis projects, has remained challenging for polyploid species, i.e., ploidy p > 2. To address this, we introduce "findGSEP," which is designed based on iterative curve fitting of k-mer frequencies. Precisely, it first disentangles up to p normal distributions by analyzing k-mer frequencies in whole genome sequencing of the focal species. Second, it computes the sizes of genomic regions related to 1∼p (homologous) chromosome(s) using each respective curve fitting, from which it infers the full polyploid and average haploid genome size. "findGSEP" can handle any level of ploidy p, and infer more accurate genome size than other well-known tools, as shown by tests using simulated and real genomic sequencing data of various species including octoploids. AVAILABILITY AND IMPLEMENTATION: "findGSEP" was implemented as a web server, which is freely available at http://146.56.237.198:3838/findGSEP/. Also, "findGSEP" was implemented as an R package for parallel processing of multiple samples. Source code and tutorial on its installation and usage is available at https://github.com/sperfu/findGSEP.
Laiyi Fu, Yanxin Xie, ShunKang Ling, Ying Wang 0065, Binzhong Wang, Hejun Du, Qinke Peng, Hequan Sun
Bioinform.1
2024 Identifying TAD-like domains on single-cell Hi-C data by graph embedding and changepoint detection
abstract
MOTIVATION: Topologically associating domains (TADs) are fundamental building blocks of 3D genome. TAD-like domains in single cells are regarded as the underlying genesis of TADs discovered in bulk cells. Understanding the organization of TAD-like domains helps to get deeper insights into their regulatory functions. Unfortunately, it remains a challenge to identify TAD-like domains on single-cell Hi-C data due to its ultra-sparsity. RESULTS: We propose scKTLD, an in silico tool for the identification of TAD-like domains on single-cell Hi-C data. It takes Hi-C contact matrix as the adjacency matrix for a graph, embeds the graph structures into a low-dimensional space with the help of sparse matrix factorization followed by spectral propagation, and the TAD-like domains can be identified using a kernel-based changepoint detection in the embedding space. The results tell that our scKTLD is superior to the other methods on the sparse contact matrices, including downsampled bulk Hi-C data as well as simulated and experimental single-cell Hi-C data. Besides, we demonstrated the conservation of TAD-like domain boundaries at single-cell level apart from heterogeneity within and across cell types, and found that the boundaries with higher frequency across single cells are more enriched for architectural proteins and chromatin marks, and they preferentially occur at TAD boundaries in bulk cells, especially at those with higher hierarchical levels. AVAILABILITY AND IMPLEMENTATION: scKTLD is freely available at https://github.com/lhqxinghun/scKTLD.
Erhu Liu, Hongqiang Lyu, Laiyi Fu, Xiaoran Yin
Bioinform.4
2024 KGRACDA: A Model Based on Knowledge Graph from Recursion and Attention Aggregation for CircRNA-Disease Association Prediction
abstract
CircRNA is closely related to human disease, so it is important to predict circRNA-disease association (CDA). However, the traditional biological detection methods have high difficulty and low accuracy, and computational methods represented by deep learning ignore the ability of the model to explicitly extract local depth information of the CDA. We propose a model based on knowledge graph from recursion and attention aggregation for circRNA-disease association prediction (KGRACDA). This model combines explicit structural features and implicit embedding information of graphs, optimizing graph embedding vectors. First, we built large-scale, multi-source heterogeneous datasets and construct a knowledge graph of multiple RNAs and diseases. After that, we use a recursive method to build multi-hop subgraphs and optimize graph attention mechanism by gating mechanism, mining local depth information. At the same time, the model uses multi-head attention mechanism to balance global and local depth features of graphs, and generate CDA prediction scores. KGRACDA surpasses other methods by capturing local and global depth information related to CDA. We update an interactive web platform HNRBase v2.0, which visualizes circRNA data, and allows users to download data and predict CDA using model.
Ying Wang 0065, Maoyuan Ma, Yanxin Xie, Qinke Peng, Hongqiang Lyu, Hequan Sun, Laiyi Fu
IEEE ACM Trans. Comput. Biol. Bioinform.7
2023 KGETCDA: an efficient representation learning framework based on knowledge graph encoder from transformer for predicting circRNA-disease associations
abstract
Recent studies have demonstrated the significant role that circRNA plays in the progression of human diseases. Identifying circRNA-disease associations (CDA) in an efficient manner can offer crucial insights into disease diagnosis. While traditional biological experiments can be time-consuming and labor-intensive, computational methods have emerged as a viable alternative in recent years. However, these methods are often limited by data sparsity and their inability to explore high-order information. In this paper, we introduce a novel method named Knowledge Graph Encoder from Transformer for predicting CDA (KGETCDA). Specifically, KGETCDA first integrates more than 10 databases to construct a large heterogeneous non-coding RNA dataset, which contains multiple relationships between circRNA, miRNA, lncRNA and disease. Then, a biological knowledge graph is created based on this dataset and Transformer-based knowledge representation learning and attentive propagation layers are applied to obtain high-quality embeddings with accurately captured high-order interaction information. Finally, multilayer perceptron is utilized to predict the matching scores of CDA based on their embeddings. Our empirical results demonstrate that KGETCDA significantly outperforms other state-of-the-art models. To enhance user experience, we have developed an interactive web-based platform named HNRBase that allows users to visualize, download data and make predictions using KGETCDA with ease. The code and datasets are publicly available at https://github.com/jinyangwu/KGETCDA.
Jinyang Wu 0001, Zhiwei Ning, Yidong Ding, Ying Wang 0065, Qinke Peng, Laiyi Fu
Briefings Bioinform.6
2023 BertNDA: A Model Based on Graph-Bert and Multi-Scale Information Fusion for ncRNA-Disease Association Prediction
abstract
Non-coding RNAs (ncRNAs) are a class of RNA molecules that lack the ability to encode proteins in human cells, but play crucial roles in various biological process. Understanding the interactions between different ncRNAs and their impact on diseases can significantly contribute to diagnosis, prevention, and treatment of diseases. However, predicting tertiary interactions between ncRNAs and diseases based on structural information in multiple scales remains a challenging task. To address this challenge, we propose a method called BertNDA, aiming to predict potential relationships between miRNAs, lncRNAs, and diseases. The framework identifies the local information through connectionless subgraph, which aggregate neighbor nodes' feature. And global information is extracted by leveraging Laplace transform of graph structures and WL (Weisfeiler-Lehman) absolute role coding. Additionally, an EMLP (Element-wise MLP) structure is designed to fuse pairwise global information. The transformer-encoder is employed as the backbone of our approach, followed by a prediction-layer to output the final correlation score. Extensive experiments demonstrate that BertNDA outperforms state-of-the-art methods in prediction assignment and exhibits significant potential for various biological applications. Moreover, we develop an online prediction platform that incorporates the prediction model, providing users with an intuitive and interactive experience. Overall, our model offers an efficient, accurate, and comprehensive tool for predicting tertiary associations between ncRNAs and diseases.
Zhiwei Ning, Jinyang Wu 0001, Yidong Ding, Ying Wang 0065, Qinke Peng, Laiyi Fu
IEEE J. Biomed. Health Informatics6
2023 LncDLSM: Identification of Long Non-Coding RNAs With Deep Learning-Based Sequence Model
abstract
Long non-coding RNAs (LncRNAs) serve a vital role in regulating gene expressions and other biological processes. Differentiation of lncRNAs from protein-coding transcripts helps researchers dig into the mechanism of lncRNA formation and its downstream regulations related to various diseases. Previous works have been proposed to identify lncRNAs, including traditional bio-sequencing and machine learning approaches. Considering the tedious work of biological characteristic-based feature extraction procedures and inevitable artifacts during bio-sequencing processes, those lncRNA detection methods are not always satisfactory. Hence, in this work, we presented lncDLSM, a deep learning-based framework differentiating lncRNA from other protein-coding transcripts without dependencies on prior biological knowledge. lncDLSM is a helpful tool for identifying lncRNAs compared with other biological feature-based machine learning methods and can be applied to other species by transfer learning achieving satisfactory results. Further experiments showed that different species display distinct boundaries among distributions corresponding to the homology and the specificity among species, respectively.
Ying Wang 0065, Hongkai Du, Yingxin Cao, Qinke Peng, Laiyi Fu
IEEE J. Biomed. Health Informatics6
2021 SAILER: scalable and accurate invariant representation learning for single-cell ATAC-seq processing and integration
abstract
MOTIVATION: Single-cell sequencing assay for transposase-accessible chromatin (scATAC-seq) provides new opportunities to dissect epigenomic heterogeneity and elucidate transcriptional regulatory mechanisms. However, computational modeling of scATAC-seq data is challenging due to its high dimension, extreme sparsity, complex dependencies and high sensitivity to confounding factors from various sources. RESULTS: Here, we propose a new deep generative model framework, named SAILER, for analyzing scATAC-seq data. SAILER aims to learn a low-dimensional nonlinear latent representation of each cell that defines its intrinsic chromatin state, invariant to extrinsic confounding factors like read depth and batch effects. SAILER adopts the conventional encoder-decoder framework to learn the latent representation but imposes additional constraints to ensure the independence of the learned representations from the confounding factors. Experimental results on both simulated and real scATAC-seq datasets demonstrate that SAILER learns better and biologically more meaningful representations of cells than other methods. Its noise-free cell embeddings bring in significant benefits in downstream analyses: clustering and imputation based on SAILER result in 6.9% and 18.5% improvements over existing methods, respectively. Moreover, because no matrix factorization is involved, SAILER can easily scale to process millions of cells. We implemented SAILER into a software package, freely available to all for large-scale scATAC-seq data analysis. AVAILABILITY AND IMPLEMENTATION: The software is publicly available at https://github.com/uci-cbcl/SAILER. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yingxin Cao, Laiyi Fu, Qinke Peng, Qing Nie, Xiaohui Xie
Bioinform.2
2020 Predicting DNA Methylation States with Hybrid Information Based Deep-Learning Model
abstract
DNA methylation plays an important role in the regulation of some biological processes. Up to now, with the development of machine learning models, there are several sequence-based deep learning models designed to predict DNA methylation states, which gain better performance than traditional methods like random forest and SVM. However, convolutional network based deep learning models that use one-hot encoding DNA sequence as input may discover limited information and cause unsatisfactory prediction performance, so more data and model structures of diverse angles should be considered. In this work, we proposed a hybrid sequence-based deep learning model with both MeDIP-seq data and Histone information to predict DNA methylated CpG states (MHCpG). We combined both MeDIP-seq data and histone modification data with sequence information and implemented convolutional network to discover sequence patterns. In addition, we used statistical data gained from previous three input data and adopted a 3-layer feedforward neuron network to extract more high-level features. We compared our method with traditional predicting methods using random forest and other previous methods like CpGenie and DeepCpG, the result showed that MHCpG exceeded the other approaches and gained more satisfactory performance.
Laiyi Fu, Qinke Peng, Ling Chai
IEEE ACM Trans. Comput. Biol. Bioinform.1