Pu-Feng Du

dblp:261/4351 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0001-9897-3932ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 9 since 2021
YearPublicationVenuePosition
2026 GraphChIAr: genome-wide super-resolution reconstruction of protein-mediated remote chromatin interactions by augmenting hi-C interaction maps with multiple ChIP-seq profiles
abstract
Protein-mediated chromatin interactions are fundamental to gene regulation. However, experimental approaches such as Chromatin Interaction Analysis by Paired-End Tag sequencing are limited by data scarcity and high cost, while existing computational models are constrained by limited resolution and challenges in effectively integrating heterogeneous genomic data. To address these issues, we propose GraphChIAr, a regression-based deep learning framework that estimates chromatin interaction strength by augmenting Hi-C contact maps with various ChIP-seq profiles and genomic sequence information. A key advantage of GraphChIAr is its super-resolution capability, enabling accurate estimations of chromatin interactions from conventional resolutions down to ultra-high near-nucleosome resolution (e.g. 200 bp). By introducing genomic shift distance in GraphChIAr, we enabled it to predict remote interactions between distant genomic loci at a genome-wide scale. Cross-referencing results demonstrate high predictive accuracy for key mediating proteins such as CTCF, highlighting the benefits of integrating complementary genomic features. Together, GraphChIAr provides an effective computational tool to augment experimental data and advance the study of 3D genome organization. The source code of GraphChIAr is available at (https://github.com/don194/GraphChIAr).
Guo-Zheng Rao, Xin-Ran Wu, Tong Xian, Bo-Qiang Wang, Pu-Feng Du
Briefings Bioinform.7
2025 scMUG: deep clustering analysis of single-cell RNA-seq data on multiple gene functional modules
abstract
Single-cell RNA sequencing (scRNA-seq) has revolutionized our understanding of cellular heterogeneity by providing gene expression data at the single-cell level. Unlike bulk RNA-seq, scRNA-seq allows identification of different cell types within a given tissue, leading to a more nuanced comprehension of cell functions. However, the analysis of scRNA-seq data presents challenges due to its sparsity and high dimensionality. Since bioinformatics plays an important role in the analysis of big data and its utility for the welfare of living beings, it has been widely applied in analyzing scRNA-seq data. To address these challenges, we introduce the scMUG computational pipeline, which incorporates gene functional module information to enhance scRNA-seq clustering analysis. The pipeline includes data preprocessing, cell representation generation, cell-cell similarity matrix construction, and clustering analysis. The scMUG pipeline also introduces a novel similarity measure that combines local density and global distribution in the latent cell representation space. As far as we can tell, this is the first attempt to integrate gene functional associations into scRNA-seq clustering analysis. We curated nine human scRNA-seq datasets to evaluate our scMUG pipeline. With the help of gene functional information and the novel similarity measure, the clustering results from scMUG pipeline present deep insights into functional relationships between gene expression patterns and cellular heterogeneity. In addition, our scMUG pipeline also presents comparable or better clustering performances than other state-of-the-art methods. All source codes of scMUG have been deposited in a GitHub repository with instructions for reproducing all results (https://github.com/degiminnal/scMUG).
De-Min Liang, Pu-Feng Du
Briefings Bioinform.2
2025 scDIFF: automatic cell type annotation using scATAC-seq data by incorporating bulk-level genomic and epigenomic information in a deep diffusive transformer
abstract
Single-cell assay for transposase-accessible chromatin sequencing (scATAC-seq) provides an opportunity to look deeply into the gene regulation mechanism at the single-cell resolution. With the rapid accumulation of scATAC-seq data, there is an urgent need for automatic cell type annotations using scATAC-seq data. Most existing methods rely on creating artificial gene activity matrix, due to the extreme sparsity nature of the scATAC-seq data. However, these methods fail to exploit the intrinsic information inherent in the scATAC-seq peaks. We present scDIFF, a diffusive transformer-based method that integrates bulk-level genomic and epigenomic information with scATAC-seq data to annotate cell types without creating artificial gene activity matrix. Our scDIFF performed constantly better than state-of-the-art methods on all 46 benchmarking pairs of reference and query datasets across different sequencing platforms. The complete implementation of scDIFF is well-documented and freely available on GitHub (https://github.com/haoyu-wangg/scDIFF).
Pu-Feng Du
Briefings Bioinform.3
2024 SilenceREIN: seeking silencers on anchors of chromatin loops by deep graph neural networks
abstract
Silencers are repressive cis-regulatory elements that play crucial roles in transcriptional regulation. Experimental methods for identifying silencers are always costly and time-consuming. Computational methods, which relies on genomic sequence features, have been introduced as alternative approaches. However, silencers do not have significant epigenomic signature. Therefore, we explore a new way to computationally identify silencers, by incorporating chromatin structural information. We propose the SilenceREIN method, which focuses on finding silencers on anchors of chromatin loops. By using graph neural networks, we extracted chromatin structural information from a regulatory element interaction network. SilenceREIN integrated the chromatin structural information with linear genomic signatures to find silencers. The predictive performance of SilenceREIN is comparable or better than other states-of-the-art methods. We performed a genome-wide scanning to systematically find silencers in human genome. Results suggest that silencers are widespread on anchors of chromatin loops. In addition, enrichment analysis of transcription factor binding motif support our prediction results. As far as we can tell, this is the first attempt to incorporate chromatin structural information in finding silencers. All datasets and source codes of SilenceREIN have been deposited in a GitHub repository (https://github.com/JianHPan/SilenceREIN).
Jian-Hua Pan, Pu-Feng Du
Briefings Bioinform.2
2023 iEssLnc: quantitative estimation of lncRNA gene essentialities with meta-path-guided random walks on the lncRNA-protein interaction network
abstract
Gene essentiality is defined as the extent to which a gene is required for the survival and reproductive success of a living system. It can vary between genetic backgrounds and environments. Essential protein coding genes have been well studied. However, the essentiality of non-coding regions is rarely reported. Most regions of human genome do not encode proteins. Determining essentialities of non-coding genes is demanded. We developed iEssLnc models, which can assign essentiality scores to lncRNA genes. As far as we know, this is the first direct quantitative estimation to the essentiality of lncRNA genes. By taking the advantage of graph neural network with meta-path-guided random walks on the lncRNA-protein interaction network, iEssLnc models can perform genome-wide screenings for essential lncRNA genes in a quantitative manner. We carried out validations and whole genome screening in the context of human cancer cell-lines and mouse genome. In comparisons to other methods, which are transferred from protein-coding genes, iEssLnc achieved better performances. Enrichment analysis indicated that iEssLnc essentiality scores clustered essential lncRNA genes with high ranks. With the screening results of iEssLnc models, we estimated the number of essential lncRNA genes in human and mouse. We performed functional analysis to find that essential lncRNA genes interact with microRNAs and cytoskeletal proteins significantly, which may be of interest in experimental life sciences. All datasets and codes of iEssLnc models have been deposited in GitHub (https://github.com/yyZhang14/iEssLnc).
Ying-Ying Zhang, De-Min Liang, Pu-Feng Du
Briefings Bioinform.3
2022 NPI-RGCNAE: Fast Predicting ncRNA-Protein Interactions Using the Relational Graph Convolutional Network Auto-Encoder
abstract
ncRNAs play important roles in a variety of biological processes by interacting with RNA-binding proteins. Therefore, identifying ncRNA-protein interactions is important to understanding the biological functions of ncRNAs. Since experimental methods to determine ncRNA-protein interactions are always costly and time-consuming, computational methods have been proposed as alternative approaches. We developed a novel method NPI-RGCNAE (predicting ncRNA-Protein Interactions by the Relational Graph Convolutional Network Auto-Encoder). With a reliable negative sample selection strategy, we applied the Relational Graph Convolutional Network encoder and the DistMult decoder to predict ncRNA-protein interactions in an accurate and efficient way. By using the 5-fold cross-validation, we found that our method achieved a comparable performance to all state-of-the-art methods. Our method requires less than 10% training time of all state-of-the-art methods. It is a more efficient choice with large datasets in practice.
Zi-Ang Shen, Pu-Feng Du
IEEE J. Biomed. Health Informatics3
2021 NPI-GNN: Predicting ncRNA-protein interactions with deep graph neural networks
abstract
Noncoding RNAs (ncRNAs) play crucial roles in many biological processes. Experimental methods for identifying ncRNA-protein interactions (NPIs) are always costly and time-consuming. Many computational approaches have been developed as alternative ways. In this work, we collected five benchmarking datasets for predicting NPIs. Based on these datasets, we evaluated and compared the prediction performances of existing machine-learning based methods. Graph neural network (GNN) is a recently developed deep learning algorithm for link predictions on complex networks, which has never been applied in predicting NPIs. We constructed a GNN-based method, which is called Noncoding RNA-Protein Interaction prediction using Graph Neural Networks (NPI-GNN), to predict NPIs. The NPI-GNN method achieved comparable performance with state-of-the-art methods in a 5-fold cross-validation. In addition, it is capable of predicting novel interactions based on network information and sequence information. We also found that insufficient sequence information does not affect the NPI-GNN prediction performance much, which makes NPI-GNN more robust than other methods. As far as we can tell, NPI-GNN is the first end-to-end GNN predictor for predicting NPIs. All benchmarking datasets in this work and all source codes of the NPI-GNN method have been deposited with documents in a GitHub repo (https://github.com/AshuiRUA/NPI-GNN).
Zi-Ang Shen, Yuan-Ke Zhou, Pu-Feng Du
Briefings Bioinform.5
2021 KNIndex: a comprehensive database of physicochemical properties for k-tuple nucleotides
abstract
With the development of high-throughput sequencing technology, the genomic sequences increased exponentially over the last decade. In order to decode these new genomic data, machine learning methods were introduced for genome annotation and analysis. Due to the requirement of most machines learning methods, the biological sequences must be represented as fixed-length digital vectors. In this representation procedure, the physicochemical properties of k-tuple nucleotides are important information. However, the values of the physicochemical properties of k-tuple nucleotides are scattered in different resources. To facilitate the studies on genomic sequences, we developed the first comprehensive database, namely KNIndex (https://knindex.pufengdu.org), for depositing and visualizing physicochemical properties of k-tuple nucleotides. Currently, the KNIndex database contains 182 properties including one for mononucleotide (DNA), 169 for dinucleotide (147 for DNA and 22 for RNA) and 12 for trinucleotide (DNA). KNIndex database also provides a user-friendly web-based interface for the users to browse, query, visualize and download the physicochemical properties of k-tuple nucleotides. With the built-in conversion and visualization functions, users are allowed to display DNA/RNA sequences as curves of multiple physicochemical properties. We wish that the KNIndex will facilitate the related studies in computational biology.
Wen-Ya Zhang, Junhai Xu, Jun Wang 0081, Yuan-Ke Zhou, Wei Chen 0064, Pu-Feng Du
Briefings Bioinform.6
2021 Predicting protein subchloroplast locations: the 10th anniversary
Pu-Feng Du
Frontiers Comput. Sci.2
2020 VisFeature: a stand-alone program for visualizing and analyzing statistical features of biological sequences
abstract
SUMMARY: Many efforts have been made in developing bioinformatics algorithms to predict functional attributes of genes and proteins from their primary sequences. One challenge in this process is to intuitively analyze and to understand the statistical features that have been selected by heuristic or iterative methods. In this paper, we developed VisFeature, which aims to be a helpful software tool that allows the users to intuitively visualize and analyze statistical features of all types of biological sequence, including DNA, RNA and proteins. VisFeature also integrates sequence data retrieval, multiple sequence alignments and statistical feature generation functions. AVAILABILITY AND IMPLEMENTATION: VisFeature is a desktop application that is implemented using JavaScript/Electron and R. The source codes of VisFeature are freely accessible from the GitHub repository (https://github.com/wangjun1996/VisFeature). The binary release, which includes an example dataset, can be freely downloaded from the same GitHub repository (https://github.com/wangjun1996/VisFeature/releases). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jun Wang 0081, Pu-Feng Du, Xin-Yu Xue, Guang-Ping Li 0005, Yuan-Ke Zhou, Hao Lin 0001, Wei Chen 0064
Bioinform.2