Fei Wang 0017

dblp:52/3194-17 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0001-5890-0448ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GATTCR: A Graph Attention Network With Multi-Feature Fusion for Peripheral Blood TCR Repertoire Classification
abstract
T cell receptor (TCR) repertoire profiling provides a promising avenue for noninvasive disease diagnostics by capturing immune signatures directly from peripheral blood. However, the high diversity and sparsity of TCR sequences pose significant challenges for robust immune state classification. In this work, we propose GATTCR, a novel framework that integrates Graph Attention Networks (GATs) with multi-feature fusion to model complex dependencies within TCR repertoires. By representing TCRs as graph nodes and incorporating biological priors-such as sequence embeddings, structural similarity, V gene usage, and clonal frequency-GATTCR enables context-aware, structure-informed representation learning. We evaluate GATTCR across a comprehensive panel of cancer and infectious disease datasets, demonstrating consistent improvements over existing methods. Notably, GATTCR achieves AUROC gains of up to +47.9% under few-shot learning scenarios, highlighting its ability to generalize from limited labeled data. Ablation studies further confirm the critical role of graph-based modeling and immunological features in driving performance gains. Overall, GATTCR offers a scalable approach for TCR repertoire analysis and paves the way for routine, noninvasive immune monitoring in precision medicine.
Hengwei Ju, Deying Kong, Yuhao Tao, Fei Wang 0017
IEEE Trans. Comput. Biol. Bioinform.4
2025 scDiformer: A Difference-Aware Transformer for Time-Series Single-Cell Gene Expression Forecasting
abstract
Modeling cellular dynamics over time is fundamental to understanding biological development and disease progression. Single-cell RNA sequencing (scRNA-seq) enables the high-resolution profiling of gene expression heterogeneity and the reconstruction of developmental trajectories. However, time-series scRNA-seq data present significant challenges due to the lack of temporal alignment inherent to destructive sampling and the limited ability of current models to capture dynamic expression patterns or extrapolate to future cell states. To address these issues, we propose scDiformer, a difference-aware Transformer framework developed for forecasting gene expression at future time points. scDiformer infers temporal couplings between cells at consecutive time points using optimal transport, enabling fine-grained trajectory modeling at the single-cell level. It further captures transcriptional dynamics by lever-aging expression-difference embeddings, which guide attention toward evolving gene expression trends. Additionally, we design a training strategy aligned with the target prediction objectives to enhance generalization under temporal shifts. Experiments on multiple real-world time-series scRNA-seq datasets demonstrate that scDiformer consistently outperforms existing methods in terms of Wasserstein distance and Pearson correlation. This framework offers a powerful tool for cell fate inference and provides valuable insights for studies in developmental biology and disease mechanisms.
Jiayi Dong, Fei Wang 0017
BIBM4
2025 Conditional Causal Representation Learning for Heterogeneous Single-cell RNA Data Integration and Prediction
abstract
Single-cell sequencing technology provides deep insights into gene activity at the individual cell level, facilitating the study of gene regulatory mechanisms. However, observed gene expression are often influenced by confounding factors such as batch effects, perturbations, and spatial position, which obscure the true gene regulatory network that governs the cell’s intrinsic state. To address these challenges, we propose scConCRL, a novel conditionally causal representation learning framework designed to extract the true gene regulatory relationships independent of confounding information. By considering both fine-grained molecular gene variables and coarse-grained latent domain variables, scConCRL not only uncovers the intrinsic biological signals but also models the complex relationships between these variables. This dual function enables the separation of genuine cellular states from domain information, providing valuable insights for downstream analyses and biological discovery. We demonstrate the effectiveness of our model on multi-domain datasets from different platforms and perturbation conditions, showing its ability to accurately disentangle confounding influences and discover novel gene relationships. Extensive comparisons across various scenarios illustrate the superior performance of scConCRL in several tasks compared to existing methods.
Jiayi Dong, Fei Wang 0017
IJCAI3
2024 stDiff: a diffusion model for imputing spatial transcriptomics through single-cell transcriptomics
abstract
Spatial transcriptomics (ST) has become a powerful tool for exploring the spatial organization of gene expression in tissues. Imaging-based methods, though offering superior spatial resolutions at the single-cell level, are limited in either the number of imaged genes or the sensitivity of gene detection. Existing approaches for enhancing ST rely on the similarity between ST cells and reference single-cell RNA sequencing (scRNA-seq) cells. In contrast, we introduce stDiff, which leverages relationships between gene expression abundance in scRNA-seq data to enhance ST. stDiff employs a conditional diffusion model, capturing gene expression abundance relationships in scRNA-seq data through two Markov processes: one introducing noise to transcriptomics data and the other denoising to recover them. The missing portion of ST is predicted by incorporating the original ST data into the denoising process. In our comprehensive performance evaluation across 16 datasets, utilizing multiple clustering and similarity metrics, stDiff stands out for its exceptional ability to preserve topological structures among cells, positioning itself as a robust solution for cell population identification. Moreover, stDiff's enhancement outcomes closely mirror the actual ST data within the batch space. Across diverse spatial expression patterns, our model accurately reconstructs them, delineating distinct spatial boundaries. This highlights stDiff's capability to unify the observed and predicted segments of ST data for subsequent analysis. We anticipate that stDiff, with its innovative approach, will contribute to advancing ST imputation methodologies.
Kongming Li, Yuhao Tao, Fei Wang 0017
Briefings Bioinform.4
2024 Deep Learning in Gene Regulatory Network Inference: A Survey
abstract
Understanding the intricate regulatory relationships among genes is crucial for comprehending the development, differentiation, and cellular response in living systems. Consequently, inferring gene regulatory networks (GRNs) based on observed data has gained significant attention as a fundamental goal in biological applications. The proliferation and diversification of available data present both opportunities and challenges in accurately inferring GRNs. Deep learning, a highly successful technique in various domains, holds promise in aiding GRN inference. Several GRN inference methods employing deep learning models have been proposed; however, the selection of an appropriate method remains a challenge for life scientists. In this survey, we provide a comprehensive analysis of 12 GRN inference methods that leverage deep learning models. We trace the evolution of these major methods and categorize them based on the types of applicable data. We delve into the core concepts and specific steps of each method, offering a detailed evaluation of their effectiveness and scalability across different scenarios. These insights enable us to make informed recommendations. Moreover, we explore the challenges faced by GRN inference methods utilizing deep learning and discuss future directions, providing valuable suggestions for the advancement of data scientists in this field.
Jiayi Dong, Fei Wang 0017
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 GROD: Joint Inference of Gene Regulatory Networks and Data Imputation in Single-Cell RNA Sequencing with Temporal Consideration
abstract
Since single-cell RNA sequencing (scRNA-seq) has revolutionized the study of cellular dynamics, the construction of gene relationships based on dynamic information has attracted much attention. However, the sparsity and dropout events inherent in scRNA-seq data present challenges for downstream analysis like Gene Regulatory Network (GRN) inference. Existing data imputation methods have yielded good results beneficial for cell clustering, but it does not adequately consider gene expression dynamics. To address this, we introduce GROD, a novel deep learning approach that simultaneously infers GRN and imputes scRNA-seq data from a temporal perspective. GROD consists of three key components: an encoder, a graph learner, and a decoder. Experimental results demonstrate that GROD outperforms existing methods in both GRN inference and data imputation tasks, providing superior accuracy in capturing gene regulatory relationships and accurately imputing missing values. By integrating these two tasks, GROD enables more accurate downstream analysis and facilitates deeper insights into cellular dynamics.
Jiayi Dong, Fei Wang 0017
BIBM2
2022 scPreGAN, a deep generative model for predicting the response of single-cell expression to perturbation
abstract
MOTIVATION: Rapid developments of single-cell RNA sequencing technologies allow study of responses to external perturbations at individual cell level. However, in many cases, it is hard to collect the perturbed cells, such as knowing the response of a cell type to the drug before actual medication to a patient. Prediction in silicon could alleviate the problem and save cost. Although several tools have been developed, their prediction accuracy leaves much room for improvement. RESULTS: In this article, we propose scPreGAN (Single-Cell data Prediction base on GAN), a deep generative model for predicting the response of single-cell expression to perturbation. ScPreGAN integrates autoencoder and generative adversarial network, the former is to extract common information of the unperturbed data and the perturbed data, the latter is to predict the perturbed data. Experiments on three real datasets show that scPreGAN outperforms three state-of-the-art methods, which can capture the complicated distribution of cell expression and generate the prediction data with the same expression abundance as the real data. AVAILABILITY AND IMPLEMENTATION: The implementation of scPreGAN is available via https://github.com/JaneJiayiDong/scPreGAN. To reproduce the results of this article, please visit https://github.com/JaneJiayiDong/scPreGAN-reproducibility. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiajie Wei, Jiayi Dong, Fei Wang 0017
Bioinform.3
2022 scSemiAE: a deep model with semi-supervised learning for single-cell transcriptomics
abstract
BACKGROUND: With the development of modern sequencing technology, hundreds of thousands of single-cell RNA-sequencing (scRNA-seq) profiles allow to explore the heterogeneity in the cell level, but it faces the challenges of high dimensions and high sparsity. Dimensionality reduction is essential for downstream analysis, such as clustering to identify cell subpopulations. Usually, dimensionality reduction follows unsupervised approach. RESULTS: In this paper, we introduce a semi-supervised dimensionality reduction method named scSemiAE, which is based on an autoencoder model. It transfers the information contained in available datasets with cell subpopulation labels to guide the search of better low-dimensional representations, which can ease further analysis. CONCLUSIONS: Experiments on five public datasets show that, scSemiAE outperforms both unsupervised and semi-supervised baselines whether the transferred information embodied in the number of labeled cells and labeled cell subpopulations is much or less.
Jiayi Dong, Fei Wang 0017
BMC Bioinform.3
2021 SSBER: removing batch effect for single-cell RNA sequencing data
abstract
BACKGROUND: With the continuous maturity of sequencing technology, different laboratories or different sequencing platforms have generated a large amount of single-cell transcriptome sequencing data for the same or different tissues. Due to batch effects and high dimensions of scRNA data, downstream analysis often faces challenges. Although a number of algorithms and tools have been proposed for removing batch effects, the current mainstream algorithms have faced the problem of data overcorrection when the cell type composition varies greatly between batches. RESULTS: In this paper, we propose a novel method named SSBER by utilizing biological prior knowledge to guide the correction, aiming to solve the problem of poor batch-effect correction when the cell type composition differs greatly between batches. CONCLUSIONS: SSBER effectively solves the above problems and outperforms other algorithms when the cell type structure among batches or distribution of cell population varies considerably, or some similar cell types exist across batches.
Fei Wang 0017
BMC Bioinform.2
2018 A study on fast calling variants from next-generation sequencing data using decision tree
abstract
BACKGROUND: The rapid development of next-generation sequencing (NGS) technology has continuously been refreshing the throughput of sequencing data. However, due to the lack of a smart tool that is both fast and accurate, the analysis task for NGS data, especially those with low-coverage, remains challenging. RESULTS: We proposed a decision-tree based variant calling algorithm. Experiments on a set of real data indicate that our algorithm achieves high accuracy and sensitivity for SNVs and indels and shows good adaptability on low-coverage data. In particular, our algorithm is obviously faster than 3 widely used tools in our experiments. CONCLUSIONS: We implemented our algorithm in a software named Fuwa and applied it together with 4 well-known variant callers, i.e., Platypus, GATK-UnifiedGenotyper, GATK-HaplotypeCaller and SAMtools, to three sequencing data sets of a well-studied sample NA12878, which were produced by whole-genome, whole-exome and low-coverage whole-genome sequencing technology respectively. We also conducted additional experiments on the WGS data of 4 newly released samples that have not been used to populate dbSNP.
Zhentang Li, Fei Wang 0017
BMC Bioinform.3
2017 SeqCNV: a novel method for identification of copy number variations in targeted next-generation sequencing data
abstract
BACKGROUND: Targeted next-generation sequencing (NGS) has been widely used as a cost-effective way to identify the genetic basis of human disorders. Copy number variations (CNVs) contribute significantly to human genomic variability, some of which can lead to disease. However, effective detection of CNVs from targeted capture sequencing data remains challenging. RESULTS: Here we present SeqCNV, a novel CNV calling method designed to use capture NGS data. SeqCNV extracts the read depth information and utilizes the maximum penalized likelihood estimation (MPLE) model to identify the copy number ratio and CNV boundary. We applied SeqCNV to both bacterial artificial clone (BAC) and human patient NGS data to identify CNVs. These CNVs were validated by array comparative genomic hybridization (aCGH). CONCLUSIONS: SeqCNV is able to robustly identify CNVs of different size using capture NGS data. Compared with other CNV-calling methods, SeqCNV shows a significant improvement in both sensitivity and specificity.
Yong Chen 0016, Ming Cao 0005, Violet Gelowani, Mingchu Xu, Smriti A. Agrawal, Yumei Li 0007, Stephen P. Daiger, Richard A. Gibbs, Fei Wang 0017, Rui Chen 0013
BMC Bioinform.11
2016 A fast read alignment method based on seed-and-vote for next generation sequencing
abstract
BACKGROUND: The next-generation of sequencing technologies, along with the development of bioinformatics, are generating a growing number of reads every day. For the convenience of further research, these reads should be aligned to the reference genome by read alignment tools. Despite the diversity of read alignment tools, most have no comprehensive advantage in both accuracy and speed. For example, BWA has comparatively high accuracy, but its speed leaves much to be desired, becoming a bottleneck while an increasing number of reads need to be aligned every day. We believe that the speed of read alignment tools still has huge room for improvement, while maintaining little to no loss in accuracy. RESULTS: Here we implement a new read alignment tool, Fast Seed-and-Vote Aligner (FSVA), which is based on seeding and voting. FSVA achieves a high accuracy close to BWA and simultaneously has a very high speed. It only requires ~10-15 CPU hours to run a whole genome read alignment, which is ~5-7 times faster than BWA. CONCLUSIONS: In some cases, reads have to be aligned in a short time. Where requirement of accuracy is not very stringent, FSVA would be a promising option. FSVA is available at https://github.com/Topwood91/FSVA.
Fei Wang 0017
BMC Bioinform.3
2011 Comparative study of computational methods to detect the correlated reaction sets in biochemical networks
abstract
Correlated reaction sets (Co-Sets) are mathematically defined modules in biochemical reaction networks which facilitate the study of biological processes by decomposing complex reaction networks into conceptually simple units. According to the degree of association, Co-Sets can be classified into three types: perfect, partial and directional. Five approaches have been developed to calculate Co-Sets, including network-based pathway analysis, Monte Carlo sampling, linear optimization, enzyme subsets and hard-coupled reaction sets. However, differences in design and implementation of these methods lead to discrepancies in the resulted Co-Sets as well as in their use in biotechnology which need careful interpretation. In this paper, we provide a comparative study of the methods for Co-Sets computing in detail from four aspects: (i) sensitivity, (ii) completeness and soundness, (iii) flexibility and (iv) scalability. By applying them to Escherichia coli core metabolic network, the differences and relationships among these methods are clearly articulated which may be useful for potential users.
Yanping Xi, Yi-Ping Phoebe Chen, Fei Wang 0017
Briefings Bioinform.4
2009 Analysis on relationship between extreme pathways and correlated reaction sets
abstract
BACKGROUND: Constraint-based modeling of reconstructed genome-scale metabolic networks has been successfully applied on several microorganisms. In constraint-based modeling, in order to characterize all allowable phenotypes, network-based pathways, such as extreme pathways and elementary flux modes, are defined. However, as the scale of metabolic network rises, the number of extreme pathways and elementary flux modes increases exponentially. Uniform random sampling solves this problem to some extent to study the contents of the available phenotypes. After uniform random sampling, correlated reaction sets can be identified by the dependencies between reactions derived from sample phenotypes. In this paper, we study the relationship between extreme pathways and correlated reaction sets. RESULTS: Correlated reaction sets are identified for E. coli core, red blood cell and Saccharomyces cerevisiae metabolic networks respectively. All extreme pathways are enumerated for the former two metabolic networks. As for Saccharomyces cerevisiae metabolic network, because of the large scale, we get a set of extreme pathways by sampling the whole extreme pathway space. In most cases, an extreme pathway covers a correlated reaction set in an 'all or none' manner, which means either all reactions in a correlated reaction set or none is used by some extreme pathway. In rare cases, besides the 'all or none' manner, a correlated reaction set may be fully covered by combination of a few extreme pathways with related function, which may bring redundancy and flexibility to improve the survivability of a cell. In a word, extreme pathways show strong complementary relationship on usage of reactions in the same correlated reaction set. CONCLUSION: Both extreme pathways and correlated reaction sets are derived from the topology information of metabolic networks. The strong relationship between correlated reaction sets and extreme pathways suggests a possible mechanism: as a controllable unit, an extreme pathway is regulated by its corresponding correlated reaction sets, and a correlated reaction set is further regulated by the organism's regulatory network.
Yanping Xi, Yi-Ping Phoebe Chen, Ming Cao 0005, Weirong Wang, Fei Wang 0017
BMC Bioinform.5