EDBT 2026 Demo / reviewers in the wild / expert
Xikang Feng
dblp:255/6821
· DBLP profile ↗
9ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0003-4029-3795ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A comprehensive benchmarking for evaluating TCR embeddings in modeling TCR-epitope interactionsabstractThe complexity of T cell receptor (TCR) sequences, particularly within the complementarity-determining region 3 (CDR3), requires efficient embedding methods for applying machine learning to immunology. While various TCR CDR3 embedding strategies have been proposed, the absence of their systematic evaluations created perplexity in the community. Here, we extracted CDR3 embedding models from 19 existing methods and benchmarked these models with four curated datasets by accessing their impact on the performance of TCR downstream tasks, including TCR-epitope binding affinity prediction, epitope-specific TCR identification, TCR clustering, and visualization analysis. We assessed these models utilizing eight downstream classifiers and five downstream clustering methods, with the performance measured by a diverse range of metrics for precision, robustness, and usability. Overall, handcrafted embeddings outperformed data-driven ones in modeling TCR-epitope interactions. To further refine our comparative findings, we developed an all-in-one TCR CDR3 embedding package comprising all evaluated embedding models. This package will assist users in easily selecting suitable embedding models for their data. Xikang Feng, Miaozhe Huo, Yongze Yang, Yuepeng Jiang, Shuaicheng Li 0001 |
Briefings Bioinform. | 1 |
| 2025 | PKDF-Net: Anticancer peptide prediction via a prior-knowledge-aware dual-path feature-entangled network
Qiangguo Jin, Ankang Wu, Leyi Wei, Hui Cui 0002, Ping Xuan, Xikang Feng, Ran Su |
Eng. Appl. Artif. Intell. | 6 |
| 2024 | DeepTAPE: Enhancing Systemic Lupus Erythematosus Diagnosis with Deep Learning Based on TCRβ CDR3 SequencesabstractSystemic Lupus Erythematosus (SLE) is a common and severe autoimmune disease driven by abnormal T cell responses, which are critical for understanding the disease’s immune pathology. Given the central role of the Complementarity Determining Region 3 (CDR3) of the TCRβ chain in T cell specificity, focused investigation into CDR3 may potentially enhance both the diagnostic accuracy and the mechanistic understanding of SLE. In this study, we developed DeepTAPE, a deep learning-based engine for predicting autoimmune diseases. It primarily uses CDR3 sequence features, supplemented by the frequency distribution of TCRβ’s V genes. This model is based on a multi-layer CNN-LSTM architecture with residual connections. DeepTAPE exhibits superior diagnostic performance for SLE when compared to existing gene-focused studies. It achieves an impressive average AUC of 97.99%, accuracy of 93.97%, and recall of 94.36%, reaffirming the diagnostic utility of CDR3 in SLE through deep learning. Building on this, we introduced a quantitative indicator based on the model, the Autoimmune Risk Score, which is positively correlated with clinical disease activity in SLE patients and assists in prognosis. Overall, DeepTAPE not only offers accurate diagnosis but also enhances our understanding of SLE. Tongfei Shen, Miaozhe Huo, Wan Nie, Kaiqi Li, Xikang Feng, Shuaicheng Li 0001 |
BIBM | 6 |
| 2024 | TSEML: A task-specific embedding-based method for few-shot classification of cancer molecular subtypesabstractMolecular subtyping of cancer is recognized as a critical and challenging upstream task for personalized therapy. Existing deep learning methods have achieved significant performance in this domain when abundant data samples are available. However, the acquisition of densely labeled samples for cancer molecular subtypes remains a significant challenge for conventional data-intensive deep learning approaches. In this work, we focus on the few-shot molecular subtype prediction problem in heterogeneous and small cancer datasets, aiming to enhance precise diagnosis and personalized treatment. We first construct a new few-shot dataset for cancer molecular subtype classification and auxiliary cancer classification, named TCGA Few-Shot, from existing publicly available datasets. To effectively leverage the relevant knowledge from both tasks, we introduce a task-specific embedding-based meta-learning framework (TSEML). TSEML leverages the synergistic strengths of a model-agnostic meta-learning (MAML) approach and a prototypical network (ProtoNet) to capture diverse and fine-grained features. Comparative experiments conducted on the TCGA FewShot dataset demonstrate that our TSEML framework achieves superior performance in addressing the problem of few-shot molecular subtype classification. Ran Su, Hui Cui 0002, Ping Xuan, Chengyan Fang, Xikang Feng, Qiangguo Jin |
BIBM | 6 |
| 2024 | MSKI-Net: Towards modality-specific knowledge interaction for glioma survival predictionabstractGliomas hold a prominent position in neurooncology due to their high malignancy and poor survival rates. Accurately predicting the prognosis and survival risk of glioma patients is crucial for clinical treatment. Recent advances in survival prediction methods emphasize the importance of integrating complementary information from diverse modalities while neglecting the significant modality gap between pathological images and genomic data. To address this issue, we propose a modality-specific knowledge interaction network (MSKI-Net), which integrates whole slide images (WSI), RNA-Seq gene expression data, and copy number variation (CNV) data for glioma survival analysis. The MSKI-Net consists of a modality-specific feature enhancement (MSFE) module, a modality-interactive cross-attention (MICA) module, and a modality-specific knowledge-guided representation learning (MSKR) module. The three modules collaborate by complementing modality-specific features with modality-agnostic knowledge to improve the learning capability of MSKI-Net. Furthermore, we construct a dataset named TCGAmm, which combines WSI, RNA-Seq, and CNV data from The Cancer Genome Atlas (TCGA) to address the issue of data scarcity. Extensive experiments demonstrate that MSKI-Net achieves superior performance in predicting the survival risk of glioma cancer. Ran Su, Hui Cui 0002, Ping Xuan, Xikang Feng, Leyi Wei, Qiangguo Jin |
BIBM | 5 |
| 2022 | Somatic variant analysis suite: copy number variation clonal visualization online platform for large-scale single-cell genomicsabstractThe recent advance of single-cell copy number variation (CNV) analysis plays an essential role in addressing intratumor heterogeneity, identifying tumor subgroups and restoring tumor-evolving trajectories at single-cell scale. Informative visualization of copy number analysis results boosts productive scientific exploration, validation and sharing. Several single-cell analysis figures have the effectiveness of visualizations for understanding single-cell genomics in published articles and software packages. However, they almost lack real-time interaction, and it is hard to reproduce them. Moreover, existing tools are time-consuming and memory-intensive when they reach large-scale single-cell throughputs. We present an online visualization platform, single-cell Somatic Variant Analysis Suite (scSVAS), for real-time interactive single-cell genomics data visualization. scSVAS is specifically designed for large-scale single-cell genomic analysis that provides an arsenal of unique functionalities. After uploading the specified input files, scSVAS deploys the online interactive visualization automatically. Users may conduct scientific discoveries, share interactive visualizations and download high-quality publication-ready figures. scSVAS provides versatile utilities for managing, investigating, sharing and publishing single-cell CNV profiles. We envision this online platform will expedite the biological understanding of cancer clonal evolution in single-cell resolution. All visualizations are publicly hosted at https://sc.deepomics.org. Lingxi Chen, Yuhao Qing, Ruikang Li, Chaohui Li, Hechen Li, Xikang Feng, Shuaicheng Li 0001 |
Briefings Bioinform. | 6 |
| 2022 | Resolving single-cell copy number profiling for large datasetsabstractThe advances of single-cell DNA sequencing (scDNA-seq) enable us to characterize the genetic heterogeneity of cancer cells. However, the high noise and low coverage of scDNA-seq impede the estimation of copy number variations (CNVs). In addition, existing tools suffer from intensive execution time and often fail on large datasets. Here, we propose SeCNV, an efficient method that leverages structural entropy, to profile the copy numbers. SeCNV adopts a local Gaussian kernel to construct a matrix, depth congruent map (DCM), capturing the similarities between any two bins along the genome. Then, SeCNV partitions the genome into segments by minimizing the structural entropy from the DCM. With the partition, SeCNV estimates the copy numbers within each segment for cells. We simulate nine datasets with various breakpoint distributions and amplitudes of noise to benchmark SeCNV. SeCNV achieves a robust performance, i.e. the F1-scores are higher than 0.95 for breakpoint detections, significantly outperforming state-of-the-art methods. SeCNV successfully processes large datasets (>50 000 cells) within 4 min, while other tools fail to finish within the time limit, i.e. 120 h. We apply SeCNV to single-nucleus sequencing datasets from two breast cancer patients and acoustic cell tagmentation sequencing datasets from eight breast cancer patients. SeCNV successfully reproduces the distinct subclones and infers tumor heterogeneity. SeCNV is available at https://github.com/deepomicslab/SeCNV. Yuwei Zhang 0005, Mengbo Wang 0001, Xikang Feng, Jianping Wang 0001, Shuaicheng Li 0001 |
Briefings Bioinform. | 4 |
| 2021 | Deep learning model reveals potential risk genes for ADHD, especially Ephrin receptor gene EPHA5abstractAttention deficit hyperactivity disorder (ADHD) is a common neurodevelopmental disorder. Although genome-wide association studies (GWAS) identify the risk ADHD-associated variants and genes with significant P-values, they may neglect the combined effect of multiple variants with insignificant P-values. Here, we proposed a convolutional neural network (CNN) to classify 1033 individuals diagnosed with ADHD from 950 healthy controls according to their genomic data. The model takes the single nucleotide polymorphism (SNP) loci of P-values $\le{1\times 10^{-3}}$, i.e. 764 loci, as inputs, and achieved an accuracy of 0.9018, AUC of 0.9570, sensitivity of 0.8980 and specificity of 0.9055. By incorporating the saliency analysis for the deep learning network, a total of 96 candidate genes were found, of which 14 genes have been reported in previous ADHD-related studies. Furthermore, joint Gene Ontology enrichment and expression Quantitative Trait Loci analysis identified a potential risk gene for ADHD, EPHA5 with a variant of rs4860671. Overall, our CNN deep learning model exhibited a high accuracy for ADHD classification and demonstrated that the deep learning model could capture variants' combining effect with insignificant P-value, while GWAS fails. To our best knowledge, our model is the first deep learning method for the classification of ADHD with SNPs data. Xikang Feng, Haimei Li, Shuaicheng Li 0001, Qiujin Qian |
Briefings Bioinform. | 2 |
| 2019 | MIRIA: a webserver for statistical, visual and meta-analysis of RNA editing data in mammalsabstractBACKGROUND: Adenosine-to-inosine RNA editing can markedly diversify the transcriptome, leading to a variety of critical molecular and biological processes in mammals. Over the past several years, researchers have developed several new pipelines and software packages to identify RNA editing sites with a focus on downstream statistical analysis and functional interpretation. RESULTS: Here, we developed a user-friendly public webserver named MIRIA that integrates statistics and visualization techniques to facilitate the comprehensive analysis of RNA editing sites data identified by the pipelines and software packages. MIRIA is unique in that provides several analytical functions, including RNA editing type statistics, genomic feature annotations, editing level statistics, genome-wide distribution of RNA editing sites, tissue-specific analysis and conservation analysis. We collected high-throughput RNA sequencing (RNA-seq) data from eight tissues across seven species as the experimental data for MIRIA and constructed an example result page. CONCLUSION: MIRIA provides both visualization and analysis of mammal RNA editing data for experimental biologists who are interested in revealing the functions of RNA editing sites. MIRIA is freely available at https://mammal.deepomics.org. Xikang Feng, Zishuai Wang, Hechen Li, Shuaicheng Li 0001 |
BMC Bioinform. | 1 |