VLDB 2026 Research / reviewers in the wild / expert
Zhengfa Xue
dblp:315/2343
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0000-8855-1285ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MRDadaptis: self-adaptive parameter configuration enhances minimal residual disease detection in heterogeneous ctDNA samplesabstractDetection of structural variations (SVs) through circulating tumor DNA (ctDNA) has become a key method for detecting minimal residual disease (MRD). However, the heterogeneity of ctDNA samples, characterized by variable limits of detection (LOD) and diverse structural variant types, significantly impacts detection stability and performance, posing persistent challenges for conventional SV detection tools such as Delly and Manta. These widely used methods require extensive manual parameter tuning, hindered by the combinatorial complexity of multiple parameters and heterogeneous sequencing data. To address this, we propose MRDadaptis, a novel SV detection tool that uniquely incorporates a self-adaptive parameter optimization mechanism. MRDadaptis distinguishes itself by integrating Bayesian optimization with meta-learning techniques to dynamically adjust detection parameters automatically, based on intrinsic features derived from the ctDNA sequencing data itself. This innovative approach not only reduces manual intervention but also effectively captures sample-specific characteristics, significantly improving detection stability, and detection performance. Extensive validation experiments using both simulated and real-world ctDNA datasets demonstrates it distinct advantages, including markedly improved average F1-scores and superior stability (reduced variance, lower RMSE, increased kurtosis). These results highlight the significant advantages of MRDadaptis in addressing sample heterogeneity, underscoring its potential to improve the accuracy and reliability of MRD detecting through ctDNA analysis. https://github.com/aAT0047/MRDadaptis.git. Xin Lai 0003, Shenjie Wang, Zhengfa Xue, Yuqian Liu, Xiaoyan Zhu 0003, Zhili Chang, Jiayin Wang 0002 |
Briefings Bioinform. | 4 |
| 2025 | PangenomeX: a graph convolutional network-based pangenome framework for unbiased population-scale genomic variation analysisabstractIn population-scale genomic variation studies based on shallow whole genome sequencing, pangenomes have become an effective tool for identifying population-specific single-nucleotide polymorphisms and indels. Extending these advantages to copy number variation (CNV), however, remains challenging due to two unresolved issues. First, current pangenome frameworks exhibit pronounced population-representation bias arising from uneven sampling across populations. As the number of samples increases, the pangenome tends to capture variations primarily from majority populations while suppressing signals from minority populations. Second, in population-scale genomic variation analyses, common but benign population-specific copy number polymorphisms (CNPs) frequently obscure pathogenic CNVs. Existing pangenome frameworks lack dedicated mechanisms for representing CNPs and CNVs, limiting their ability to distinguish pathogenic CNVs from benign, population-specific CNPs. In this study, we present PangenomeX, a graph-convolutional pangenome framework tailored for low-coverage, population-scale CNV analysis. To address CNP representation, we embed known CNPs as prior knowledge into the pangenome graph and construct a CNV relationship network guided by a phylogenetic tree. A graph convolutional network (GCN) then learns the interactions between CNV and CNP nodes. To mitigate population-representation bias, the GCN aggregates information from only one- and two-hop neighborhoods, preserving local population context while preventing majority group signals from dominating. Evaluation on simulated cohorts and 561 real samples shows that PangenomeX distinguishes pathogenic CNVs from common population CNPs markedly better than existing methods. Overall, PangenomeX offers a methodological blueprint for large-cohort variant screening and provides a practical path for bringing graph-based genomics into clinical practice. Zhengfa Xue, Yu Wang 0069, Xuwen Wang, Jiajing Yuan, Jingyu Zeng, Huanhuan Zhu, Jiayin Wang 0002 |
Briefings Bioinform. | 1 |
| 2025 | ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencingabstractMOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome‑wide coverage of only 0.1-5×, sWGS data display a pronounced zero‑inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several‑fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy‑number gains (false positives), and true deletions often become indistinguishable from pervasive zero‑coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv. Zhengfa Xue, Jingyu Zeng, Xuwen Wang, Jiajing Yuan, Xin Lai 0003, Yu Wang 0069, Huanhuan Zhu, Jiayin Wang 0002 |
Bioinform. | 1 |
| 2024 | Enhancing Dental Implant Risk Prediction with an Interpretable Multi-Instance Learning ModelabstractVariations in clinical and biological factors often lead to differing risks of dental implant failure among patients, even when undergoing similar procedures. Traditional predictive models often oversimplify outcomes into binary classifications, lacking the interpretability needed for accurate risk assessment and effective patient stratification. To address this challenge, this paper presents a novel multi-instance learning (MIL) framework that incorporates a Hosmer-Lemeshow-based loss function for implant failure risk assessment based on patient-specific clinical features. The framework also identifies the optimal threshold for key features, such as bone density, to enable robust patient stratification. The effectiveness label for each patient was constructed first and sampled patients into 600 and 1,000 groups. Results demonstrate that the proposed framework captures inter-patient variability with improved statistical calibration and enhanced risk stratification, offering practical insights to guide clinical decision-making. This study highlights the scalability and versatility of the proposed framework, bridging computational methodologies and practical applications in personalized implantology. Furthermore, the approach provides a transparent and effective tool for risk assessment, with potential applications in broader clinical stratification problems. Yuqian Liu, Almonzer Salah Nooraldaim, Zhengfa Xue, Jiayin Wang 0002 |
BIBM | 3 |
| 2024 | NIPT-PG: empowering non-invasive prenatal testing to learn from population genomics through an incremental pan-genomic approachabstractNon-invasive prenatal testing (NIPT) is a quite popular approach for detecting fetal genomic aneuploidies. However, due to the limitations on sequencing read length and coverage, NIPT suffers a bottleneck on further improving performance and conducting earlier detection. The errors mainly come from reference biases and population polymorphism. To break this bottleneck, we proposed NIPT-PG, which enables the NIPT algorithm to learn from population data. A pan-genome model is introduced to incorporate variant and polymorphic loci information from tested population. Subsequently, we proposed a sequence-to-graph alignment method, which considers the read mis-match rates during the mapping process, and an indexing method using hash indexing and adjacency lists to accelerate the read alignment process. Finally, by integrating multi-source aligned read and polymorphic sites across the pan-genome, NIPT-PG obtains a more accurate z-score, thereby improving the accuracy of chromosomal aneuploidy detection. We tested NIPT-PG on two simulated datasets and 745 real-world cell-free DNA sequencing data sets from pregnant women. Results demonstrate that NIPT-PG outperforms the standard z-score test. Furthermore, combining experimental and theoretical analyses, we demonstrate the probably approximately correct learnability of NIPT-PG. In summary, NIPT-PG provides a new perspective for fetal chromosomal aneuploidies detection. NIPT-PG may have broad applications in clinical testing, and its detection results can serve as a reference for false positive samples approaching the critical threshold. Zhengfa Xue, Aifen Zhou, Xiaoyan Zhu 0003, Huanhuan Zhu, Jiayin Wang 0002 |
Briefings Bioinform. | 1 |
| 2022 | MVE-FLK: A multi-task legal judgment prediction via multi-view encoder fusing legal keywords
Shuxin Yang, Suxin Tong, Guixiang Zhu, Jie Cao 0001, Youquan Wang, Zhengfa Xue |
Knowl. Based Syst. | 6 |