Shenjie Wang

dblp:81/9432 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 12 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PScnv: personalized self-normalizing CNV detection with a hierarchical multi-phase framework
abstract
MOTIVATION: Accurate detection of copy number variations (CNVs) from targeted panel sequencing remains challenging due to limited genomic coverage and pronounced sample-specific biases. Existing normalization strategies, including baseline-cohort, matched-control, and single-sample approaches, often struggle to balance noise suppression with adaptability, leading to inconsistent performance across heterogeneous samples. RESULTS: We present PScnv, a personalized self-normalizing framework for robust CNV detection from panel sequencing data. PScnv integrates a pre-built panel-of-normals (PoN) with sample-intrinsic stable chromosomes through ridge-regression normalization to generate individualized log2 ratio profiles with reduced systematic variation. CNVs are then identified using a hierarchical multi-phase segmentation pipeline incorporating z-score pre-partitioning, kernel-based correction, and circular binary segmentation. In 139 clinical tumor samples with orthogonal FISH validation at MET, ERBB2, and MTAP, PScnv showed improved accuracy and robustness over existing methods that do not require patient-matched normal samples, provided that a pre-built PoN cohort is available. AVAILABILITY: Source code is available for academic use at https://github.com/lvws/PScnv.
Xuwen Wang, Zhili Chang, Wansheng Lv, Akhatov Akmal, Xamidov Munis, Xunbiao Liu, Shenjie Wang, Xiaoyan Zhu 0003, Chong Du, Shuqun Zhang, Jiayin Wang 0002
Bioinform.7
2025 EMcnv: enhancing CNV detection performance through ensemble strategies with heterogeneous meta-graph neural networks
abstract
Copy number variation (CNV) is a crucial biomarker for many complex traits and diseases. Although numerous CNV detection tools are available, no single method consistently achieves optimal performance across diverse sequencing samples, as each tool has distinct advantages and limitations. Therefore, integrating the strengths of these tools to improve CNV detection accuracy is both a promising strategy and a significant challenge. To address this, we propose EMcnv, a novel deep ensemble framework based on meta-learning. EMcnv combines multiple CNV detection strategies through a three-step approach: (i) leveraging meta-learning and meta-path heterogeneous graphs, employing Relational Graph Convolutional Networks as a specific model within the Heterogeneous Graph Neural Networks framework to develop a probabilistic weight meta-model that ensembles various CNV detection strategies; (ii) assigning probabilistic weights to calls from different CNV detection tools and aggregating them into weighted CNV regions (CNVRs); (iii) refining Copy number variations based on weighted CNVRs. We conducted comprehensive experiments on both simulated and real sequencing data using benchmark datasets. The results demonstrate that EMcnv significantly outperforms popular existing methods, underscoring its superiority and importance in CNV detection. To support further research, the source code is available for academic use at https://github.com/Sherwin-xjtu/EMcnv.
Xuwen Wang, Zhili Chang, Yuqian Liu, Shenjie Wang, Xiaoyan Zhu 0003, Jiayin Wang 0002
Briefings Bioinform.4
2025 MRDadaptis: self-adaptive parameter configuration enhances minimal residual disease detection in heterogeneous ctDNA samples
abstract
Detection of structural variations (SVs) through circulating tumor DNA (ctDNA) has become a key method for detecting minimal residual disease (MRD). However, the heterogeneity of ctDNA samples, characterized by variable limits of detection (LOD) and diverse structural variant types, significantly impacts detection stability and performance, posing persistent challenges for conventional SV detection tools such as Delly and Manta. These widely used methods require extensive manual parameter tuning, hindered by the combinatorial complexity of multiple parameters and heterogeneous sequencing data. To address this, we propose MRDadaptis, a novel SV detection tool that uniquely incorporates a self-adaptive parameter optimization mechanism. MRDadaptis distinguishes itself by integrating Bayesian optimization with meta-learning techniques to dynamically adjust detection parameters automatically, based on intrinsic features derived from the ctDNA sequencing data itself. This innovative approach not only reduces manual intervention but also effectively captures sample-specific characteristics, significantly improving detection stability, and detection performance. Extensive validation experiments using both simulated and real-world ctDNA datasets demonstrates it distinct advantages, including markedly improved average F1-scores and superior stability (reduced variance, lower RMSE, increased kurtosis). These results highlight the significant advantages of MRDadaptis in addressing sample heterogeneity, underscoring its potential to improve the accuracy and reliability of MRD detecting through ctDNA analysis. https://github.com/aAT0047/MRDadaptis.git.
Xin Lai 0003, Shenjie Wang, Zhengfa Xue, Yuqian Liu, Xiaoyan Zhu 0003, Zhili Chang, Jiayin Wang 0002
Briefings Bioinform.3
2025 TMBquant: an explainable AI-powered caller advancing tumor mutation burden quantification across heterogeneous samples
abstract
Accurate tumor mutation burden (TMB) quantification is critical for immunotherapy stratification, yet remains challenging due to variability across sequencing platforms, tumor heterogeneity, and variant calling pipelines. Here, we introduce TMBquant, an explainable AI-powered caller designed to optimize TMB estimation through dynamic feature selection, ensemble learning, and automated strategy adaptation. Built upon the H2O AutoML framework, TMBquant integrates variant features, minimizes classification errors, and enhances both accuracy and stability across diverse datasets. We benchmarked TMBquant against nine widely used variant callers, including traditional tools (e.g. Mutect2, VarScan2, Strelka2) and recent AI-based methods (DeepSomatic, Octopus), using 706 whole-exome sequencing tumor-control pairs. To evaluate clinical relevance, we further assessed TMBquant through survival analyses across immunotherapy-treated cohorts of non-small cell lung cancer (NSCLC), nasopharyngeal carcinoma (NPC), and the two NSCLC subtypes: lung adenocarcinoma and lung squamous cell carcinoma. In each cohort, TMBquant consistently achieved the highest hazard ratios, demonstrating superior patient stratification compared to all other methods. Importantly, TMBquant maintained robust predictive performance across both high-TMB (NSCLC) and low-TMB (NPC) settings, highlighting its generalizability across cancer types with distinct biological characteristics. These findings establish TMBquant as a reliable, reproducible, and clinically actionable tool for precision oncology. The software is open source and freely available at https://github.com/SomaticCaller/SomaticCaller. To enhance reproducibility, we provide detailed usage instructions and representative code snippets for TMBquant in the Methods section (see Code Availability).
Shenjie Wang, Xiaoyan Zhu 0003, Xuwen Wang, Yuqian Liu, Minchao Zhao, Zhili Chang, Shuanying Yang, Jiayin Wang 0002
Briefings Bioinform.1
2025 MRDagent: iterative and adaptive parameter optimization for stable ctDNA-based MRD detection in heterogeneous samples
abstract
MOTIVATION: Minimal residual disease (MRD) as critical biomarker for cancer prognosis and management plays a crucial role in improving patient outcomes. However, detecting MRD via next-generation sequencing-based circulating tumor DNA variant calling remains unstable due to the extremely low variant allele frequency and significant inter- and intra-sample heterogeneity. Although parameter optimization can theoretically enhance the detection performance of variants, achieving stable MRD detection remains challenging due to three key factors: (i) the necessity for individualized parameter tuning across numerous heterogeneous genomic intervals within each sample, (ii) the tightly interdependent parameter requirements across different stages of variant detection workflows, and (iii) the limitations of current automated parameter optimization methods. RESULTS: In this study, we propose MRDagent, a novel variant detection tool designed specifically for MRD detection. MRDagent incorporates an iterative and self-adaptive optimization framework capable of handling unknown objectives, varying constraints, and highly coupled parameters across stages. A key innovation of MRDagent is the integration of a convolutional neural network-based meta-model, trained on historical data to enable rapid parameter prediction. This significantly enhances computational efficiency and generalization performance. Extensive evaluations on simulated and real-world datasets demonstrate MRDagent's superior and stable performance, providing an efficient, reliable solution for MRD detection in clinical and high-throughput research applications. AVAILABILITY AND IMPLEMENTATION: MRDagent is freely available at https://github.com/aAT0047/MRDagent.git. The corresponding dataset and software archive are available at Zenodo: https://doi.org/10.5281/zenodo.15458496.
Xin Lai 0003, Shenjie Wang, Yuqian Liu, Xiaoyan Zhu 0003, Jiayin Wang 0002
Bioinform.3
2025 MRDtarget: A heuristic Gaussian approach for optimizing targeted capture regions to enhance Minimal Residual Disease detection
abstract
Molecular residual disease (MRD) detection, initially developed for hematologic malignancies, has become a critical biomarker for monitoring solid tumors. MRD detection primarily relies on circulating tumor DNA (ctDNA) analysis using next-generation sequencing, offering high sensitivity and broad genomic coverage. However, challenges remain in designing cost-effective panels that maximize mutation detection while maintaining biological relevance. Fixed panels often lack sufficient patient-specific mutation coverage, while WES-based personalized MRD assays, despite their high sensitivity, are costly and less accessible. We developed a tumor comprehensive genomic profiling (CGP)-informed personalized MRD assay to detect tumor-derived mutations, which allowed us to design patient-specific personalized panels and meanwhile, provide a cost-effective alternative to whole exome sequencing (WES). To address these limitations, we developed MRDtarget, a heuristic multivariate Gaussian model-based targeted capture region selection method. By expanding beyond traditional hotspot regions, MRDtarget optimizes variant tracking for MRD detection, significantly improving sensitivity. Using a Bayesian inference-based heuristic approach, MRDtarget integrates multi-feature informativeness rates to identify optimal genomic regions for capture. Experimental results demonstrate that MRDtarget enables the detection of more variants per patient. This study underscores the importance of rational panel design to improve MRD sensitivity and provides a novel approach to enhance precision diagnostics and treatment for solid tumor patients.
Xuwen Wang, Yanfang Guan, Xin Lai 0003, Wuqiang Cao, Xiaoyan Zhu 0003, Xiaoling Zeng, Yuqian Liu, Shenjie Wang, Ruoyu Liu, Shuanying Yang, Jiayin Wang 0002
PLoS Comput. Biol.9
2025 Cross-Task Collaborative Meta-Learning for Cold-Start Recommendations
abstract
Optimizer-based meta-learning, specifically model-agnostic meta-learning (MAML), has emerged as a powerful tool for tackling the cold-start recommendation problem. In these meta-learning-based methods, recommendations for individual users are typically treated as separate tasks and learned independently. However, this task-by-task learning paradigm presents several observable limitations. First, learning one task at a time ignores inter-task correlations, i.e., collaborative signals, which limits the meta-model's receptive field and prevents it from leveraging valuable shared information, ultimately leading to subpar performance. Second, the meta-model is susceptible to the task distribution, i.e., the varied preference distributions among different users, which in turn introduces biases and inconsistencies, resulting in a less robust model that may perform well on certain user groups while underperforming on others. In this paper, we explore the correlations among different tasks in cold-start recommendations and develop a novel strategy termed cross-task collaborative meta-learning (CCML). More specifically, we propose a collaborative task sampling module designed to mitigate the adverse impact of irrelevant tasks during meta-model learning. This module adaptively identifies tasks that are both similar and beneficial to the primary task, ensuring that the meta-model learns from relevant and supportive information. Additionally, to harness collaborative information across relevant tasks, we introduce a bi-level cross-task meta-training strategy. This strategy leverages multi-task learning to capture collaborative knowledge simultaneously and enhance user profiling with pertinent information. Extensive experiments on four public benchmark datasets demonstrate the advantages of CCML over many state-of-the-art cold-start recommendation methods. Our results show significant improvements in recommendation accuracy and robustness, highlighting the potential of cross-task collaboration in enhancing meta-learning-based recommender systems. The code is available athttps://anonymous.4open.science/r/CCML-F064.
Yantong Du, Rui Chen 0012, Qiaoyu Tan, Qilong Han, Shenjie Wang, Xiangyu Zhao 0001
IEEE Trans. Knowl. Data Eng.5
2024 Multi-Objective Policy Monitoring Method for Epidemic Control
abstract
In the face of emerging infectious diseases such as COVID-19, timely government intervention is crucial, as swift policy actions can effectively prevent greater losses. However, policymakers often need to balance multiple conflicting objectives. The lack of high-quality data and suitable analytical tools poses significant challenges for policy evaluation, especially in multi-objective decision-making, where accurately assessing the impact of interventions becomes even more difficult. To address this issue, this paper proposes a real-time data-driven policy monitoring method for dynamically tracking the effects of policy interventions. We introduce a new non-parametric one-sided test control chart, leveraging the interpretability and ease of implementation of control charts to monitor risk levels across various policies. Experimental results demonstrate the effectiveness of this method in policy monitoring.
Xin Lai 0003, Rundong Fan, Ruoyu Liu, Jiayin Wang 0002, Xiaoyan Zhu 0003, Yuqian Liu, Xuwen Wang, Shenjie Wang
BIBM8
2024 An Enhanced Multiple Correction Method with Limited Independent Effective SNPs
abstract
In genome-wide association studies (GWAS) and candidate gene studies (CGS), appropriate multiple testing correction methods can effectively address linkage disequilibrium (LD) blocks and are crucial for controlling the family-wise error rate (FWER) while ensuring the reliability of results. In our recent research, we observed that when the number of independent effective Single Nucleotide Polymorphisms (SNPs) is relatively small, existing multiple testing correction methods struggle to effectively control the FWER at 0.05 and lack robustness. This study proposes an enhanced method that builds upon the Moskvina and Schmidt approach by incorporating additional SNP correlation information, allowing for precise control of the FWER at 0.05 with limited independent effective SNPs. We also leverage Monte Carlo integration on a GPU to accelerate the computation of significance thresholds. Our method was evaluated through both simulation studies and real genotype datasets, demonstrating superior performance and robustness compared to existing methods.
Xin Lai 0003, Xiaohai Yang, Jiayin Wang 0002, Xiaoyan Zhu 0003, Yuqian Liu, Ruoyu Liu, Xuwen Wang, Shenjie Wang
BIBM8
2024 RMComBat: A Batch Effect Correction Algorithm for Repeated Measurement Sequencing Data to Prevent Overcorrection
abstract
Batch effects, caused by non-biological variations such as differences in laboratory conditions, reagent lots, or personnel, are a substantial source of noise in gene expression data. Accurately correcting these effects is crucial for valid biological inferences. However, the majority of existing batch effect correction algorithms are prone to overcorrection, where biologically meaningful signals are mistakenly identified as noise, especially in repeated measurement studies where time is confounded with batch. The failure to accurately distinguish between batch-related and biologically relevant variation leads to a loss of critical biological information. This paper presents RMComBat, an enhancement of the widely-used ComBat framework, which addresses this limitation by replacing the general linear model with a linear mixed-effects model. RMComBat incorporates subject-specific random intercepts to correct for sample correlation and enhance the preservation of biological signals. We tested RMComBat and several popular algorithms on simulated and real repeated measurement gene expression datasets, evaluating their performance through visual inspections and quantitative metrics. Results indicate that although most algorithms can reduce batch effects, they often do so at the cost of removing true biological signals. RMComBat demonstrates superior performance in preventing overcorrection, providing a more balanced and biologically informative correction in repeated measurement studies, so making it a valuable tool for improving the accuracy of gene expression analyses.
Yuqian Liu, Zhaoxing Wei, Jiayin Wang 0002, Xiaoyan Zhu 0003, Ruoyu Liu, Xuwen Wang, Shenjie Wang, Xin Lai 0003
BIBM7
2024 Correction of Read Biases Induced by Complex Reference Genome Regions for Improving Copy Number Variation Detection Using a Gaussian Mixture Model
abstract
Copy number variations are crucial in cancer research, but their detection through next-generation sequencing is often hindered by read biases, particularly in complex genomic regions. Existing bias-correction methods address common issues like GC content but often fail in regions with repetitive sequences or segmental duplications, leading to false-positive CNVs. We propose refMask, a hybrid Gaussian model-based method that dynamically identifies low-confidence regions in the reference genome, correcting read biases and improving CNV detection accuracy. By integrating features from hg38 and T2T genomes, refMask tailors a custom blacklist for each sequencing sample, enhancing the reliability of CNV detection across diverse conditions. Our method provides a more accurate and flexible solution compared to current fixed blacklists, offering improved performance in challenging genomic regions.
Xuwen Wang, Zhili Chang, Shenjie Wang, Ruoyu Liu, Yuqian Liu, Xiaoyan Zhu 0003, Xin Lai 0003, Shuanying Yang, Jiayin Wang 0002
BIBM3
2024 TMBstable: a variant caller controls performance variation across heterogeneous sequencing samples
abstract
In cancer genomics, variant calling has advanced, but traditional mean accuracy evaluations are inadequate for biomarkers like tumor mutation burden, which vary significantly across samples, affecting immunotherapy patient selection and threshold settings. In this study, we introduce TMBstable, an innovative method that dynamically selects optimal variant calling strategies for specific genomic regions using a meta-learning framework, distinguishing it from traditional callers with uniform sample-wide strategies. The process begins with segmenting the sample into windows and extracting meta-features for clustering, followed by using a pre-trained meta-model to select suitable algorithms for each cluster, thereby addressing strategy-sample mismatches, reducing performance fluctuations and ensuring consistent performance across various samples. We evaluated TMBstable using both simulated and real non-small cell lung cancer and nasopharyngeal carcinoma samples, comparing it with advanced callers. The assessment, focusing on stability measures, such as the variance and coefficient of variation in false positive rate, false negative rate, precision and recall, involved 300 simulated and 106 real tumor samples. Benchmark results showed TMBstable's superior stability with the lowest variance and coefficient of variation across performance metrics, highlighting its effectiveness in analyzing the counting-based biomarker. The TMBstable algorithm can be accessed at https://github.com/hello-json/TMBstable for academic usage only.
Shenjie Wang, Xiaoyan Zhu 0003, Xuwen Wang, Yuqian Liu, Minchao Zhao, Zhili Chang, Jiayin Wang 0002
Briefings Bioinform.1
2022 PEcnv: accurate and efficient detection of copy number variations of various lengths
abstract
Copy number variation (CNV) is a class of key biomarkers in many complex traits and diseases. Detecting CNV from sequencing data is a substantial bioinformatics problem and a standard requirement in clinical practice. Although many proposed CNV detection approaches exist, the core statistical model at their foundation is weakened by two critical computational issues: (i) identifying the optimal setting on the sliding window and (ii) correcting for bias and noise. We designed a statistical process model to overcome these limitations by calculating regional read depths via an exponentially weighted moving average strategy. A one-run detection of CNVs of various lengths is then achieved by a dynamic sliding window, whose size is self-adopted according to the weighted averages. We also designed a novel bias/noise reduction model, accompanied by the moving average, which can handle complicated patterns and extend training data. This model, called PEcnv, accurately detects CNVs ranging from kb-scale to chromosome-arm level. The model performance was validated with simulation samples and real samples. Comparative analysis showed that PEcnv outperforms current popular approaches. Notably, PEcnv provided considerable advantages in detecting small CNVs (1 kb-1 Mb) in panel sequencing data. Thus, PEcnv fills the gap left by existing methods focusing on large CNVs. PEcnv may have broad applications in clinical testing where panel sequencing is the dominant strategy. Availability and implementation: Source code is freely available at https://github.com/Sherwin-xjtu/PEcnv.
Xuwen Wang, Ruoyu Liu, Xin Lai 0003, Yuqian Liu, Shenjie Wang, Xuanping Zhang, Jiayin Wang 0002
Briefings Bioinform.6
2019 GSDcreator: An Efficient and Comprehensive Simulator for Genarating NGS Data with Population Genetic Information
abstract
In recent decades, NGS data analysis has become a major research field in bioinformatics, which presents great advantages in many application scenarios. Many algorithms and software were designed for analyzing the NGS data, while simulation datasets are urgently needed for testing software and optimizing their parameter configurations. Thus, a series of NGS data simulators have been published. However, the existing simulators cannot satisfy the requirements from many specific scenarios. First, they do not support many newly discovered variations. Second, complex structural variations are difficult to generate. In addition, along with the increase of population data, it is urgent to increase population information simulation. In this paper, we propose GSDcreator, a comprehensive NGS simulator that overcome the three weaknesses mentioned above. It can produce all known types of variation, where the complex of variations are also supported. Furthermore, it can capture many important real data features including population polymorphism, insert size distribution, adjacent site depth distribution, overall depth distribution, quality score distribution, amplification bias, sequencing errors and so on. It's highlighted that 1000 Genomes Project Database is taken as a reference and integrates population genetic information to simulate population polymorphism. To test the performance, we did a lot of experiments and found that simulated data produced by GSDcreator are quit mimic to the real sequencing data.
Shenjie Wang, Jiayin Wang 0002, Xuanping Zhang, Xuwen Wang, Xiaoyan Zhu 0003, Xin Lai 0003
BIBM1
2019 FilterLAP: Filtering False-positive Mutation Calls via a Label Propagation Framework
abstract
Benefiting from the recent advantages of genomic sequencing, detecting genomic mutations becomes a routine work in precise diagnoses and treatments for cancers. In clinical practices, many factors, such as tumor purity, clonal structure, etc., interfere the performance of calling mutations. The computational pipelines prefer to sensitively report the candidate calls, while a filter is applied for removing the false-positive calls. The existing filters rely on the whole genome/exome sequencing data, which can provide sufficient samples for training the filters. However, the gene-panel sequencing is more popular in clinical practices, but there is no practical filter for limited training samples. In light of this, we develop a semi-learning filter for gene-panel sequencing data, FilterLAP, which implemented via a label propagation framework. Given few labeled samples with a set of unlabeled ones, its basic idea is to predict the label information of unlabeled nodes from the label information of labeled nodes, and establishes a complete graph model by using the relationship between samples, by combining transductive inference with label propagation algorithm. For each node in the network, tags are propagated to adjacent nodes according to similarity and the probability distribution of similar nodes tends to be similar and can be divided into a class. We perform multiple sets of experiments on gene-panel sequencing data captured from Illumina platform. FilterLAP outperforms on both SNV and INDEL filtering, where the AUCs reach 0.90-0.97, and the average accuracies on overall mutation calls are over 90%. Comparing to GATK hard filters, FilterLAP present a 5% improvement on accuracy. These results demonstrate that the proposed method can better reduce the false positive mutation calls on gene-panel sequencing data. In addition, it is stable and efficient, which can be used as a practical tool for mutation call filtering for gene-panel sequencing data.
Xuwen Wang, Xiaoyan Zhu 0003, Shenjie Wang, Xuanping Zhang, Xin Lai 0003, Jiayin Wang 0002
BIBM4
2010 Instantaneously companding baseband SC low-pass filter and ADC for 802.1 la/g WLAN receiver
abstract
To handle the 12dB peak-to-average-power ratio (PAPR) of OFDM signals in a 802.11a/g WLAN receiver baseband, an instantaneously companding system consisting of a 5-order low pass SC filter and a 10-bit pipeline ADC is presented. The filter cut-off and clock frequencies are 10MHz and 100MHz respectively and the ADC sampling frequency is 25MS/s. The filter provides the compressed output directly to the ADC and the signal expansion is done in the digital domain, which eliminates the need of an analog expansion amplifier. For a 12dB increase in dynamic range, it is estimated that the filter consumes 3.7 times less power than a conventional filter while the dynamic range required from the ADC is reduced by 12dB due to companding by a factor of 4. The filter and the ADC are designed to be implemented in a 1.2V, IBM 0.13μm CMOS process and the total power consumption is 75mW.
Shenjie Wang, Vaibhav Maheshwari, Wouter A. Serdijn
ISCAS1