Xuwen Wang

dblp:134/4618 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 PScnv: personalized self-normalizing CNV detection with a hierarchical multi-phase framework
abstract
MOTIVATION: Accurate detection of copy number variations (CNVs) from targeted panel sequencing remains challenging due to limited genomic coverage and pronounced sample-specific biases. Existing normalization strategies, including baseline-cohort, matched-control, and single-sample approaches, often struggle to balance noise suppression with adaptability, leading to inconsistent performance across heterogeneous samples. RESULTS: We present PScnv, a personalized self-normalizing framework for robust CNV detection from panel sequencing data. PScnv integrates a pre-built panel-of-normals (PoN) with sample-intrinsic stable chromosomes through ridge-regression normalization to generate individualized log2 ratio profiles with reduced systematic variation. CNVs are then identified using a hierarchical multi-phase segmentation pipeline incorporating z-score pre-partitioning, kernel-based correction, and circular binary segmentation. In 139 clinical tumor samples with orthogonal FISH validation at MET, ERBB2, and MTAP, PScnv showed improved accuracy and robustness over existing methods that do not require patient-matched normal samples, provided that a pre-built PoN cohort is available. AVAILABILITY: Source code is available for academic use at https://github.com/lvws/PScnv.
Xuwen Wang, Zhili Chang, Wansheng Lv, Akhatov Akmal, Xamidov Munis, Xunbiao Liu, Shenjie Wang, Xiaoyan Zhu 0003, Chong Du, Shuqun Zhang, Jiayin Wang 0002
Bioinform.1
2025 Neuropathvision: a Knowledge Graph-Enhanced LLM Agent for Patient-Centered Interpretation of Neuropathology Reports
abstract
Neuropathology reports are written for clinicians, making them difficult for patients to understand. Recent advances in large language models (LLMs) suggest a path to plain-language explanations for lay readers. In this study, we present NeuropathVision, a knowledge graph-enhanced LLM agent that generates patient-centered interpretations of neuropathology reports. The agent combines semantic retrieval from authoritative sources with a glioma-focused domain knowledge graph capturing entities and relations. It applies prompt-level style constraints such as clarity, transparent uncertainty, and empathy. To evaluate NeuropathVision, a nine-item patient-centered readability assessment form was developed covering four dimensions: Accuracy, Clarity, Empathy & Respect, and Treatment & Prognosis. Twenty cases were rated independently by two senior neuropathologists. All interpretations met the prespecified pass criteria, and the pooled mean score was 96.78. Exact agreement on item-level ratings was 86.7 %. Clarity and empathy scores were high, reflecting LLM strengths in following style guidance. The knowledge graph contributed structured anchors and consistency checks that supported accuracy. These results indicate the feasibility and acceptability of NeuropathVision for producing patient-centered neuropathology interpretations. Future work will add statementlevel citations and evaluate patient-reported outcomes.
Xuwen Wang, Xuchao Lang, Yuxiao Fu
BIBM1
2025 EMcnv: enhancing CNV detection performance through ensemble strategies with heterogeneous meta-graph neural networks
abstract
Copy number variation (CNV) is a crucial biomarker for many complex traits and diseases. Although numerous CNV detection tools are available, no single method consistently achieves optimal performance across diverse sequencing samples, as each tool has distinct advantages and limitations. Therefore, integrating the strengths of these tools to improve CNV detection accuracy is both a promising strategy and a significant challenge. To address this, we propose EMcnv, a novel deep ensemble framework based on meta-learning. EMcnv combines multiple CNV detection strategies through a three-step approach: (i) leveraging meta-learning and meta-path heterogeneous graphs, employing Relational Graph Convolutional Networks as a specific model within the Heterogeneous Graph Neural Networks framework to develop a probabilistic weight meta-model that ensembles various CNV detection strategies; (ii) assigning probabilistic weights to calls from different CNV detection tools and aggregating them into weighted CNV regions (CNVRs); (iii) refining Copy number variations based on weighted CNVRs. We conducted comprehensive experiments on both simulated and real sequencing data using benchmark datasets. The results demonstrate that EMcnv significantly outperforms popular existing methods, underscoring its superiority and importance in CNV detection. To support further research, the source code is available for academic use at https://github.com/Sherwin-xjtu/EMcnv.
Xuwen Wang, Zhili Chang, Yuqian Liu, Shenjie Wang, Xiaoyan Zhu 0003, Jiayin Wang 0002
Briefings Bioinform.1
2025 TMBquant: an explainable AI-powered caller advancing tumor mutation burden quantification across heterogeneous samples
abstract
Accurate tumor mutation burden (TMB) quantification is critical for immunotherapy stratification, yet remains challenging due to variability across sequencing platforms, tumor heterogeneity, and variant calling pipelines. Here, we introduce TMBquant, an explainable AI-powered caller designed to optimize TMB estimation through dynamic feature selection, ensemble learning, and automated strategy adaptation. Built upon the H2O AutoML framework, TMBquant integrates variant features, minimizes classification errors, and enhances both accuracy and stability across diverse datasets. We benchmarked TMBquant against nine widely used variant callers, including traditional tools (e.g. Mutect2, VarScan2, Strelka2) and recent AI-based methods (DeepSomatic, Octopus), using 706 whole-exome sequencing tumor-control pairs. To evaluate clinical relevance, we further assessed TMBquant through survival analyses across immunotherapy-treated cohorts of non-small cell lung cancer (NSCLC), nasopharyngeal carcinoma (NPC), and the two NSCLC subtypes: lung adenocarcinoma and lung squamous cell carcinoma. In each cohort, TMBquant consistently achieved the highest hazard ratios, demonstrating superior patient stratification compared to all other methods. Importantly, TMBquant maintained robust predictive performance across both high-TMB (NSCLC) and low-TMB (NPC) settings, highlighting its generalizability across cancer types with distinct biological characteristics. These findings establish TMBquant as a reliable, reproducible, and clinically actionable tool for precision oncology. The software is open source and freely available at https://github.com/SomaticCaller/SomaticCaller. To enhance reproducibility, we provide detailed usage instructions and representative code snippets for TMBquant in the Methods section (see Code Availability).
Shenjie Wang, Xiaoyan Zhu 0003, Xuwen Wang, Yuqian Liu, Minchao Zhao, Zhili Chang, Shuanying Yang, Jiayin Wang 0002
Briefings Bioinform.4
2025 PangenomeX: a graph convolutional network-based pangenome framework for unbiased population-scale genomic variation analysis
abstract
In population-scale genomic variation studies based on shallow whole genome sequencing, pangenomes have become an effective tool for identifying population-specific single-nucleotide polymorphisms and indels. Extending these advantages to copy number variation (CNV), however, remains challenging due to two unresolved issues. First, current pangenome frameworks exhibit pronounced population-representation bias arising from uneven sampling across populations. As the number of samples increases, the pangenome tends to capture variations primarily from majority populations while suppressing signals from minority populations. Second, in population-scale genomic variation analyses, common but benign population-specific copy number polymorphisms (CNPs) frequently obscure pathogenic CNVs. Existing pangenome frameworks lack dedicated mechanisms for representing CNPs and CNVs, limiting their ability to distinguish pathogenic CNVs from benign, population-specific CNPs. In this study, we present PangenomeX, a graph-convolutional pangenome framework tailored for low-coverage, population-scale CNV analysis. To address CNP representation, we embed known CNPs as prior knowledge into the pangenome graph and construct a CNV relationship network guided by a phylogenetic tree. A graph convolutional network (GCN) then learns the interactions between CNV and CNP nodes. To mitigate population-representation bias, the GCN aggregates information from only one- and two-hop neighborhoods, preserving local population context while preventing majority group signals from dominating. Evaluation on simulated cohorts and 561 real samples shows that PangenomeX distinguishes pathogenic CNVs from common population CNPs markedly better than existing methods. Overall, PangenomeX offers a methodological blueprint for large-cohort variant screening and provides a practical path for bringing graph-based genomics into clinical practice.
Zhengfa Xue, Yu Wang 0069, Xuwen Wang, Jiajing Yuan, Jingyu Zeng, Huanhuan Zhu, Jiayin Wang 0002
Briefings Bioinform.3
2025 ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencing
abstract
MOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome‑wide coverage of only 0.1-5×, sWGS data display a pronounced zero‑inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several‑fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy‑number gains (false positives), and true deletions often become indistinguishable from pervasive zero‑coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv.
Zhengfa Xue, Jingyu Zeng, Xuwen Wang, Jiajing Yuan, Xin Lai 0003, Yu Wang 0069, Huanhuan Zhu, Jiayin Wang 0002
Bioinform.3
2025 MedScaleRE-PF: a prompt-based framework with retrieval-augmented generation, chain-of-thought, and self-verification for scale-specific relation extraction in Chinese medical literature
abstract
Large language models have shown promise in biomedical natural language processing, yet their use in extracting structured knowledge from medical scales remains limited. This study introduces MedScaleRE-PF, a novel prompting framework designed for relation extraction in Chinese medical scale texts. The framework combines few-shot in-context learning with retrieval-augmented generation, chain-of-thought prompting, and self-verification strategies to improve contextual understanding and factual consistency. We constructed the CMedS-RE dataset, consisting of 606 full-text articles with 19,051 sentences, 29,359 annotated entities, and 7217 relation instances. Experiments were conducted on two tasks: relational triple extraction (RTE) and relation classification (RC). We evaluated both single-step and multi-step prompting, along with four self-verification strategies: direct (D-SV), stepwise (S-CoT-SV), relation-specific (R-CoT-SV), and stepwise relation-specific (SR-CoT-SV). The best results were achieved with single-step prompting and the R-CoT-SV strategy, yielding F1 scores of 42.58 % for RTE under the 32-shot setting and 65.42 % for RC under the 8-shot setting. Compared to a RAG-only baseline, this configuration improved F1 by 7.59 % on RTE and 1.07 % on RC. Additional experiments demonstrated strong performance under annotation-scarce conditions, achieving 46.99 % F1 on RTE with 20 training articles and 59.87 % on RC with 50 articles. Ablation and error analyses further confirmed that task-specific prompt structure and verification design significantly impact performance under few-shot conditions. MedScaleRE-PF also showed consistent results across multiple LLMs, confirming its stability and generalizability. These findings highlight the effectiveness of combining simple prompting and CoT-inspired verification in domain-specific information extraction. MedScaleRE-PF offers a flexible and structured approach for mining medical scale knowledge and supports prompt-based development in biomedical applications.
Zhenli Chen, Jiao Li 0001, Qinglong Peng, Xuwen Wang, Shan Cong, Liu Shen, Siyue Pu
Inf. Process. Manag.8
2025 MRDtarget: A heuristic Gaussian approach for optimizing targeted capture regions to enhance Minimal Residual Disease detection
abstract
Molecular residual disease (MRD) detection, initially developed for hematologic malignancies, has become a critical biomarker for monitoring solid tumors. MRD detection primarily relies on circulating tumor DNA (ctDNA) analysis using next-generation sequencing, offering high sensitivity and broad genomic coverage. However, challenges remain in designing cost-effective panels that maximize mutation detection while maintaining biological relevance. Fixed panels often lack sufficient patient-specific mutation coverage, while WES-based personalized MRD assays, despite their high sensitivity, are costly and less accessible. We developed a tumor comprehensive genomic profiling (CGP)-informed personalized MRD assay to detect tumor-derived mutations, which allowed us to design patient-specific personalized panels and meanwhile, provide a cost-effective alternative to whole exome sequencing (WES). To address these limitations, we developed MRDtarget, a heuristic multivariate Gaussian model-based targeted capture region selection method. By expanding beyond traditional hotspot regions, MRDtarget optimizes variant tracking for MRD detection, significantly improving sensitivity. Using a Bayesian inference-based heuristic approach, MRDtarget integrates multi-feature informativeness rates to identify optimal genomic regions for capture. Experimental results demonstrate that MRDtarget enables the detection of more variants per patient. This study underscores the importance of rational panel design to improve MRD sensitivity and provides a novel approach to enhance precision diagnostics and treatment for solid tumor patients.
Xuwen Wang, Yanfang Guan, Xin Lai 0003, Wuqiang Cao, Xiaoyan Zhu 0003, Xiaoling Zeng, Yuqian Liu, Shenjie Wang, Ruoyu Liu, Shuanying Yang, Jiayin Wang 0002
PLoS Comput. Biol.1
2024 Multi-Objective Policy Monitoring Method for Epidemic Control
abstract
In the face of emerging infectious diseases such as COVID-19, timely government intervention is crucial, as swift policy actions can effectively prevent greater losses. However, policymakers often need to balance multiple conflicting objectives. The lack of high-quality data and suitable analytical tools poses significant challenges for policy evaluation, especially in multi-objective decision-making, where accurately assessing the impact of interventions becomes even more difficult. To address this issue, this paper proposes a real-time data-driven policy monitoring method for dynamically tracking the effects of policy interventions. We introduce a new non-parametric one-sided test control chart, leveraging the interpretability and ease of implementation of control charts to monitor risk levels across various policies. Experimental results demonstrate the effectiveness of this method in policy monitoring.
Xin Lai 0003, Rundong Fan, Ruoyu Liu, Jiayin Wang 0002, Xiaoyan Zhu 0003, Yuqian Liu, Xuwen Wang, Shenjie Wang
BIBM7
2024 An Enhanced Multiple Correction Method with Limited Independent Effective SNPs
abstract
In genome-wide association studies (GWAS) and candidate gene studies (CGS), appropriate multiple testing correction methods can effectively address linkage disequilibrium (LD) blocks and are crucial for controlling the family-wise error rate (FWER) while ensuring the reliability of results. In our recent research, we observed that when the number of independent effective Single Nucleotide Polymorphisms (SNPs) is relatively small, existing multiple testing correction methods struggle to effectively control the FWER at 0.05 and lack robustness. This study proposes an enhanced method that builds upon the Moskvina and Schmidt approach by incorporating additional SNP correlation information, allowing for precise control of the FWER at 0.05 with limited independent effective SNPs. We also leverage Monte Carlo integration on a GPU to accelerate the computation of significance thresholds. Our method was evaluated through both simulation studies and real genotype datasets, demonstrating superior performance and robustness compared to existing methods.
Xin Lai 0003, Xiaohai Yang, Jiayin Wang 0002, Xiaoyan Zhu 0003, Yuqian Liu, Ruoyu Liu, Xuwen Wang, Shenjie Wang
BIBM7
2024 RMComBat: A Batch Effect Correction Algorithm for Repeated Measurement Sequencing Data to Prevent Overcorrection
abstract
Batch effects, caused by non-biological variations such as differences in laboratory conditions, reagent lots, or personnel, are a substantial source of noise in gene expression data. Accurately correcting these effects is crucial for valid biological inferences. However, the majority of existing batch effect correction algorithms are prone to overcorrection, where biologically meaningful signals are mistakenly identified as noise, especially in repeated measurement studies where time is confounded with batch. The failure to accurately distinguish between batch-related and biologically relevant variation leads to a loss of critical biological information. This paper presents RMComBat, an enhancement of the widely-used ComBat framework, which addresses this limitation by replacing the general linear model with a linear mixed-effects model. RMComBat incorporates subject-specific random intercepts to correct for sample correlation and enhance the preservation of biological signals. We tested RMComBat and several popular algorithms on simulated and real repeated measurement gene expression datasets, evaluating their performance through visual inspections and quantitative metrics. Results indicate that although most algorithms can reduce batch effects, they often do so at the cost of removing true biological signals. RMComBat demonstrates superior performance in preventing overcorrection, providing a more balanced and biologically informative correction in repeated measurement studies, so making it a valuable tool for improving the accuracy of gene expression analyses.
Yuqian Liu, Zhaoxing Wei, Jiayin Wang 0002, Xiaoyan Zhu 0003, Ruoyu Liu, Xuwen Wang, Shenjie Wang, Xin Lai 0003
BIBM6
2024 Enabling Adaptive CNV Detection through A Novel Predictive Control Framework
abstract
Accurate detection of copy number variations (CNVs) from sequencing data is crucial in many complex traits and diseases research. Although many CNV detection algorithms have been developed, challenges in precisely identifying CNVs persist. The core statistical model of these algorithms cannot self-adjust, which limits their adaptability to heterogeneous samples and reduces detection accuracy. address this challenge, we reframed the CNV detection problem as a quality control issue and incorporated adaptive mechanisms. We developed adapCNV, a novel adaptive CNV detection framework that integrates machine learning with optimization control. This framework enables dynamic adaptation of primary parameters based on sample features. We defined a quantifiable metric, RD fluctuation values, to assess signal characteristics when the algorithm accurately detects CNVs. We then employed machine learning techniques extract features from panel sequencing data, select initial parameter values for samples, and determine optimal RD fluctuation values. By adopting adaptive model predictive control (AMPC), adapCNV performs optimizations within rolling window. It dynamically adjusts the primary parameters based on error feedback from RD fluctuation values. This adaptive control strategy enables dynamic adjustment automatically match the characteristics of panel sequencing samples, significantly enhancing overall detection quality. The performance of this framework was validated with simulated data. Comparative analysis demonstrated that the proposed method outperforms the baseline approach, particularly in detecting small CNVs. The adapCNV framework is particularly suitable for panel sequencing, which may have broad applications in clinical practice. This novel approach from quality control perspective introduces a new paradigm for CNV detection.
Yuqian Liu, Jiajing Yuan, Xiaoyan Zhu 0003, Xin Lai 0003, Ruoyu Liu, Xuwen Wang, Jiayin Wang 0002
BIBM6
2024 Correction of Read Biases Induced by Complex Reference Genome Regions for Improving Copy Number Variation Detection Using a Gaussian Mixture Model
abstract
Copy number variations are crucial in cancer research, but their detection through next-generation sequencing is often hindered by read biases, particularly in complex genomic regions. Existing bias-correction methods address common issues like GC content but often fail in regions with repetitive sequences or segmental duplications, leading to false-positive CNVs. We propose refMask, a hybrid Gaussian model-based method that dynamically identifies low-confidence regions in the reference genome, correcting read biases and improving CNV detection accuracy. By integrating features from hg38 and T2T genomes, refMask tailors a custom blacklist for each sequencing sample, enhancing the reliability of CNV detection across diverse conditions. Our method provides a more accurate and flexible solution compared to current fixed blacklists, offering improved performance in challenging genomic regions.
Xuwen Wang, Zhili Chang, Shenjie Wang, Ruoyu Liu, Yuqian Liu, Xiaoyan Zhu 0003, Xin Lai 0003, Shuanying Yang, Jiayin Wang 0002
BIBM1
2024 TMBstable: a variant caller controls performance variation across heterogeneous sequencing samples
abstract
In cancer genomics, variant calling has advanced, but traditional mean accuracy evaluations are inadequate for biomarkers like tumor mutation burden, which vary significantly across samples, affecting immunotherapy patient selection and threshold settings. In this study, we introduce TMBstable, an innovative method that dynamically selects optimal variant calling strategies for specific genomic regions using a meta-learning framework, distinguishing it from traditional callers with uniform sample-wide strategies. The process begins with segmenting the sample into windows and extracting meta-features for clustering, followed by using a pre-trained meta-model to select suitable algorithms for each cluster, thereby addressing strategy-sample mismatches, reducing performance fluctuations and ensuring consistent performance across various samples. We evaluated TMBstable using both simulated and real non-small cell lung cancer and nasopharyngeal carcinoma samples, comparing it with advanced callers. The assessment, focusing on stability measures, such as the variance and coefficient of variation in false positive rate, false negative rate, precision and recall, involved 300 simulated and 106 real tumor samples. Benchmark results showed TMBstable's superior stability with the lowest variance and coefficient of variation across performance metrics, highlighting its effectiveness in analyzing the counting-based biomarker. The TMBstable algorithm can be accessed at https://github.com/hello-json/TMBstable for academic usage only.
Shenjie Wang, Xiaoyan Zhu 0003, Xuwen Wang, Yuqian Liu, Minchao Zhao, Zhili Chang, Jiayin Wang 0002
Briefings Bioinform.3
2022 PEcnv: accurate and efficient detection of copy number variations of various lengths
abstract
Copy number variation (CNV) is a class of key biomarkers in many complex traits and diseases. Detecting CNV from sequencing data is a substantial bioinformatics problem and a standard requirement in clinical practice. Although many proposed CNV detection approaches exist, the core statistical model at their foundation is weakened by two critical computational issues: (i) identifying the optimal setting on the sliding window and (ii) correcting for bias and noise. We designed a statistical process model to overcome these limitations by calculating regional read depths via an exponentially weighted moving average strategy. A one-run detection of CNVs of various lengths is then achieved by a dynamic sliding window, whose size is self-adopted according to the weighted averages. We also designed a novel bias/noise reduction model, accompanied by the moving average, which can handle complicated patterns and extend training data. This model, called PEcnv, accurately detects CNVs ranging from kb-scale to chromosome-arm level. The model performance was validated with simulation samples and real samples. Comparative analysis showed that PEcnv outperforms current popular approaches. Notably, PEcnv provided considerable advantages in detecting small CNVs (1 kb-1 Mb) in panel sequencing data. Thus, PEcnv fills the gap left by existing methods focusing on large CNVs. PEcnv may have broad applications in clinical testing where panel sequencing is the dominant strategy. Availability and implementation: Source code is freely available at https://github.com/Sherwin-xjtu/PEcnv.
Xuwen Wang, Ruoyu Liu, Xin Lai 0003, Yuqian Liu, Shenjie Wang, Xuanping Zhang, Jiayin Wang 0002
Briefings Bioinform.1
2019 GSDcreator: An Efficient and Comprehensive Simulator for Genarating NGS Data with Population Genetic Information
abstract
In recent decades, NGS data analysis has become a major research field in bioinformatics, which presents great advantages in many application scenarios. Many algorithms and software were designed for analyzing the NGS data, while simulation datasets are urgently needed for testing software and optimizing their parameter configurations. Thus, a series of NGS data simulators have been published. However, the existing simulators cannot satisfy the requirements from many specific scenarios. First, they do not support many newly discovered variations. Second, complex structural variations are difficult to generate. In addition, along with the increase of population data, it is urgent to increase population information simulation. In this paper, we propose GSDcreator, a comprehensive NGS simulator that overcome the three weaknesses mentioned above. It can produce all known types of variation, where the complex of variations are also supported. Furthermore, it can capture many important real data features including population polymorphism, insert size distribution, adjacent site depth distribution, overall depth distribution, quality score distribution, amplification bias, sequencing errors and so on. It's highlighted that 1000 Genomes Project Database is taken as a reference and integrates population genetic information to simulate population polymorphism. To test the performance, we did a lot of experiments and found that simulated data produced by GSDcreator are quit mimic to the real sequencing data.
Shenjie Wang, Jiayin Wang 0002, Xuanping Zhang, Xuwen Wang, Xiaoyan Zhu 0003, Xin Lai 0003
BIBM5
2019 FilterLAP: Filtering False-positive Mutation Calls via a Label Propagation Framework
abstract
Benefiting from the recent advantages of genomic sequencing, detecting genomic mutations becomes a routine work in precise diagnoses and treatments for cancers. In clinical practices, many factors, such as tumor purity, clonal structure, etc., interfere the performance of calling mutations. The computational pipelines prefer to sensitively report the candidate calls, while a filter is applied for removing the false-positive calls. The existing filters rely on the whole genome/exome sequencing data, which can provide sufficient samples for training the filters. However, the gene-panel sequencing is more popular in clinical practices, but there is no practical filter for limited training samples. In light of this, we develop a semi-learning filter for gene-panel sequencing data, FilterLAP, which implemented via a label propagation framework. Given few labeled samples with a set of unlabeled ones, its basic idea is to predict the label information of unlabeled nodes from the label information of labeled nodes, and establishes a complete graph model by using the relationship between samples, by combining transductive inference with label propagation algorithm. For each node in the network, tags are propagated to adjacent nodes according to similarity and the probability distribution of similar nodes tends to be similar and can be divided into a class. We perform multiple sets of experiments on gene-panel sequencing data captured from Illumina platform. FilterLAP outperforms on both SNV and INDEL filtering, where the AUCs reach 0.90-0.97, and the average accuracies on overall mutation calls are over 90%. Comparing to GATK hard filters, FilterLAP present a 5% improvement on accuracy. These results demonstrate that the proposed method can better reduce the false positive mutation calls on gene-panel sequencing data. In addition, it is stable and efficient, which can be used as a practical tool for mutation call filtering for gene-panel sequencing data.
Xuwen Wang, Xiaoyan Zhu 0003, Shenjie Wang, Xuanping Zhang, Xin Lai 0003, Jiayin Wang 0002
BIBM1
2019 Joints Relation Inference Network for Skeleton-Based Action Recognition
abstract
Recently, the graph convolutional networks based methods have achieved remarkable performance in skeleton-based action recognition. However, current methods do not make full use of the topology of the skeleton graph, because they set and fixed the relation of skeleton joints or adjacency matrix for all input samples manually. In order to solve the problem, we design an end-to-end architecture consisting of joints relation inference network (JRIN) and skeleton graph convolutional network (SGCN). JRIN can aggregate spatial-temporal feature of every two joints globally, then infer the optimal relation between every two joints automatically. These relations of joints is quantified as the optimal adjacency matrices. Then SGCN will take the optimal matrices to do the final action recognition. During training, special initialization and alternate training strategies are proposed to optimize the model. Our method achieves very competitive performance on widely used NTU-RGB+D and Kinetics datasets.
Fanfan Ye, Huiming Tang, Xuwen Wang
ICIP3
2019 farPPI: a webserver for accurate prediction of protein-ligand binding structures for small-molecule PPI inhibitors by MM/PB(GB)SA methods
abstract
SUMMARY: Protein-protein interactions (PPIs) have been regarded as an attractive emerging class of therapeutic targets for the development of new treatments. Computational approaches, especially molecular docking, have been extensively employed to predict the binding structures of PPI-inhibitors or discover novel small molecule PPI inhibitors. However, due to the relatively 'undruggable' features of PPI interfaces, accurate predictions of the binding structures for ligands towards PPI targets are quite challenging for most docking algorithms. Here, we constructed a non-redundant pose ranking benchmark dataset for small-molecule PPI inhibitors, which contains 900 binding poses for 184 protein-ligand complexes. Then, we evaluated the performance of MM/PB(GB)SA approaches to identify the correct binding poses for PPI inhibitors, including two Prime MM/GBSA procedures from the Schrödinger suite and seven different MM/PB(GB)SA procedures from the Amber package. Our results showed that MM/PBSA outperformed the Glide SP scoring function (success rate of 58.6%) and MM/GBSA in most cases, especially the PB3 procedure which could achieve an overall success rate of ∼74%. Moreover, the GB6 procedure (success rate of 68.9%) performed much better than the other MM/GBSA procedures, highlighting the excellent potential of the GBNSR6 implicit solvation model for pose ranking. Finally, we developed the webserver of Fast Amber Rescoring for PPI Inhibitors (farPPI), which offers a freely available service to rescore the docking poses for PPI inhibitors by using the MM/PB(GB)SA methods. AVAILABILITY AND IMPLEMENTATION: farPPI web server is freely available at http://cadd.zju.edu.cn/farppi/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhe Wang 0041, Xuwen Wang, Youyong Li, Tailong Lei, Ercheng Wang, Dan Li 0013, Yu Kang 0002, Feng Zhu 0004, Tingjun Hou
Bioinform.2
2015 Cross-lingual Pseudo Relevance Feedback Based on Weak Relevant Topic Alignment
Xuwen Wang, Xiaojie Wang 0006, Junlian Li
PACLIC1