EDBT 2026 Demo / reviewers in the wild / expert
Xin Lai 0003
dblp:24/8076-3
· DBLP profile ↗
26ranked-venue papers
2as first author
23since 2021 · last 2026
0000-0003-3850-5680ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts ModelsabstractSparse Mixture-of-Experts (SMoE) architectures have enabled a new frontier in scaling Large Language Models (LLMs), offering superior performance by activating only a fraction of their total parameters during inference. However, their practical deployment is severely hampered by substantial static memory overhead, as all experts must be loaded into memory. Existing post-training pruning methods, while reducing model size, often derive their pruning criteria from a single, general-purpose corpus. This leads to a critical limitation: a catastrophic performance degradation when the pruned model is applied to other domains, necessitating a costly re-pruning for each new domain. To address this generalization gap, we introduce Mosaic Pruning (MoP). The core idea of MoP is to construct a functionally comprehensive set of experts through a structured ``cluster-then-select" process. This process leverages a similarity metric that captures expert performance across different task domains to functionally cluster the experts, and subsequently selects the most representative expert from each cluster based on our proposed Activation Variability Score. Unlike methods that optimize for a single corpus, our proposed Mosaic Pruning ensures that the pruned model retains a functionally complementary set of experts, much like the tiles of a mosaic that together form a complete picture of the original model's capabilities, enabling it to handle diverse downstream tasks.Extensive experiments on various MoE models demonstrate the superiority of our approach. MoP significantly outperforms prior work, achieving a 7.24\% gain on general tasks and 8.92\% on specialized tasks like math reasoning and code generation. Mingkuan Zhao, Shuangyong Song, Xiaoyan Zhu 0003, Xin Lai 0003, Jiayin Wang 0002 |
AAAI | 5 |
| 2026 | Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-offabstractThe design of Large Language Models (LLMs) has long been hampered by a fundamental conflict within their core attention mechanism: its remarkable expressivity is built upon a computational complexity of O(H·N²) that grows quadratically with the context size (N) and linearly with the number of heads (H). This standard implementation harbors significant computational redundancy, as all heads independently compute attention over the same sequence space. Existing sparse methods, meanwhile, often trade information integrity for computational efficiency. To resolve this efficiency-performance trade-off, we propose SPAttention, whose core contribution is the introduction of a new paradigm we term Principled Structural Sparsity. SPAttention does not merely drop connections but instead reorganizes the computational task by partitioning the total attention workload into balanced, non-overlapping distance bands, assigning each head a unique segment. This approach transforms the multi-head attention mechanism from H independent O(N²) computations into a single, collaborative O(N²) computation, fundamentally reducing complexity by a factor of H. The structured inductive bias compels functional specialization among heads, enabling a more efficient allocation of computational resources from redundant modeling to distinct dependencies across the entire sequence span. Extensive empirical validation on the OLMoE-1B-7B and 0.25B-1.75B model series demonstrates that while delivering an approximately two-fold increase in training throughput, its performance is on par with standard dense attention, even surpassing it on select key metrics, while consistently outperforming representative sparse attention methods including Longformer, Reformer, and BigBird across all evaluation metrics. Our work demonstrates that thoughtfully designed structural sparsity can serve as an effective inductive bias that simultaneously improves both computational efficiency and model performance, opening a new avenue for the architectural design of next-generation, high-performance LLMs. Mingkuan Zhao, Jiayin Wang 0002, Xin Lai 0003, Tianchen Huang, Yuheng Min, Xiaoyan Zhu 0003 |
AAAI | 4 |
| 2026 | UF-CDDFM: A unified framework for code defect detection using multi-modal inputs and few-shot learningabstractContext: The detection of code defects is foundational to modern software development and maintenance, playing a critical role in ensuring software quality and security. However, as software systems grow in scale and complexity, the limitations of traditional static analysis and conventional machine learning techniques have become increasingly evident. These methods rely heavily on intricate, manual feature engineering and fail to capture dynamic runtime behavior, resulting in suboptimal accuracy and elevated error rates. Objective: To address these deficiencies, we propose UF-CDDFM, a unified framework for code defect detection that integrates multi-modal inputs, active learning, and state-of-the-art few-shot learning techniques. We aim to improve detection performance, reduce feature selection complexity and sample bias through active learning, and maintain practical efficiency in real-world development contexts. Methods: UF-CDDFM employs parallel encoding of source code, code annotations, and abstract syntax trees (ASTs) using large language models (LLMs) alongside multilayer perceptrons (MLPs) to derive robust, high-fidelity representations of code. To streamline feature selection and mitigate sample bias, an active learning component is introduced for automated identification of high-quality features. Addressing the pervasive challenge of data scarcity, we incorporate two complementary few-shot learning strategies-MAML for small-scale datasets and LEO for larger-scale settings to enhance overall generalization capability. Results: Empirical evaluations demonstrate that UF-CDDFM consistently outperforms existing methods, establishing new state-of-the-art detection rates: 72.04% for defect detection and 95.23% for clone detection. Crucially, these gains are achieved within resource-constrained computational environments, which highlights the practicality of the method. Conclusion: By fusing multi-modal code representations, active learning, and adaptive few-shot learning techniques, UF-CDDFM delivers significant improvements in detection accuracy and computational efficiency. This work offers a new paradigm for robust, scalable, and practical code defect and clone detection in modern software engineering. Xianglu Zhou, Tianxiang Cui, Xiaoyan Zhu 0003, Jiayin Wang 0002, Xin Lai 0003 |
Inf. Softw. Technol. | 5 |
| 2026 | Reaching Software Quality for Bioinformatics Applications: How Far Are We?abstractWith the rapid advancements in medicine, biology, and information technology, their deep integration has given rise to the emerging field of bioinformatics. In this process, high–throughput technologies such as genomics, transcriptomics, and proteomics have generated massive volumes of biological data. The biological significance of these data heavily relies on bioinformatics software for analysis and processing. Therefore, it is crucial for both scientific research and clinical applications to ensure the quality of bioinformatics software and avoiding errors or hidden defects. However, to date, no dedicated study has systematically analyzed the quality of bioinformatics software. we conduct a comprehensive empirical study that aggregates, synthesizes, and analyzes findings from 167 bioinformatics software projects. Following the Preferred Reporting Items for Systematic Review and Meta–Analysis (PRISMA) protocol, we extract and evaluate quality–related data to answer our research questions (RQs). Our analysis reveals several key findings. The quality of bioinformatics software requires significant improvement, with an average defect density approximately 11.8× higher than that of general-purpose software. Additionally, unlike traditional software domains, a considerable proportion of defects in bioinformatics software are related to annotations. These issues can lead developers to overlook potential security vulnerabilities or make incorrect fixes, thereby increasing the cost and complexity of subsequent code maintenance. Based on these findings, we further discuss the challenges faced by bioinformatics software and propose potential solutions. This paper lays a foundation for further research on software quality in the bioinformatics domain and offers actionable insights for researchers and practitioners alike. Xiaoyan Zhu 0003, Xin Lai 0003, Xin Lian, Hangyu Cheng, Jiayin Wang 0002 |
IEEE Trans. Software Eng. | 3 |
| 2025 | Multi-Label Ranking Loss Minimization for Matrix CompletionabstractThe common matrix completion methods minimize the rank of the matrix to be completed in addition to the Hamming loss between the incomplete and completed matrices. The rank of matrix measures the linear relation among the vectors of matrix, which may introduce ambiguity for data recovery. To cope with this issue, we extend multi-label ranking loss into matrix completion, and employ multi-label ranking loss minimization (MLRM) in this paper to exploit the relative correlation among matrix vectors. In MLRM, the original incomplete matrix is converted into a pairwise ranking matrix, and the approximation on this newly generated matrix can be viewed as a surrogate of multi-label ranking loss to replace the Hamming loss pattern in the existing methods. Extensive experiments demonstrate that MLRM outperforms the state-of-the-art matrix completion methods in varies of applications, including movie recommendation, drug-target interaction prediction and multi-label learning. Jiaxuan Li 0001, Xiaoyan Zhu 0003, Hongrui Wang 0004, Yu Zhang 0203, Xin Lai 0003, Jiayin Wang 0002 |
AAAI | 5 |
| 2025 | Cross-Project Defect Prediction Based on Feature Fusion and Local Domain AdaptationabstractCross-project defect prediction (CPDP) is hindered by distribution shifts between source and target projects, so models that excel in within-project software defect prediction (WPDP) often degrade across projects. We propose FLDP, which couples (i) local subset alignment selecting similar source-target file pairs via three file-level metrics and aligning only those subsets with (ii) sequence-graph feature fusion, where TLSTM encodes token sequences and TGCN encodes AST structure into a unified representation. Across 10 transfers on 7 projects, FLDP consistently outperforms classical and recent CPDP baselines in AUC/F1/MCC. Ablation shows both local alignment and fusion are necessary for the gains, and our analysis of selection metrics offers practical guidance for applying CPDP in heterogeneous settings. Xianglu Zhou, Xiaoyan Zhu 0003, Yu Wang 0069, Jiayin Wang 0002, Xin Lai 0003 |
APSEC | 5 |
| 2025 | THOR: a TMB heterogeneity-adaptive optimization model predicts immunotherapy response using clonal genomic features in group-structured dataabstractWith the increasing number of indications for immune checkpoint inhibitors in early and advanced cancers, the prospect of a tumor-agnostic biomarker to prioritize patients is compelling. Tumor mutation burden (TMB) is a widely endorsed biomarker that quantifies nonsynonymous mutations within tumor DNA, essential for neoantigen production, which, in turn, correlates with the immune response and guides decision-making. However, the general clinical application of TMB-relying on simple mutational counts targeted at a single endpoint-does not adequately capture the complex clonal structure of tumors nor the multifaceted nature of prognostic indicators. This recognition has spurred the exploration of sophisticated high-dimensional regression techniques. Unfortunately, the limited cohort sizes in immunotherapy trials have hindered the full potential of these advanced methods. Our approach considers patient subgroups as related yet distinct entities, enabling precise tailoring and refinement to address subgroup-specific dynamics. Given the deficiencies and the constraints, we introduce a TMB heterogeneity-optimized regression (THOR). This innovative model enhances the predictive capabilities of TMB by integrating tumor clonality and a diverse spectrum of clinical endpoints, further augmented by fusion techniques across subgroups to facilitate robust data sharing and interpretation. Our simulations validate THOR's superiority in parameter estimation for statistical inference. Clinically, we assess the utility of THOR in a structured cohort of 238 cancer patients undergoing immunotherapy, supplemented by 2212 patients across 19 subgroups from public datasets. The forecast of the responses and comparison of survival hazards demonstrate that THOR significantly enhances patient stratification and prognostic predictions by incorporating complex immunogenetic biology and subgroup-specific dynamics. Yanfang Guan, Xin Lai 0003, Yuqian Liu, Zhili Chang, Quan Wang 0004, Jian Zhao 0034, Shuanying Yang, Jiayin Wang 0002 |
Briefings Bioinform. | 3 |
| 2025 | MRDadaptis: self-adaptive parameter configuration enhances minimal residual disease detection in heterogeneous ctDNA samplesabstractDetection of structural variations (SVs) through circulating tumor DNA (ctDNA) has become a key method for detecting minimal residual disease (MRD). However, the heterogeneity of ctDNA samples, characterized by variable limits of detection (LOD) and diverse structural variant types, significantly impacts detection stability and performance, posing persistent challenges for conventional SV detection tools such as Delly and Manta. These widely used methods require extensive manual parameter tuning, hindered by the combinatorial complexity of multiple parameters and heterogeneous sequencing data. To address this, we propose MRDadaptis, a novel SV detection tool that uniquely incorporates a self-adaptive parameter optimization mechanism. MRDadaptis distinguishes itself by integrating Bayesian optimization with meta-learning techniques to dynamically adjust detection parameters automatically, based on intrinsic features derived from the ctDNA sequencing data itself. This innovative approach not only reduces manual intervention but also effectively captures sample-specific characteristics, significantly improving detection stability, and detection performance. Extensive validation experiments using both simulated and real-world ctDNA datasets demonstrates it distinct advantages, including markedly improved average F1-scores and superior stability (reduced variance, lower RMSE, increased kurtosis). These results highlight the significant advantages of MRDadaptis in addressing sample heterogeneity, underscoring its potential to improve the accuracy and reliability of MRD detecting through ctDNA analysis. https://github.com/aAT0047/MRDadaptis.git. Xin Lai 0003, Shenjie Wang, Zhengfa Xue, Yuqian Liu, Xiaoyan Zhu 0003, Zhili Chang, Jiayin Wang 0002 |
Briefings Bioinform. | 2 |
| 2025 | TMBclaw: tumor clone-aware graph learning improves immunotherapy response prediction across heterogeneous cohortsabstractImmune checkpoint inhibitors (ICIs) have emerged as a cornerstone of modern oncology, necessitating the development of robust biomarkers for optimizing patient stratification and treatment selection. While tumor mutation burden (TMB) has demonstrated prognostic value, conventional quantification methods based on mutation counts fail to reflect immunogenic neoantigen presentation due to intratumoral clonal heterogeneity. Recent efforts have focused on mutation subsets derived from tumor clonality, yet the complex interactions among clones remain a significant obstacle to accurate prognosis. This challenge is further exacerbated by the inherent constraints of limited cohort sizes in clinical studies, which severely compromise model generalizability across heterogeneous cohorts. Therefore, we propose TMBclaw (Tumor Mutation Burden-based Clonal attention with Laplacian Adaptive Weighting), a graph-regularized multi-task learning framework for immunotherapy response prediction. TMBclaw establishes unified integration of group-structured cohorts while enabling cross-cohort knowledge transfer and clonal relationship exploration. For clinical validation, we utilized four cohorts of 238 patients with non-small-cell lung cancer (NSCLC), melanoma, or nasopharyngeal carcinoma treated with ICIs, along with external multicenter validation cohorts (N = 1433) of melanoma and NSCLC patients from public datasets. Comparative analyses demonstrate that TMBclaw significantly outperforms conventional methods in prognostic accuracy and risk stratification. Through systematic quantification of clonal dynamics and discriminative identification of driver clones, TMBclaw shows potential to improve understanding of tumor heterogeneity and provides interpretable insights into the immunotherapy process. Xiaoyan Zhu 0003, Zhili Chang, Xin Lai 0003, Jiayin Wang 0002 |
Briefings Bioinform. | 8 |
| 2025 | MRDagent: iterative and adaptive parameter optimization for stable ctDNA-based MRD detection in heterogeneous samplesabstractMOTIVATION: Minimal residual disease (MRD) as critical biomarker for cancer prognosis and management plays a crucial role in improving patient outcomes. However, detecting MRD via next-generation sequencing-based circulating tumor DNA variant calling remains unstable due to the extremely low variant allele frequency and significant inter- and intra-sample heterogeneity. Although parameter optimization can theoretically enhance the detection performance of variants, achieving stable MRD detection remains challenging due to three key factors: (i) the necessity for individualized parameter tuning across numerous heterogeneous genomic intervals within each sample, (ii) the tightly interdependent parameter requirements across different stages of variant detection workflows, and (iii) the limitations of current automated parameter optimization methods. RESULTS: In this study, we propose MRDagent, a novel variant detection tool designed specifically for MRD detection. MRDagent incorporates an iterative and self-adaptive optimization framework capable of handling unknown objectives, varying constraints, and highly coupled parameters across stages. A key innovation of MRDagent is the integration of a convolutional neural network-based meta-model, trained on historical data to enable rapid parameter prediction. This significantly enhances computational efficiency and generalization performance. Extensive evaluations on simulated and real-world datasets demonstrate MRDagent's superior and stable performance, providing an efficient, reliable solution for MRD detection in clinical and high-throughput research applications. AVAILABILITY AND IMPLEMENTATION: MRDagent is freely available at https://github.com/aAT0047/MRDagent.git. The corresponding dataset and software archive are available at Zenodo: https://doi.org/10.5281/zenodo.15458496. Xin Lai 0003, Shenjie Wang, Yuqian Liu, Xiaoyan Zhu 0003, Jiayin Wang 0002 |
Bioinform. | 2 |
| 2025 | ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencingabstractMOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome‑wide coverage of only 0.1-5×, sWGS data display a pronounced zero‑inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several‑fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy‑number gains (false positives), and true deletions often become indistinguishable from pervasive zero‑coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv. Zhengfa Xue, Jingyu Zeng, Xuwen Wang, Jiajing Yuan, Xin Lai 0003, Yu Wang 0069, Huanhuan Zhu, Jiayin Wang 0002 |
Bioinform. | 6 |
| 2025 | MRDtarget: A heuristic Gaussian approach for optimizing targeted capture regions to enhance Minimal Residual Disease detectionabstractMolecular residual disease (MRD) detection, initially developed for hematologic malignancies, has become a critical biomarker for monitoring solid tumors. MRD detection primarily relies on circulating tumor DNA (ctDNA) analysis using next-generation sequencing, offering high sensitivity and broad genomic coverage. However, challenges remain in designing cost-effective panels that maximize mutation detection while maintaining biological relevance. Fixed panels often lack sufficient patient-specific mutation coverage, while WES-based personalized MRD assays, despite their high sensitivity, are costly and less accessible. We developed a tumor comprehensive genomic profiling (CGP)-informed personalized MRD assay to detect tumor-derived mutations, which allowed us to design patient-specific personalized panels and meanwhile, provide a cost-effective alternative to whole exome sequencing (WES). To address these limitations, we developed MRDtarget, a heuristic multivariate Gaussian model-based targeted capture region selection method. By expanding beyond traditional hotspot regions, MRDtarget optimizes variant tracking for MRD detection, significantly improving sensitivity. Using a Bayesian inference-based heuristic approach, MRDtarget integrates multi-feature informativeness rates to identify optimal genomic regions for capture. Experimental results demonstrate that MRDtarget enables the detection of more variants per patient. This study underscores the importance of rational panel design to improve MRD sensitivity and provides a novel approach to enhance precision diagnostics and treatment for solid tumor patients. Xuwen Wang, Yanfang Guan, Xin Lai 0003, Wuqiang Cao, Xiaoyan Zhu 0003, Xiaoling Zeng, Yuqian Liu, Shenjie Wang, Ruoyu Liu, Shuanying Yang, Jiayin Wang 0002 |
PLoS Comput. Biol. | 4 |
| 2024 | Multi-Objective Policy Monitoring Method for Epidemic ControlabstractIn the face of emerging infectious diseases such as COVID-19, timely government intervention is crucial, as swift policy actions can effectively prevent greater losses. However, policymakers often need to balance multiple conflicting objectives. The lack of high-quality data and suitable analytical tools poses significant challenges for policy evaluation, especially in multi-objective decision-making, where accurately assessing the impact of interventions becomes even more difficult. To address this issue, this paper proposes a real-time data-driven policy monitoring method for dynamically tracking the effects of policy interventions. We introduce a new non-parametric one-sided test control chart, leveraging the interpretability and ease of implementation of control charts to monitor risk levels across various policies. Experimental results demonstrate the effectiveness of this method in policy monitoring. Xin Lai 0003, Rundong Fan, Ruoyu Liu, Jiayin Wang 0002, Xiaoyan Zhu 0003, Yuqian Liu, Xuwen Wang, Shenjie Wang |
BIBM | 1 |
| 2024 | An Enhanced Multiple Correction Method with Limited Independent Effective SNPsabstractIn genome-wide association studies (GWAS) and candidate gene studies (CGS), appropriate multiple testing correction methods can effectively address linkage disequilibrium (LD) blocks and are crucial for controlling the family-wise error rate (FWER) while ensuring the reliability of results. In our recent research, we observed that when the number of independent effective Single Nucleotide Polymorphisms (SNPs) is relatively small, existing multiple testing correction methods struggle to effectively control the FWER at 0.05 and lack robustness. This study proposes an enhanced method that builds upon the Moskvina and Schmidt approach by incorporating additional SNP correlation information, allowing for precise control of the FWER at 0.05 with limited independent effective SNPs. We also leverage Monte Carlo integration on a GPU to accelerate the computation of significance thresholds. Our method was evaluated through both simulation studies and real genotype datasets, demonstrating superior performance and robustness compared to existing methods. Xin Lai 0003, Xiaohai Yang, Jiayin Wang 0002, Xiaoyan Zhu 0003, Yuqian Liu, Ruoyu Liu, Xuwen Wang, Shenjie Wang |
BIBM | 1 |
| 2024 | LMR-EWMA: A LASSO-based Multivariate Residual Control Chart for Monitoring Rare Health-Related EventsabstractMonitoring rare health-related events using control charts is crucial for timely detecting potential changes in healthcare scenarios. For example, sequentially testing the level of changes in infectious disease patient numbers helps prepare before an epidemic. Unlike general health-related events, the observation of rare ones often involves an excess of zeros, making it more appropriate to use the zero-inflated Poisson (ZIP) distribution rather than the classical Poisson. Although residual-based charts have attracted significant attention in this field, few studies have explored how to appropriately select residuals with different advantages in the complex situations like healthcare scenarios. Therefore, in this paper, we propose LMR-EWMA, a least absolute shrinkage and selection operator (LASSO)-based multivariate residual exponentially weighted moving average (EWMA) control chart, to automatically select the optimal residuals for monitoring changes (i.e., shifts) in the number of rare health-related events. Additionally, we have innovatively designed a bi-directional moving mechanism to address the limitation of current research in distinguishing the practical significance of shifts. Experimental results on three simulation cases and two real datasets demonstrate that LMR-EWMA outperforms existing charts in monitoring performance. Ruoyu Liu, Jiayin Wang 0002, Xiaoyan Zhu 0003, Yuqian Liu, Shuanying Yang, Xin Lai 0003 |
BIBM | 7 |
| 2024 | RMComBat: A Batch Effect Correction Algorithm for Repeated Measurement Sequencing Data to Prevent OvercorrectionabstractBatch effects, caused by non-biological variations such as differences in laboratory conditions, reagent lots, or personnel, are a substantial source of noise in gene expression data. Accurately correcting these effects is crucial for valid biological inferences. However, the majority of existing batch effect correction algorithms are prone to overcorrection, where biologically meaningful signals are mistakenly identified as noise, especially in repeated measurement studies where time is confounded with batch. The failure to accurately distinguish between batch-related and biologically relevant variation leads to a loss of critical biological information. This paper presents RMComBat, an enhancement of the widely-used ComBat framework, which addresses this limitation by replacing the general linear model with a linear mixed-effects model. RMComBat incorporates subject-specific random intercepts to correct for sample correlation and enhance the preservation of biological signals. We tested RMComBat and several popular algorithms on simulated and real repeated measurement gene expression datasets, evaluating their performance through visual inspections and quantitative metrics. Results indicate that although most algorithms can reduce batch effects, they often do so at the cost of removing true biological signals. RMComBat demonstrates superior performance in preventing overcorrection, providing a more balanced and biologically informative correction in repeated measurement studies, so making it a valuable tool for improving the accuracy of gene expression analyses. Yuqian Liu, Zhaoxing Wei, Jiayin Wang 0002, Xiaoyan Zhu 0003, Ruoyu Liu, Xuwen Wang, Shenjie Wang, Xin Lai 0003 |
BIBM | 8 |
| 2024 | Enabling Adaptive CNV Detection through A Novel Predictive Control FrameworkabstractAccurate detection of copy number variations (CNVs) from sequencing data is crucial in many complex traits and diseases research. Although many CNV detection algorithms have been developed, challenges in precisely identifying CNVs persist. The core statistical model of these algorithms cannot self-adjust, which limits their adaptability to heterogeneous samples and reduces detection accuracy. address this challenge, we reframed the CNV detection problem as a quality control issue and incorporated adaptive mechanisms. We developed adapCNV, a novel adaptive CNV detection framework that integrates machine learning with optimization control. This framework enables dynamic adaptation of primary parameters based on sample features. We defined a quantifiable metric, RD fluctuation values, to assess signal characteristics when the algorithm accurately detects CNVs. We then employed machine learning techniques extract features from panel sequencing data, select initial parameter values for samples, and determine optimal RD fluctuation values. By adopting adaptive model predictive control (AMPC), adapCNV performs optimizations within rolling window. It dynamically adjusts the primary parameters based on error feedback from RD fluctuation values. This adaptive control strategy enables dynamic adjustment automatically match the characteristics of panel sequencing samples, significantly enhancing overall detection quality. The performance of this framework was validated with simulated data. Comparative analysis demonstrated that the proposed method outperforms the baseline approach, particularly in detecting small CNVs. The adapCNV framework is particularly suitable for panel sequencing, which may have broad applications in clinical practice. This novel approach from quality control perspective introduces a new paradigm for CNV detection. Yuqian Liu, Jiajing Yuan, Xiaoyan Zhu 0003, Xin Lai 0003, Ruoyu Liu, Xuwen Wang, Jiayin Wang 0002 |
BIBM | 4 |
| 2024 | Correction of Read Biases Induced by Complex Reference Genome Regions for Improving Copy Number Variation Detection Using a Gaussian Mixture ModelabstractCopy number variations are crucial in cancer research, but their detection through next-generation sequencing is often hindered by read biases, particularly in complex genomic regions. Existing bias-correction methods address common issues like GC content but often fail in regions with repetitive sequences or segmental duplications, leading to false-positive CNVs. We propose refMask, a hybrid Gaussian model-based method that dynamically identifies low-confidence regions in the reference genome, correcting read biases and improving CNV detection accuracy. By integrating features from hg38 and T2T genomes, refMask tailors a custom blacklist for each sequencing sample, enhancing the reliability of CNV detection across diverse conditions. Our method provides a more accurate and flexible solution compared to current fixed blacklists, offering improved performance in challenging genomic regions. Xuwen Wang, Zhili Chang, Shenjie Wang, Ruoyu Liu, Yuqian Liu, Xiaoyan Zhu 0003, Xin Lai 0003, Shuanying Yang, Jiayin Wang 0002 |
BIBM | 8 |
| 2024 | TCSR: Self-attention with time and category for session-based recommendationabstractAbstract Session‐based recommendation that uses sequence of items clicked by anonymous users to make recommendations has drawn the attention of many researchers, and a lot of approaches have been proposed. However, there are still problems that have not been well addressed: (1) Time information is either ignored or exploited with a fixed time span and granularity, which fails to understand the personalized interest transfer pattern of users with different clicking speeds; (2) Category information is either omitted or considered independent of the items, which defies the fact that the relationships between categories and items are helpful for the recommendation. To solve these problems, we propose a new session‐based recommendation method, TCSR (self‐attention with time and category for session‐based recommendation). TCSR uses a non‐linear normalized time embedding to perceive user interest transfer patterns on variable granularity and employs a heterogeneous SAN to make full use of both items and categories. Moreover, a cross‐recommendation unit is adapted to adjust recommendations on the item and category sides. Extensive experiments on four real datasets show that TCSR significantly outperforms state‐of‐the‐art approaches. Xiaoyan Zhu 0003, Yu Zhang 0203, Jiaxuan Li 0001, Jiayin Wang 0002, Xin Lai 0003 |
Comput. Intell. | 5 |
| 2024 | A novel instance-based method for cross-project just-in-time defect predictionabstractSummary Cross‐project (CP) just‐in‐time software defect prediction (JIT‐SDP) uses CP data to overcome initial data scarcity for training high‐performing JIT‐SDP classifiers in the early stages of software projects. The primary challenge faced by JIT‐SDP in a cross‐project context lies in the distinct distributions between training and test data. To tackle this issue, we select source data instances that closely resemble target data for building classifiers. Software datasets commonly exhibit a class imbalance problem, where the ratio of the defective class to the clean class is notably low. This imbalance typically diminishes classifier performance. In this study, we propose an instance selection method utilizing kernel mean matching (ISKMM) that addresses both knowledge transfer and class imbalance in cross‐project defect prediction (CPDP). The method employs the kernel mean matching (KMM) technique to assess the similarity between training and target data. It selects instances with high similarity, retains them, and resamples the data based on similarity weighting to mitigate the class imbalance problem. Our experiments, conducted on 10 open‐source projects, reveal that the ISKMM algorithm outperforms existing CP single‐source software defect prediction (SDP) algorithms. Moreover, when employing the proposed algorithm, defect predictors constructed from cross‐project data demonstrate an overall performance comparable to predictors learned from within‐project data. Xiaoyan Zhu 0003, Jiayin Wang 0002, Xin Lai 0003 |
Softw. Pract. Exp. | 4 |
| 2023 | A Control Chart Method for Simultaneously Monitoring the Average Level and Stability of Surgical QualityabstractGood and stable surgical quality is of great significance to ensure the life safety of patients and spare patients unnecessary health burdens. The variable life-adjusted display (VLAD) is a popular assessment method for surgical quality and some VLAD-based control charts have been proposed to motivate the quality improvement. However, existing charts can only monitor the average level of surgical quality by detecting the changes of VLAD’s mean, but lack a mechanism to monitor the stability based on VLAD’s variance. The volatile surgical quality which can hardly be considered good will make existing charts provide delay alarms. Therefore, in this paper, we propose a risk-adjusted exponential weighted moving average (EWMA) control chart to monitor the mean and variance of VLAD simultaneously, named MVV-EWMA. Firstly, we give the explicit form of VLAD’s variance. Then, the EWMA statistics for the mean and variance are respectively constructed and integrated by the generalized likelihood ratio test. Moreover, two auxiliary mechanisms are adopted to enhance MVV-EWMA’s monitoring ability. Both the results of simulation and case study show that MVV-EWMA is not only able to effectively monitor the stability of surgical quality, but also has the best performance compared to existing chart. Ruoyu Liu, Xin Lai 0003, Jiayin Wang 0002, Paul B. S. Lai, Ka Chun Chong |
BIBM | 2 |
| 2022 | PEcnv: accurate and efficient detection of copy number variations of various lengthsabstractCopy number variation (CNV) is a class of key biomarkers in many complex traits and diseases. Detecting CNV from sequencing data is a substantial bioinformatics problem and a standard requirement in clinical practice. Although many proposed CNV detection approaches exist, the core statistical model at their foundation is weakened by two critical computational issues: (i) identifying the optimal setting on the sliding window and (ii) correcting for bias and noise. We designed a statistical process model to overcome these limitations by calculating regional read depths via an exponentially weighted moving average strategy. A one-run detection of CNVs of various lengths is then achieved by a dynamic sliding window, whose size is self-adopted according to the weighted averages. We also designed a novel bias/noise reduction model, accompanied by the moving average, which can handle complicated patterns and extend training data. This model, called PEcnv, accurately detects CNVs ranging from kb-scale to chromosome-arm level. The model performance was validated with simulation samples and real samples. Comparative analysis showed that PEcnv outperforms current popular approaches. Notably, PEcnv provided considerable advantages in detecting small CNVs (1 kb-1 Mb) in panel sequencing data. Thus, PEcnv fills the gap left by existing methods focusing on large CNVs. PEcnv may have broad applications in clinical testing where panel sequencing is the dominant strategy. Availability and implementation: Source code is freely available at https://github.com/Sherwin-xjtu/PEcnv. Xuwen Wang, Ruoyu Liu, Xin Lai 0003, Yuqian Liu, Shenjie Wang, Xuanping Zhang, Jiayin Wang 0002 |
Briefings Bioinform. | 4 |
| 2021 | Ensemble of ML-KNN for classification algorithm recommendation
Xiaoyan Zhu 0003, Chenzhen Ying, Jiayin Wang 0002, Jiaxuan Li 0001, Xin Lai 0003, Guangtao Wang |
Knowl. Based Syst. | 5 |
| 2019 | An Artificial Fish Swarm Algorithm for Identifying Associations between Multiple Variants and Multiple PhenotypesabstractIdentifying associations between genomic variants and phenotypes has always been an interesting research field of population genetics, which is of great significance for studying the pathogenesis of complex diseases and supporting clinical assistant decision making. Nowadays, many identification methods have been proposed to find the associations between variants and phenotypes, such as GWAS and pheWAS, and have made excellent achievements in pathological research and clinical practice. However, the existing methods only focus on single phenotype-multiple variants or single variant-multiple phenotypes, but not on multiple variants-multiple phenotypes. In the view of the fact that complex diseases often have several subtypes which differ greatly in variants and phenotypes, focusing only on single variant or single phenotype is far from enough and limits the ability of identification of those methods. Therefore, we propose a heuristic method with an AFSA framework on the solution space to identify associations between multiple variants and multiple phenotypes. In our method, each fish carries two logic trees that respectively represent the associations between variants and the associations between phenotypes. The logic trees will be iteratively updated to find a better solution according to the preset update strategies. When the iteration stop condition is reached, the algorithm will stop and output the optimal fish. The logical expression represented by the logic trees carried by the optimal fish is the associations we find. We validated the proposed method on the simulation data generated by hapgen2 and PhenotypeSimulator, and took the ratio of the number of people that can be explained by the found logical expression as the index to evaluate the performance, which was called Coverage. We conducted 9 groups of experiments, each of which was different in the number of variants and phenotypes. The best Coverage of was from the group including 500 variants and 10 phenotypes, which reached 72.12%, and the worst result is from the group including 100 variants and 20 phenotypes, 31.73%. We also exhausted the simulation data to find the optimal logical expression and several most important logic rules to evaluate the results obtained by the method. Ruoyu Liu, Xin Lai 0003, Xuanping Zhang, Xiaoyan Zhu 0003, Jiayin Wang 0002 |
BIBM | 2 |
| 2019 | GSDcreator: An Efficient and Comprehensive Simulator for Genarating NGS Data with Population Genetic InformationabstractIn recent decades, NGS data analysis has become a major research field in bioinformatics, which presents great advantages in many application scenarios. Many algorithms and software were designed for analyzing the NGS data, while simulation datasets are urgently needed for testing software and optimizing their parameter configurations. Thus, a series of NGS data simulators have been published. However, the existing simulators cannot satisfy the requirements from many specific scenarios. First, they do not support many newly discovered variations. Second, complex structural variations are difficult to generate. In addition, along with the increase of population data, it is urgent to increase population information simulation. In this paper, we propose GSDcreator, a comprehensive NGS simulator that overcome the three weaknesses mentioned above. It can produce all known types of variation, where the complex of variations are also supported. Furthermore, it can capture many important real data features including population polymorphism, insert size distribution, adjacent site depth distribution, overall depth distribution, quality score distribution, amplification bias, sequencing errors and so on. It's highlighted that 1000 Genomes Project Database is taken as a reference and integrates population genetic information to simulate population polymorphism. To test the performance, we did a lot of experiments and found that simulated data produced by GSDcreator are quit mimic to the real sequencing data. Shenjie Wang, Jiayin Wang 0002, Xuanping Zhang, Xuwen Wang, Xiaoyan Zhu 0003, Xin Lai 0003 |
BIBM | 7 |
| 2019 | FilterLAP: Filtering False-positive Mutation Calls via a Label Propagation FrameworkabstractBenefiting from the recent advantages of genomic sequencing, detecting genomic mutations becomes a routine work in precise diagnoses and treatments for cancers. In clinical practices, many factors, such as tumor purity, clonal structure, etc., interfere the performance of calling mutations. The computational pipelines prefer to sensitively report the candidate calls, while a filter is applied for removing the false-positive calls. The existing filters rely on the whole genome/exome sequencing data, which can provide sufficient samples for training the filters. However, the gene-panel sequencing is more popular in clinical practices, but there is no practical filter for limited training samples. In light of this, we develop a semi-learning filter for gene-panel sequencing data, FilterLAP, which implemented via a label propagation framework. Given few labeled samples with a set of unlabeled ones, its basic idea is to predict the label information of unlabeled nodes from the label information of labeled nodes, and establishes a complete graph model by using the relationship between samples, by combining transductive inference with label propagation algorithm. For each node in the network, tags are propagated to adjacent nodes according to similarity and the probability distribution of similar nodes tends to be similar and can be divided into a class. We perform multiple sets of experiments on gene-panel sequencing data captured from Illumina platform. FilterLAP outperforms on both SNV and INDEL filtering, where the AUCs reach 0.90-0.97, and the average accuracies on overall mutation calls are over 90%. Comparing to GATK hard filters, FilterLAP present a 5% improvement on accuracy. These results demonstrate that the proposed method can better reduce the false positive mutation calls on gene-panel sequencing data. In addition, it is stable and efficient, which can be used as a practical tool for mutation call filtering for gene-panel sequencing data. Xuwen Wang, Xiaoyan Zhu 0003, Shenjie Wang, Xuanping Zhang, Xin Lai 0003, Jiayin Wang 0002 |
BIBM | 6 |