EDBT 2026 Demo / reviewers in the wild / expert
Xiguo Yuan
dblp:86/7617
· DBLP profile ↗
19ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-0822-5189ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hi-Enhancer: a two-stage framework for prediction and localization of enhancers based on Blending-KAN and Stacking-Auto modelsabstractMOTIVATION: Gene expression plays a crucial role in cell function, and enhancers can regulate gene expression precisely. Therefore, accurate prediction of enhancers is particularly critical. However, existing prediction methods have low accuracy or rely on fixed multiple epigenetic signals, which may not always be available. RESULTS: We propose a two-stage framework that accurately predicts enhancers by flexibly combining multiple epigenetic signals. In the first stage, we designed a Blending-KAN model, which integrates the results of various base classifiers and employs Kolmogorov-Arnold Networks (KAN) as a meta-classifier to predict enhancers based on flexible combinations of multiple epigenetic signals. In the second stage, we developed a Stacking-Auto model, which extracted sequence features using DNABERT-2 and located the enhancers based on the Stacking strategy and AutoGluon framework. The accuracy of the Blending-KAN model reached 99.69 ± 0.11% when five epigenetic signals were used. In cross-cell line prediction, the accuracy was more significant than or equal to 93.72%. With Gaussian noise, it still maintains an accuracy of 98.74 ± 0.03%. In the second stage, the accuracy of the Stacking-Auto model is 80.50%, which is better than the existing 17 methods. The results show that our models can be flexibly used to predict and locate enhancers utilizing a combination of multiple epigenetic signals. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/emanlee/Hi-Enhancer and https://doi.org/10.6084/m9.figshare.29262158.v1. Rong Fei, Juntao Zou, Xiguo Yuan, Saurav Mallik, Xinhong Hei 0001, Lei Wang 0029 |
Bioinform. | 5 |
| 2026 | BiGAM-Net: A bilateral gated fusion framework for imbalanced breast imaging
A. K. Alvi Haque, Xiguo Yuan |
Expert Syst. Appl. | 4 |
| 2026 | Representation learning for 12-lead ECGs via dual-view conditional diffusion and lead-aware attention
Fanyi Yang, Xiguo Yuan |
Inf. Process. Manag. | 4 |
| 2025 | COPCNVBD: An Integrated Approach for Somatic Copy Number Variation and Breakpoint Detection Using Whole Genome Sequencing DataabstractRead depth (RD) signals anomaly-based copy number variation (CNV) detection methods using whole genome sequencing data are affected by the measurement scale and parameters, and the breakpoint of CNVs is easily influenced by the sizes of windows in preparing RD signals. In this study, we propose an integrated approach for somatic CNV and breakpoint detection, COPCNVBD, which builds an improved Copula-Based Outlier Detector (COPOD) anomaly detection method to infer approximate locations of CNVs without hyper-parameters. Then, COPCNVBD first regards the precise CNV breakpoint location as the image boundary detection and makes use of the pair-end mapping (PEM) reads information to design a CNV breakpoint identification strategy. We use simulation datasets with different tumor purity and coverage settings and real samples to demonstrate the performance of COPCNVBD and compare it with peer-popular tools. Simulation results show that COPCNVBD has a superior comprehensive performance even on low coverage data with low tumor purity. Results on three real cancer samples show that our proposed COPCNVBD is able to detect moderate CNVs and has a high consistency. The proposed COPCNVBD can be as a tool for the analysis of CNVs in the genome. Yajie Yan, Xiguo Yuan |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | LDSSNV: A Linkage Disequilibrium-Based Method for the Detection of Somatic Single-Nucleotide VariantsabstractSingle nucleotide variants (SNVs) are very common in human genome and pose a significant effect on cellular proliferation and tumorigenesis in various cancers. Somatic variant and germline variant are the two forms of SNVs. They are the major drivers of inherited diseases and acquired tumors respectively. A reasonable analysis of the next generation sequencing data profiles from cancer genomes could provide crucial information for cancer diagnosis and treatment. Accurate detection of SNVs and distinguishing the two forms are still considered challenging tasks in cancer analysis. Herein, we propose a new approach, LDSSNV, to detect somatic SNVs without matched normal samples. LDSSNV predicts SNVs by training the XGboost classifier on a concise combination of features and distinguishes the two forms based on linkage disequilibrium which is a trait between germline mutations. LDSSNV provides two modes to distinguish the somatic variants from germline variants, the single-mode and multiple-mode by respectively using a single tumor sample and multiple tumor samples. The performance of the proposed method is assessed on both simulation data and real sequencing datasets. The analysis shows that the LDSSNV method outperforms competing methods and can become a robust and reliable tool for analyzing tumor genome variation. Jingfen Lan, Wenxiang Chen, Ganggang Yin, Haque A. K. Alvi, Kun Xie 0011, Qiang Yu 0003, Xiguo Yuan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2022 | Greedy-mRMR: An emotion recognition algorithm based on EEG using greedy algorithmabstractElectroencephalography (EEG)-based emotion recognition methods mostly adopt features as many as possible to achieve high performance. However, this is not feasible and flexible for portable or wearable helmet devices because of the limited computing resource and requirement of real-time performance. Aiming at improving EEG emotion recognition performance with features as few as possible, this paper firstly designed a feature called Relative Intensity Ratio Entropy (RIRE), then proposed a novel feature selection method based on greedy algorithm and Max-Relevance and Min-Redundancy(mRMR), termed as Greedy-mRMR. Greedy algorithm is adopted to take top features in the ranking list, and mRMR algorithm is used to select features that are most relevant and minimum redundant. Greedy-mRMR algorithm also uses dynamic termination threshold to guarantee recognition accuracy in each iteration. Experiments were carried on DEAP dataset. SVM classifier achieved average classification accuracy of 91.16% in valence and 91.70% in arousal, by extracting RIRE feature and selecting with Greedy-mRMR algorithm. Experimental results show that RIRE is suitable for EEG emotion recognition and Greedy-mRMR outperforms state-of-the-art methods. Liying Yang 0001, Si Chao, Dunhui Liu, Xiguo Yuan |
BIBM | 5 |
| 2021 | NeuroTIS: Enhancing the prediction of translation initiation sites in mRNA sequences via a hybrid dependency network and deep learning framework
Xiguo Yuan, Zongzhen He |
Knowl. Based Syst. | 3 |
| 2021 | A Local Outlier Factor-Based Detection of Copy Number Variations From NGS DataabstractCopy number variation (CNV) is a major type of genomic structural variations that play an important role in human disorders. Next generation sequencing (NGS) has fueled the advancement in algorithm design to detect CNVs at base-pair resolution. However, accurate detection of CNVs of low amplitudes remains a challenging task. This paper proposes a new computational method, CNV-LOF, to identify CNVs of full-range amplitudes from NGS data. CNV-LOF is distinctly different from traditional methods, which mainly consider aberrations from a global perspective and rely on some assumed distribution of NGS read depths. In contrast, CNV-LOF takes a local view on the read depths and assigns an outlier factor to each genome segment. With the outlier factor profile, CNV-LOF uses a boxplot procedure to declare CNVs without the reliance of any distribution assumptions. Simulation experiments indicate that CNV-LOF outperforms five existing methods with respect to F1-measure, sensitivity, and precision. CNV-LOF is further validated on real sequencing samples, yielding highly consistent results with peer methods. CNV-LOF is able to detect CNVs of low and moderate amplitudes where the other existing methods fail, and it is expected to become a routine approach for the discovery of novel CNVs on whole sequencing genome. Xiguo Yuan, Junping Li, Jianing Xi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | STIC: Predicting Single Nucleotide Variants and Tumor Purity in Cancer GenomeabstractSingle nucleotide variant (SNV) plays an important role in cellular proliferation and tumorigenesis in various types of human cancer. Next-generation sequencing (NGS) has provided high-throughput data at an unprecedented resolution to predict SNVs. Currently, there exist many computational methods for either germline or somatic SNV discovery from NGS data, but very few of them are versatile enough to adapt to any situations. In the absence of matched normal samples, the prediction of somatic SNVs from single-tumor samples becomes considerably challenging, especially when the tumor purity is unknown. Here, we propose a new approach, STIC, to predict somatic SNVs and estimate tumor purity from NGS data without matched normal samples. The main features of STIC include: (1) extracting a set of SNV-relevant features on each site and training the BP neural network algorithm on the features to predict SNVs; (2) creating an iterative process to distinguish somatic SNVs from germline ones by disturbing allele frequency; and (3) establishing a reasonable relationship between tumor purity and allele frequencies of somatic SNVs to accurately estimate the purity. We quantitatively evaluate the performance of STIC on both simulation and real sequencing datasets, the results of which indicate that STIC outperforms competing methods. Xiguo Yuan, Haiyong Zhao, Liying Yang 0001, Shuzhen Wang, Jianing Xi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | ERINS: Novel Sequence Insertion Detection by Constructing an Extended ReferenceabstractNext generation sequencing technology has led to the development of methods for the detection of novel sequence insertions (nsINS). Multiple signatures from short reads are usually extracted to improve nsINS detection performance. However, characterization of nsINSs larger than the mean insert size is still challenging. This article presents a new method, ERINS, to detect nsINS contents and genotypes of full spectrum range size. It integrates the features of structural variations and mapping states of split reads to find nsINS breakpoints, and then adopts a left-most mapping strategy to infer nsINS content by iteratively extending the standard reference at each breakpoint. Finally, it realigns all reads to the extended reference and infers nsINS genotypes through statistical testing on read counts. We test and validate the performance of ERINS on simulation and real sequencing datasets. The simulation experimental results demonstrate that it outperforms several peer methods with respect to sensitivity and precision. The real data application indicates that ERINS obtains high consistent results with those of previously reported and detects nsINSs over 200 base pairs that many other methods fail. In conclusion, ERINS can be used as a supplement to existing tools and will become a routine approach for characterizing nsINSs. Xiguo Yuan, Xiangyan Xu, Haiyong Zhao, Junbo Duan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | CNV_IFTV: An Isolation Forest and Total Variation-Based Detection of CNVs from Short-Read Sequencing DataabstractAccurate detection of copy number variations (CNVs) from short-read sequencing data is challenging due to the uneven distribution of reads and the unbalanced amplitudes of gains and losses. The direct use of read depths to measure CNVs tends to limit performance. Thus, robust computational approaches equipped with appropriate statistics are required to detect CNV regions and boundaries. This study proposes a new method called CNV_IFTV to address this need. CNV_IFTV assigns an anomaly score to each genome bin through a collection of isolation trees. The trees are trained based on isolation forest algorithm through conducting subsampling from measured read depths. With the anomaly scores, CNV_IFTV uses a total variation model to smooth adjacent bins, leading to a denoised score profile. Finally, a statistical model is established to test the denoised scores for calling CNVs. CNV_IFTV is tested on both simulated and real data in comparison to several peer methods. The results indicate that the proposed method outperforms the peer methods. CNV_IFTV is a reliable tool for detecting CNVs from short-read sequencing data even for low-level coverage and tumor purity. The detection results on tumor samples can aid to evaluate known cancer genes and to predict target drugs for disease diagnosis. Xiguo Yuan, Jianing Xi, Liying Yang 0001, Junliang Shang, Junbo Duan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | Inferring subgroup-specific driver genes from heterogeneous cancer samples via subspace learning with subgroup indicationabstractMOTIVATION: Detecting driver genes from gene mutation data is a fundamental task for tumorigenesis research. Due to the fact that cancer is a heterogeneous disease with various subgroups, subgroup-specific driver genes are the key factors in the development of precision medicine for heterogeneous cancer. However, the existing driver gene detection methods are not designed to identify subgroup specificities of their detected driver genes, and therefore cannot indicate which group of patients is associated with the detected driver genes, which is difficult to provide specifically clinical guidance for individual patients. RESULTS: By incorporating the subspace learning framework, we propose a novel bioinformatics method called DriverSub, which can efficiently predict subgroup-specific driver genes in the situation where the subgroup annotations are not available. When evaluated by simulation datasets with known ground truth and compared with existing methods, DriverSub yields the best prediction of driver genes and the inference of their related subgroups. When we apply DriverSub on the mutation data of real heterogeneous cancers, we can observe that the predicted results of DriverSub are highly enriched for experimentally validated known driver genes. Moreover, the subgroups inferred by DriverSub are significantly associated with the annotated molecular subgroups, indicating its capability of predicting subgroup-specific driver genes. AVAILABILITY AND IMPLEMENTATION: The source code is publicly available at https://github.com/JianingXi/DriverSub. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jianing Xi, Xiguo Yuan, Ao Li 0001, Xuelong Li 0001, Qinghua Huang |
Bioinform. | 2 |
| 2020 | Comparative study of whole exome sequencing-based copy number variation detection toolsabstractBACKGROUND: With the rapid development of whole exome sequencing (WES), an increasing number of tools are being proposed for copy number variation (CNV) detection based on this technique. However, no comprehensive guide is available for the use of these tools in clinical settings, which renders them inapplicable in practice. To resolve this problem, in this study, we evaluated the performances of four WES-based CNV tools, and established a guideline for the recommendation of a suitable tool according to the application requirements. RESULTS: In this study, first, we selected four WES-based CNV detection tools: CoNIFER, cn.MOPS, CNVkit and exomeCopy. Then, we evaluated their performances in terms of three aspects: sensitivity and specificity, overlapping consistency and computational costs. From this evaluation, we obtained four main results: (1) The sensitivity increases and subsequently stabilizes as the coverage or CNV size increases, while the specificity decreases. (2) CoNIFER performs better for CNV insertions than for CNV deletions, while the remaining tools exhibit the opposite trend. (3) CoNIFER, cn.MOPS and CNVkit realize satisfactory overlapping consistency, which indicates their results are trustworthy. (4) CoNIFER has the best space complexity and cn.MOPS has the best time complexity among these four tools. Finally, we established a guideline for tools' usage according to these results. CONCLUSION: No available tool performs excellently under all conditions; however, some tools perform excellently in some scenarios. Users can obtain a CNV tool recommendation from our paper according to the targeted CNV size, the CNV type or computational costs of their projects, as presented in Table 1, which is helpful even for users with limited knowledge of computer science. Lanling Zhao, Xiguo Yuan, Junbo Duan |
BMC Bioinform. | 3 |
| 2020 | CONDEL: Detecting Copy Number Variation and Genotyping Deletion Zygosity from Single Tumor Samples Using Sequence DataabstractCharacterizing copy number variations (CNVs) from sequenced genomes is a both feasible and cost-effective way to search for driver genes in cancer diagnosis. A number of existing algorithms for CNV detection only explored part of the features underlying sequence data and copy number structures, resulting in limited performance. Here, we describe CONDEL, a method for detecting CNVs from single tumor samples using high-throughput sequence data. CONDEL utilizes a novel statistic in combination with a peel-off scheme to assess the statistical significance of genome bins, and adopts a Bayesian approach to infer copy number gains, losses, and deletion zygosity based on statistical mixture models. We compare CONDEL to six peer methods on a large number of simulation datasets, showing improved performance in terms of true positive and false positive rates, and further validate CONDEL on three real datasets derived from the 1000 Genomes Project and the EGA archive. CONDEL obtained higher consistent results in comparison with other three single sample-based methods, and exclusively identified a number of CNVs that were previously associated with cancers. We conclude that CONDEL is a powerful tool for detecting copy number variations on single tumor samples even if these are sequenced at low-coverage. Xiguo Yuan, Liying Yang 0001, Junbo Duan, Meihong Gao |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | SVSR: A Program to Simulate Structural Variations and Generate Sequencing Reads for Multiple PlatformsabstractStructural variation accounts for a major fraction of mutations in the human genome and confers susceptibility to complex diseases. Next generation sequencing along with the rapid development of computational methods provides a cost-effective procedure to detect such variations. Simulation of structural variations and sequencing reads with real characteristics is essential for benchmarking the computational methods. Here, we develop a new program, SVSR, to simulate five types of structural variations (indels, tandem duplication, CNVs, inversions, and translocations) and SNPs for the human genome and to generate sequencing reads with features from popular platforms (Illumina, SOLiD, 454, and Ion Torrent). We adopt a selection model trained from real data to predict copy number states, starting from the first site of a particular genome to the end. Furthermore, we utilize references of microbial genomes to produce insertion fragments and design probabilistic models to imitate inversions and translocations. Moreover, we create platform-specific errors and base quality profiles to generate normal, tumor, or normal-tumor mixture reads. Experimental results show that SVSR could capture more features that are realistic and generate datasets with satisfactory quality scores. SVSR is able to evaluate the performance of structural variation detection methods and guide the development of new computational methods. Xiguo Yuan, Meihong Gao, Junbo Duan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2018 | A new differential evolution algorithm for solving multimodal optimization problems with high dimensionality
Shouheng Tuo, Xiguo Yuan, Longquan Yong |
Soft Comput. | 3 |
| 2017 | Analysis of breast cancer subtypes by AP-ISA biclusteringabstractBACKGROUND: Gene expression profiling has led to the definition of breast cancer molecular subtypes: Basal-like, HER2-enriched, LuminalA, LuminalB and Normal-like. Different subtypes exhibit diverse responses to treatment. In the past years, several traditional clustering algorithms have been applied to analyze gene expression profiling. However, accurate identification of breast cancer subtypes, especially within highly variable LuminalA subtype, remains a challenge. Furthermore, the relationship between DNA methylation and expression level in different breast cancer subtypes is not clear. RESULTS: In this study, a modified ISA biclustering algorithm, termed AP-ISA, was proposed to identify breast cancer subtypes. Comparing with ISA, AP-ISA provides the optimized strategy to select seeds and thresholds in the circumstance that prior knowledge is absent. Experimental results on 574 breast cancer samples were evaluated using clinical ER/PR information, PAM50 subtypes and the results of five peer to peer methods. One remarkable point in the experiment is that, AP-ISA divided the expression profiles of the luminal samples into four distinct classes. Enrichment analysis and methylation analysis showed obvious distinction among the four subgroups. Tumor variability within the Luminal subtype is observed in the experiments, which could contribute to the development of novel directed therapies. CONCLUSIONS: Aiming at breast cancer subtype classification, a novel biclustering algorithm AP-ISA is proposed in this paper. AP-ISA classifies breast cancer into seven subtypes and we argue that there are four subtypes in luminal samples. Comparison with other methods validates the effectiveness of AP-ISA. New genes that would be useful for targeted treatment of breast cancer were also obtained in this study. Liying Yang 0001, Yunyan Shen, Xiguo Yuan, Jianhua Wei |
BMC Bioinform. | 3 |
| 2014 | AISAIC: a software suite for accurate identification of significant aberrations in cancersabstractUNLABELLED: Accurate identification of significant aberrations in cancers (AISAIC) is a systematic effort to discover potential cancer-driving genes such as oncogenes and tumor suppressors. Two major confounding factors against this goal are the normal cell contamination and random background aberrations in tumor samples. We describe a Java AISAIC package that provides comprehensive analytic functions and graphic user interface for integrating two statistically principled in silico approaches to address the aforementioned challenges in DNA copy number analyses. In addition, the package provides a command-line interface for users with scripting and programming needs to incorporate or extend AISAIC to their customized analysis pipelines. This open-source multiplatform software offers several attractive features: (i) it implements a user friendly complete pipeline from processing raw data to reporting analytic results; (ii) it detects deletion types directly from copy number signals using a Bayes hypothesis test; (iii) it estimates the fraction of normal contamination for each sample; (iv) it produces unbiased null distribution of random background alterations by iterative aberration-exclusive permutations; and (v) it identifies significant consensus regions and the percentage of homozygous/hemizygous deletions across multiple samples. AISAIC also provides users with a parallel computing option to leverage ubiquitous multicore machines. AVAILABILITY AND IMPLEMENTATION: AISAIC is available as a Java application, with a user's guide and source code, at https://code.google.com/p/aisaic/. Bai Zhang, Xuchu Hou, Xiguo Yuan, Ie-Ming Shih, Robert Clarke, Roger R. Wang, Subha Madhavan, Yue Joseph Wang, Guoqiang Yu |
Bioinform. | 3 |
| 2007 | Investigating Novel Immune-Inspired Multi-agent Systems for Anomaly DetectionabstractDue to the biological immune system applied to the field of computer security, immunological scientists have made much development for anomaly detection systems. However, there are still a number of significant hurdles to prevent it from solving real-world problems efficiently, such as the high false positive and false negative errors. In order to present a more feasible anomaly detection system, we outline multi-agent systems (MAS) to design an artificial immune system inspired by a novel immune theory- danger theory, following an appropriate evaluation tool (DCs) for network packets and a suitable mechanism of communication between agents. We set up two kinds of immune responses logically on both host layer and network layer to the coming intruders for the purpose of mitigating the damage and infection. We hope that this system will eventually become more powerful as a distributed immune system, based on the sound immunological concepts. Haidong Fu, Xiguo Yuan, Xiaolong Zhang 0002 |
APSCC | 2 |