VLDB 2026 Research / reviewers in the wild / expert
Xiaodan Fan
dblp:95/7112
· DBLP profile ↗
28ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-2744-9030ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | gSV: a general structural variant detector using the third-generation sequencing dataabstractStructural variants (SVs) are major contributors to genome diversity and disease susceptibility, particularly in cancer. Although third-generation sequencing technologies have substantially improved SV detection sensitivity, accurate detection of complex SVs remains challenging due to fragmented and heterogeneous alignment signals, as well as the dependence of many existing methods on predefined variant models. In this paper, we propose gSV, a general SV detector that integrates alignment-based and assembly-based approaches with the maximum exact match strategy, with particular emphasis on resolving SVs with complex or atypical alignment signatures. Without predefined assumptions about SV types, gSV captures diverse variant signals, enabling the detection of SVs that are usually missed by conventional tools. Benchmarking using both simulated datasets and real long-read sequencing data demonstrates that gSV achieves improved sensitivity and overall detection performance compared with current state-of-the-art SV callers, particularly for simple and complex SV events with complex alignment patterns. Unique SV discoveries in four breast cancer cell lines, particularly in cancer-associated genes, demonstrate the potential biological relevance of gSV-enabled discoveries. Furthermore, analysis of a breast cancer cohort from the Chinese population highlights the utility of gSV for population-scale genomic studies. Collectively, gSV provides a unified framework for comprehensive SV discovery in both research and clinical genomics settings. Jingyu Hao, Jiandong Shi, Sheng Lian, Zhen Zhang 0016, Yongyi Luo, Taobo Hu, Toyotaka Ishibashi, De-Peng Wang, Xiaodan Fan, Weichuan Yu |
Briefings Bioinform. | 10 |
| 2026 | RoBep: a region-oriented deep learning model for B-cell epitope predictionabstractMOTIVATION: Accurate in silico identification of B-cell epitope residues is crucial for antibody design and structure-guided vaccine development. Although recent protein language models and structure-aware methods can capture spatial information of tertiary structure when generating residue embeddings, most existing epitope predictors use these embeddings to perform classification for individual residues one by one, without enforcing spatial continuity for reported epitope residues. Such methods often result in biologically implausible predictions because B-cell epitope residues always cluster together on the antigen surface. RESULTS: We present RoBep, a region-oriented B-cell epitope predictor that explicitly models the spatial clustering of epitope residues. RoBep introduces a novel region constraint mechanism and combines the advanced protein language model ESM-Cambrian with an equivariant graph neural network. Our method outperforms existing structure-based methods on the benchmark dataset, demonstrating improvements of 26%, 45%, 13%, and 43% in F1, Matthews correlation coefficient, area under the precision-recall curve, and AUROC0.1, respectively. In addition to residue-level predictions, RoBep can also provide antibody-antigen binding regions. Importantly, the predicted epitope residues are ensured to be spatially compact, enhancing biological plausibility and practical relevance for immunotherapeutic design. AVAILABILITY AND IMPLEMENTATION: A user-friendly website for using RoBep is provided at https://huggingface.co/spaces/NielTT/RoBep. All datasets, source code used in this work, and implementation instructions of the website are publicly available at https://github.com/YitaoXU/RoBep. Guanyun Wei, Jingying Zhou, Yuanhua Huang, Weichuan Yu, Zhixiang Lin, Xiaodan Fan |
Bioinform. | 8 |
| 2026 | Interpretable integration of unpaired multi-omics for Alzheimer's diagnosis via cross-modal transformer reconstructionabstractAlzheimer's disease (AD) is a progressive neurodegenerative disorder with limited diagnostic tools and poorly understood molecular underpinnings. Although multi-omics technologies hold promise for early detection, integrating unpaired transcriptomic and epigenetic data remains a major challenge due to modality heterogeneity and small sample sizes. We present AE-Trans, an interpretable dual-channel Transformer framework that aligns RNA and DNA methylation data through cross-modal reconstruction and multi-head attention. AE-Trans achieves superior performance on prefrontal cortex datasets (accuracy = 0.9736, AUC = 0.9910) and demonstrates strong generalizability to external regions temporal cortex cohorts across brain regions (accuracy = 0.7389, AUC = 0.8432). To validate the performance within the same brain region, we tested AE-Trans on an external unpaired multi-omics dataset from the prefrontal cortex. Additionally, we validated the model on a paired multi-omics dataset to assess whether it could achieve good results in real-world scenarios. In the unpaired dataset from the external same brain region, AE-Trans achieved an accuracy of (accuracy = 0.87) and AUC of (AUC = 0.94), while in the real-world paired multi-omics dataset, the accuracy was (accuracy = 0.88) and AUC was (AUC = 0.93). These results demonstrate that AE-Trans not only validates well on external unpaired datasets, but also generalizes effectively to real-world multi-omics paired datasets, highlighting its robustness in practical applications. Through counterfactual integrated gradients, we identified key features associated with immune regulation, hormonal signaling, and neuronal metabolism. These were validated via pathway enrichment and logistic regression (AUC = 0.9749), confirming the biological relevance of model-derived markers. Furthermore, AE-Trans generalized well to two independent RNA datasets, where latent representations not only improved classification (AUCs = 0.92 and 0.89) but also stratified patients into subgroups with significantly different prognoses. These results highlight AE-Trans as a robust and explainable tool for multi-omics integration, supporting early diagnosis, biomarker discovery, and individualized risk prediction in Alzheimer's disease. Danfeng Du, Xiaodan Fan, Changshui Chen, Bowei Yan |
PLoS Comput. Biol. | 5 |
| 2026 | Iterative optimal transport for multimodal image registration
Mengyu Li 0001, Cheng Meng, Xiaodan Fan |
Pattern Recognit. | 3 |
| 2025 | RBPtool: A Deep Language Model Framework for Multi-Resolution RBP-RNA Binding Prediction and RNA Molecule DesignabstractJiyue Jiang, Yitao Xu, Zikang Wang, Yihan Ye, Yanruisheng Shao, Yuheng Shan, Jiuming Wang, Xiaodan Fan, Jiao Yuan, Yu Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jiyue Jiang, Zikang Wang, Yihan Ye, Yanruisheng Shao, Yuheng Shan, Jiuming Wang, Xiaodan Fan, Jiao Yuan, Yu Li 0006 |
EMNLP | 8 |
| 2025 | Double optimal transport for differential gene regulatory network inference with unpaired samplesabstractMOTIVATION: Inferring differential gene regulatory networks (GRNs) between different conditions from gene expression profiles remains a significant challenge. Current GRN inference approaches are limited by either scalability in large networks or accuracy in high-dimensional scenarios. Furthermore, most existing methods require paired samples for comparative GRN analyses. RESULTS: To overcome these challenges, we model gene regulation as a distribution transportation problem and propose an efficient and effective method, called double optimal transport (OT), for reconstructing differential GRNs from the perspective of optimal transport theory, applicable to unpaired samples. Double OT is a novel two-level OT framework. It first aligns unpaired samples by solving a partial OT problem at the sample level, and then infers GRNs from the aligned samples by solving a robust OT problem at the gene level. Comprehensive simulation studies demonstrate the superior efficiency and efficacy of double OT in different scales of networks compared to state-of-the-art methods. We also apply the proposed method to a gastric cancer dataset, identifying the proto-oncogene MET as a central node in the gastric cancer GRN. Its crucial role in early oncogenesis and potential as a therapeutic target further validate our approach and enhance our understanding of the regulatory mechanisms of gastric cancer. AVAILABILITY AND IMPLEMENTATION: A Python library that implements the proposed method is available at https://github.com/Mengyu8042/ot-grn. Mengyu Li 0001, Bencong Zhu, Cheng Meng, Xiaodan Fan |
Bioinform. | 4 |
| 2024 | NetMIM: network-based multi-omics integration with block missingness for biomarker selection and disease outcome predictionabstractCompared with analyzing omics data from a single platform, an integrative analysis of multi-omics data provides a more comprehensive understanding of the regulatory relationships among biological features associated with complex diseases. However, most existing frameworks for integrative analysis overlook two crucial aspects of multi-omics data. Firstly, they neglect the known dependencies among biological features that exist in highly credible biological databases. Secondly, most existing integrative frameworks just simply remove the subjects without full omics data to handle block missingness, resulting in decreasing statistical power. To overcome these issues, we propose a network-based integrative Bayesian framework for biomarker selection and disease outcome prediction based on multi-omics data. Our framework utilizes Dirac spike-and-slab variable selection prior to identifying a small subset of biomarkers. The incorporation of gene pathway information improves the interpretability of feature selection. Furthermore, with the strategy in the FBM (stand for "full Bayesian model with missingness") model where missing omics data are augmented via a mechanistic model, our framework handles block missingness in multi-omics data via a data augmentation approach. The real application illustrates that our approach, which incorporates existing gene pathway information and includes subjects without DNA methylation data, results in more interpretable feature selection results and more accurate predictions. Bencong Zhu, Zhen Zhang 0016, Suet Yi Leung, Xiaodan Fan |
Briefings Bioinform. | 4 |
| 2023 | A Bayesian model for identifying cancer subtypes from paired methylation profilesabstractAberrant DNA methylation is the most common molecular lesion that is crucial for the occurrence and development of cancer, but has thus far been underappreciated as a clinical tool for cancer classification, diagnosis or as a guide for therapeutic decisions. Partly, this has been due to a lack of proven algorithms that can use methylation data to stratify patients into clinically relevant risk groups and subtypes that are of prognostic importance. Here, we proposed a novel Bayesian model to capture the methylation signatures of different subtypes from paired normal and tumor methylation array data. Application of our model to synthetic and empirical data showed high clustering accuracy, and was able to identify the possible epigenetic cause of a cancer subtype. April S. Chan, Suet Yi Leung, Xiaodan Fan |
Briefings Bioinform. | 5 |
| 2023 | A Bayesian approach to estimate MHC-peptide binding thresholdabstractMajor histocompatibility complex (MHC)-peptide binding is a critical step in enabling a peptide to serve as an antigen for T-cell recognition. Accurate prediction of this binding can facilitate various applications in immunotherapy. While many existing methods offer good predictive power for the binding affinity of a peptide to a specific MHC, few models attempt to infer the binding threshold that distinguishes binding sequences. These models often rely on experience-based ad hoc criteria, such as 500 or 1000nM. However, different MHCs may have different binding thresholds. As such, there is a need for an automatic, data-driven method to determine an accurate binding threshold. In this study, we proposed a Bayesian model that jointly infers core locations (binding sites), the binding affinity and the binding threshold. Our model provided the posterior distribution of the binding threshold, enabling accurate determination of an appropriate threshold for each MHC. To evaluate the performance of our method under different scenarios, we conducted simulation studies with varying dominant levels of motif distributions and proportions of random sequences. These simulation studies showed desirable estimation accuracy and robustness of our model. Additionally, when applied to real data, our results outperformed commonly used thresholds. Ye-Fan Hu, Jian-Dong Huang, Xiaodan Fan |
Briefings Bioinform. | 4 |
| 2023 | Automatic block-wise genotype-phenotype association detection based on hidden Markov modelabstractBACKGROUND: For detecting genotype-phenotype association from case-control single nucleotide polymorphism (SNP) data, one class of methods relies on testing each genomic variant site individually. However, this approach ignores the tendency for associated variant sites to be spatially clustered instead of uniformly distributed along the genome. Therefore, a more recent class of methods looks for blocks of influential variant sites. Unfortunately, existing such methods either assume prior knowledge of the blocks, or rely on ad hoc moving windows. A principled method is needed to automatically detect genomic variant blocks which are associated with the phenotype. RESULTS: In this paper, we introduce an automatic block-wise Genome-Wide Association Study (GWAS) method based on Hidden Markov model. Using case-control SNP data as input, our method detects the number of blocks associated with the phenotype and the locations of the blocks. Correspondingly, the minor allele of each variate site will be classified as having negative influence, no influence or positive influence on the phenotype. We evaluated our method using both datasets simulated from our model and datasets from a block model different from ours, and compared the performance with other methods. These included both simple methods based on the Fisher's exact test, applied site-by-site, as well as more complex methods built into the recent Zoom-Focus Algorithm. Across all simulations, our method consistently outperformed the comparisons. CONCLUSIONS: With its demonstrated better performance, we expect our algorithm for detecting influential variant sites may help find more accurate signals across a wide range of case-control GWAS. Jin Du, Chaojie Wang 0008, Shanjun Mao, Bencong Zhu, Xiaodan Fan |
BMC Bioinform. | 7 |
| 2023 | Revealing Free Energy Landscape From MD Data via Conditional Angle Partition TreeabstractDeciphering the free energy landscape of biomolecular structure space is crucial for understanding many complex molecular processes, such as protein-protein interaction, RNA folding, and protein folding. A major source of current dynamic structure data is Molecular Dynamics (MD) simulations. Several methods have been proposed to investigate the free energy landscape from MD data, but all of them rely on the assumption that kinetic similarity is associated with global geometric similarity, which may lead to unsatisfactory results. In this paper, we proposed a new method called Conditional Angle Partition Tree to reveal the hierarchical free energy landscape by correlating local geometric similarity with kinetic similarity. Its application on the benchmark alanine dipeptide MD data showed a much better performance than existing methods in exploring and understanding the free energy landscape. We also applied it to the MD data of Villin HP35. Our results are more reasonable on various aspects than those from other methods and very informative on the hierarchical structure of its energy landscape. Hangjin Jiang, Wing Hung Wong, Xiaodan Fan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2023 | Family-Specific Training Improves Linear B Cell Epitope Prediction for Emerging VirusesabstractThe rational design of vaccines and antibody-based therapeutics against newly emerging viruses relies on B cell epitopes mainly. To predict the B cell epitopes of a novel virus, several algorithms have been developed. While most existing algorithms are trained on a dataset in which B cell epitopes are classified as 'Positive' or 'Negative'. However, we found that training on such data contaminates the target pattern of specific viruses, leading to inaccurate predictions in some cases. In this paper, we introduce a novel framework for predicting linear B cell epitopes of novel viruses by exclusively using highly similar viruses for training data. We employed kernel regression based on seropositive rates, which are the percentages of seropositive samples among the population, to predict the potential epitopes. To assess our method, we conducted simulations and utilized two real-world datasets. Our method significantly outperformed other existing methods on the testing data of four viruses with seropositive rates. Also, our strategy showed a better prediction in a larger dataset from the IEDB. Thus, a novel framework providing better linear B cell prediction of newly emerging viruses is established, which will benefit the rational design of vaccines and antibody-based therapeutics in the future. Ye-Fan Hu, Jin Du, Bao-Zhong Zhang, Thomas Yau, Xiaodan Fan, Jian-Dong Huang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2022 | High-dimensional correlation matrix estimation for general continuous data with Bagging techniqueabstractAbstract High-dimensional covariance matrix estimation plays a central role in multivariate statistical analysis. It is well-known that the sample covariance matrix is singular when the sample size is smaller than the dimension of the variable, but the covariance estimate must be positive-definite. This motivates some modifications of the sample covariance matrix to preserve its efficient estimation of pairwise covariance. In this paper, we modify the sample correlation matrix using the Bagging technique. The proposed Bagging estimator is flexible for general continuous data. Under some mild conditions, we show theoretically that the Bagging estimator can ensure positive-definiteness with probability one in finite samples. We also prove the consistency of the bootstrap estimator of Pearson correlation and the consistency of our Bagging estimator when the dimension p is fixed. Simulation results and a real application are provided to demonstrate that our method strikes a better balance between RMSE and likelihood, and is more robust, than other existing estimators. Chaojie Wang 0008, Jin Du, Xiaodan Fan |
Mach. Learn. | 3 |
| 2021 | Deep6mA: A deep learning framework for exploring similar patterns in DNA N6-methyladenine sites across different speciesabstractN6-methyladenine (6mA) is an important DNA modification form associated with a wide range of biological processes. Identifying accurately 6mA sites on a genomic scale is crucial for under-standing of 6mA's biological functions. However, the existing experimental techniques for detecting 6mA sites are cost-ineffective, which implies the great need of developing new computational methods for this problem. In this paper, we developed, without requiring any prior knowledge of 6mA and manually crafted sequence features, a deep learning framework named Deep6mA to identify DNA 6mA sites, and its performance is superior to other DNA 6mA prediction tools. Specifically, the 5-fold cross-validation on a benchmark dataset of rice gives the sensitivity and specificity of Deep6mA as 92.96% and 95.06%, respectively, and the overall prediction accuracy is 94%. Importantly, we find that the sequences with 6mA sites share similar patterns across different species. The model trained with rice data predicts well the 6mA sites of other three species: Arabidopsis thaliana, Fragaria vesca and Rosa chinensis with a prediction accuracy over 90%. In addition, we find that (1) 6mA tends to occur at GAGG motifs, which means the sequence near the 6mA site may be conservative; (2) 6mA is enriched in the TATA box of the promoter, which may be the main source of its regulating downstream gene expression. Zutan Li, Hangjin Jiang, Lingpeng Kong, Yuanyuan Chen 0014, Kun Lang, Xiaodan Fan, Liang-Yun Zhang, Cong Pian |
PLoS Comput. Biol. | 6 |
| 2020 | miR+Pathway: the integration and visualization of miRNA and KEGG pathwaysabstractmiRNAs represent a type of noncoding small molecule RNA. Many studies have shown that miRNAs are widely involved in the regulation of various pathways. The key to fully understanding the regulatory function of miRNAs is the determination of the pathways in which the miRNAs participate. However, the major pathway databases such as KEGG only include information regarding protein-coding genes. Here, we redesigned a pathway database (called miR+Pathway) by integrating and visualizing the 8882 human experimentally validated miRNA-target interactions (MTIs) and 150 KEGG pathways. This database is freely accessible at http://www.insect-genome.com/miR-pathway. Researchers can intuitively determine the pathways and the genes in the pathways that are regulated by miRNAs as well as the miRNAs that target the pathways. To determine the pathways in which targets of a certain miRNA or multiple miRNAs are enriched, we performed a KEGG analysis miRNAs by using the hypergeometric test. In addition, miR+Pathway provides information regarding MTIs, PubMed IDs and the experimental verification method. Users can retrieve pathways regulated by an miRNA or a gene by inputting its names. Cong Pian, Guang-Le Zhang, Libin Gao, Xiaodan Fan |
Briefings Bioinform. | 4 |
| 2020 | Can ODE gene regulatory models neglect time lag or measurement scaling?abstractMOTIVATION: Many ordinary differential equation (ODE) models have been introduced to replace linear regression models for inferring gene regulatory relationships from time-course gene expression data. But, since the observed data are usually not direct measurements of the gene products or there is an unknown time lag in gene regulation, it is problematic to directly apply traditional ODE models or linear regression models. RESULTS: We introduce a lagged ODE model to infer lagged gene regulatory relationships from time-course measurements, which are modeled as linear transformation of the gene products. A time-course microarray dataset from a yeast cell-cycle study is used for simulation assessment of the methods and real data analysis. The results show that our method, by considering both time lag and measurement scaling, performs much better than other linear and ODE models. It indicates the necessity of explicitly modeling the time lag and measurement scaling in ODE gene regulatory models. AVAILABILITY AND IMPLEMENTATION: R code is available at https://www.sta.cuhk.edu.hk/xfan/share/lagODE.zip. Jie Hu 0035, Huihui Qin, Xiaodan Fan |
Bioinform. | 3 |
| 2020 | MM-6mAPred: identifying DNA N6-methyladenine sites based on Markov modelabstractMOTIVATION: Recent studies have shown that DNA N6-methyladenine (6mA) plays an important role in epigenetic modification of eukaryotic organisms. It has been found that 6mA is closely related to embryonic development, stress response and so on. Developing a new algorithm to quickly and accurately identify 6mA sites in genomes is important for explore their biological functions. RESULTS: In this paper, we proposed a new classification method called MM-6mAPred based on a Markov model which makes use of the transition probability between adjacent nucleotides to identify 6mA site. The sensitivity and specificity of our method are 89.32% and 90.11%, respectively. The overall accuracy of our method is 89.72%, which is 6.59% higher than that of the previous method i6mA-Pred. It indicated that, compared with the 41 nucleotide chemical properties used by i6mA-Pred, the transition probability between adjacent nucleotides can capture more discriminant sequence information. AVAILABILITY AND IMPLEMENTATION: The web server of MM-6mAPred is freely accessible at http://www.insect-genome.com/MM-6mAPred/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Cong Pian, Guang-Le Zhang, Xiaodan Fan |
Bioinform. | 4 |
| 2020 | SOMM4mC: a second-order Markov model for DNA N4-methylcytosine site prediction in six speciesabstractMOTIVATION: DNA N4-methylcytosine (4mC) modification is an important epigenetic modification in prokaryotic DNA due to its role in regulating DNA replication and protecting the host DNA against degradation. An efficient algorithm to identify 4mC sites is needed for downstream analyses. RESULTS: In this study, we propose a new prediction method named SOMM4mC based on a second-order Markov model, which makes use of the transition probability between adjacent nucleotides to identify 4mC sites. The results show that the first-order and second-order Markov model are superior to the three existing algorithms in all six species (Caenorhabditis elegans, Drosophila melanogaster, Arabidopsis thaliana, Escherichia coli, Geoalkalibacter subterruneus and Geobacter pickeringii) where benchmark datasets are available. However, the classification performance of SOMM4mC is more outstanding than that of first-order Markov model. Especially, for E.coli and C.elegans, the overall accuracy of SOMM4mC are 91.8% and 87.6%, which are 8.5% and 6.1% higher than those of the latest method 4mcPred-SVM, respectively. This shows that more discriminant sequence information is captured by SOMM4mC through the dependency between adjacent nucleotides. AVAILABILITY AND IMPLEMENTATION: The web server of SOMM4mC is freely accessible at www.insect-genome.com/SOMM4mC. CONTACT: [email protected] or [email protected]. Jiali Yang, Kun Lang, Guang-Le Zhang, Xiaodan Fan, Yuanyuan Chen 0014, Cong Pian |
Bioinform. | 4 |
| 2018 | Guest Editorial for Special Section on the Sixth National Conference on Bioinformatics and System Biology of ChinaabstractThe three papers in this special section were presented at the Sixth National Conference on Bioinformatics and System Biology of China that was held in Nanjing, China, on October 6-9, 2014. Xiaodan Fan, Xinglai Ji, Rui Jiang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | Detect differentially methylated regions using non-homogeneous hidden Markov model for methylation array dataabstractMOTIVATION: DNA methylation is an important epigenetic mechanism in gene regulation and the detection of differentially methylated regions (DMRs) is enthralling for many disease studies. There are several aspects that we can improve over existing DMR detection methods: (i) methylation statuses of nearby CpG sites are highly correlated, but this fact has seldom been modelled rigorously due to the uneven spacing; (ii) it is practically important to be able to handle both paired and unpaired samples; and (iii) the capability to detect DMRs from a single pair of samples is demanded. RESULTS: We present DMRMark (DMR detection based on non-homogeneous hidden Markov model), a novel Bayesian framework for detecting DMRs from methylation array data. It combines the constrained Gaussian mixture model that incorporates the biological knowledge with the non-homogeneous hidden Markov model that models spatial correlation. Unlike existing methods, our DMR detection is achieved without predefined boundaries or decision windows. Furthermore, our method can detect DMRs from a single pair of samples and can also incorporate unpaired samples. Both simulation studies and real datasets from The Cancer Genome Atlas showed the significant improvement of DMRMark over other methods. AVAILABILITY AND IMPLEMENTATION: DMRMark is freely available as an R package at the CRAN R package repository. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Linghao Shen, Shuo-Yen Robert Li, Xiaodan Fan |
Bioinform. | 4 |
| 2017 | BPP: a sequence-based algorithm for branch point predictionabstractMOTIVATION: Although high-throughput sequencing methods have been proposed to identify splicing branch points in the human genome, these methods can only detect a small fraction of the branch points subject to the sequencing depth, experimental cost and the expression level of the mRNA. An accurate computational model for branch point prediction is therefore an ongoing objective in human genome research. RESULTS: We here propose a novel branch point prediction algorithm that utilizes information on the branch point sequence and the polypyrimidine tract. Using experimentally validated data, we demonstrate that our proposed method outperforms existing methods. Availability and implementation: https://github.com/zhqingit/BPP. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qing Zhang 0019, Xiaodan Fan, Yejun Wang, Ming-an Sun, Jianlin Shao, Dianjing Guo |
Bioinform. | 2 |
| 2016 | Variable selection and prediction of clinical outcome with multiply-imputed data via Bayesian model averagingabstractMultiple imputation (MI) is increasingly used to deal with missing data in medical studies, whilst variable selection and prediction on multiply-imputed data is an area under intense research in statistics. A commonly used strategy is to select a single top model based on the Rubin's rules (RR). However, such approaches do not take the model uncertainty into consideration, which might lead to over-confident inferences. In this paper, we extended the Bayesian model averaging method to perform variable selection and prediction under multiple imputation (MI-BMA), which takes into account the uncertainties originated from both the missing data and the model selection. We applied the MI-BMA method to simulated datasets as well as a real data set from a prospective cohort, and demonstrated the advantage of our method as compared with the classical RR stepwise method. Guozhi Jiang, Claudia H. Tam, Andrea O. Y. Luk, Alice P. S. Kong, Wing-Yee So, Juliana C. Chan, Ronald C. Ma, Xiaodan Fan |
BIBM | 8 |
| 2015 | A Resampling Based Clustering Algorithm for Replicated Gene Expression DataabstractIn gene expression data analysis, clustering is a fruitful exploratory technique to reveal the underlying molecular mechanism by identifying groups of co-expressed genes. To reduce the noise, usually multiple experimental replicates are performed. An integrative analysis of the full replicate data, instead of reducing the data to the mean profile, carries the promise of yielding more precise and robust clusters. In this paper, we propose a novel resampling based clustering algorithm for genes with replicated expression measurements. Assuming those replicates are exchangeable, we formulate the problem in the bootstrap framework, and aim to infer the consensus clustering based on the bootstrap samples of replicates. In our approach, we adopt the mixed effect model to accommodate the heterogeneous variances and implement a quasi-MCMC algorithm to conduct statistical inference. Experiments demonstrate that by taking advantage of the full replicate data, our algorithm produces more reliable clusters and has robust performance in diverse scenarios, especially when the data is subject to multiple sources of variance. Xiaodan Fan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2012 | Short adjacent repeat identification based on Chemical Reaction OptimizationabstractThe analysis of short tandem repeats (STRs) in DNA sequences has become an attractive method for determining the genetic profile of an individual. Here we focus on a more general and practical issue named short adjacent repeats identification problem (SARIP), which is extended from STR by allowing short gaps between neighboring units. Presently, the best available solution to SARIP is BASARD, which uses Markov chain Monte Carlo algorithms to determine the posterior estimate. However, the computational complexity and the tendency to get stuck in a local mode lower the efficiency of BASARD and impede its wide application. In this paper, we prove that SARIP is NP-hard, and we also solve it with Chemical Reaction Optimization (CRO), a recently developed metaheuristic approach. CRO mimics the interactions of molecules in a chemical reaction and it can explore the solution space efficiently to find the optimal or near optimal solution(s). We test the CRO algorithm with both synthetic and real data, and compare its performance in mode searching with BASARD. Simulation results show that CRO enjoys dozens of times, or even a hundred times shorter computational time compared with BASARD. It is also demonstrated that CRO can obtain the global optima most of the time. Moreover, CRO is more stable in different runs, which is of great importance in practical use. Thus, CRO is by far the best method on SARIP. Albert Y. S. Lam, Victor O. K. Li, Qiwei Li 0001, Xiaodan Fan |
IEEE Congress on Evolutionary Computation | 5 |
| 2011 | An MCMC algorithm for detecting short adjacent repeats shared by multiple sequencesabstractMOTIVATION: Repeats detection problems are traditionally formulated as string matching or signal processing problems. They cannot readily handle gaps between repeat units and are incapable of detecting repeat patterns shared by multiple sequences. This study detects short adjacent repeats with interunit insertions from multiple sequences. For biological sequences, such studies can shed light on molecular structure, biological function and evolution. RESULTS: The task of detecting short adjacent repeats is formulated as a statistical inference problem by using a probabilistic generative model. An Markov chain Monte Carlo algorithm is proposed to infer the parameters in a de novo fashion. Its applications on synthetic and real biological data show that the new method not only has a competitive edge over existing methods, but also can provide a way to study the structure and the evolution of repeat-containing genes. AVAILABILITY: The related C++ source code and datasets are available at http://ihome.cuhk.edu.hk/%7Eb118998/share/BASARD.zip. CONTACT: [email protected] Qiwei Li 0001, Xiaodan Fan, Tong Liang, Shuo-Yen Robert Li |
Bioinform. | 2 |
| 2010 | An automatic procedure to search highly repetitive sequences in genome as fluorescence in situ hybridization probes and its application on Brachypodium distachyonabstractFluorescence in situ hybridization (FISH) is a powerful technique that localizes specific DNA sequences on chromosomes for use in physical and genetic maps assembling, genetic counselling, species identification, etc. Highly repetitive sequences are considered to be suitable FISH probes that can avoid many potential problems of using unique sequences as FISH probes. The distinct chromosomal distributions of these highly repetitive sequences are also ideal for labelling purposes such as karyotyping. In this paper, we present an automatic computational procedure for searching highly repetitive sequences from a whole genome as FISH probes, as well as an experimental protocol to use them in FISH analysis. We successfully applied the method on the newly released genome of Brachypodium distachyon (Brachypodium) and produced satisfactory results of FISH experiment. Qiwei Li 0001, Tong Liang, Xiaodan Fan, Weichang Yu 0002, Shuo-Yen Robert Li |
BIBM | 3 |
| 2010 | An Evolutionary Monte Carlo algorithm for identifying short adjacent repeats in multiple sequencesabstractEvolutionary Monte Carlo (EMC) algorithm is an effective and powerful method to sample complicated distributions. Short adjacent repeats identification problem (SARIP), i.e., searching for the common sequence pattern in multiple DNA sequences, is considered as one of the key challenges in the field of bioinformatics. A recently proposed Markov chain Monte Carlo (MCMC) algorithm has demonstrated its effectiveness in solving SARIP. However, high computation time and inevitable local optima hinder its wide application. In this paper, we apply EMC to parallelize the MCMC algorithm to solve SARIP. Our proposed EMC scheme is implemented on a parallel platform and the simulation results show that, compared with the conventional MCMC algorithm, EMC not only improves the quality of final solution but also reduces the computation time. Qiwei Li 0001, Xiaodan Fan, Victor O. K. Li, Shuo-Yen Robert Li |
BIBM | 3 |
| 2007 | Statistical power of phylo-HMM for evolutionarily conserved element detectionabstractBACKGROUND: An important goal of comparative genomics is the identification of functional elements through conservation analysis. Phylo-HMM was recently introduced to detect conserved elements based on multiple genome alignments, but the method has not been rigorously evaluated. RESULTS: We report here a simulation study to investigate the power of phylo-HMM. We show that the power of the phylo-HMM approach depends on many factors, the most important being the number of species-specific genomes used and evolutionary distances between pairs of species. This finding is consistent with results reported by other groups for simpler comparative genomics models. In addition, the conservation ratio of conserved elements and the expected length of the conserved elements are also major factors. In contrast, the influence of the topology and the nucleotide substitution model are relatively minor factors. CONCLUSION: Our results provide for general guidelines on how to select the number of genomes and their evolutionary distance in comparative genomics studies, as well as the level of power we can expect under different parameter settings. Xiaodan Fan, Eric E. Schadt, Jun S. Liu |
BMC Bioinform. | 1 |