Ling-Yun Wu

dblp:75/6006 · DBLP profile ↗
← Back
20ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0001-9487-0215ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 Adaptive segmentation and optimization for efficient centralizer selection and placement in oilfield drilling operations
Ling-Yun Wu
Expert Syst. Appl.5
2024 Incorporating network diffusion and peak location information for better single-cell ATAC-seq data analysis
abstract
Single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data provided new insights into the understanding of epigenetic heterogeneity and transcriptional regulation. With the increasing abundance of dataset resources, there is an urgent need to extract more useful information through high-quality data analysis methods specifically designed for scATAC-seq. However, analyzing scATAC-seq data poses challenges due to its near binarization, high sparsity and ultra-high dimensionality properties. Here, we proposed a novel network diffusion-based computational method to comprehensively analyze scATAC-seq data, named Single-Cell ATAC-seq Analysis via Network Refinement with Peaks Location Information (SCARP). SCARP formulates the Network Refinement diffusion method under the graph theory framework to aggregate information from different network orders, effectively compensating for missing signals in the scATAC-seq data. By incorporating distance information between adjacent peaks on the genome, SCARP also contributes to depicting the co-accessibility of peaks. These two innovations empower SCARP to obtain lower-dimensional representations for both cells and peaks more effectively. We have demonstrated through sufficient experiments that SCARP facilitated superior analyses of scATAC-seq data. Specifically, SCARP exhibited outstanding cell clustering performance, enabling better elucidation of cell heterogeneity and the discovery of new biologically significant cell subpopulations. Additionally, SCARP was also instrumental in portraying co-accessibility relationships of accessible regions and providing new insight into transcriptional regulation. Consequently, SCARP identified genes that were involved in key Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways related to diseases and predicted reliable cis-regulatory interactions. To sum up, our studies suggested that SCARP is a promising tool to comprehensively analyze the scATAC-seq data.
Jiating Yu, Jiacheng Leng, Zhichao Hou, Duanchen Sun, Ling-Yun Wu
Briefings Bioinform.5
2024 Reverse network diffusion to remove indirect noise for better inference of gene regulatory networks
abstract
MOTIVATION: Gene regulatory networks (GRNs) are vital tools for delineating regulatory relationships between transcription factors and their target genes. The boom in computational biology and various biotechnologies has made inferring GRNs from multi-omics data a hot topic. However, when networks are constructed from gene expression data, they often suffer from false-positive problem due to the transitive effects of correlation. The presence of spurious noise edges obscures the real gene interactions, which makes downstream analyses, such as detecting gene function modules and predicting disease-related genes, difficult and inefficient. Therefore, there is an urgent and compelling need to develop network denoising methods to improve the accuracy of GRN inference. RESULTS: In this study, we proposed a novel network denoising method named REverse Network Diffusion On Random walks (RENDOR). RENDOR is designed to enhance the accuracy of GRNs afflicted by indirect effects. RENDOR takes noisy networks as input, models higher-order indirect interactions between genes by transitive closure, eliminates false-positive effects using the inverse network diffusion method, and produces refined networks as output. We conducted a comparative assessment of GRN inference accuracy before and after denoising on simulated networks and real GRNs. Our results emphasized that the network derived from RENDOR more accurately and effectively captures gene interactions. This study demonstrates the significance of removing network indirect noise and highlights the effectiveness of the proposed method in enhancing the signal-to-noise ratio of noisy networks. AVAILABILITY AND IMPLEMENTATION: The R package RENDOR is provided at https://github.com/Wu-Lab/RENDOR and other source code and data are available at https://github.com/Wu-Lab/RENDOR-reproduce.
Jiating Yu, Jiacheng Leng, Duanchen Sun, Ling-Yun Wu
Bioinform.5
2023 PathExpSurv: pathway expansion for explainable survival analysis and disease gene discovery
abstract
BACKGROUND: In the field of biology and medicine, the interpretability and accuracy are both important when designing predictive models. The interpretability of many machine learning models such as neural networks is still a challenge. Recently, many researchers utilized prior information such as biological pathways to develop neural networks-based methods, so as to provide some insights and interpretability for the models. However, the prior biological knowledge may be incomplete and there still exists some unknown information to be explored. RESULTS: We proposed a novel method, named PathExpSurv, to gain an insight into the black-box model of neural network for cancer survival analysis. We demonstrated that PathExpSurv could not only incorporate the known prior information into the model, but also explore the unknown possible expansion to the existing pathways. We performed downstream analyses based on the expanded pathways and successfully identified some key genes associated with the diseases and original pathways. CONCLUSIONS: Our proposed PathExpSurv is a novel, effective and interpretable method for survival analysis. It has great utility and value in medical diagnosis and offers a promising framework for biological research.
Zhichao Hou, Jiacheng Leng, Jiating Yu, Zheng Xia, Ling-Yun Wu
BMC Bioinform.5
2022 Interaction-based transcriptome analysis via differential network inference
abstract
Gene-based transcriptome analysis, such as differential expression analysis, can identify the key factors causing disease production, cell differentiation and other biological processes. However, this is not enough because basic life activities are mainly driven by the interactions between genes. Although there have been already many differential network inference methods for identifying the differential gene interactions, currently, most studies still only use the information of nodes in the network for downstream analyses. To investigate the insight into differential gene interactions, we should perform interaction-based transcriptome analysis (IBTA) instead of gene-based analysis after obtaining the differential networks. In this paper, we illustrated a workflow of IBTA by developing a Co-hub Differential Network inference (CDN) algorithm, and a novel interaction-based metric, pivot APC2. We confirmed the superior performance of CDN through simulation experiments compared with other popular differential network inference algorithms. Furthermore, three case studies are given using colorectal cancer, COVID-19 and triple-negative breast cancer datasets to demonstrate the ability of our interaction-based analytical process to uncover causative mechanisms.
Jiacheng Leng, Ling-Yun Wu
Briefings Bioinform.2
2022 hDirect-MAP: projection-free single-cell modeling of response to checkpoint immunotherapy
abstract
There is a lack of robust generalizable predictive biomarkers of response to immune checkpoint blockade in multiple types of cancer. We develop hDirect-MAP, an algorithm that maps T cells into a shared high-dimensional (HD) expression space of diverse T cell functional signatures in which cells group by the common T cell phenotypes rather than dimensional reduced features or a distorted view of these features. Using projection-free single-cell modeling, hDirect-MAP first removed a large group of cells that did not contribute to response and then clearly distinguished T cells into response-specific subpopulations that were defined by critical T cell functional markers of strong differential expression patterns. We found that these grouped cells cannot be distinguished by dimensional-reduction algorithms but are blended by diluted expression patterns. Moreover, these identified response-specific T cell subpopulations enabled a generalizable prediction by their HD metrics. Tested using five single-cell RNA-seq or mass cytometry datasets from basal cell carcinoma, squamous cell carcinoma and melanoma, hDirect-MAP demonstrated common response-specific T cell phenotypes that defined a generalizable and accurate predictive biomarker.
Ningbo Zheng, Wenzhong Yang, Rui-Sheng Wang, Ling-Yun Wu, Lance D. Miller, Timothy Pardee, Pierre L. Triozzi, Hui-Wen Lo, Kounosuke Watabe, Stephen T. C. Wong, Boris C. Pasche, Guangxu Jin
Briefings Bioinform.7
2022 Importance-Penalized Joint Graphical Lasso (IPJGL): differential network inference via GGMs
abstract
MOTIVATION: Differential network inference is a fundamental and challenging problem to reveal gene interactions and regulation relationships under different conditions. Many algorithms have been developed for this problem; however, they do not consider the differences between the importance of genes, which may not fit the real-world situation. Different genes have different mutation probabilities, and the vital genes associated with basic life activities have less fault tolerance to mutation. Equally treating all genes may bias the results of differential network inference. Thus, it is necessary to consider the importance of genes in the models of differential network inference. RESULTS: Based on the Gaussian graphical model with adaptive gene importance regularization, we develop a novel Importance-Penalized Joint Graphical Lasso method (IPJGL) for differential network inference. The presented method is validated by the simulation experiments as well as the real datasets. Furthermore, to precisely evaluate the results of differential network inference, we propose a new metric named APC2 for the differential levels of gene pairs. We apply IPJGL to analyze the TCGA colorectal and breast cancer datasets and find some candidate cancer genes with significant survival analysis results, including SOST for colorectal cancer and RBBP8 for breast cancer. We also conduct further analysis based on the interactions in the Reactome database and confirm the utility of our method. AVAILABILITY AND IMPLEMENTATION: R source code of Importance-Penalized Joint Graphical Lasso is freely available at https://github.com/Wu-Lab/IPJGL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiacheng Leng, Ling-Yun Wu
Bioinform.2
2019 Discovering cooperative biomarkers for heterogeneous complex disease diagnoses
abstract
Biomarkers with high reproducibility and accurate prediction performance can contribute to comprehending the underlying pathogenesis of related complex diseases and further facilitate disease diagnosis and therapy. Techniques integrating gene expression profiles and biological networks for the identification of network-based disease biomarkers are receiving increasing interest. The biomarkers for heterogeneous diseases often exhibit strong cooperative effects, which implies that a set of genes may achieve more accurate outcome prediction than any single gene. In this study, we evaluated various biomarker identification methods that consider gene cooperative effects implicitly or explicitly, and proposed the gene cooperation network to explicitly model the cooperative effects of gene combinations. The gene cooperation network-enhanced method, named as MarkRank, achieves superior performance compared with traditional biomarker identification methods in both simulation studies and real data sets. The biomarkers identified by MarkRank not only have a better prediction accuracy but also have stronger topological relationships in the biological network and exhibit high specificity associated with the related diseases. Furthermore, the top genes identified by MarkRank involve crucial biological processes of related diseases and give a good prioritization for known disease genes. In conclusion, MarkRank suggests that explicit modeling of gene cooperative effects can greatly improve biomarker identification for complex diseases, especially for diseases with high heterogeneity.
Duanchen Sun, Xian-Wen Ren, Eszter Ari, Tamás Korcsmáros, Peter Csermely, Ling-Yun Wu
Briefings Bioinform.6
2017 Structure alignment-based classification of RNA-binding pockets reveals regional RNA recognition motifs on protein surfaces
abstract
BACKGROUND: Many critical biological processes are strongly related to protein-RNA interactions. Revealing the protein structure motifs for RNA-binding will provide valuable information for deciphering protein-RNA recognition mechanisms and benefit complementary structural design in bioengineering. RNA-binding events often take place at pockets on protein surfaces. The structural classification of local binding pockets determines the major patterns of RNA recognition. RESULTS: In this work, we provide a novel framework for systematically identifying the structure motifs of protein-RNA binding sites in the form of pockets on regional protein surfaces via a structure alignment-based method. We first construct a similarity network of RNA-binding pockets based on a non-sequential-order structure alignment method for local structure alignment. By using network community decomposition, the RNA-binding pockets on protein surfaces are clustered into groups with structural similarity. With a multiple structure alignment strategy, the consensus RNA-binding pockets in each group are identified. The crucial recognition patterns, as well as the protein-RNA binding motifs, are then identified and analyzed. CONCLUSIONS: Large-scale RNA-binding pockets on protein surfaces are grouped by measuring their structural similarities. This similarity network-based framework provides a convenient method for modeling the structural relationships of functional pockets. The local structural patterns identified serve as structure motifs for the recognition with RNA on protein surfaces.
Zhi-Ping Liu, Shutang Liu, Ruitang Chen, Xiaopeng Huang, Ling-Yun Wu
BMC Bioinform.5
2014 Discovery of co-occurring driver pathways in cancer
abstract
BACKGROUND: It has been widely realized that pathways rather than individual genes govern the course of carcinogenesis. Therefore, discovering driver pathways is becoming an important step to understand the molecular mechanisms underlying cancer and design efficient treatments for cancer patients. Previous studies have focused mainly on observation of the alterations in cancer genomes at the individual gene or single pathway level. However, a great deal of evidence has indicated that multiple pathways often function cooperatively in carcinogenesis and other key biological processes. RESULTS: In this study, an exact mathematical programming method was proposed to de novo identify co-occurring mutated driver pathways (CoMDP) in carcinogenesis without any prior information beyond mutation profiles. Two possible properties of mutations that occurred in cooperative pathways were exploited to achieve this: (1) each individual pathway has high coverage and high exclusivity; and (2) the mutations between the pair of pathways showed statistically significant co-occurrence. The efficiency of CoMDP was validated first by testing on simulated data and comparing it with a previous method. Then CoMDP was applied to several real biological data including glioblastoma, lung adenocarcinoma, and ovarian carcinoma datasets. The discovered co-occurring driver pathways were here found to be involved in several key biological processes, such as cell survival and protein synthesis. Moreover, CoMDP was modified to (1) identify an extra pathway co-occurring with a known pathway and (2) detect multiple significant co-occurring driver pathways for carcinogenesis. CONCLUSIONS: The present method can be used to identify gene sets with more biological relevance than the ones currently used for the discovery of single driver pathways.
Junhua Zhang 0002, Ling-Yun Wu, Xiang-Sun Zhang
BMC Bioinform.2
2012 Efficient methods for identifying mutated driver pathways in cancer
abstract
MOTIVATION: The first step for clinical diagnostics, prognostics and targeted therapeutics of cancer is to comprehensively understand its molecular mechanisms. Large-scale cancer genomics projects are providing a large volume of data about genomic, epigenomic and gene expression aberrations in multiple cancer types. One of the remaining challenges is to identify driver mutations, driver genes and driver pathways promoting cancer proliferation and filter out the unfunctional and passenger ones. RESULTS: In this study, we propose two methods to solve the so-called maximum weight submatrix problem, which is designed to de novo identify mutated driver pathways from mutation data in cancer. The first one is an exact method that can be helpful for assessing other approximate or/and heuristic algorithms. The second one is a stochastic and flexible method that can be employed to incorporate other types of information to improve the first method. Particularly, we propose an integrative model to combine mutation and expression data. We first apply our methods onto simulated data to show their efficiency. We further apply the proposed methods onto several real biological datasets, such as the mutation profiles of 74 head and neck squamous cell carcinomas samples, 90 glioblastoma tumor samples and 313 ovarian carcinoma samples. The gene expression profiles were also considered for the later two data. The results show that our integrative model can identify more biologically relevant gene sets. We have implemented all these methods and made a package called mutated driver pathway finder, which can be easily used for other researchers. AVAILABILITY: A MATLAB package of MDPFinder is available at http://zhangroup.aporc.org/ShiHuaZhang. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Junfei Zhao, Ling-Yun Wu, Xiang-Sun Zhang
Bioinform.3
2011 An efficient network querying method based on conditional random fields
abstract
MOTIVATION: A large amount of biomolecular network data for multiple species have been generated by high-throughput experimental techniques, including undirected and directed networks such as protein-protein interaction networks, gene regulatory networks and metabolic networks. There are many conserved functionally similar modules and pathways among multiple biomolecular networks in different species; therefore, it is important to analyze the similarity between the biomolecular networks. Network querying approaches aim at efficiently discovering the similar subnetworks among different species. However, many existing methods only partially solve this problem. RESULTS: In this article, a novel approach for network querying problem based on conditional random fields (CRFs) model is presented, which can handle both undirected and directed networks, acyclic and cyclic networks and any number of insertions/deletions. The CRF method is fast and can query pathways in a large network in seconds using a PC. To evaluate the CRF method, extensive computational experiments are conducted on the simulated and real data, and the results are compared with the existing network querying methods. All results show that the CRF method is very useful and efficient to find the conserved functionally similar modules and pathways in multiple biomolecular networks.
Ling-Yun Wu, Xiang-Sun Zhang
Bioinform.2
2010 Prediction of protein-RNA binding sites by a random forest method with combined features
abstract
MOTIVATION: Protein-RNA interactions play a key role in a number of biological processes, such as protein synthesis, mRNA processing, mRNA assembly, ribosome function and eukaryotic spliceosomes. As a result, a reliable identification of RNA binding site of a protein is important for functional annotation and site-directed mutagenesis. Accumulated data of experimental protein-RNA interactions reveal that a RNA binding residue with different neighbor amino acids often exhibits different preferences for its RNA partners, which in turn can be assessed by the interacting interdependence of the amino acid fragment and RNA nucleotide. RESULTS: In this work, we propose a novel classification method to identify the RNA binding sites in proteins by combining a new interacting feature (interaction propensity) with other sequence- and structure-based features. Specifically, the interaction propensity represents a binding specificity of a protein residue to the interacting RNA nucleotide by considering its two-side neighborhood in a protein residue triplet. The sequence as well as the structure-based features of the residues are combined together to discriminate the interaction propensity of amino acids with RNA. We predict RNA interacting residues in proteins by implementing a well-built random forest classifier. The experiments show that our method is able to detect the annotated protein-RNA interaction sites in a high accuracy. Our method achieves an accuracy of 84.5%, F-measure of 0.85 and AUC of 0.92 prediction of the RNA binding residues for a dataset containing 205 non-homologous RNA binding proteins, and also outperforms several existing RNA binding residue predictors, such as RNABindR, BindN, RNAProB and PPRint, and some alternative machine learning methods, such as support vector machine, naive Bayes and neural network in the comparison study. Furthermore, we provide some biological insights into the roles of sequences and structures in protein-RNA interactions by both evaluating the importance of features for their contributions in predictive accuracy and analyzing the binding patterns of interacting residues. AVAILABILITY: All the source data and code are available at http://www.aporc.org/doc/wiki/PRNA or http://www.sysbio.ac.cn/datatools.asp CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhi-Ping Liu, Ling-Yun Wu, Yong Wang 0001, Xiang-Sun Zhang, Luonan Chen
Bioinform.2
2009 Conditional random pattern algorithm for LOH inference and segmentation
abstract
MOTIVATION: Loss of heterozygosity (LOH) is one of the most important mechanisms in the tumor evolution. LOH can be detected from the genotypes of the tumor samples with or without paired normal samples. In paired sample cases, LOH detection for informative single nucleotide polymorphisms (SNPs) is straightforward if there is no genotyping error. But genotyping errors are always unavoidable, and there are about 70% non-informative SNPs whose LOH status can only be inferred from the neighboring informative SNPs. RESULTS: This article presents a novel LOH inference and segmentation algorithm based on the conditional random pattern (CRP) model. The new model explicitly considers the distance between two neighboring SNPs, as well as the genotyping error rate and the heterozygous rate. This new method is tested on the simulated and real data of the Affymetrix Human Mapping 500K SNP arrays. The experimental results show that the CRP method outperforms the conventional methods based on the hidden Markov model (HMM). AVAILABILITY: Software is available upon request.
Ling-Yun Wu, Xiaobo Zhou 0001, Fuhai Li 0001, Xiaorong Yang, Chung-Che Jeff Chang, Stephen T. C. Wong
Bioinform.1
2009 Evaluating Protein Similarity from Coarse Structures
abstract
To unscramble the relationship between protein function and protein structure, it is essential to assess the protein similarity from different aspects. Although many methods have been proposed for protein structure alignment or comparison, alternative similarity measures are still strongly demanded due to the requirement of fast screening and query in large-scale structure databases. In this paper, we first formulate a novel representation of a protein structure, i.e., Feature Sequence of Surface (FSS). Then, a new score scheme is developed to measure the similarity between two representations. To verify the proposed method, numerical experiments are conducted in four different protein data sets. We also classify SARS coronavirus to verify the effectiveness of the new method. Furthermore, preliminary results of fast classification of the whole CATH v2.5.1 database based on the new macrostructure similarity are given as a pilot study. We demonstrate that the proposed approach to measure the similarities between protein structures is simple to implement, computationally efficient, and surprisingly fast. In addition, the method itself provides a new and quantitative tool to view a protein structure.
Yong Wang 0001, Ling-Yun Wu, Zhong-Wei Zhan, Xiang-Sun Zhang, Luonan Chen
IEEE ACM Trans. Comput. Biol. Bioinform.2
2007 Predicting gene ontology functions from protein's regional surface structures
abstract
BACKGROUND: Annotation of protein functions is an important task in the post-genomic era. Most early approaches for this task exploit only the sequence or global structure information. However, protein surfaces are believed to be crucial to protein functions because they are the main interfaces to facilitate biological interactions. Recently, several databases related to structural surfaces, such as pockets and cavities, have been constructed with a comprehensive library of identified surface structures. For example, CASTp provides identification and measurements of surface accessible pockets as well as interior inaccessible cavities. RESULTS: A novel method was proposed to predict the Gene Ontology (GO) functions of proteins from the pocket similarity network, which is constructed according to the structure similarities of pockets. The statistics of the networks were presented to explore the relationship between the similar pockets and GO functions of proteins. Cross-validation experiments were conducted to evaluate the performance of the proposed method. Results and codes are available at: http://zhangroup.aporc.org/bioinfo/PSN/. CONCLUSION: The computational results demonstrate that the proposed method based on the pocket similarity network is effective and efficient for predicting GO functions of proteins in terms of both computational complexity and prediction accuracy. The proposed method revealed strong relationship between small surface patterns (or pockets) and GO functions, which can be further used to identify active sites or functional motifs. The high quality performance of the prediction method together with the statistics also indicates that pockets play essential roles in biological interactions or the GO functions. Moreover, in addition to pockets, the proposed network framework can also be used for adopting other protein spatial surface patterns to predict the protein functions.
Zhi-Ping Liu, Ling-Yun Wu, Yong Wang 0001, Luonan Chen, Xiang-Sun Zhang
BMC Bioinform.2
2007 Analysis on multi-domain cooperation for predicting protein-protein interactions
abstract
BACKGROUND: Domains are the basic functional units of proteins. It is believed that protein-protein interactions are realized through domain interactions. Revealing multi-domain cooperation can provide deep insights into the essential mechanism of protein-protein interactions at the domain level and be further exploited to improve the accuracy of protein interaction prediction. RESULTS: In this paper, we aim to identify cooperative domains for protein interactions by extending two-domain interactions to multi-domain interactions. Based on the high-throughput experimental data from multiple organisms with different reliabilities, the interactions of domains were inferred by a Linear Programming algorithm with Multi-domain pairs (LPM) and an Association Probabilistic Method with Multi-domain pairs (APMM). Experimental results demonstrate that our approach not only can find cooperative domains effectively but also has a higher accuracy for predicting protein interaction than the existing methods. Cooperative domains, including strongly cooperative domains and superdomains, were detected from major interaction databases MIPS and DIP, and many of them were verified by physical interactions from the crystal structures of protein complexes in PDB which provide intuitive evidences for such cooperation. Comparison experiments in terms of protein/domain interaction prediction justified the benefit of considering multi-domain cooperation. CONCLUSION: From the computational viewpoint, this paper gives a general framework to predict protein interactions in a more accurate manner by considering the information of both multi-domains and multiple organisms, which can also be applied to identify cooperative domains, to reconstruct large complexes and further to annotate functions of domains. Supplementary information and software are provided in http://intelligent.eic.osaka-sandai.ac.jp/chenen/MDCinfer.htm and http://zhangroup.aporc.org/bioinfo/MDCinfer.
Rui-Sheng Wang, Yong Wang 0001, Ling-Yun Wu, Xiang-Sun Zhang, Luonan Chen
BMC Bioinform.3
2006 Automatic Classification of Protein Structures Based on Convex Hull Representation by Integrated Neural Network
Yong Wang 0001, Ling-Yun Wu, Xiang-Sun Zhang, Luonan Chen
TAMC2
2005 Haplotype reconstruction from SNP fragments by minimum error correction
abstract
MOTIVATION: Haplotype reconstruction based on aligned single nucleotide polymorphism (SNP) fragments is to infer a pair of haplotypes from localized polymorphism data gathered through short genome fragment assembly. An important computational model of this problem is the minimum error correction (MEC) model, which has been mentioned in several literatures. The model retrieves a pair of haplotypes by correcting minimum number of SNPs in given genome fragments coming from an individual's DNA. RESULTS: In the first part of this paper, an exact algorithm for the MEC model is presented. Owing to the NP-hardness of the MEC model, we also design a genetic algorithm (GA). The designed GA is intended to solve large size problems and has very good performance. The strength and weakness of the MEC model are shown using experimental results on real data and simulation data. In the second part of this paper, to improve the MEC model for haplotype reconstruction, a new computational model is proposed, which simultaneously employs genotype information of an individual in the process of SNP correction, and is called MEC with genotype information (shortly, MEC/GI). Computational results on extensive datasets show that the new model has much higher accuracy in haplotype reconstruction than the pure MEC model.
Rui-Sheng Wang, Ling-Yun Wu, Zhen-Ping Li, Xiang-Sun Zhang
Bioinform.2
2003 Reconstruction of DNA sequencing by hybridization
abstract
MOTIVATION: It is widely recognized that the hybridization process is prone to errors and that the future of DNA sequencing by hybridization is predicated on the ability to successfully cope with such errors. However, the occurrence of hybridization errors results in the computational difficulty of the reconstruction of DNA sequencing by hybridization. The reconstruction problem of DNA sequencing by hybridization with errors is a strongly NP-hard problem. So far the problem has not been solved well. RESULTS: In this paper, a new approach is presented to solve the reconstruction problem of DNA sequencing by hybridization, which realizes the computational part of the SBH experiment. The proposed algorithm accepts both the negative and positive errors. The computational experiments show that the algorithm behaves satisfactorily, especially for the case with k-tuple repetitions and positive errors.
Ling-Yun Wu, Xiang-Sun Zhang
Bioinform.2