Hong-Wen Deng

dblp:02/1248 · DBLP profile ↗
← Back
35ranked-venue papers
0as first author
7since 2021 · last 2022
0000-0002-0387-8818ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 35 · 7 since 2021
YearPublicationVenuePosition
2022 BERT6mA: prediction of DNA N6-methyladenine site using deep learning-based approaches
abstract
N6-methyladenine (6mA) is associated with important roles in DNA replication, DNA repair, transcription, regulation of gene expression. Several experimental methods were used to identify DNA modifications. However, these experimental methods are costly and time-consuming. To detect the 6mA and complement these shortcomings of experimental methods, we proposed a novel, deep leaning approach called BERT6mA. To compare the BERT6mA with other deep learning approaches, we used the benchmark datasets including 11 species. The BERT6mA presented the highest AUCs in eight species in independent tests. Furthermore, BERT6mA showed higher and comparable performance with the state-of-the-art models while the BERT6mA showed poor performances in a few species with a small sample size. To overcome this issue, pretraining and fine-tuning between two species were applied to the BERT6mA. The pretrained and fine-tuned models on specific species presented higher performances than other models even for the species with a small sample size. In addition to the prediction, we analyzed the attention weights generated by BERT6mA to reveal how the BERT6mA model extracts critical features responsible for the 6mA prediction. To facilitate biological sciences, the BERT6mA online web server and its source codes are freely accessible at https://github.com/kuratahiroyuki/BERT6mA.git, respectively.
Sho Tsukiyama, Md. Mehedi Hasan 0002, Hong-Wen Deng, Hiroyuki Kurata
Briefings Bioinform.3
2022 Effective Cancer Subtype and Stage Prediction via Dropfeature-DNNs
abstract
Precise cancer subtype and/or stage prediction is instrumental for cancer diagnosis, treatment and management. However, most of the existing methods based on genomic profiles suffer from issues such as overfitting, high computational complexity and selected features (i.e., genes) not directly related to forecast precision. These deficiencies are largely due to the nature of "high dimensionality and small sample size" inherent in molecular data, and such a nature is often deemed as an obstacle to the application of deep learning, e.g., deep neural networks (DNNs), to biomedicine and cancer research. In this paper, we propose a DNN-based algorithm coupled with a new embedded feature selection technique, named Dropfeature-DNNs, to address these issues. Dropfeature-DNNs can discard some irrelevant features (i.e., genes) when training DNNs, and we formulate Dropfeature-DNNs as an iterative AUC optimization problem. As such, an "optimal" feature subset that contains meaningful genes for accurate tumor subtype and/or stage prediction can be obtained when the AUC optimization converges in the training stage. Since the feature subset and AUC optimizations are synchronous with the training phase of DNNs, model complexity and computational cost are simultaneously reduced. Rigorous feature subset convergence analysis and error bound inference provide a solid theoretical foundation for the proposed method. Extensive empirical comparisons to benchmark methods further demonstrate the efficacy of Dropfeature-DNNs in cancer subtype and/or stage prediction using HDSS gene expression data from multiple cancer types.
Zhong Chen 0003, Wensheng Zhang 0005, Hong-Wen Deng, Kun Zhang 0012
IEEE ACM Trans. Comput. Biol. Bioinform.3
2021 NeuroPred-FRL: an interpretable prediction model for identifying neuropeptide using feature representation learning
abstract
Neuropeptides (NPs) are the most versatile neurotransmitters in the immune systems that regulate various central anxious hormones. An efficient and effective bioinformatics tool for rapid and accurate large-scale identification of NPs is critical in immunoinformatics, which is indispensable for basic research and drug development. Although a few NP prediction tools have been developed, it is mandatory to improve their NPs' prediction performances. In this study, we have developed a machine learning-based meta-predictor called NeuroPred-FRL by employing the feature representation learning approach. First, we generated 66 optimal baseline models by employing 11 different encodings, six different classifiers and a two-step feature selection approach. The predicted probability scores of NPs based on the 66 baseline models were combined to be deemed as the input feature vector. Second, in order to enhance the feature representation ability, we applied the two-step feature selection approach to optimize the 66-D probability feature vector and then inputted the optimal one into a random forest classifier for the final meta-model (NeuroPred-FRL) construction. Benchmarking experiments based on both cross-validation and independent tests indicate that the NeuroPred-FRL achieves a superior prediction performance of NPs compared with the other state-of-the-art predictors. We believe that the proposed NeuroPred-FRL can serve as a powerful tool for large-scale identification of NPs, facilitating the characterization of their functional mechanisms and expediting their applications in clinical therapy. Moreover, we interpreted some model mechanisms of NeuroPred-FRL by leveraging the robust SHapley Additive exPlanation algorithm.
Md. Mehedi Hasan 0002, Md. Ashad Alam, Watshara Shoombuatong, Hong-Wen Deng, Balachandran Manavalan, Hiroyuki Kurata
Briefings Bioinform.4
2021 Combining artificial intelligence: deep learning with Hi-C data to predict the functional effects of non-coding variants
abstract
MOTIVATION: Although genome-wide association studies (GWASs) have identified thousands of variants for various traits, the causal variants and the mechanisms underlying the significant loci are largely unknown. In this study, we aim to predict non-coding variants that may functionally affect translation initiation through long-range chromatin interaction. RESULTS: By incorporating the Hi-C data, we propose a novel and powerful deep learning model of artificial intelligence to classify interacting and non-interacting fragment pairs and predict the functional effects of sequence alteration of single nucleotide on chromatin interaction and thus on gene expression. The changes in chromatin interaction probability between the reference sequence and the altered sequence reflect the degree of functional impact for the variant. The model was effective and efficient with the classification of interacting and non-interacting fragment pairs. The predicted causal SNPs that had a larger impact on chromatin interaction were more likely to be identified by GWAS and eQTL analyses. We demonstrate that an integrative approach combining artificial intelligence-deep learning with high throughput experimental evidence of chromatin interaction leads to prioritizing the functional variants in disease- and phenotype-related loci and thus will greatly expedite uncover of the biological mechanism underlying the association identified in genomic studies. AVAILABILITY AND IMPLEMENTATION: Source code used in data preparing and model training is available at the GitHub website (https://github.com/biocai/DeepHiC). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiang-He Meng, Hong-Mei Xiao, Hong-Wen Deng
Bioinform.3
2021 A gene-level methylome-wide association analysis identifies novel Alzheimer's disease genes
abstract
Abstract Motivation Transcriptome-wide association studies (TWAS) have successfully facilitated the discovery of novel genetic risk loci for many complex traits, including late-onset Alzheimer’s disease (AD). However, most existing TWAS methods rely only on gene expression and ignore epigenetic modification (i.e. DNA methylation) and functional regulatory information (i.e. enhancer-promoter interactions), both of which contribute significantly to the genetic basis of AD. Results We develop a novel gene-level association testing method that integrates genetically regulated DNA methylation and enhancer–target gene pairs with genome-wide association study (GWAS) summary results. Through simulations, we show that our approach, referred to as the CMO (cross methylome omnibus) test, yielded well controlled type I error rates and achieved much higher statistical power than competing methods under a wide range of scenarios. Furthermore, compared with TWAS, CMO identified an average of 124% more associations when analyzing several brain imaging-related GWAS results. By analyzing to date the largest AD GWAS of 71 880 cases and 383 378 controls, CMO identified six novel loci for AD, which have been ignored by competing methods. Availabilityand implementation The data used in this work were obtained from the following publicly available datasets: IGAP1, GWAX, UK Biobank, a 2019 meta-analyzed AD GWAS results and a imaging-derived phenotype GWAS results. The data resources are summarized in Supplementary Table S7. We used the publicly available software and tools for competing methods. All codes used to generate results that are reported in this manuscript and software for our newly proposed method CMO are available at https://github.com/ChongWuLab/CMO. Supplementary information Supplementary data are available at Bioinformatics online.
Jonathan Bradley, Hong-Wen Deng
Bioinform.5
2021 A generalized kernel machine approach to identify higher-order composite effects in multi-view datasets, with application to adolescent brain development and osteoporosis
Md. Ashad Alam, Chuan Qiu, Hui Shen 0007, Yu-Ping Wang 0002, Hong-Wen Deng
J. Biomed. Informatics5
2021 A Joint Analysis of Multi-Paradigm fMRI Data With Its Application to Cognitive Study
abstract
With the development of neuroimaging techniques, a growing amount of multi-modal brain imaging data are collected, facilitating comprehensive study of the brain. In this paper, we jointly analyzed functional magnetic resonance imaging (fMRI) collected under different paradigms in order to understand cognitive behaviors of an individual. To this end, we proposed a novel multi-view learning algorithm called structure-enforced collaborative regression (SCoRe) to extract co-expressed discriminative brain regions under the guidance of anatomical structure of the brain. An advantage of SCoRe over its predecessor collaborative regression (CoRe) lies in its incorporation of group structures in the brain imaging data, which makes the model biologically more meaningful. Results from real data analysis has confirmed that by incorporating prior knowledge of brain structure, SCoRe can deliver better prediction performance and is less sensitive to hyper-parameters than CoRe. After validation with simulation experiments, we applied SCoRe to fMRI data collected from the Philadelphia Neurodevelopmental Cohort and adopted the scores from the wide range achievement test (WRAT) to evaluate an individual's cognitive skills. We located 14 relevant brain regions that can efficiently predict WRAT scores and these brain regions were further confirmed by other independent studies.
Yuntong Bai, Yun Gong, Jianchao Bai, Jingyu Liu 0001, Hong-Wen Deng, Vince D. Calhoun, Yu-Ping Wang 0002
IEEE Trans. Medical Imaging5
2020 A novel computational strategy for DNA methylation imputation using mixture regression model (MRM)
abstract
BACKGROUND: DNA methylation is an important heritable epigenetic mark that plays a crucial role in transcriptional regulation and the pathogenesis of various human disorders. The commonly used DNA methylation measurement approaches, e.g., Illumina Infinium HumanMethylation-27 and -450 BeadChip arrays (27 K and 450 K arrays) and reduced representation bisulfite sequencing (RRBS), only cover a small proportion of the total CpG sites in the human genome, which considerably limited the scope of the DNA methylation analysis in those studies. RESULTS: We proposed a new computational strategy to impute the methylation value at the unmeasured CpG sites using the mixture of regression model (MRM) of radial basis functions, integrating information of neighboring CpGs and the similarities in local methylation patterns across subjects and across multiple genomic regions. Our method achieved a better imputation accuracy over a set of competing methods on both simulated and empirical data, particularly when the missing rate is high. By applying MRM to an RRBS dataset from subjects with low versus high bone mineral density (BMD), we recovered methylation values of ~ 300 K CpGs in the promoter regions of chromosome 17 and identified some novel differentially methylated CpGs that are significantly associated with BMD. CONCLUSIONS: Our method is well applicable to the numerous methylation studies. By expanding the coverage of the methylation dataset to unmeasured sites, it can significantly enhance the discovery of novel differential methylation signals and thus reveal the mechanisms underlying various human disorders/traits.
Fangtang Yu, Chao Xu 0014, Hong-Wen Deng, Hui Shen 0007
BMC Bioinform.3
2019 PCA-based GRS analysis enhances the effectiveness for genetic correlation detection
abstract
Genetic risk score (GRS, also known as polygenic risk score) analysis is an increasingly popular method for exploring genetic architectures and relationships of complex diseases. However, complex diseases are usually measured by multiple correlated phenotypes. Analyzing each disease phenotype individually is likely to reduce statistical power due to multiple testing correction. In order to conquer the disadvantage, we proposed a principal component analysis (PCA)-based GRS analysis approach. Extensive simulation studies were conducted to compare the performance of PCA-based GRS analysis and traditional GRS analysis approach. Simulation results observed significantly improved performance of PCA-based GRS analysis compared to traditional GRS analysis under various scenarios. For the sake of verification, we also applied both PCA-based GRS analysis and traditional GRS analysis to a real Caucasian genome-wide association study (GWAS) data of bone geometry. Real data analysis results further confirmed the improved performance of PCA-based GRS analysis. Given that GWAS have flourished in the past decades, our approach may help researchers to explore the genetic architectures and relationships of complex diseases or traits.
Yujie Ning, Feng Zhang 0044, Miao Ding, Yan Wen 0003, Mengnan Lu, Jingyan Sun, Menglu Wu, Bolun Cheng, Mei Ma, Shiqiang Cheng, Hui Shen 0007, Qing Tian 0004, Xiong Guo, Hong-Wen Deng
Briefings Bioinform.18
2018 EPS-LASSO: test for high-dimensional regression under extreme phenotype sampling of continuous traits
abstract
Motivation: Extreme phenotype sampling (EPS) is a broadly-used design to identify candidate genetic factors contributing to the variation of quantitative traits. By enriching the signals in extreme phenotypic samples, EPS can boost the association power compared to random sampling. Most existing statistical methods for EPS examine the genetic factors individually, despite many quantitative traits have multiple genetic factors underlying their variation. It is desirable to model the joint effects of genetic factors, which may increase the power and identify novel quantitative trait loci under EPS. The joint analysis of genetic data in high-dimensional situations requires specialized techniques, e.g. the least absolute shrinkage and selection operator (LASSO). Although there are extensive research and application related to LASSO, the statistical inference and testing for the sparse model under EPS remain unknown. Results: We propose a novel sparse model (EPS-LASSO) with hypothesis test for high-dimensional regression under EPS based on a decorrelated score function. The comprehensive simulation shows EPS-LASSO outperforms existing methods with stable type I error and FDR control. EPS-LASSO can provide a consistent power for both low- and high-dimensional situations compared with the other methods dealing with high-dimensional situations. The power of EPS-LASSO is close to other low-dimensional methods when the causal effect sizes are small and is superior when the effects are large. Applying EPS-LASSO to a transcriptome-wide gene expression study for obesity reveals 10 significant body mass index associated genes. Our results indicate that EPS-LASSO is an effective method for EPS data analysis, which can account for correlated predictors. Availability and implementation: The source code is available at https://github.com/xu1912/EPSLASSO. Supplementary information: Supplementary data are available at Bioinformatics online.
Chao Xu 0014, Jian Fang 0001, Hui Shen 0007, Yu-Ping Wang 0002, Hong-Wen Deng
Bioinform.5
2018 A Sparse Regression Method for Group-Wise Feature Selection with False Discovery Rate Control
abstract
The method of Sorted L-One Penalized Estimation, or SLOPE, is a sparse regression method recently introduced by Bogdan et. al. [1] . It can be used to identify significant predictor variables in a linear model that may have more unknown parameters than observations. When the correlations between predictor variables are small, the SLOPE method is shown to successfully control the false discovery rate (the expected proportion of the irrelevant among all selected predictors) at a user specified level. However, the requirement for nearly uncorrelated predictors is too restrictive for genomic data, as demonstrated in our recent study [2] by an application of SLOPE to realistic simulated DNA sequence data. A possible solution is to divide the predictor variables into nearly uncorrelated groups, and to modify the procedure to select entire groups with an overall significant group effect, rather than individual predictors. Following this motivation, we extend SLOPE in the spirit of Group LASSO to Group SLOPE, a method that can handle group structures between the predictor variables, which are ubiquitous in real genomic data. Our theoretical results show that Group SLOPE controls the group-wise false discovery rate (gFDR), when groups are orthogonal to each other. For use in non-orthogonal settings, we propose two types of Monte Carlo based heuristics, which lead to gFDR control with Group SLOPE in simulations based on real SNP data. As an illustration of the merits of this method, an application of Group SLOPE to a dataset from the Framingham Heart Study results in the identification of some known DNA sequence regions associated with bone health, as well as some new candidate regions. The novel methods are implemented in the R package grpSLOPEMC , which is publicly available at https://github.com/agisga/grpSLOPEMC.
Alexej Gossmann, Shaolong Cao, Damian Brzyski, Hong-Wen Deng, Yu-Ping Wang 0002
IEEE ACM Trans. Comput. Biol. Bioinform.5
2018 Joint Detection of Associations Between DNA Methylation and Gene Expression From Multiple Cancers
abstract
DNA methylation plays an important role in the development of various cancers mainly through the regulation on gene expression. Hence, the study on the relation between DNA methylation and gene expression is of particular interest to understand cancers. Recently, an increasing number of datasets are available from multiple cancers, which makes it possible to study both the similarity and difference of genomic alterations across multiple tumor types. However, most of the existing pan-cancer analysis methods perform simple aggregations, which may overlook the heterogeneity of the interactions. In this paper, we propose a novel method to jointly detect complex associations between DNA methylation and gene expression levels from multiple cancers. The main idea is to apply joint sparse canonical correlation analysis to detect a small set of methylated sites, which are associated with another set of genes either shared across cancers or specific to a particular group (group-specific) of cancers. These methylated sites and genes form a complex module with strong multivariate correlations. We further introduced a joint sparse precision matrix estimation method to identify driver methylation-gene pairs in the module. These pairs are characterized by significant partial correlations, which may imply high functional impacts and contribute to complementary information to the main step. We apply our method to The Cancer Genome Atlas(TCGA) datasets with 1166 samples from four cancers. The results reveal significant shared and groupspecific interactions between DNA methylation and gene expression levels. To promote reproducible research, the Matlab code is available at https://sites.google.com/site/jianfang86/jointTCGA.
Jian Fang 0001, Ji-Gang Zhang, Hong-Wen Deng, Yu-Ping Wang 0002
IEEE J. Biomed. Health Informatics3
2018 Fast and Accurate Detection of Complex Imaging Genetics Associations Based on Greedy Projected Distance Correlation
abstract
Recent advances in imaging genetics produce large amounts of data including functional MRI images, single nucleotide polymorphisms (SNPs), and cognitive assessments. Understanding the complex interactions among these heterogeneous and complementary data has the potential to help with diagnosis and prevention of mental disorders. However, limited efforts have been made due to the high dimensionality, group structure, and mixed type of these data. In this paper we present a novel method to detect conditional associations between imaging genetics data. We use projected distance correlation to build a conditional dependency graph among high-dimensional mixed data, then use multiple testing to detect significant group level associations (e.g., ROI-gene). In addition, we introduce a scalable algorithm based on orthogonal greedy algorithm, yielding the greedy projected distance correlation (G-PDC). This can reduce the computational cost, which is critical for analyzing large-volume of imaging genomics data. The results from our simulations demonstrate a higher degree of accuracy with GPDC than distance correlation, Pearson's correlation and partial correlation, especially when the correlation is nonlinear. Finally, we apply our method to the Philadelphia Neurodevelopmental data cohort with 866 samples including fMRI images and SNP profiles. The results uncover several statistically significant and biologically interesting interactions, which are further validated with many existing studies. The Matlab code is available at https://sites.google.com/site/jianfang86/gPDC.
Jian Fang 0001, Chao Xu 0014, Pascal Zille, Dongdong Lin, Hong-Wen Deng, Vince D. Calhoun, Yu-Ping Wang 0002
IEEE Trans. Medical Imaging5
2017 Tissue-specific pathway association analysis using genome-wide association study summaries
abstract
MOTIVATION: Pathway association analysis has made great achievements in elucidating the genetic basis of human complex diseases. However, current pathway association analysis approaches fail to consider tissue-specificity. RESULTS: We developed a tissue-specific pathway interaction enrichment analysis algorithm (TPIEA). TPIEA was applied to two large Caucasian and Chinese genome-wide association study summary datasets of bone mineral density (BMD). TPIEA identified several significant pathways for BMD [false discovery rate (FDR) < 0.05], such as KEGG FOCAL ADHESION and KEGG AXON GUIDANCE, which had been demonstrated to be involved in the development of osteoporosis. We also compared the performance of TPIEA and classical pathway enrichment analysis, and TPIEA presented improved performance in recognizing disease relevant pathways. TPIEA may help to fill the gap of classic pathway association analysis approaches by considering tissue specificity. AVAILABILITY AND IMPLEMENTATION: The online web tool of TPIEA is available at https://sourceforge.net/projects/tpieav1/files CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online.
Jingcan Hao, Qianrui Fan, Awen He, Yan Wen 0003, Xiong Guo, Cuiyan Wu, Tielin Yang, Hui Shen 0007, Xiangding Chen, Qing Tian 0004, Lijun Tan, Hong-Wen Deng, Feng Zhang 0034
Bioinform.15
2017 Comparison of statistical methods for subnetwork detection in the integration of gene expression and protein interaction network
abstract
BACKGROUND: With the advancement of high-throughput technologies and enrichment of popular public databases, more and more research focuses of bioinformatics research have been on computational integration of network and gene expression profiles for extracting context-dependent active subnetworks. Many methods for subnetwork searching have been developed. Scoring and searching algorithms present a range of computational considerations and implementations. The primary goal of present study is to comprehensively evaluate the performance of different subnetwork detection methods. Eleven popular methods were selected for comprehensive comparison. RESULTS: First, taking into account the dependence of genes given a protein-protein interaction (PPI) network, we simulated microarray gene expression data under case and control conditions. Then each method was applied to the simulated data for subnetwork identification. Second, a large microarray data set of prostate cancer was used to assess the practical performance of each method. Using both simulation studies and a real data application, we evaluated the performance of different methods in terms of recall and precision. CONCLUSIONS: jActiveModules, PinnacleZ and WMAXC performed well in identifying subnetwork with relative high precision and recall. BioNet performed very well only in precision. As none of methods outperformed other methods overall, users should choose an appropriate method based on the purposes of their studies.
Hao He 0002, Dongdong Lin, Ji-Gang Zhang, Yu-Ping Wang 0002, Hong-Wen Deng
BMC Bioinform.5
2016 Unified tests for fine-scale mapping and identifying sparse high-dimensional sequence associations
abstract
MOTIVATION: In searching for genetic variants for complex diseases with deep sequencing data, genomic marker sets of high-dimensional genotypic data and sparse functional variants are quite common. Existing sequence association tests are incapable of identifying such marker sets or individual causal loci, although they appeared powerful to identify small marker sets with dense functional variants. In sequence association studies of admixed individuals, cryptic relatedness and population structure are known to confound the association analyses. METHOD: We here propose a unified marker wise test (uFineMap) to accurately localize causal loci and a unified high-dimensional set based test (uHDSet) to identify high-dimensional sparse associations in deep sequencing genomic data of multi-ethnic individuals with random relatedness. These two novel tests are based on scaled sparse linear mixed regressions with Lp (0 < p < 1) norm regularization. They jointly adjust for cryptic relatedness, population structure and other confounders to prevent false discoveries and improve statistical power for identifying promising individual markers and marker sets that harbor functional genetic variants of a complex trait. RESULTS: With large scale simulation data and real data analyses, the proposed tests appropriately controlled Type I error rates and appeared to be more powerful than several prominent methods. We illustrated their practical utilities by the applications to DNA sequence data of Framingham Heart Study for osteoporosis. The proposed tests identified 11 novel significant genes that were missed by the prominent famSKAT and GEMMA. In particular, four out of six most significant pathways identified by the uHDSet but missed by famSKAT have been reported to be related to BMD or osteoporosis in the literature. AVAILABILITY AND IMPLEMENTATION: The computational toolkit is available for academic use: https://sites.google.com/site/shaolongscode/home/uhdset CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shaolong Cao, Huaizhen Qin, Alexej Gossmann, Hong-Wen Deng, Yu-Ping Wang 0002
Bioinform.4
2016 An integrative imputation method based on multi-omics datasets
abstract
BACKGROUND: Integrative analysis of multi-omics data is becoming increasingly important to unravel functional mechanisms of complex diseases. However, the currently available multi-omics datasets inevitably suffer from missing values due to technical limitations and various constrains in experiments. These missing values severely hinder integrative analysis of multi-omics data. Current imputation methods mainly focus on using single omics data while ignoring biological interconnections and information imbedded in multi-omics data sets. RESULTS: In this study, a novel multi-omics imputation method was proposed to integrate multiple correlated omics datasets for improving the imputation accuracy. Our method was designed to: 1) combine the estimates of missing value from individual omics data itself as well as from other omics, and 2) simultaneously impute multiple missing omics datasets by an iterative algorithm. We compared our method with five imputation methods using single omics data at different noise levels, sample sizes and data missing rates. The results demonstrated the advantage and efficiency of our method, consistently in terms of the imputation error and the recovery of mRNA-miRNA network structure. CONCLUSIONS: We concluded that our proposed imputation method can utilize more biological information to minimize the imputation error and thus can improve the performance of downstream analysis such as genetic regulatory network construction.
Dongdong Lin, Ji-Gang Zhang, Chao Xu 0014, Hong-Wen Deng, Yu-Ping Wang 0002
BMC Bioinform.5
2015 SWGDT: A sliding window-based genotype dependence testing tool for genome-wide susceptibility gene scan
Yan Wen 0003, Jingcan Hao, Xiong Guo, Cuiyan Wu, Tielin Yang, Hui Shen 0007, Xiangding Chen, Lijun Tan, Qing Tian 0004, Hong-Wen Deng, Feng Zhang 0034
J. Biomed. Informatics12
2014 FISH: fast and accurate diploid genotype imputation via segmental hidden Markov model
abstract
MOTIVATION: Fast and accurate genotype imputation is necessary for facilitating gene-mapping studies, especially with the ever increasing numbers of both common and rare variants generated by high-throughput-sequencing experiments. However, most of the existing imputation approaches suffer from either inaccurate results or heavy computational demand. RESULTS: In this article, aiming to perform fast and accurate genotype-imputation analysis, we propose a novel, fast and yet accurate method to impute diploid genotypes. Specifically, we extend a hidden Markov model that is widely used to describe haplotype structures. But we model hidden states onto single reference haplotypes rather than onto pairs of haplotypes. Consequently the computational complexity is linear to size of reference haplotypes. We further develop an algorithm 'merge-and-recover (MAR)' to speed up the calculation. Working on compact representation of segmental reference haplotypes, the MAR algorithm always calculates an exact form of transition probabilities regardless of partition of segments. Both simulation studies and real-data analyses demonstrated that our proposed method was comparable to most of the existing popular methods in terms of imputation accuracy, but was much more efficient in terms of computation. The MAR algorithm can further speed up the calculation by several folds without loss of accuracy. The proposed method will be useful in large-scale imputation studies with a large number of reference subjects. AVAILABILITY: The implemented multi-threading software FISH is freely available for academic use at https://sites.google.com/site/lzhanghomepage/FISH.
Lei Zhang 0041, Yu-Fang Pei, Xiaoying Fu, Yu-Ping Wang 0002, Hong-Wen Deng
Bioinform.6
2014 FISH: fast and accurate diploid genotype imputation via segmental hidden Markov model
abstract
Bioinformatics (2014) 30 (13), 1876–1883 The author would like to note an affiliation correction. The first affiliation address was mistakenly stated as ‘School of Public Health, Xi'an Jiaotong University, Shaanxi, China’, while the correct one should be ‘Center for Genetic Epidemiology and Genomics, School of Public Health, Soochow University, Jiangsu, P. R. China’.
Lei Zhang 0041, Yu-Fang Pei, Xiaoying Fu, Yu-Ping Wang 0002, Hong-Wen Deng
Bioinform.6
2014 Critical assessment of coalescent simulators in modeling recombination hotspots in genomic sequences
abstract
BACKGROUND: Coalescent simulation is pivotal for understanding population evolutionary models and demographic histories, as well as for developing novel analytical methods for genetic association studies for DNA sequence data. A plethora of coalescent simulators are developed, but selecting the most appropriate program remains challenging. RESULTS: We extensively compared performances of five widely used coalescent simulators - Hudson's ms, msHOT, MaCS, Simcoal2, and fastsimcoal, to provide a practical guide considering three crucial factors, 1) speed, 2) scalability and 3) recombination hotspot position and intensity accuracy. Although ms represents a popular standard coalescent simulator, it lacks the ability to simulate sequences with recombination hotspots. An extended program msHOT has compensated for the deficiency of ms by incorporating recombination hotspots and gene conversion events at arbitrarily chosen locations and intensities, but remains limited in simulating long stretches of DNA sequences. Simcoal2, based on a discrete generation-by-generation approach, could simulate more complex demographic scenarios, but runs comparatively slow. MaCS and fastsimcoal, both built on fast, modified sequential Markov coalescent algorithms to approximate standard coalescent, are much more efficient whilst keeping salient features of msHOT and Simcoal2, respectively. Our simulations demonstrate that they are more advantageous over other programs for a spectrum of evolutionary models. To validate recombination hotspots, LDhat 2.2 rhomap package, sequenceLDhot and Haploview were compared for hotspot detection, and sequenceLDhot exhibited the best performance based on both real and simulated data. CONCLUSIONS: While ms remains an excellent choice for general coalescent simulations of DNA sequences, MaCS and fastsimcoal are much more scalable and flexible in simulating a variety of demographic events under different recombination hotspot models. Furthermore, sequenceLDhot appears to give the most optimal performance in detecting and validating cross-over hotspots.
Hong-Wen Deng, Tianhua Niu
BMC Bioinform.2
2013 Modeling exome sequencing data with generalized Gaussian distribution with application to copy number variation detection
abstract
Exome sequencing provides us an effective way to discover genetic factors that might be associated with phenotypes for complex diseases. Compared with the whole-genome sequencing, exome sequencing can satisfy the high sequencing coverage requirement while under the limited budge constraint. However, due to the nature that exons are distributed sparsely along the genome, and the technical variability between samples, the analysis of exome sequencing data is complicated and direct utilization of current whole-genome sequencing targeted methods yields wrong results. In this paper, we propose a novel model to represent the exome sequencing data. Under this model, we show that the technical variability as well as random sequencing error follow the generalized Gaussian distribution. Based on this observation, we propose a method to detect the copy number variation. Studies on real data from 1000 Genomes Projects validate the proposed algorithm.
Junbo Duan, Mingxi Wan, Hong-Wen Deng, Yu-Ping Wang 0002
BIBM3
2013 Network-based investigation of genetic modules associated with functional brain networks in schizophrenia
abstract
We developed a new sparse multivariate regression method, collaborative sparse reduced rank regression(C-sRRR) for detecting genetic networks associated with brain functional networks in schizophrenia (SZ). Our study: 1) introduced both genetic and brain network structure to group single nucleotide polymorphism (SNP) and voxels simultaneously for utilizing the interacting effects implied in both features; 2) used collaborative sparse group lasso to perform genetic variants selection and nuclear norm penalty to address the interrelationship among voxels; 3) developed an efficient algorithm for solving the non-smooth optimization. In real data analysis, we constructed 8605 genetic sub-networks (modules) from 722177 SNPs with a median module size of 9. A functional brain network was extracted which also showed significant discriminative characteristics between SZ and healthy controls. A sub sampling strategy was applied to identify 57 highly ranked genes from 14 high-ranking modules. 14 of them are SZ susceptibility genes and 6 genes were consistent with the findings in previous study.
Dongdong Lin, Hao He 0002, Hong-Wen Deng, Vince D. Calhoun, Yu-Ping Wang 0002
BIBM4
2013 CNV-TV: A robust method to discover copy number variation from short sequencing reads
abstract
BACKGROUND: Copy number variation (CNV) is an important structural variation (SV) in human genome. Various studies have shown that CNVs are associated with complex diseases. Traditional CNV detection methods such as fluorescence in situ hybridization (FISH) and array comparative genomic hybridization (aCGH) suffer from low resolution. The next generation sequencing (NGS) technique promises a higher resolution detection of CNVs and several methods were recently proposed for realizing such a promise. However, the performances of these methods are not robust under some conditions, e.g., some of them may fail to detect CNVs of short sizes. There has been a strong demand for reliable detection of CNVs from high resolution NGS data. RESULTS: A novel and robust method to detect CNV from short sequencing reads is proposed in this study. The detection of CNV is modeled as a change-point detection from the read depth (RD) signal derived from the NGS, which is fitted with a total variation (TV) penalized least squares model. The performance (e.g., sensitivity and specificity) of the proposed approach are evaluated by comparison with several recently published methods on both simulated and real data from the 1000 Genomes Project. CONCLUSION: The experimental results showed that both the true positive rate and false positive rate of the proposed detection method do not change significantly for CNVs with different copy numbers and lengthes, when compared with several existing methods. Therefore, our proposed approach results in a more reliable detection of CNVs than the existing methods.
Junbo Duan, Ji-Gang Zhang, Hong-Wen Deng, Yu-Ping Wang 0002
BMC Bioinform.3
2013 Group sparse canonical correlation analysis for genomic data integration
abstract
BACKGROUND: The emergence of high-throughput genomic datasets from different sources and platforms (e.g., gene expression, single nucleotide polymorphisms (SNP), and copy number variation (CNV)) has greatly enhanced our understandings of the interplay of these genomic factors as well as their influences on the complex diseases. It is challenging to explore the relationship between these different types of genomic data sets. In this paper, we focus on a multivariate statistical method, canonical correlation analysis (CCA) method for this problem. Conventional CCA method does not work effectively if the number of data samples is significantly less than that of biomarkers, which is a typical case for genomic data (e.g., SNPs). Sparse CCA (sCCA) methods were introduced to overcome such difficulty, mostly using penalizations with l-1 norm (CCA-l1) or the combination of l-1and l-2 norm (CCA-elastic net). However, they overlook the structural or group effect within genomic data in the analysis, which often exist and are important (e.g., SNPs spanning a gene interact and work together as a group). RESULTS: We propose a new group sparse CCA method (CCA-sparse group) along with an effective numerical algorithm to study the mutual relationship between two different types of genomic data (i.e., SNP and gene expression). We then extend the model to a more general formulation that can include the existing sCCA models. We apply the model to feature/variable selection from two data sets and compare our group sparse CCA method with existing sCCA methods on both simulation and two real datasets (human gliomas data and NCI60 data). We use a graphical representation of the samples with a pair of canonical variates to demonstrate the discriminating characteristic of the selected features. Pathway analysis is further performed for biological interpretation of those features. CONCLUSIONS: The CCA-sparse group method incorporates group effects of features into the correlation analysis while performs individual feature selection simultaneously. It outperforms the two sCCA methods (CCA-l1 and CCA-group) by identifying the correlated features with more true positives while controlling total discordance at a lower level on the simulated data, even if the group effect does not exist or there are irrelevant features grouped with true correlated features. Compared with our proposed CCA-group sparse models, CCA-l1 tends to select less true correlated features while CCA-group inclines to select more redundant features.
Dongdong Lin, Ji-Gang Zhang, Vince D. Calhoun, Hong-Wen Deng, Yu-Ping Wang 0002
BMC Bioinform.5
2011 Comparative studies of de novo assembly tools for next-generation sequencing technologies
abstract
MOTIVATION: Several new de novo assembly tools have been developed recently to assemble short sequencing reads generated by next-generation sequencing platforms. However, the performance of these tools under various conditions has not been fully investigated, and sufficient information is not currently available for informed decisions to be made regarding the tool that would be most likely to produce the best performance under a specific set of conditions. RESULTS: We studied and compared the performance of commonly used de novo assembly tools specifically designed for next-generation sequencing data, including SSAKE, VCAKE, Euler-sr, Edena, Velvet, ABySS and SOAPdenovo. Tools were compared using several performance criteria, including N50 length, sequence coverage and assembly accuracy. Various properties of read data, including single-end/paired-end, sequence GC content, depth of coverage and base calling error rates, were investigated for their effects on the performance of different assembly tools. We also compared the computation time and memory usage of these seven tools. Based on the results of our comparison, the relative performance of individual tools are summarized and tentative guidelines for optimal selection of different assembly tools, under different conditions, are provided.
Jian Li 0012, Hui Shen 0007, Lei Zhang 0041, Christopher J. Papasian, Hong-Wen Deng
Bioinform.6
2011 Integrated Analysis of Gene Expression and Copy Number Data on Gene Shaving Using Independent Component Analysis
abstract
DNA microarray gene expression and microarray-based comparative genomic hybridization (aCGH) have been widely used for biomedical discovery. Because of the large number of genes and the complex nature of biological networks, various analysis methods have been proposed. One such method is "gene shaving," a procedure which identifies subsets of the genes with coherent expression patterns and large variation across samples. Since combining genomic information from multiple sources can improve classification and prediction of diseases, in this paper we proposed a new method, "ICA gene shaving" (ICA, independent component analysis), for jointly analyzing gene expression and copy number data. First we used ICA to analyze joint measurements, gene expression and copy number, of a biological system and project the data onto statistically independent biological processes. Next, we used these results to identify patterns of variation in the data and then applied an iterative shaving method. We investigated the properties of our proposed method by analyzing both simulated and real data. We demonstrated that the robustness of our method to noise using simulated data. Using breast cancer data, we showed that our method is superior to the Generalized Singular Value Decomposition (GSVD) gene shaving method for identifying genes associated with breast cancer.
Jinhua Sheng, Hong-Wen Deng, Vince D. Calhoun, Yu-Ping Wang 0002
IEEE ACM Trans. Comput. Biol. Bioinform.2
2009 Bayesian robust analysis for genetic architecture of quantitative traits
abstract
MOTIVATION: In most quantitative trait locus (QTL) mapping studies, phenotypes are assumed to follow normal distributions. Deviations from this assumption may affect the accuracy of QTL detection and lead to detection of spurious QTLs. To improve the robustness of QTL mapping methods, we replaced the normal distribution for residuals in multiple interacting QTL models with the normal/independent distributions that are a class of symmetric and long-tailed distributions and are able to accommodate residual outliers. Subsequently, we developed a Bayesian robust analysis strategy for dissecting genetic architecture of quantitative traits and for mapping genome-wide interacting QTLs in line crosses. RESULTS: Through computer simulations, we showed that our strategy had a similar power for QTL detection compared with traditional methods assuming normal-distributed traits, but had a substantially increased power for non-normal phenotypes. When this strategy was applied to a group of traits associated with physical/chemical characteristics and quality in rice, more main and epistatic QTLs were detected than traditional Bayesian model analyses under the normal assumption.
Runqing Yang, Jian Li 0012, Hong-Wen Deng
Bioinform.4
2009 A new permutation strategy of pathway-based approach for genome-wide association study
abstract
BACKGROUND: Recently introduced pathway-based approach is promising and advantageous to improve the efficiency of analyzing genome-wide association scan (GWAS) data to identify disease variants by jointly considering variants of the genes that belong to the same biological pathway. However, the current available pathway-based approaches for analyzing GWAS have limited power and efficiency. RESULTS: We proposed a new and efficient permutation strategy based on SNP randomization for determining significance in pathway analysis of GWAS. The developed permutation strategy was evaluated and compared to two previously available methods, i.e. sample permutation and gene permutation, through simulation studies and a study on a real dataset. Results showed that the proposed permutation strategy is more powerful and efficient with greatly reducing the computational complexity. CONCLUSION: Our findings indicate the improved performance of SNP permutation and thus render pathway-based analysis of GWAS more applicable and attractive.
Yan-Fang Guo, Jian Li 0012, Li-Shu Zhang, Hong-Wen Deng
BMC Bioinform.5
2008 HAPSIMU: a genetic simulation platform for population-based association studies
abstract
BACKGROUND: Population structure is an important cause leading to inconsistent results in population-based association studies (PBAS) of human diseases. Various statistical methods have been proposed to reduce the negative impact of population structure on PBAS. Due to lack of structural information in real populations, it is difficult to evaluate the impact of population structure on PBAS in real populations. RESULTS: We developed a genetic simulation platform, HAPSIMU, based on real haplotype data from the HapMap ENCODE project. This platform can simulate heterogeneous populations with various known and controllable structures under the continuous migration model or the discrete model. Moreover, both qualitative and quantitative traits can be simulated using additive genetic model with various genetic parameters designated by users. CONCLUSION: HAPSIMU provides a common genetic simulation platform to evaluate the impact of population structure on PBAS, and compare the relative performance of various population structure identification and PBAS methods.
Jianfeng Liu 0003, Jie Chen 0008, Hong-Wen Deng
BMC Bioinform.4
2007 Gene selection for classification of microarray data based on the Bayes error
abstract
ABSTRACT: BACKGROUND: With DNA microarray data, selecting a compact subset of discriminative genes from thousands of genes is a critical step for accurate classification of phenotypes for, e.g., disease diagnosis. Several widely used gene selection methods often select top-ranked genes according to their individual discriminative power in classifying samples into distinct categories, without considering correlations among genes. A limitation of these gene selection methods is that they may result in gene sets with some redundancy and yield an unnecessary large number of candidate genes for classification analyses. Some latest studies show that incorporating gene to gene correlations into gene selection can remove redundant genes and improve classification accuracy. RESULTS: In this study, we propose a new method, Based Bayes error Filter (BBF), to select relevant genes and remove redundant genes in classification analyses of microarray data. The effectiveness and accuracy of this method is demonstrated through analyses of five publicly available microarray datasets. The results show that our gene selection method is capable of achieving better accuracies than previous studies, while being able to effectively select relevant genes, remove redundant genes and obtain efficient and small gene sets for sample classification purposes. CONCLUSION: The proposed method can effectively identify a compact set of genes with high classification accuracy. This study also indicates that application of the Bayes error is a feasible and effective wayfor removing redundant genes in gene selection.
Ji-Gang Zhang, Hong-Wen Deng
BMC Bioinform.2
2006 JADE: a distributed Java application for deleterious genomic mutation (DGM) estimation
abstract
SUMMARY: The characterization of deleterious genomic mutation (DGM) is of central significance for evolutionary biology and genetic studies. Fitness moment method has been developed to efficiently characterize DGM from natural population directly. In order to enable researchers to employ this method for theoretical and empirical research on characterizing DGM, we here present a distributed Java Application for DGM Estimation (JADE). AVAILABILITY: http://orclinux.creighton.edu/DGM/index.htm.
Miao-Xin Li, Yan-Fang Guo, H.-Y. Deng, Hong-Wen Deng
Bioinform.5
2005 PhD: a web database application for phenotype data management
abstract
Summary: A database application has been developed for phenotype data management employing the Entity-Attribute-Value (EAV) model. By applying the EAV model, this application allows users to manage arbitrary phenotypes and customize data entry forms; therefore, it is suitable for different and multi-center projects. Availability: http://apps.sbri.org/gpdb (Beta version). Contacts: [email protected] Supplementary information: http://apps.sbri.org/gpdb/supp.htm
Miao-Xin Li, H.-Y. Deng, P. E. Duffy, Hong-Wen Deng
Bioinform.5
2005 Hotelling's T2 multivariate profiling for detecting differential expression in microarrays
abstract
The most widely used statistical methods for finding differentially expressed genes (DEGs) are essentially univariate. In this study, we present a new T(2) statistic for analyzing microarray data. We implemented our method using a multiple forward search (MFS) algorithm that is designed for selecting a subset of feature vectors in high-dimensional microarray datasets. The proposed T2 statistic is a corollary to that originally developed for multivariate analyses and possesses two prominent statistical properties. First, our method takes into account multidimensional structure of microarray data. The utilization of the information hidden in gene interactions allows for finding genes whose differential expressions are not marginally detectable in univariate testing methods. Second, the statistic has a close relationship to discriminant analyses for classification of gene expression patterns. Our search algorithm sequentially maximizes gene expression difference/distance between two groups of genes. Including such a set of DEGs into initial feature variables may increase the power of classification rules. We validated our method by using a spike-in HGU95 dataset from Affymetrix. The utility of the new method was demonstrated by application to the analyses of gene expression patterns in human liver cancers and breast cancers. Extensive bioinformatics analyses and cross-validation of DEGs identified in the application datasets showed the significant advantages of our new algorithm.
Hong-Wen Deng
Bioinform.4
2005 SNPP: automating large-scale SNP genotype data management
abstract
UNLABELLED: To manage high-throughput single nucleotide polymorphism (SNP) genotyping data efficiently, we developed a dynamic general database management system-SNPP (SNP Processor). It provides several functions, including data importing with comparison, Mendelian inheritance check within pedigrees, data compiling and exporting. Furthermore, SNPP may generate files for repeat genotyping and transform them into files that can be executed by a liquid handling system. AVAILABILITY: http://orclinux.creighton.edu/snpp/ CONTACT: [email protected]
Miao-Xin Li, Yan-Fang Guo, Fu-Hua Xu, Hong-Wen Deng
Bioinform.6