Yong Wang 0001

dblp:84/2694-1 · DBLP profile ↗
← Back
32ranked-venue papers
6as first author
4since 2021 · last 2024
0000-0003-0695-5273ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 28 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 2Theory of computation · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Predicting RNA polymerase II transcriptional elongation pausing and associated histone code
abstract
RNA Polymerase II (Pol II) transcriptional elongation pausing is an integral part of the dynamic regulation of gene transcription in the genome of metazoans. It plays a pivotal role in many vital biological processes and disease progression. However, experimentally measuring genome-wide Pol II pausing is technically challenging and the precise governing mechanism underlying this process is not fully understood. Here, we develop RP3 (RNA Polymerase II Pausing Prediction), a network regularized logistic regression machine learning method, to predict Pol II pausing events by integrating genome sequence, histone modification, gene expression, chromatin accessibility, and protein-protein interaction data. RP3 can accurately predict Pol II pausing in diverse cellular contexts and unveil the transcription factors that are associated with the Pol II pausing machinery. Furthermore, we utilize a forward feature selection framework to systematically identify the combination of histone modification signals associated with Pol II pausing. RP3 is freely available at https://github.com/AMSSwanglab/RP3.
Lixin Ren, Wanbiao Ma, Yong Wang 0001
Briefings Bioinform.3
2024 ClusterMatch aligns single-cell RNA-sequencing data at the multi-scale cluster level via stable matching
abstract
MOTIVATION: Unsupervised clustering of single-cell RNA sequencing (scRNA-seq) data holds the promise of characterizing known and novel cell type in various biological and clinical contexts. However, intrinsic multi-scale clustering resolutions poses challenges to deal with multiple sources of variability in the high-dimensional and noisy data. RESULTS: We present ClusterMatch, a stable match optimization model to align scRNA-seq data at the cluster level. In one hand, ClusterMatch leverages the mutual correspondence by canonical correlation analysis and multi-scale Louvain clustering algorithms to identify cluster with optimized resolutions. In the other hand, it utilizes stable matching framework to align scRNA-seq data in the latent space while maintaining interpretability with overlapped marker gene set. Through extensive experiments, we demonstrate the efficacy of ClusterMatch in data integration, cell type annotation, and cross-species/timepoint alignment scenarios. Our results show ClusterMatch's ability to utilize both global and local information of scRNA-seq data, sets the appropriate resolution of multi-scale clustering, and offers interpretability by utilizing marker genes. AVAILABILITY AND IMPLEMENTATION: The code of ClusterMatch software is freely available at https://github.com/AMSSwanglab/ClusterMatch.
Teer Ba, Lirong Zhang, Caixia Gao, Yong Wang 0001
Bioinform.5
2022 Annotating regulatory elements by heterogeneous network embedding
abstract
MOTIVATION: Regulatory elements (REs), such as enhancers and promoters, are known as regulatory sequences functional in a heterogeneous regulatory network to control gene expression by recruiting transcription regulators and carrying genetic variants in a context specific way. Annotating those REs relies on costly and labor-intensive next-generation sequencing and RNA-guided editing technologies in many cellular contexts. RESULTS: We propose a systematic Gene Ontology Annotation method for Regulatory Elements (RE-GOA) by leveraging the powerful word embedding in natural language processing. We first assemble a heterogeneous network by integrating context specific regulations, protein-protein interactions and gene ontology (GO) terms. Then we perform network embedding and associate regulatory elements with GO terms by assessing their similarity in a low dimensional vector space. With three applications, we show that RE-GOA outperforms existing methods in annotating TFs' binding sites from ChIP-seq data, in functional enrichment analysis of differentially accessible peaks from ATAC-seq data, and in revealing genetic correlation among phenotypes from their GWAS summary statistics data. AVAILABILITY AND IMPLEMENTATION: The source code and the systematic RE annotation for human and mouse are available at https://github.com/AMSSwanglab/RE-GOA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yurun Lu, Zhan-Ying Feng, Songmao Zhang, Yong Wang 0001
Bioinform.4
2021 MIMIC: an optimization method to identify cell type-specific marker panel for cell sorting
abstract
Multi-omics data allow us to select a small set of informative markers for the discrimination of specific cell types and study of cellular heterogeneity. However, it is often challenging to choose an optimal marker panel from the high-dimensional molecular profiles for a large amount of cell types. Here, we propose a method called Mixed Integer programming Model to Identify Cell type-specific marker panel (MIMIC). MIMIC maintains the hierarchical topology among different cell types and simultaneously maximizes the specificity of a fixed number of selected markers. MIMIC was benchmarked on the mouse ENCODE RNA-seq dataset, with 29 diverse tissues, for 43 surface markers (SMs) and 1345 transcription factors (TFs). MIMIC could select biologically meaningful markers and is robust for different accuracy criteria. It shows advantages over the standard single gene-based approaches and widely used dimensional reduction methods, such as multidimensional scaling and t-SNE, both in accuracy and in biological interpretation. Furthermore, the combination of SMs and TFs achieves better specificity than SMs or TFs alone. Applying MIMIC to a large collection of 641 RNA-seq samples covering 231 cell types identifies a panel of TFs and SMs that reveal the modularity of cell type association networks. Finally, the scalability of MIMIC is demonstrated by selecting enhancer markers from mouse ENCODE data. MIMIC is freely available at https://github.com/MengZou1/MIMIC.
Meng Zou, Zhana Duren, Qiuyue Yuan, Henry Li, Andrew Paul Hutchins, Wing Hung Wong, Yong Wang 0001
Briefings Bioinform.7
2020 scTIM: seeking cell-type-indicative marker from single cell RNA-seq data by consensus optimization
abstract
MOTIVATION: Single cell RNA-seq data offers us new resource and resolution to study cell type identity and its conversion. However, data analyses are challenging in dealing with noise, sparsity and poor annotation at single cell resolution. Detecting cell-type-indicative markers is promising to help denoising, clustering and cell type annotation. RESULTS: We developed a new method, scTIM, to reveal cell-type-indicative markers. scTIM is based on a multi-objective optimization framework to simultaneously maximize gene specificity by considering gene-cell relationship, maximize gene's ability to reconstruct cell-cell relationship and minimize gene redundancy by considering gene-gene relationship. Furthermore, consensus optimization is introduced for robust solution. Experimental results on three diverse single cell RNA-seq datasets show scTIM's advantages in identifying cell types (clustering), annotating cell types and reconstructing cell development trajectory. Applying scTIM to the large-scale mouse cell atlas data identifies critical markers for 15 tissues as 'mouse cell marker atlas', which allows us to investigate identities of different tissues and subtle cell types within a tissue. scTIM will serve as a useful method for single cell RNA-seq data mining. AVAILABILITY AND IMPLEMENTATION: scTIM is freely available at https://github.com/Frank-Orwell/scTIM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhan-Ying Feng, Xianwen Ren, Yining Yin, Chutian Huang, Yong Wang 0001
Bioinform.7
2020 ELF: Extract Landmark Features By Optimizing Topology Maintenance, Redundancy, and Specificity
abstract
Feature selection is the process of selecting a subset of landmark features for model construction when there are many features and a comparatively few samples. The far-reaching development technologies such as biological sequencing at single cell level make feature selection a more challenging work. The difficulty lies in four facts: those features measured are in high dimension and with noise; dropouts make the data much sparse; many features are either redundant or irrelevant; and samples are not well-labeled in the experiments. Here, we propose a new model called ELF (Extract Landmark Features) to address the above challenges. ELF aims to simultaneously maximize topology maintenance to keep the pairwise relationships among samples, minimize feature redundancy to diversify the features, and maximize feature specificity to make every selected feature more representative. This makes ELF a nonlinear combinatorial optimization. To solve this difficult problem, we propose a heuristic algorithm based on greedy strategy. We show ELF's outstanding performance on two single cell RNA-seq datasets. One is the direct reprogramming from mouse embryonic fibroblasts to induced neuron and the other is hepatoblast differentiation. ELF is able to choose only hundreds of landmark genes to maintain the cells' correlativity. Topology maintenance, redundancy removal, and specificity each plays its important role in selecting landmark features and revealing cells' biological functions. In addition, ELF can be generally applied in other scenarios. We demonstrate that ELF can reveal pivotal pixel in writing region and human face in two public image datasets. We believe that ELF is a useful tool to obtain more interpretable results by revealing key features while clustering the samples well.
Zhan-Ying Feng, Yong Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2019 Systems biology intertwines with single cell and AI
abstract
A report of the 12th International Conference on Systems Biology (ISB2018), 18-21 August, Guiyang, China.
Yong Wang 0001, Xiang-Sun Zhang, Luonan Chen
BMC Bioinform.1
2019 Associating lncRNAs with small molecules via bilevel optimization reveals cancer-related lncRNAs
abstract
Long noncoding RNA (lncRNA) transcripts have emerging impacts in cancer studies, which suggests their potential as novel therapeutic agents. However, the molecular mechanism behind their treatment effects is still unclear. Here, we designed a computational model to Associate LncRNAs with Anti-Cancer Drugs (ALACD) based on a bilevel optimization model, which optimized the gene signature overlap in the upper level and imputed the missing lncRNA-gene association in the lower level. ALACD predicts genes coexpressed with lncRNAs mean while matching drug's gene signatures. This model allows us to borrow the target gene information of small molecules to understand the mechanisms of action of lncRNAs and their roles in cancer. The ALACD model was systematically applied to the 10 cancer types in The Cancer Genome Atlas (TCGA) that had matched lncRNA and mRNA expression data. Cancer type-specific lncRNAs and associated drugs were identified. These lncRNAs show significantly different expression levels in cancer patients. Follow-up functional and molecular pathway analysis suggest the gene signatures bridging drugs and lncRNAs are closely related to cancer development. Importantly, patient survival information and evidence from the literature suggest that the lncRNAs and drug-lncRNA associations identified by the ALACD model can provide an alternative choice for cancer targeting treatment and potential cancer pognostic biomarkers. The ALACD model is freely available at https://github.com/wangyc82/ALACD-v1.
Yongcui Wang, Shilong Chen, Luonan Chen, Yong Wang 0001
PLoS Comput. Biol.4
2019 A Semisupervised Classification Approach for Multidomain Networks With Domain Selection
abstract
Multidomain network classification has attracted significant attention in data integration and machine learning, which can enhance network classification or prediction performance by integrating information from different sources. Despite the previous success, existing multidomain network learning methods usually assume that different views are available for the same set of instances, and thus, they seek a consistent classification result for all domains. However, in many real-world problems, each domain has its specific instance set, and one instance in one domain may correspond to multiple instances in another domain. Moreover, due to the rapid growth of data sources, different domains may not be relevant to each other, which asks for selecting domains relevant to the target/focused domain. A key challenge under this setting is how to achieve accurate prediction by integrating different data representations without losing data information. In this paper, we propose a semisupervised classification approach for a multidomain network based on label propagation, i.e., multidomain classification with domain selection (MCS), which can deal with the cross-domain information and different instance sets in domains. In particular, with sparse weight properties, the proposed MCS can automatically identify those domains relevant to our target domain by assigning them higher weights than the other irrelevant domains. This not only significantly improves a classification accuracy but also helps to obtain optimal network partition for the target domain. From the theoretical viewpoint, we equivalently decompose MCS into two simpler subproblems with analytical solutions, which can be efficiently solved by their computational procedures. Extensive experimental results on both synthetic and real-world data sets empirically demonstrate the advantages of the proposed approach in terms of both prediction performance and domain selection ability.
Chuan Chen 0001, Jingxue Xin, Yong Wang 0001, Luonan Chen, Michael Kwok-Po Ng
IEEE Trans. Neural Networks Learn. Syst.3
2016 Computational probing protein-protein interactions targeting small molecules
abstract
MOTIVATION: With the booming of interactome studies, a lot of interactions can be measured in a high throughput way and large scale datasets are available. It is becoming apparent that many different types of interactions can be potential drug targets. Compared with inhibition of a single protein, inhibition of protein-protein interaction (PPI) is promising to improve the specificity with fewer adverse side-effects. Also it greatly broadens the drug target search space, which makes the drug target discovery difficult. Computational methods are highly desired to efficiently provide candidates for further experiments and hold the promise to greatly accelerate the discovery of novel drug targets. RESULTS: Here, we propose a machine learning method to predict PPI targets in a genomic-wide scale. Specifically, we develop a computational method, named as PrePPItar, to Predict PPIs as drug targets by uncovering the potential associations between drugs and PPIs. First, we survey the databases and manually construct a gold-standard positive dataset for drug and PPI interactions. This effort leads to a dataset with 227 associations among 63 PPIs and 113 FDA-approved drugs and allows us to build models to learn the association rules from the data. Second, we characterize drugs by profiling in chemical structure, drug ATC-code annotation, and side-effect space and represent PPI similarity by a symmetrical S-kernel based on protein amino acid sequence. Then the drugs and PPIs are correlated by Kronecker product kernel. Finally, a support vector machine (SVM), is trained to predict novel associations between drugs and PPIs. We validate our PrePPItar method on the well-established gold-standard dataset by cross-validation. We find that all chemical structure, drug ATC-code, and side-effect information are predictive for PPI target. Moreover, we can increase the PPI target prediction coverage by integrating multiple data sources. Follow-up database search and pathway analysis indicate that our new predictions are worthy of future experimental validation. CONCLUSION: In conclusion, PrePPItar can serve as a useful tool for PPI target discovery and provides a general heterogeneous data integrative framework. AVAILABILITY AND IMPLEMENTATION: PrePPItar is available at http://doc.aporc.org/wiki/PrePPItar. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yong-Cui Wang, Shi-Long Chen, Naiyang Deng, Yong Wang 0001
Bioinform.4
2015 NCC-AUC: an AUC optimization method to identify multi-biomarker panel for cancer prognosis from genomic and clinical data
abstract
MOTIVATION: In prognosis and survival studies, an important goal is to identify multi-biomarker panels with predictive power using molecular characteristics or clinical observations. Such analysis is often challenged by censored, small-sample-size, but high-dimensional genomic profiles or clinical data. Therefore, sophisticated models and algorithms are in pressing need. RESULTS: In this study, we propose a novel Area Under Curve (AUC) optimization method for multi-biomarker panel identification named Nearest Centroid Classifier for AUC optimization (NCC-AUC). Our method is motived by the connection between AUC score for classification accuracy evaluation and Harrell's concordance index in survival analysis. This connection allows us to convert the survival time regression problem to a binary classification problem. Then an optimization model is formulated to directly maximize AUC and meanwhile minimize the number of selected features to construct a predictor in the nearest centroid classifier framework. NCC-AUC shows its great performance by validating both in genomic data of breast cancer and clinical data of stage IB Non-Small-Cell Lung Cancer (NSCLC). For the genomic data, NCC-AUC outperforms Support Vector Machine (SVM) and Support Vector Machine-based Recursive Feature Elimination (SVM-RFE) in classification accuracy. It tends to select a multi-biomarker panel with low average redundancy and enriched biological meanings. Also NCC-AUC is more significant in separation of low and high risk cohorts than widely used Cox model (Cox proportional-hazards regression model) and L1-Cox model (L1 penalized in Cox model). These performance gains of NCC-AUC are quite robust across 5 subtypes of breast cancer. Further in an independent clinical data, NCC-AUC outperforms SVM and SVM-RFE in predictive accuracy and is consistently better than Cox model and L1-Cox model in grouping patients into high and low risk categories. CONCLUSION: In summary, NCC-AUC provides a rigorous optimization framework to systematically reveal multi-biomarker panel from genomic and clinical data. It can serve as a useful tool to identify prognostic biomarkers for survival analysis. AVAILABILITY AND IMPLEMENTATION: NCC-AUC is available at http://doc.aporc.org/wiki/NCC-AUC. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Meng Zou, Zhaoqi Liu, Xiang-Sun Zhang, Yong Wang 0001
Bioinform.4
2013 Network predicting drug's anatomical therapeutic chemical code
abstract
MOTIVATION: Discovering drug's Anatomical Therapeutic Chemical (ATC) classification rules at molecular level is of vital importance to understand a vast majority of drugs action. However, few studies attempt to annotate drug's potential ATC-codes by computational approaches. RESULTS: Here, we introduce drug-target network to computationally predict drug's ATC-codes and propose a novel method named NetPredATC. Starting from the assumption that drugs with similar chemical structures or target proteins share common ATC-codes, our method, NetPredATC, aims to assign drug's potential ATC-codes by integrating chemical structures and target proteins. Specifically, we first construct a gold-standard positive dataset from drugs' ATC-code annotation databases. Then we characterize ATC-code and drug by their similarity profiles and define kernel function to correlate them. Finally, we use a kernel method, support vector machine, to automatically predict drug's ATC-codes. Our method was validated on four drug datasets with various target proteins, including enzymes, ion channels, G-protein couple receptors and nuclear receptors. We found that both drug's chemical structure and target protein are predictive, and target protein information has better accuracy. Further integrating these two data sources revealed more experimentally validated ATC-codes for drugs. We extensively compared our NetPredATC with SuperPred, which is a chemical similarity-only based method. Experimental results showed that our NetPredATC outperforms SuperPred not only in predictive coverage but also in accuracy. In addition, database search and functional annotation analysis support that our novel predictions are worthy of future experimental validation. CONCLUSION: In conclusion, our new method, NetPredATC, can predict drug's ATC-codes more accurately by incorporating drug-target network and integrating data, which will promote drug mechanism understanding and drug repositioning and discovery. AVAILABILITY: NetPredATC is available at http://doc.aporc.org/wiki/NetPredATC. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yong-Cui Wang, Shi-Long Chen, Naiyang Deng, Yong Wang 0001
Bioinform.4
2012 A unified computational model for revealing and predicting subtle subtypes of cancers
abstract
BACKGROUND: Gene expression profiling technologies have gradually become a community standard tool for clinical applications. For example, gene expression data has been analyzed to reveal novel disease subtypes (class discovery) and assign particular samples to well-defined classes (class prediction). In the past decade, many effective methods have been proposed for individual applications. However, there is still a pressing need for a unified framework that can reveal the complicated relationships between samples. RESULTS: We propose a novel convex optimization model to perform class discovery and class prediction in a unified framework. An efficient algorithm is designed and software named OTCC (Optimization Tool for Clustering and Classification) is developed. Comparison in a simulated dataset shows that our method outperforms the existing methods. We then applied OTCC to acute leukemia and breast cancer datasets. The results demonstrate that our method not only can reveal the subtle structures underlying those cancer gene expression data but also can accurately predict the class labels of unknown cancer samples. Therefore, our method holds the promise to identify novel cancer subtypes and improve diagnosis. CONCLUSIONS: We propose a unified computational framework for class discovery and class prediction to facilitate the discovery and prediction of subtle subtypes of cancers. Our method can be generally applied to multiple types of measurements, e.g., gene expression profiling, proteomic measuring, and recent next-generation sequencing, since it only requires the similarities among samples as input.
Xian-Wen Ren, Yong Wang 0001, Xiang-Sun Zhang
BMC Bioinform.2
2011 Improving accuracy of protein-protein interaction prediction by considering the converse problem for sequence representation
abstract
BACKGROUND: With the development of genome-sequencing technologies, protein sequences are readily obtained by translating the measured mRNAs. Therefore predicting protein-protein interactions from the sequences is of great demand. The reason lies in the fact that identifying protein-protein interactions is becoming a bottleneck for eventually understanding the functions of proteins, especially for those organisms barely characterized. Although a few methods have been proposed, the converse problem, if the features used extract sufficient and unbiased information from protein sequences, is almost untouched. RESULTS: In this study, we interrogate this problem theoretically by an optimization scheme. Motivated by the theoretical investigation, we find novel encoding methods for both protein sequences and protein pairs. Our new methods exploit sufficiently the information of protein sequences and reduce artificial bias and computational cost. Thus, it significantly outperforms the available methods regarding sensitivity, specificity, precision, and recall with cross-validation evaluation and reaches ~80% and ~90% accuracy in Escherichia coli and Saccharomyces cerevisiae respectively. Our findings here hold important implication for other sequence-based prediction tasks because representation of biological sequence is always the first step in computational biology. CONCLUSIONS: By considering the converse problem, we propose new representation methods for both protein sequences and protein pairs. The results show that our method significantly improves the accuracy of protein-protein interaction predictions.
Xian-Wen Ren, Yong-Cui Wang, Yong Wang 0001, Xiang-Sun Zhang, Naiyang Deng
BMC Bioinform.3
2010 Prediction of protein-RNA binding sites by a random forest method with combined features
abstract
MOTIVATION: Protein-RNA interactions play a key role in a number of biological processes, such as protein synthesis, mRNA processing, mRNA assembly, ribosome function and eukaryotic spliceosomes. As a result, a reliable identification of RNA binding site of a protein is important for functional annotation and site-directed mutagenesis. Accumulated data of experimental protein-RNA interactions reveal that a RNA binding residue with different neighbor amino acids often exhibits different preferences for its RNA partners, which in turn can be assessed by the interacting interdependence of the amino acid fragment and RNA nucleotide. RESULTS: In this work, we propose a novel classification method to identify the RNA binding sites in proteins by combining a new interacting feature (interaction propensity) with other sequence- and structure-based features. Specifically, the interaction propensity represents a binding specificity of a protein residue to the interacting RNA nucleotide by considering its two-side neighborhood in a protein residue triplet. The sequence as well as the structure-based features of the residues are combined together to discriminate the interaction propensity of amino acids with RNA. We predict RNA interacting residues in proteins by implementing a well-built random forest classifier. The experiments show that our method is able to detect the annotated protein-RNA interaction sites in a high accuracy. Our method achieves an accuracy of 84.5%, F-measure of 0.85 and AUC of 0.92 prediction of the RNA binding residues for a dataset containing 205 non-homologous RNA binding proteins, and also outperforms several existing RNA binding residue predictors, such as RNABindR, BindN, RNAProB and PPRint, and some alternative machine learning methods, such as support vector machine, naive Bayes and neural network in the comparison study. Furthermore, we provide some biological insights into the roles of sequences and structures in protein-RNA interactions by both evaluating the importance of features for their contributions in predictive accuracy and analyzing the binding patterns of interacting residues. AVAILABILITY: All the source data and code are available at http://www.aporc.org/doc/wiki/PRNA or http://www.sysbio.ac.cn/datatools.asp CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhi-Ping Liu, Ling-Yun Wu, Yong Wang 0001, Xiang-Sun Zhang, Luonan Chen
Bioinform.3
2009 Dynamically dysfunctional protein interactions in the development of Alzheimer's disease
abstract
Alzheimer's disease usually causes dementia in the old people and the symptom progression of the disease phenotype displays certain patterns. One possible reason is that the nerve cells in the brains of the patients degenerate at different stages. Here, we analyze the dynamics of disease progression based on its biomolecular network. We develop a novel computational method to integrate an ensemble protein network and the hippocampal gene expression data. Specifically, we construct the induced dynamical pathways which present particular characteristics at different disease stages from the control to disease samples. Based on the network-based method, we reveal that the active pathways tend to be more complicated during the development of disease. Also we find that the disease proteins performing important functions are always located in the cooperations of the identified pathways. These results also demonstrate that the network-based analysis can provide knowledge and evidences on the dynamics and pathological pathways of the complex Alzheimer's disease.
Zhi-Ping Liu, Yong Wang 0001, Tieqiao Wen, Xiang-Sun Zhang, Weiming Xia, Luonan Chen
SMC2
2009 Disease-Aging Network Reveals Significant Roles of Aging Genes in Connecting Genetic Diseases
abstract
One of the challenging problems in biology and medicine is exploring the underlying mechanisms of genetic diseases. Recent studies suggest that the relationship between genetic diseases and the aging process is important in understanding the molecular mechanisms of complex diseases. Although some intricate associations have been investigated for a long time, the studies are still in their early stages. In this paper, we construct a human disease-aging network to study the relationship among aging genes and genetic disease genes. Specifically, we integrate human protein-protein interactions (PPIs), disease-gene associations, aging-gene associations, and physiological system-based genetic disease classification information in a single graph-theoretic framework and find that (1) human disease genes are much closer to aging genes than expected by chance; and (2) diseases can be categorized into two types according to their relationships with aging. Type I diseases have their genes significantly close to aging genes, while type II diseases do not. Furthermore, we examine the topological characters of the disease-aging network from a systems perspective. Theoretical results reveal that the genes of type I diseases are in a central position of a PPI network while type II are not; (3) more importantly, we define an asymmetric closeness based on the PPI network to describe relationships between diseases, and find that aging genes make a significant contribution to associations among diseases, especially among type I diseases. In conclusion, the network-based study provides not only evidence for the intricate relationship between the aging process and genetic diseases, but also biological implications for prying into the nature of human diseases.
Yong Wang 0001, Luonan Chen, Xiang-Sun Zhang
PLoS Comput. Biol.3
2009 Evaluating Protein Similarity from Coarse Structures
abstract
To unscramble the relationship between protein function and protein structure, it is essential to assess the protein similarity from different aspects. Although many methods have been proposed for protein structure alignment or comparison, alternative similarity measures are still strongly demanded due to the requirement of fast screening and query in large-scale structure databases. In this paper, we first formulate a novel representation of a protein structure, i.e., Feature Sequence of Surface (FSS). Then, a new score scheme is developed to measure the similarity between two representations. To verify the proposed method, numerical experiments are conducted in four different protein data sets. We also classify SARS coronavirus to verify the effectiveness of the new method. Furthermore, preliminary results of fast classification of the whole CATH v2.5.1 database based on the new macrostructure similarity are given as a pilot study. We demonstrate that the proposed approach to measure the similarities between protein structures is simple to implement, computationally efficient, and surprisingly fast. In addition, the method itself provides a new and quantitative tool to view a protein structure.
Yong Wang 0001, Ling-Yun Wu, Zhong-Wei Zhan, Xiang-Sun Zhang, Luonan Chen
IEEE ACM Trans. Comput. Biol. Bioinform.1
2008 Gene function prediction using labeled and unlabeled data
abstract
BACKGROUND: In general, gene function prediction can be formalized as a classification problem based on machine learning technique. Usually, both labeled positive and negative samples are needed to train the classifier. For the problem of gene function prediction, however, the available information is only about positive samples. In other words, we know which genes have the function of interested, while it is generally unclear which genes do not have the function, i.e. the negative samples. If all the genes outside of the target functional family are seen as negative samples, the imbalanced problem will arise because there are only a relatively small number of genes annotated in each family. Furthermore, the classifier may be degraded by the false negatives in the heuristically generated negative samples. RESULTS: In this paper, we present a new technique, namely Annotating Genes with Positive Samples (AGPS), for defining negative samples in gene function prediction. With the defined negative samples, it is straightforward to predict the functions of unknown genes. In addition, the AGPS algorithm is able to integrate various kinds of data sources to predict gene functions in a reliable and accurate manner. With the one-class and two-class Support Vector Machines as the core learning algorithm, the AGPS algorithm shows good performances for function prediction on yeast genes. CONCLUSION: We proposed a new method for defining negative samples in gene function prediction. Experimental results on yeast genes show that AGPS yields good performances on both training and test sets. In addition, the overlapping between prediction results and GO annotations on unknown genes also demonstrates the effectiveness of the proposed method.
Xing-Ming Zhao, Yong Wang 0001, Luonan Chen, Kazuyuki Aihara
BMC Bioinform.2
2008 Clustering complex networks and biological networks by nonnegative matrix factorization with various similarity measures
Rui-Sheng Wang, Yong Wang 0001, Xiang-Sun Zhang, Luonan Chen
Neurocomputing3
2007 A novel approach based on the combination image of fraction image and normalized mnf image to urban land use/cover mapping
abstract
Urban land use/cover mapping is very important and it is the base and foundation of further urban analysis and research. Whereas urban land use/cover mapping of using medium spatial resolution remotely sensed images presents numerous challenges due to the intensive heterogeneity of urban landscapes. In order to solve the above challenges and improve the accuracy of urban land cover/use mapping, we proposed a novel approach, which produced firstly combination image based on fraction image and normalized MNF image and then performed decision tree classification and gained urban land use/cover mapping at last. An ETM+ image was used as data source and Nanjing City, China was selected as study area. The accuracy of classification result was validated using IKONOS images and was compared with the other three classification schemes. Results show that this decision tree classification based on the combination image is superior to the other classification schemes.
Dafang Zhuang, Yong Wang 0001
IGARSS5
2007 Alignment of molecular networks by integer quadratic programming
abstract
MOTIVATION: With more and more data on molecular networks (e.g. protein interaction networks, gene regulatory networks and metabolic networks) available, the discovery of conserved patterns or signaling pathways by comparing various kinds of networks among different species or within a species becomes an increasingly important problem. However, most of the conventional approaches either restrict comparative analysis to special structures, such as pathways, or adopt heuristic algorithms due to computational burden. RESULTS: In this article, to find the conserved substructures, we develop an efficient algorithm for aligning molecular networks based on both molecule similarity and architecture similarity, by using integer quadratic programming (IQP). Such an IQP can be relaxed into the corresponding quadratic programming (QP) which almost always ensures an integer solution, thereby making molecular network alignment tractable without any approximation. The proposed framework is very flexible and can be applied to many kinds of molecular networks including weighted and unweighted, directed and undirected networks with or without loops. AVAILABILITY: Matlab code and data are available from http://zhangroup.aporc.org/bioinfo/MNAligner or http://intelligent.eic.osaka-sandai.ac.jp/chenen/software/MNAligner, or upon request from authors. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zhen-Ping Li, Yong Wang 0001, Xiang-Sun Zhang, Luonan Chen
Bioinform.3
2007 Inferring transcriptional regulatory networks from high-throughput data
abstract
MOTIVATION: Inferring the relationships between transcription factors (TFs) and their targets has utmost importance for understanding the complex regulatory mechanisms in cellular systems. However, the transcription factor activities (TFAs) cannot be measured directly by standard microarray experiment owing to various post-translational modifications. In particular, cooperative mechanism and combinatorial control are common in gene regulation, e.g. TFs usually recruit other proteins cooperatively to facilitate transcriptional reaction processes. RESULTS: In this article, we propose a novel method for inferring transcriptional regulatory networks (TRN) from gene expression data based on protein transcription complexes and mass action law. With gene expression data and TFAs estimated from transcription complex information, the inference of TRN is formulated as a linear programming (LP) problem which has a globally optimal solution in terms of L(1) norm error. The proposed method not only can easily incorporate ChIP-Chip data as prior knowledge, but also can integrate multiple gene expression datasets from different experiments simultaneously. A unique feature of our method is to take into account protein cooperation in transcription process. We tested our method by using both synthetic data and several experimental datasets in yeast. The extensive results illustrate the effectiveness of the proposed method for predicting transcription regulatory relationships between TFs with co-regulators and target genes.
Rui-Sheng Wang, Yong Wang 0001, Xiang-Sun Zhang, Luonan Chen
Bioinform.2
2007 Predicting gene ontology functions from protein's regional surface structures
abstract
BACKGROUND: Annotation of protein functions is an important task in the post-genomic era. Most early approaches for this task exploit only the sequence or global structure information. However, protein surfaces are believed to be crucial to protein functions because they are the main interfaces to facilitate biological interactions. Recently, several databases related to structural surfaces, such as pockets and cavities, have been constructed with a comprehensive library of identified surface structures. For example, CASTp provides identification and measurements of surface accessible pockets as well as interior inaccessible cavities. RESULTS: A novel method was proposed to predict the Gene Ontology (GO) functions of proteins from the pocket similarity network, which is constructed according to the structure similarities of pockets. The statistics of the networks were presented to explore the relationship between the similar pockets and GO functions of proteins. Cross-validation experiments were conducted to evaluate the performance of the proposed method. Results and codes are available at: http://zhangroup.aporc.org/bioinfo/PSN/. CONCLUSION: The computational results demonstrate that the proposed method based on the pocket similarity network is effective and efficient for predicting GO functions of proteins in terms of both computational complexity and prediction accuracy. The proposed method revealed strong relationship between small surface patterns (or pockets) and GO functions, which can be further used to identify active sites or functional motifs. The high quality performance of the prediction method together with the statistics also indicates that pockets play essential roles in biological interactions or the GO functions. Moreover, in addition to pockets, the proposed network framework can also be used for adopting other protein spatial surface patterns to predict the protein functions.
Zhi-Ping Liu, Ling-Yun Wu, Yong Wang 0001, Luonan Chen, Xiang-Sun Zhang
BMC Bioinform.3
2007 Analysis on multi-domain cooperation for predicting protein-protein interactions
abstract
BACKGROUND: Domains are the basic functional units of proteins. It is believed that protein-protein interactions are realized through domain interactions. Revealing multi-domain cooperation can provide deep insights into the essential mechanism of protein-protein interactions at the domain level and be further exploited to improve the accuracy of protein interaction prediction. RESULTS: In this paper, we aim to identify cooperative domains for protein interactions by extending two-domain interactions to multi-domain interactions. Based on the high-throughput experimental data from multiple organisms with different reliabilities, the interactions of domains were inferred by a Linear Programming algorithm with Multi-domain pairs (LPM) and an Association Probabilistic Method with Multi-domain pairs (APMM). Experimental results demonstrate that our approach not only can find cooperative domains effectively but also has a higher accuracy for predicting protein interaction than the existing methods. Cooperative domains, including strongly cooperative domains and superdomains, were detected from major interaction databases MIPS and DIP, and many of them were verified by physical interactions from the crystal structures of protein complexes in PDB which provide intuitive evidences for such cooperation. Comparison experiments in terms of protein/domain interaction prediction justified the benefit of considering multi-domain cooperation. CONCLUSION: From the computational viewpoint, this paper gives a general framework to predict protein interactions in a more accurate manner by considering the information of both multi-domains and multiple organisms, which can also be applied to identify cooperative domains, to reconstruct large complexes and further to annotate functions of domains. Supplementary information and software are provided in http://intelligent.eic.osaka-sandai.ac.jp/chenen/MDCinfer.htm and http://zhangroup.aporc.org/bioinfo/MDCinfer.
Rui-Sheng Wang, Yong Wang 0001, Ling-Yun Wu, Xiang-Sun Zhang, Luonan Chen
BMC Bioinform.2
2006 Supervised Inference of Gene Regulatory Networks by Linear Programming
Yong Wang 0001, Trupti Joshi, Dong Xu 0002, Xiang-Sun Zhang, Luonan Chen
ICIC (3)1
2006 Automatic Classification of Protein Structures Based on Convex Hull Representation by Integrated Neural Network
Yong Wang 0001, Ling-Yun Wu, Xiang-Sun Zhang, Luonan Chen
TAMC1
2006 Protein Structure Comparison Based on a Measure of Information Discrepancy
Zikai Wu, Yong Wang 0001, En-Min Feng, Jin-Cheng Zhao
TAMC2
2006 Inferring gene regulatory networks from multiple microarray datasets
abstract
MOTIVATION: Microarray gene expression data has increasingly become the common data source that can provide insights into biological processes at a system-wide level. One of the major problems with microarrays is that a dataset consists of relatively few time points with respect to a large number of genes, which makes the problem of inferring gene regulatory network an ill-posed one. On the other hand, gene expression data generated by different groups worldwide are increasingly accumulated on many species and can be accessed from public databases or individual websites, although each experiment has only a limited number of time-points. RESULTS: This paper proposes a novel method to combine multiple time-course microarray datasets from different conditions for inferring gene regulatory networks. The proposed method is called GNR (Gene Network Reconstruction tool) which is based on linear programming and a decomposition procedure. The method theoretically ensures the derivation of the most consistent network structure with respect to all of the datasets, thereby not only significantly alleviating the problem of data scarcity but also remarkably improving the prediction reliability. We tested GNR using both simulated data and experimental data in yeast and Arabidopsis. The result demonstrates the effectiveness of GNR in terms of predicting new gene regulatory relationship in yeast and Arabidopsis. AVAILABILITY: The software is available from http://zhangorup.aporc.org/bioinfo/grninfer/, http://digbio.missouri.edu/grninfer/ and http://intelligent.eic.osaka-sandai.ac.jp or upon request from the authors.
Yong Wang 0001, Trupti Joshi, Xiang-Sun Zhang, Dong Xu 0002, Luonan Chen
Bioinform.1
2005 A comparison of university of Maryland 1km land cover dataset and a land cover dataset in China
abstract
Land cover plays important roles in the understanding of the physical, chemical, biological and anthropological process in the earth system sciences. Land cover map at scales from local to global has been produced using the remote sensing data by visual interpretation or automatic classification methods during the past several decades. University of Maryland (UMd) land cover dataset is a global land cover dataset based on remote sensing method produced in recent years. The UMd approach employed a supervised decision tree method to classify global land cover types. This paper makes a comparison of this UMd land cover dataset with a Chinese land cover dataset. Firstly, we present a method to compare land cover datasets produced at different time based on invariant reliable land unit. Secondly we compare UMd land cover dataset with Chinese land cover dataset (CLCD) based on above method. Finally we analyze the possible factors affecting the differences among the land cover classification datasets. The comparison results demonstrate that most of the land surface in China was identified as different types in these two datasets. For example, UMd maps 51.1% of the deciduous needleleaf forest units in CLCD to the mixed forest. The classification scheme and method used in these datasets are the most likely reasons to explain the differences between them.
Dafang Zhuang, Runhe Shi, Yong Wang 0001, Xinfang Yu
IGARSS5
2005 Monitoring the biophysical composition changes of urban environments using normalized spectral mixture analysis
abstract
Analyzing and monitoring biophysical composition changes of urban environments is an important issue due to the rapid urbanization in China. Unfortunately, the estimation of urban composition using moderate spatial resolution satellite images presents great challenges due to the diversity of urban landscapes. In this paper, we performed NSMA proposed by Wu (2004) to analyze biophysical composition changes of Nanjing city, East China of the two dates using TM and ETM+ images. Three percentage composition images corresponding to the three endmembers (vegetation, impervious surface and water or shade) were acquired by applying a fully constrained SMA to the normalized images. The fraction changes of three kinds of composition were systematically analyzed through associating them with urban growth of Nanjing city. The accuracy of composition images was validated using IKONOS images. Results show that NSMA is able to monitor urban biophysical composition changes effectively and has a considerable accuracy.
Dafang Zhuang, Runhe Shi, Yong Wang 0001
IGARSS4
2005 A WebGIS-based system model of vehicle monitoring central platform
abstract
The vehicle monitoring central platform is establishing a uniform information portal on vehicle monitoring in China. Rather than elaborating sophisticated vehicle monitoring and management approaches, the WebGIS-based system model is striving to develop a technical umbrella for the distributed and Web-based vehicle management systems, which employs several technologies, including communication technology - SMS (short message services), vehicle position technology - GPS (global positioning systems), and information display and matching technology - GIS (geographic information systems). The model presents prototype application for vehicle monitoring central platform, which provides an enabling technology to directives and city management for online dissemination of vehicles information. The model's architecture and its designing considerations are introduced in this paper. The systematic framework and structure, the data management components and the program tools to develop final applications are also explained. The prototype system shows this prototype model is an efficient and feasible one.
Yong Wang 0001, Dafang Zhuang, Runhe Shi
IGARSS1