Dong Wang 0011

dblp:40/3934-11 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Deep learning-driven survival prediction in pan-cancer studies by integrating multimodal histology-genomic data
abstract
Accurate cancer prognosis is essential for personalized clinical management, guiding treatment strategies and predicting patient survival. Conventional methods, which depend on the subjective evaluation of histopathological features, exhibit significant inter-observer variability and limited predictive power. To overcome these limitations, we developed cross-attention transformer-based multimodal fusion network (CATfusion), a deep learning framework that integrates multimodal histology-genomic data for comprehensive cancer survival prediction. By employing self-supervised learning strategy with TabAE for feature extraction and utilizing cross-attention mechanisms to fuse diverse data types, including mRNA-seq, miRNA-seq, copy number variation, DNA methylation variation, mutation data, and histopathological images. By successfully integrating this multi-tiered patient information, CATfusion has become an advanced survival prediction model to utilize the most diverse data types across various cancer types. CATfusion's architecture, which includes a bidirectional multimodal attention mechanism and self-attention block, is adept at synchronizing the learning and integration of representations from various modalities. CATfusion achieves superior predictive performance over traditional and unimodal models, as demonstrated by enhanced C-index and survival area under the curve scores. The model's high accuracy in stratifying patients into distinct risk groups is a boon for personalized medicine, enabling tailored treatment plans. Moreover, CATfusion's interpretability, enabled by attention-based visualization, offers insights into the biological underpinnings of cancer prognosis, underscoring its potential as a transformative tool in oncology.
Yongfei Hu, Dong Wang 0011
Briefings Bioinform.6
2023 LncReader: identification of dual functional long noncoding RNAs using a multi-head self-attention mechanism
abstract
Long noncoding ribonucleic acids (RNAs; LncRNAs) endowed with both protein-coding and noncoding functions are referred to as 'dual functional lncRNAs'. Recently, dual functional lncRNAs have been intensively studied and identified as involved in various fundamental cellular processes. However, apart from time-consuming and cell-type-specific experiments, there is virtually no in silico method for predicting the identity of dual functional lncRNAs. Here, we developed a deep-learning model with a multi-head self-attention mechanism, LncReader, to identify dual functional lncRNAs. Our data demonstrated that LncReader showed multiple advantages compared to various classical machine learning methods using benchmark datasets from our previously reported cncRNAdb project. Moreover, to obtain independent in-house datasets for robust testing, mass spectrometry proteomics combined with RNA-seq and Ribo-seq were applied in four leukaemia cell lines, which further confirmed that LncReader achieved the best performance compared to other tools. Therefore, LncReader provides an accurate and practical tool that enables fast dual functional lncRNA identification.
Bohao Zou, Manman He, Yongfei Hu, Yiying Dou, Tianyu Cui, Puwen Tan, Shaobin Li, Shuan Rao, Sixi Liu, Kaican Cai, Dong Wang 0011
Briefings Bioinform.13
2021 Design powerful predictor for mRNA subcellular location prediction in Homo sapiens
abstract
Messenger RNAs (mRNAs) shoulder special responsibilities that transmit genetic code from DNA to discrete locations in the cytoplasm. The locating process of mRNA might provide spatial and temporal regulation of mRNA and protein functions. The situ hybridization and quantitative transcriptomics analysis could provide detail information about mRNA subcellular localization; however, they are time consuming and expensive. It is highly desired to develop computational tools for timely and effectively predicting mRNA subcellular location. In this work, by using binomial distribution and one-way analysis of variance, the optimal nonamer composition was obtained to represent mRNA sequences. Subsequently, a predictor based on support vector machine was developed to identify the mRNA subcellular localization. In 5-fold cross-validation, results showed that the accuracy is 90.12% for Homo sapiens (H. sapiens). The predictor may provide a reference for the study of mRNA localization mechanisms and mRNA translocation strategies. An online web server was established based on our models, which is available at http://lin-group.cn/server/iLoc-mRNA/.
Zhao-Yue Zhang 0002, Hui Ding 0005, Dong Wang 0011, Wei Chen 0064, Hao Lin 0001
Briefings Bioinform.4
2021 Cellinker: a platform of ligand-receptor interactions for intercellular communication analysis
abstract
MOTIVATION: Ligand-receptor (L-R) interactions mediate cell adhesion, recognition and communication and play essential roles in physiological and pathological signaling. With the rapid development of single-cell RNA sequencing (scRNA-seq) technologies, systematically decoding the intercellular communication network involving L-R interactions has become a focus of research. Therefore, construction of a comprehensive, high-confidence and well-organized resource to retrieve L-R interactions in order to study the functional effects of cell-cell communications would be of great value. RESULTS: In this study, we developed Cellinker, a manually curated resource of literature-supported L-R interactions that play roles in cell-cell communication. We aimed to provide a useful platform for studies on cell-cell communication mediated by L-R interactions. The current version of Cellinker documents over 3,700 human and 3,200 mouse L-R protein-protein interactions (PPIs) and embeds a practical and convenient webserver with which researchers can decode intercellular communications based on scRNA-seq data. And over 400 endogenous small molecule (sMOL) related L-R interactions were collected as well. Moreover, to help with research on coronavirus (CoV) infection, Cellinker collects information on 16 L-R PPIs involved in CoV-human interactions (including 12 L-R PPIs involved in SARS-CoV-2 infection). In summary, Cellinker provides a user-friendly interface for querying, browsing and visualizing L-R interactions as well as a practical and convenient web tool for inferring intercellular communications based on scRNA-seq data. We believe this platform could promote intercellular communication research and accelerate the development of related algorithms for scRNA-seq studies. AVAILABILITY: Cellinker is available at http://www.rna-society.org/cellinker/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yang Zhang 0125, Jing Wang 0004, Bohao Zou, Linhui Yao, Kechen Chen, Lin Ning 0002, Bingyi Wu, Dong Wang 0011
Bioinform.11
2019 RIscoper: a tool for RNA-RNA interaction extraction from the literature
abstract
MOTIVATION: Numerous experimental and computational studies in the biomedical literature have provided considerable amounts of data on diverse RNA-RNA interactions (RRIs). However, few text mining systems for RRIs information extraction are available. RESULTS: RNA Interactome Scoper (RIscoper) represents the first tool for full-scale RNA interactome scanning and was developed for extracting RRIs from the literature based on the N-gram model. Notably, a reliable RRI corpus was integrated in RIscoper, and more than 13 300 manually curated sentences with RRI information were recruited. RIscoper allows users to upload full texts or abstracts, and provides an online search tool that is connected with PubMed (PMID and keyword input), and these capabilities are useful for biologists. RIscoper has a strong performance (90.4% precision and 93.9% recall), integrates natural language processing techniques and has a reliable RRI corpus. AVAILABILITY AND IMPLEMENTATION: The standalone software and web server of RIscoper are freely available at www.rna-society.org/riscoper/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yang Zhang 0125, Jinxurong Yang, Jiayi Yin, Yuncong Zhang, Zhixi Yun, Lin Ning 0002, Feng-Biao Guo, Yongshuai Jiang, Hao Lin 0001, Dong Wang 0011, Jian Huang 0004
Bioinform.13
2019 Exploiting locational and topological overlap model to identify modules in protein interaction networks
abstract
BACKGROUND: Clustering molecular network is a typical method in system biology, which is effective in predicting protein complexes or functional modules. However, few studies have realized that biological molecules are spatial-temporally regulated to form a dynamic cellular network and only a subset of interactions take place at the same location in cells. RESULTS: In this study, considering the subcellular localization of proteins, we first construct a co-localization human protein interaction network (PIN) and systematically investigate the relationship between subcellular localization and biological functions. After that, we propose a Locational and Topological Overlap Model (LTOM) to preprocess the co-localization PIN to identify functional modules. LTOM requires the topological overlaps, the common partners shared by two proteins, to be annotated in the same localization as the two proteins. We observed the model has better correspondence with the reference protein complexes and shows more relevance to cancers based on both human and yeast datasets and two clustering algorithms, ClusterONE and MCL. CONCLUSION: Taking into consideration of protein localization and topological overlap can improve the performance of module detection from protein interaction networks.
Lixin Cheng, Dong Wang 0011, Kwong-Sak Leung
BMC Bioinform.3
2018 iLoc-lncRNA: predict the subcellular location of lncRNAs by incorporating octamer composition into general PseKNC
abstract
Motivation: Long non-coding RNAs (lncRNAs) are a class of RNA molecules with more than 200 nucleotides. They have important functions in cell development and metabolism, such as genetic markers, genome rearrangements, chromatin modifications, cell cycle regulation, transcription and translation. Their functions are generally closely related to their localization in the cell. Therefore, knowledge about their subcellular locations can provide very useful clues or preliminary insight into their biological functions. Although biochemical experiments could determine the localization of lncRNAs in a cell, they are both time-consuming and expensive. Therefore, it is highly desirable to develop bioinformatics tools for fast and effective identification of their subcellular locations. Results: We developed a sequence-based bioinformatics tool called 'iLoc-lncRNA' to predict the subcellular locations of LncRNAs by incorporating the 8-tuple nucleotide features into the general PseKNC (Pseudo K-tuple Nucleotide Composition) via the binomial distribution approach. Rigorous jackknife tests have shown that the overall accuracy achieved by the new predictor on a stringent benchmark dataset is 86.72%, which is over 20% higher than that by the existing state-of-the-art predictor evaluated on the same tests. Availability and implementation: A user-friendly webserver has been established at http://lin-group.cn/server/iLoc-LncRNA, by which users can easily obtain their desired results. Supplementary information: Supplementary data are available at Bioinformatics online.
Zhen-Dong Su 0001, Zhao-Yue Zhang 0002, Ya-Wei Zhao, Dong Wang 0011, Wei Chen 0064, Kuo-Chen Chou, Hao Lin 0001
Bioinform.5
2009 Edge-based scoring and searching method for identifying condition-responsive protein-protein interaction sub-network
abstract
Bioinformatics 23(16), 2121–2128. We would like to correct the author list in this manuscript. We apologize to the two authors Lei Wang and Shaoqi Rao who were missed from the published version of the manuscript. The corrected author list is: Zheng Guo, Lei Wang, Yongjin Li, Xue Gong, Chen Yao, Wencai Ma, Dong Wang, Yanhhui Li, Jing Zhu, Min Zhang, Da Yang, Shaoqi Rao and Jing Wang
Zheng Guo 0002, Yongjin Li, Wencai Ma, Dong Wang 0011, Jing Zhu 0004, Min Zhang 0009, Da Yang 0003, Shaoqi Rao, Jing Wang 0004
Bioinform.7
2009 Evaluating reproducibility of differential expression discoveries in microarray studies by considering correlated molecular changes
abstract
MOTIVATION: According to current consistency metrics such as percentage of overlapping genes (POG), lists of differentially expressed genes (DEGs) detected from different microarray studies for a complex disease are often highly inconsistent. This irreproducibility problem also exists in other high-throughput post-genomic areas such as proteomics and metabolism. A complex disease is often characterized with many coordinated molecular changes, which should be considered when evaluating the reproducibility of discovery lists from different studies. RESULTS: We proposed metrics percentage of overlapping genes-related (POGR) and normalized POGR (nPOGR) to evaluate the consistency between two DEG lists for a complex disease, considering correlated molecular changes rather than only counting gene overlaps between the lists. Based on microarray datasets of three diseases, we showed that though the POG scores for DEG lists from different studies for each disease are extremely low, the POGR and nPOGR scores can be rather high, suggesting that the apparently inconsistent DEG lists may be highly reproducible in the sense that they are actually significantly correlated. Observing different discovery results for a disease by the POGR and nPOGR scores will obviously reduce the uncertainty of the microarray studies. The proposed metrics could also be applicable in many other high-throughput post-genomic areas.
Min Zhang 0009, Lin Zhang 0057, Jinfeng Zou, Jing Wang 0004, Dong Wang 0011, Chenguang Wang 0004, Zheng Guo 0002
Bioinform.8
2008 Gaining confidence in biological interpretation of the microarray data: the functional consistence of the significant GO categories
abstract
MOTIVATION: In microarray studies, numerous tools are available for functional enrichment analysis based on GO categories. Most of these tools, due to their requirement of a prior threshold for designating genes as differentially expressed genes (DEGs), are categorized as threshold-dependent methods that often suffer from a major criticism on their changing results with different thresholds. RESULTS: In the present article, by considering the inherent correlation structure of the GO categories, a continuous measure based on semantic similarity of GO categories is proposed to investigate the functional consistence (or stability) of threshold-dependent methods. The results from several datasets show when simply counting overlapping categories between two groups, the significant category groups selected under different DEG thresholds are seemingly very different. However, based on the semantic similarity measure proposed in this article, the results are rather functionally consistent for a wide range of DEG thresholds. Moreover, we find that the functional consistence of gene lists ranked by SAM metric behaves relatively robust against changing DEG thresholds. AVAILABILITY: Source code in R is available on request from the authors.
Da Yang 0003, Min Zhang 0009, Jing Zhu 0004, Wencai Ma, Jing Wang 0004, Dong Wang 0011, Zheng Guo 0002, Baofeng Yang
Bioinform.10
2008 Apparently low reproducibility of true differential expression discoveries in microarray studies
abstract
MOTIVATION: Differentially expressed gene (DEG) lists detected from different microarray studies for a same disease are often highly inconsistent. Even in technical replicate tests using identical samples, DEG detection still shows very low reproducibility. It is often believed that current small microarray studies will largely introduce false discoveries. RESULTS: Based on a statistical model, we show that even in technical replicate tests using identical samples, it is highly likely that the selected DEG lists will be very inconsistent in the presence of small measurement variations. Therefore, the apparently low reproducibility of DEG detection from current technical replicate tests does not indicate low quality of microarray technology. We also demonstrate that heterogeneous biological variations existing in real cancer data will further reduce the overall reproducibility of DEG detection. Nevertheless, in small subsamples from both simulated and real data, the actual false discovery rate (FDR) for each DEG list tends to be low, suggesting that each separately determined list may comprise mostly true DEGs. Rather than simply counting the overlaps of the discovery lists from different studies for a complex disease, novel metrics are needed for evaluating the reproducibility of discoveries characterized with correlated molecular changes. Supplementaty information: Supplementary data are available at Bioinformatics online.
Min Zhang 0009, Zheng Guo 0002, Jinfeng Zou, Lin Zhang 0057, Dong Wang 0011, Da Yang 0003, Jing Zhu 0004, Xia Li 0004
Bioinform.7
2007 Edge-based scoring and searching method for identifying condition-responsive protein-protein interaction sub-network
abstract
MOTIVATION: Current high-throughput protein-protein interaction (PPI) data do not provide information about the condition(s) under which the interactions occur. Thus, the identification of condition-responsive PPI sub-networks is of great importance for investigating how a living cell adapts to changing environments. RESULTS: In this article, we propose a novel edge-based scoring and searching approach to extract a PPI sub-network responsive to conditions related to some investigated gene expression profiles. Using this approach, what we constructed is a sub-network connected by the selected edges (interactions), instead of only a set of vertices (proteins) as in previous works. Furthermore, we suggest a systematic approach to evaluate the biological relevance of the identified responsive sub-network by its ability of capturing condition-relevant functional modules. We apply the proposed method to analyze a human prostate cancer dataset and a yeast cell cycle dataset. The results demonstrate that the edge-based method is able to efficiently capture relevant protein interaction behaviors under the investigated conditions. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zheng Guo 0002, Yongjin Li, Wencai Ma, Dong Wang 0011, Jing Zhu 0004, Min Zhang 0009, Da Yang 0003, Jing Wang 0004
Bioinform.6
2006 Effects of replacing the unreliable cDNA microarray measurements on the disease classification based on gene expression profiles and functional modules
abstract
MOTIVATION: Microarrays datasets frequently contain a large number of missing values (MVs), which need to be estimated and replaced for subsequent data mining. The focus of the paper is to study the effects of different MV treatments for cDNA microarray data on disease classification analysis. RESULTS: By analyzing five datasets, we demonstrate that among three kinds of classifiers evaluated in this study, support vector machine (SVM) classifiers are robust to varied MV imputation methods [e.g. replacing MVs by zero, K nearest-neighbor (KNN) imputation algorithm, local least square imputation and Bayesian principal component analysis], while the classification and regression tree classifiers are sensitive in terms of classification accuracy. The KNNclassifiers built on differentially expressed genes (DEGs) are robust to the varied MV treatments, but the performances of the KNN classifiers based on all measured genes can be significantly deteriorated when imputing MVs for genes with larger missing rate (MR) (e.g. MR > 5%). Generally, while replacing MVs by zero performs relatively poor, the other imputation algorithms have little difference in affecting classification performances of the SVM or KNN classifiers. We further demonstrate the power and feasibility of our recently proposed functional expression profile (FEP) approach as means to handle microarray data with MVs. The FEPs, which are derived from the functional modules that are enriched with sets of DEGs and thus can be consistently identified under varied MV treatments, achieve precise disease classification with better biological interpretation. We conclude that the choice of MV treatments should be determined in context of the later approaches used for disease classification. The suggested exclusion criterion of ignoring the genes with larger MR (e.g. >5%), while justifiable for some classifiers such as KNN classifiers, might not be considered as a general rule for all classifiers.
Dong Wang 0011, Yingli Lv, Zheng Guo 0002, Xia Li 0004, Jing Zhu 0004, Da Yang 0003, Jianzhen Xu, Chenguang Wang 0004, Shaoqi Rao, Baofeng Yang
Bioinform.1